Strategies for disaster recovery with IBM Storage Protect

IBM Storage Protect provides several ways to recover the server if the database or storage pools fail.

Automatic failover for disaster recovery

Automatic failover is an operation that switches to a standby system if a software, hardware, or network interruption occurs. Automatic failover is used with node replication to recover data after a system failure. Figure 1 shows the IBM Storage Protect automatic failover process.

Figure 1. Automatic failover process
Illustration shows automatic failover process

Automatic failover for data recovery occurs if the source replication server is unavailable because of a disaster or a system outage. During normal operations, when the client accesses a source replication server, the client receives connection information for the target replication server. The client node stores the failover connection information in the client options file.

During client restore operations, the server automatically changes clients from the source replication server to the target replication server and back again. Only one server per node can be used for failover protection at any time. When a new client operation is started, the client attempts to connect to the source replication server. The client resumes operations on the source server if the source replication server is available.

To use automatic failover for replicated client nodes, the source replication server, the target replication server, and the client must be at the V7.1 level or later. If any of the servers are at an earlier level, automatic failover is disabled and you must rely on a manual failover process.

Recovery of IBM Storage Protect components

The server database, recovery log, and storage pools are critical to the operation of IBM Storage Protect and must be protected. If the database is unusable, the entire server is unavailable and recovering data that is managed by the server might be difficult or impossible.

Even without the database, fragments of data or complete files might be read from storage pool volumes that are not encrypted and security can be compromised. Therefore, you must always back up the database. Also, always encrypt sensitive data by using the client or the storage device, unless the storage media is physically secured.

IBM Storage Protect provides several data protection methods, which include backing up storage pools and the database. For example, you can define schedules so that the following operations occur:
  • After the initial full backup of your storage pools, incremental storage pool backups are run every night.
  • Incremental database backups are run every night.
  • Full database backups are run once a week.
For tape-based environments, you can use disaster recovery manager (DRM) to assist you in many of the tasks that are associated with protecting and recovering data. DRM is available with IBM Storage Protect Extended Edition.

Preventive actions for recovery

Recovery is based on the following preventive actions:
  • Mirroring, by which the server maintains a copy of the active log
  • Backing up the database
  • Backing up the storage pools
  • Auditing storage pools for damaged files and recovery of damaged files when necessary
  • Backing up the device configuration and volume history files
  • Validating the data in storage pools by using cyclic redundancy checking
  • Storing the cert.kdb file in a safe place to ensure that the Secure Sockets Layer (SSL) is secure
If you are using tape for storage, you can also create a disaster recovery plan to guide you through the recovery process by using DRM. You can use the disaster recovery plan for audit purposes to certify the recoverability of the server. The disaster recovery methods of DRM are based on taking the following actions:
  • Creating a disaster recovery plan file for the server
  • Backing up server data to tape
  • Sending the server backup data to a remote site or to another server
  • Storing client system information
  • Defining and tracking the storage media that is used for storing and recovering client data
If you take preventive actions for recovery, you should have backups for the following files and directories. The files and directories that are required to restore a server and its operations:
  • Server options file (dsmserv.opt)
  • Device configuration file (for example, devconf.dat)
  • Volume history file (for example, volhist.dat)
  • Master encryption key files (dsmkeydb.kdb or dsmkeydb.sth)
  • Server certificate and private key files (cert.kbd or cert.sth)