How to perform disaster recovery with storage replication¶
Active-passive disaster recovery with storage replication requires advance preparation: you must set up storage replication between a primary and secondary LXD deployment. For details, see How to set up storage replication.
If the primary deployment becomes unavailable, you can follow these steps to fail over to the secondary deployment, and then fail back when the primary deployment comes back online.
Promote the secondary location¶
If the primary location becomes unreachable, the secondary location can be promoted to become the new source of truth. The method to promote the secondary storage array depends on the storage vendor. For links to vendor guides, see: Set up storage replication.
Important
Promoting the secondary array might result in data loss if there is data on the primary location that has not been replicated. Consult the storage vendor’s documentation for further information.
Recover resources¶
After the secondary storage array has been promoted, you can start recovering the workload. Run the steps in How to recover LXD database records for instances and custom volumes on the secondary LXD deployment.
When prompted to choose the pools to scan for unknown volumes, select the storage pool that was configured during the replication setup.
The instances and custom storage volumes are then recovered on the secondary LXD deployment. Use lxc start to bring up the instances that were originally running on the primary deployment.
Add missing storage pool¶
If the LXD storage pool at the secondary location exists only in the storage array and has not yet been created in LXD (as described in Storage pool), you must recover it first.
Use the lxc storage create command to add the storage pool. This works for both single and clustered LXD deployments. For more information, see: Create a storage pool.
Recover Ceph RBD pool¶
LXD’s Ceph RBD driver uses a placeholder volume to reserve the storage pool and ensure it isn’t used more than once. For replication, this behavior can be ignored because the replicated pool must be recovered at the secondary location. To allow this, set source.recover to ignore the placeholder volume if it was also replicated to the secondary location.
When creating the storage pool in a LXD cluster, make sure to add the source.recover=true setting when creating the pending storage pools per cluster member as this setting is cluster member specific.
Fail back to the primary location¶
Once the primary location is back online, the storage layer ensures data consistency because the secondary storage array now acts as the source of truth and no longer receives updates from the primary array. As long as this replication flow is not reversed, the running instances and custom volumes on the secondary location are protected.
Warning
Network collisions might occur if the primary location comes back online and LXD automatically starts up any instances. This issue is outside the scope of storage replication, but you must take appropriate measures to prevent such conflicts.
Service failback to the primary location can be performed in two ways. In both cases, the operations on the storage layer are identical, but the correct approach depends on the state of the instances and custom volumes on the secondary location:
Shut down the resources on the secondary location and bring them back up on the primary This approach requires that the configuration of the recovered instances and volumes on the secondary location has not been modified in any way. Any modifications would not be reflected in the database of the primary LXD deployment and might cause unexpected side effects.
Set up a fresh deployment of LXD on the primary location and repeat the steps outlined in How to recover LXD database records for instances and custom volumes. This approach repeats the same process performed for the initial disaster recovery, but in reverse.
After choosing an approach, demote the storage array at the secondary location and promote the array at the primary location. Refer to Set up storage replication for details on how to perform these actions.
Finally, either bring up the instances on the primary deployment using lxc start, or recover them first to make them known again to the primary deployment before starting them.