How to set up storage replication¶
You can prepare for active-passive disaster recovery by setting up a secondary LXD deployment in a different location that can take over running workloads if the primary deployment (a non-clustered LXD server or an entire cluster) goes offline or becomes unreachable.
If such an incident occurs, you can rely on the storage layer that replicates all instances and custom volumes to the secondary location. You can then consolidate the storage layer and recover the resources to make them available to your secondary deployment (see How to recover LXD database records for instances and custom volumes).
This requires not only two separate LXD deployments, but also storage replication configuration for the respective storage array. Configure each deployment to operate independently and use only its own co-located storage array.
Note
Recovery with storage replication is only possible when using remote Storage drivers that support volume recovery (see Driver and feature comparison). Configuring replication on the storage array is out of scope for LXD and highly dependent on how each vendor implements replication.
This how-to guide focuses on the steps performed within LXD and mentions storage array requirements where applicable.
Set up entities at each location¶
Before you can set up storage replication, you must set up the required entities at each location.
Storage pool¶
Ensure that both the primary and secondary LXD deployments have a storage pool on their respective storage arrays that can later be used for replication.
If you need to create a storage pool at either location, see: Create a storage pool.
Networks and profiles¶
You might also want to set up other entities, such as networks and profiles, in advance on the secondary location. This way, in the event of a disaster, you can focus on recovering the volumes.
When performing How to recover LXD database records for instances and custom volumes, LXD checks if the required entities are present and notifies you if anything is missing. The recovery does not create these entities.
Set up storage replication¶
Replication must be configured outside of LXD, according to the concepts and constructs by the storage vendor.
The following links lead to replication setup guides published by various storage vendors:
Ceph RBD: RBD mirroring
Dell PowerFlex: Introduction to Replication
For Ceph RBD, LXD replicators can also manage the mirroring, failover, and failback of a project. See How to set up replicators.
Once you have configured the connection between the primary and secondary storage arrays, follow the relevant storage vendor’s steps to set up the actual replication of volumes.
Some vendors (such as Dell) use a concept called replication consistency group (RCG), which allows consistent replication of a group of volumes. An RCG can contain an instance’s volume along with all of its attached custom volumes. Other vendors might use different concepts.
Known storage array limitations¶
When setting up replication, consider the following limitations:
PowerFlex¶
- Cannot replicate and recover volumes with snapshots
In PowerFlex, a volume’s snapshot appears as its own volume but is still logically connected to its parent volume (vTree). When replicating a volume inside an RCG, its snapshots are not replicated; this causes inconsistencies on the secondary location. A volume’s snapshot can be replicated but will be placed inside a new vTree, losing the logical relation to its parent volume. During recovery, LXD notices this inconsistency and raises an error.
Ceph RBD¶
- Cannot use journaling mode
On Ceph RBD storage arrays, it’s possible to configure mirroring using either journaling or snapshot mode. However, with LXD, only snapshot mode is supported. This is because the volumes need to be mapped to the host for read access during recovery, which might not be possible due to missing kernel features.
Verify replication¶
After setting up storage replication, confirm that the primary location’s volumes are successfully replicated to the secondary location.
Important
Check replication regularly. If the replication fails to run, you are at risk of losing data whenever the primary location experiences an outage.