How to perform disaster recovery with replicators

Active-passive disaster recovery with replicators requires advance preparation: you must set up replicators on the primary cluster to regularly copy instances to a secondary cluster.

If the primary cluster becomes unavailable, you can promote your secondary cluster to take over workloads. Then, after the primary cluster comes back online, you can synchronize the clusters and resume the original replication setup.

Important

Changing replica modes does not redirect application traffic. You must manage traffic separately from LXD.

Fail over to the secondary cluster

If the primary cluster becomes unavailable, you can manually fail over to the secondary cluster.

On the secondary cluster, promote the replica project to leader mode:

lxc project promote-replica <project_name>

If the primary cluster is unreachable, promotion proceeds automatically without requiring validation. Use --force only to promote a standby project when the primary cluster is still reachable and its project is still in leader mode:

lxc project promote-replica <project_name> --force

For a planned switchover rather than a failover, do not use --force; instead, demote the leader first, then promote the standby, as described in Replicators.

Caution

Forced promotion of a standby project can create a split-brain risk: if the target of replication is still in leader mode, then both projects are writable. Any instances created, started, or modified independently on each side during that window will diverge and cannot be automatically reconciled by the next replicator run. Only force-promote in this situation when you understand and accept that risk; for a planned switchover, demote the leader first instead.

After promotion, the project on the secondary cluster becomes writable. Start the instances to resume your workloads:

lxc start --all --project <project_name>

You can then verify that the instances start successfully.

Run this command to confirm that all instances have state RUNNING:

lxc list --project <project_name>

You can also run the following command to check that custom volumes have been mounted:

lxc exec <instance_name> -- df -h

Instances may fail to boot for application-related reasons, or may require a strict boot order, which LXD does not orchestrate automatically. If a virtual machine fails to boot, you can attach its root volume to another virtual machine to investigate its contents.

Refer to the troubleshooting guides if you encounter any issues.

Important

Verifying that an instance is RUNNING does not confirm application readiness. You must manage application validation separately from LXD.

Fail back to the primary cluster

After failover, the secondary cluster manages workloads with a project in leader mode. As a result, when the primary cluster comes back online, the replica projects will be out of sync. Scheduled replicator runs on the primary cluster will fail because both projects are in leader mode. To fail back to the primary cluster, you must synchronize the projects, restore the original replica modes, and resume the original replication direction.

1. Verify cluster restoration

Once the primary cluster comes back online, you can confirm the health of the cluster.

Run this command to ensure that all cluster members are back online:

lxc cluster list

2. Synchronize projects

You can run the replicator on the primary cluster in “restore” mode to synchronize the project on the primary cluster with the project on the secondary cluster. The replicator only runs in “restore” mode if all instances in the original source project are stopped, to prevent partial restoration of instances.

Run this command to stop running instances:

lxc stop <instance_name> [<instance_name>...] --force --project <project_name>

Demote the project on the primary cluster to standby mode:

lxc project demote-replica <project_name>

Though demotion does not require contact with the other project, a project can only be demoted to standby if replica.cluster is set, to specify the source of replication data. If this key is unset, use --force to demote the project:

lxc project demote-replica <project_name> --force

On the primary cluster, run the replicator in “restore” mode to pull data from the secondary cluster:

lxc replicator run <replicator_name> --restore

Restore mode uses the instance list from the secondary cluster as the authoritative source. Any instances created on the secondary cluster during the failover period are also created on the primary cluster automatically.

The project on the primary cluster is now a standby replica of the project on the secondary cluster (the replicator remains on the primary cluster).

3. Resume original replication direction

To restore the original setup, in which the primary cluster replicates to the secondary cluster, stop any running instances in the project on the secondary cluster. Next, demote the project on the secondary cluster back to standby mode:

lxc project demote-replica <project_name>

Finally, promote the project on the primary cluster back to leader mode. When a project is promoted from standby to leader, LXD uses the replica.cluster key to identify the corresponding replica project and confirm that it is no longer in leader mode. If this key is unset, set it to the name of the cluster link used for replication; then, promote the project:

lxc project promote-replica <project_name>

Your original active-passive disaster recovery setup is now restored. You can restart your instances on the primary cluster and resume your scheduled replicator runs.