How to perform disaster recovery with replicators

Active-passive disaster recovery with replicators requires advance preparation: you must set up replicators on the primary cluster to regularly copy instances to a secondary cluster.

If the primary cluster becomes unavailable, you can promote your secondary cluster to take over workloads. Then, after the primary cluster comes back online, you can synchronize the clusters and resume the original replication setup.

Important

Changing replica modes does not redirect application traffic. You must manage traffic separately from LXD.

Fail over to the secondary cluster

If the primary cluster becomes unavailable, you can manually fail over to the secondary cluster.

On the secondary cluster, promote the replica project to leader mode:

lxc project promote-replica <project_name>

If the primary cluster is unreachable, promotion proceeds automatically without requiring validation. Use --force only to promote a standby project when the primary cluster is still reachable and its project is still in leader mode:

lxc project promote-replica <project_name> --force

For a planned switchover rather than a failover, do not use --force; instead, demote the leader first, then promote the standby, as described in Replicators.

Important

If the project is mirrored with Ceph RBD, LXD also promotes the project’s volumes in Ceph, and different rules apply:

  • If the primary cluster is unavailable, you must force the promotion. The primary cluster cannot demote its volumes while it is unavailable, so Ceph refuses a promotion without force.

  • If the primary cluster is available, do not force the promotion. A forced promotion makes the volumes on the two clusters diverge. The volumes on the primary cluster must then be discarded and copied again, and the changes that did not reach the secondary cluster are lost. Instead, stop the instances, run the replicator one more time, and demote the project on the primary cluster. Then promote the project on the secondary cluster without force.

Caution

Forced promotion of a standby project can create a split-brain risk: if the target of replication is still in leader mode, then both projects are writable. Any instances created, started, or modified independently on each side during that window will diverge and cannot be automatically reconciled by the next replicator run. Only force-promote in this situation when you understand and accept that risk; for a planned switchover, demote the leader first instead.

After promotion, the project on the secondary cluster becomes writable. Start the instances to resume your workloads:

lxc start --all --project <project_name>

You can then verify that the instances start successfully.

Run this command to confirm that all instances have state RUNNING:

lxc list --project <project_name>

You can also run the following command to check that custom volumes have been mounted:

lxc exec <instance_name> -- df -h

Instances may fail to boot for application-related reasons, or may require a strict boot order, which LXD does not orchestrate automatically. If a virtual machine fails to boot, you can attach its root volume to another virtual machine to investigate its contents.

Refer to the troubleshooting guides if you encounter any issues.

Important

Verifying that an instance is RUNNING does not confirm application readiness. You must manage application validation separately from LXD.

Fail back to the primary cluster

After failover, the secondary cluster manages workloads with a project in leader mode. As a result, when the primary cluster comes back online, the replica projects will be out of sync. Scheduled replicator runs on the primary cluster will fail because both projects are in leader mode. To fail back to the primary cluster, you must synchronize the projects, restore the original replica modes, and resume the original replication direction.

1. Verify cluster restoration

Once the primary cluster comes back online, you can confirm the health of the cluster.

Run this command to ensure that all cluster members are back online:

lxc cluster list

2. Synchronize projects

How you synchronize the projects depends on whether the project is mirrored with Ceph RBD.

Without Ceph RBD mirroring

You can run the replicator on the primary cluster in “restore” mode to synchronize the project on the primary cluster with the project on the secondary cluster. The replicator only runs in “restore” mode if all instances in the original source project are stopped, to prevent partial restoration of instances.

Run this command to stop running instances:

lxc stop <instance_name> [<instance_name>...] --force --project <project_name>

Demote the project on the primary cluster to standby mode:

lxc project demote-replica <project_name>

Though demotion does not require contact with the other project, a project can only be demoted to standby if replica.cluster is set, to specify the source of replication data. If this key is unset, use --force to demote the project:

lxc project demote-replica <project_name> --force

On the primary cluster, run the replicator in “restore” mode to pull data from the secondary cluster:

lxc replicator run <replicator_name> --restore

Restore mode uses the instance list from the secondary cluster as the authoritative source. Any instances created on the secondary cluster during the failover period are also created on the primary cluster automatically.

The project on the primary cluster is now a standby replica of the project on the secondary cluster (the replicator remains on the primary cluster).

With Ceph RBD mirroring

A replicator cannot run in “restore” mode for a mirrored project. To synchronize the projects, you must discard the volumes on the primary cluster, let Ceph copy volumes from the secondary to the primary cluster, and then replicate the instance and volume records from the secondary cluster.

When the primary cluster comes back online, its project is still in leader mode. LXD might also have restarted the instances that were running when the cluster became unavailable.

  1. On the primary cluster, stop all instances in the project:

    lxc stop --all --force --project <project_name>
    
  2. On the primary cluster, make sure that replica.cluster is set to the cluster link of the secondary cluster, and demote the project:

    lxc project set <project_name> replica.cluster=<secondary_cluster_link_name>
    lxc project demote-replica <project_name>
    

    This makes the project’s volumes on the primary cluster read-only.

  3. Wait until Ceph reports that the volumes have diverged from the volumes on the secondary cluster. To check, run the following command against the primary cluster’s Ceph cluster:

    rbd mirror pool status <osd_pool_name> --verbose
    

    Within about a minute of the demotion, the project’s volumes show the state up+error with the description split-brain.

  4. On the primary cluster, demote the project again:

    lxc project demote-replica <project_name>
    

    LXD discards every volume of the project that Ceph reports as split-brain, and Ceph copies it again in full from the secondary cluster. Changes that did not reach the secondary cluster before the failover are lost. The second demotion is needed only after a forced promotion, because the volumes on the primary cluster then have writes that never reached the secondary cluster. A planned switchover demotes the project before the promotion, so its volumes do not diverge and no second demotion is needed.

  5. Wait until Ceph has copied the volumes. To check the progress, run the rbd mirror pool status command again. The project’s volumes are ready when their state is up+replaying.

  6. On the secondary cluster, create a replicator that targets the primary cluster, and run it:

    lxc replicator create <replicator_name> cluster=<primary_cluster_link_name> --project <project_name>
    lxc replicator run <replicator_name> --project <project_name>
    

    The instances on the secondary cluster can keep running. If the run fails because Ceph has not finished copying the volumes, wait and run the replicator again. If the run fails because it reports a volume as split-brain, demote the project on the primary cluster again.

The project on the primary cluster is now a standby replica of the project on the secondary cluster.

Note

If you deleted an instance or a custom volume on the secondary cluster while the primary cluster was unavailable, Ceph removes its volume from the primary cluster when it copies the volumes again, but the record remains in LXD. Delete the record on the primary cluster yourself.

3. Resume original replication direction

To restore the original setup, in which the primary cluster replicates to the secondary cluster, stop any running instances in the project on the secondary cluster.

If the project is mirrored with Ceph RBD, run the replicator on the secondary cluster one more time after you have stopped the instances, so that the final state reaches the primary cluster:

lxc replicator run <replicator_name> --project <project_name>

The demotion of a mirrored project fails if any instance in the project is still running.

Next, demote the project on the secondary cluster back to standby mode:

lxc project demote-replica <project_name>

Finally, promote the project on the primary cluster back to leader mode. When a project is promoted from standby to leader, LXD uses the replica.cluster key to identify the corresponding replica project and confirm that it is no longer in leader mode. If this key is unset, set it to the name of the cluster link used for replication; then, promote the project:

lxc project promote-replica <project_name>

If the project is mirrored with Ceph RBD and the promotion fails because the demotion has not reached the primary cluster yet, wait a moment and try again.

Your original active-passive disaster recovery setup is now restored. You can restart your instances on the primary cluster and resume your scheduled replicator runs.