<a id="exp-disaster-recovery"></a>

# Disaster recovery

LXD provides support for different approaches to active-passive disaster recovery: a strategy that allows for the restoration of workloads on a secondary cluster after a disaster.

#### NOTE
Backups of standalone LXD servers can also protect against data loss.
For details about different backup methods, see [How to back up a LXD server](https://canonical.com/lxd/docs/latest/backup/index.html.md#backups).

<a id="exp-disaster-recovery-concepts"></a>

## Disaster recovery concepts

Active-passive disaster recovery
: This strategy requires the deployment of two LXD clusters with the same infrastructure and configuration.
  Under this setup, a primary (active) cluster manages workloads, and a secondary (passive) cluster only becomes active if the primary cluster fails.
  Preparation for active-passive disaster recovery requires the periodic replication of data from the primary cluster on the secondary cluster.
  Infrastructure as code tools () can also facilitate the consistent deployment and configuration of LXD and required infrastructure.
  The [LXD Terraform provider](https://registry.terraform.io/providers/terraform-lxd/lxd/latest/docs), for example, is an IaC tool that manages LXD resources.

Recovery point objective ()
: The maximum acceptable amount of time since the last data recovery point.
  The RPO determines expectations for an acceptable loss of data.

Recovery time objective ()
: The maximum acceptable delay between the interruption of services and restoration of service.
  The RTO determines an acceptable length of time for service downtime.

<a id="exp-disaster-recovery-monitoring"></a>

## Monitoring

LXD collects metrics and streams events.
Setting up external observability systems in advance to collect this data can assist with disaster detection and evaluation.
Refer to the [metrics reference](https://canonical.com/lxd/docs/latest/reference/provided_metrics/index.html.md#provided-metrics) and [events reference](https://canonical.com/lxd/docs/latest/events/index.html.md#events) for details about the data produced by LXD.
For example, [Replicator metrics](https://canonical.com/lxd/docs/latest/reference/provided_metrics/index.html.md#replicator-metrics) can be used to determine the current RPO of clusters that use replicators for data replication.
For information about how to gather and store metrics and logs, see [How to monitor metrics](https://canonical.com/lxd/docs/latest/metrics/index.html.md#metrics), [How to send logs to Loki](https://canonical.com/lxd/docs/latest/howto/logs_loki/index.html.md#logs-loki), and [Set up a Grafana dashboard](https://canonical.com/lxd/docs/latest/howto/grafana/index.html.md#grafana).

<a id="exp-disaster-recovery-approaches"></a>

## Active-passive disaster recovery

LXD supports two approaches to active-passive disaster recovery:

LXD replicators
: Replicators are LXD entities that use cluster links to periodically copy instances from a project on one cluster to a project on another cluster.
  In the event of a disaster at the primary location, you can manage failover to the secondary cluster through LXD.
  Once the primary cluster comes back online, you can use replicators to restore the original replication direction.

Storage replication
: Data replication at the storage layer is possible when using remote storage drivers that support volume recovery.
  In this setup, you must configure replication separately from LXD, through the remote storage provider.
  Likewise, in the event of a disaster, you must manage promotion of the secondary cluster through the remote storage provider.

|                           | Replicators                                                                                                       | Storage replication                                                     |
|---------------------------|-------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------|
| **Level**                 | LXD instance layer                                                                                                | Storage array layer                                                     |
| **Data replication**      | Incremental instance refresh over cluster links                                                                   | Vendor storage replication, such as Ceph RBD mirroring or PowerFlex RCG |
| **Scheduling**            | Controlled by LXD ([`schedule`](https://canonical.com/lxd/docs/latest/reference/replicator_config/index.html.md#replicator-conf:schedule) config key) | Controlled by the storage vendor                                        |
| **Requires cluster link** | Yes                                                                                                               | No                                                                      |
| **Recovery method**       | Promote standby project with LXD CLI or UI                                                                        | Promote storage array, then use the `lxd recover` command               |
| **Snapshot support**      | Automatic pre-replication snapshots                                                                               | Depends on storage vendor                                               |

<a id="exp-replicators"></a>

## Replicators

You can use LXD replicators to manage replication, failover, and failback end-to-end with LXD, without dependency on a specific storage backend.

<a id="exp-replicators-concepts"></a>

### Leader and standby projects

Replication is configured at the project level. You must create projects with the same name on both clusters, and then configure the replica mode of each project:

- `leader`: The project is writable.
  Instances in this project are the source of replication.
  The replicator runs from the cluster with this project.
- `standby`: Instances in this project are replicas, kept in sync by the replicator.
  New instances cannot be created directly in this project, and existing instances cannot be started.
  The project must be promoted to `leader` during a failover before instances can be started.
- (empty): The project is not part of any replication setup.
  This is the default for new projects.

The replica mode is not a configuration key. For details about the dedicated CLI commands and UI processes you must use to manage the replica mode, refer to [How to perform disaster recovery with replicators](https://canonical.com/lxd/docs/latest/howto/disaster_recovery_replicators/index.html.md#howto-replicators-dr).

Clearing the replica mode of a standby project must be forced, because it drops the record of which cluster was replicating into it. Once the replica mode is unset, the project could then be promoted without checking that the cluster replicating into the project has stepped down. If that cluster’s project is still in `leader` mode, both projects become writable at the same time, and changes made independently on each side diverge and cannot be reconciled by the next replicator run.

The [`replica.cluster`](https://canonical.com/lxd/docs/latest/reference/projects/index.html.md#project-replica:replica.cluster) configuration key identifies the cluster link that is allowed to push replication data into a standby project. It is required on the standby project, and it must also be set on the leader project if you intend to fail over and later return to the original replication direction: after a failover the original leader becomes a standby, and it can only be promoted back to leader once LXD can identify the cluster it was replicating with.

A project can only be promoted or demoted if it takes part in a replication topology. Demotion requires the `replica.cluster` key, because a standby project cannot accept replication data without it. Promoting a project that has no replica mode set requires at least one replicator, and promoting a standby requires the `replica.cluster` key, which is what identifies the cluster whose project must have stepped down first. Forced promotion and demotion override these checks.

To swap the roles of two clusters in a planned switchover, demote the current leader first, then promote the standby. Demoting first is always safe: the topology is briefly left without a leader, which only pauses writes; promoting first, however, would allow both clusters to accept writes at the same time.

The leader project pushes its instances to the standby project over the cluster link.
The standby project mirrors the leader at the time of the last replicator run.

<a id="exp-replicators-how"></a>

### How replication works

When a replicator runs, LXD performs an incremental refresh of every instance in the leader project to the standby project, together with the custom storage volumes attached only to that instance. Instances and volumes that do not yet exist on the standby are created; existing ones are updated to match the leader’s current state. A custom volume attached to more than one instance is not replicated and must exist on the standby before the instances using it can be replicated. A volume attached through a profile is replicated when one instance alone uses it. Because a profile device must point at an existing volume, create the volume on the standby and add the device to the standby’s copy of the profile before the first run; the run then refreshes it. Without the device the run refuses the instance before any data is sent. Custom volumes are only replicated when the project has `features.storage.volumes=true`; projects that inherit volumes from the default project have no project-local custom volumes to replicate.

Before each refresh, LXD creates a point-in-time snapshot of each instance on the leader, capturing its root disk and the custom volumes attached only to it at the same moment. This provides a consistent rollback point on the source cluster in case anything goes wrong during replication, and the volume snapshots travel to the standby with the volumes. The exception is instances that already have a [`snapshots.schedule`](https://canonical.com/lxd/docs/latest/reference/instance_options/index.html.md#instance-snapshots:snapshots.schedule) configured and no custom volume attached: their scheduled snapshots already provide point-in-time history, so LXD skips the extra snapshot to avoid redundancy. Scheduled snapshots capture the root disk alone, so an instance with custom volumes always gets a snapshot here.

Replication can be triggered manually through the LXD CLI or UI, or scheduled automatically using a cron expression in the [`schedule`](https://canonical.com/lxd/docs/latest/reference/replicator_config/index.html.md#replicator-conf:schedule) configuration key.

<a id="exp-replicators-failover"></a>

### Failover and recovery

If the leader cluster fails, the standby project can be promoted through the LXD CLI or UI. This makes the project writable and allows instances to be started. If the leader cluster is unreachable, validation against it is skipped automatically. Forced promotion skips all validation without attempting to connect, which is useful when the leader is known to be down. Do not use it for a planned switchover: demote the leader first, then promote the standby.

When the primary cluster comes back online, you can synchronize the projects by running the replicator in “restore” mode; then you can demote the project on the secondary cluster and promote the project on the primary cluster to return the projects to their original roles. In restore mode, the instance list on the secondary cluster is used as the authoritative source: instances that were created on the secondary cluster after failover are also created on the recovering cluster, not just the instances that existed before the failure.

See [How to perform disaster recovery with replicators](https://canonical.com/lxd/docs/latest/howto/disaster_recovery_replicators/index.html.md#howto-replicators-dr) for step-by-step instructions.

<a id="exp-storage-replication"></a>

## Storage replication

Replication at the storage array layer is possible with remote storage drivers that support volume recovery.
You must configure replication outside of LXD, through a process that depends on the remote storage vendor.

For detailed instructions and information about storage providers that support storage replication, refer to [How to perform disaster recovery with storage replication](https://canonical.com/lxd/docs/latest/howto/disaster_recovery_replication/index.html.md#disaster-recovery-replication).

<a id="exp-storage-failover"></a>

### Failover with storage replication

In the event of a disaster, you must manage promotion of the secondary cluster outside of LXD, through the remote storage provider.
Once you have promoted the secondary cluster, you can use the LXD recovery tool (`lxd recover`) to recover instances and custom volumes from the replicated data.
Infrastructure as code () can facilitate redeployment of the original resources and configuration.

For instructions on how to use the recovery tool, see [How to recover LXD database records for instances and custom volumes](https://canonical.com/lxd/docs/latest/howto/database_recovery/index.html.md#disaster-recovery).

<a id="exp-disaster-recovery-lxd-recover"></a>

## LXD database recovery with `lxd recover`

The LXD recovery tool, `lxd recover`, facilitates recovery of LXD instances and custom volumes when a LXD database is lost or corrupted.
You can also use this tool to recover resources when [performing disaster recovery with storage replication](https://canonical.com/lxd/docs/latest/howto/disaster_recovery_replication/index.html.md#disaster-recovery-replication).

When you run `lxd recover`, the recovery tool scans all storage pools that exist in the database and identifies missing volumes that can be recovered.
The tool also scans volumes on the known storage pools and, in the process, may discover additional storage pools that exist on disk but are missing from the LXD database.
In such cases, the tool prints information about the storage pools so that you can re-create their database records manually.
Concrete recovery examples for each storage driver can be found in [Recover a storage pool](https://canonical.com/lxd/docs/latest/howto/storage_pools/index.html.md#howto-storage-pools-recover).
The tool then mounts any unmounted storage pools and continues scanning for volumes that may be associated with LXD.

Through this scan, the recovery tool can identify some custom volumes by name.
Some [remote storage drivers](https://canonical.com/lxd/docs/latest/reference/storage_drivers/index.html.md#storage-drivers-remote), however, such as the [PowerFlex](https://canonical.com/lxd/docs/latest/reference/storage_powerflex/index.html.md#storage-powerflex), [PowerStore](https://canonical.com/lxd/docs/latest/reference/storage_powerstore/index.html.md#storage-powerstore), and [Pure](https://canonical.com/lxd/docs/latest/reference/storage_pure/index.html.md#storage-pure) drivers, use transformed volume names, and the recovery tool is unable to discover these volumes from their name alone.
Instead, these volumes can only be discovered if they are attached to an instance.

LXD maintains a `backup.yaml` file in each instance’s storage volume, which contains all necessary information to recover a given instance.
The recovery tool compares the `backup.yaml` file with what is actually on disk (such as matching snapshots) and, if this consistency check passes, re-creates the database records.
The tool can also use the `backup.yaml` file to gather information about profiles, storage pools, and attached devices (such as storage volumes with transformed names).
Based on this information, the tool will prompt you to re-create missing entities, but it will not display information about how those entities were configured, unless they are storage pools.
For example, if an instance used a bridge network attached to a profile, the tool will notify you that the two entities are missing from the database, but it will not direct you to attach the network to the profile.

## Related topics

How-to guides:

- [How to set up replicators](https://canonical.com/lxd/docs/latest/howto/replicators_create/index.html.md#howto-replicators-setup)
- [How to manage replicators](https://canonical.com/lxd/docs/latest/howto/replicators_manage/index.html.md#howto-replicators-manage)
- [How to perform disaster recovery with replicators](https://canonical.com/lxd/docs/latest/howto/disaster_recovery_replicators/index.html.md#howto-replicators-dr)
- [How to perform disaster recovery with storage replication](https://canonical.com/lxd/docs/latest/howto/disaster_recovery_replication/index.html.md#disaster-recovery-replication)
- [How to recover LXD database records for instances and custom volumes](https://canonical.com/lxd/docs/latest/howto/database_recovery/index.html.md#disaster-recovery)
- [How to monitor metrics](https://canonical.com/lxd/docs/latest/metrics/index.html.md#metrics)
- [How to send logs to Loki](https://canonical.com/lxd/docs/latest/howto/logs_loki/index.html.md#logs-loki)
- [Set up a Grafana dashboard](https://canonical.com/lxd/docs/latest/howto/grafana/index.html.md#grafana)

Reference:

- [Replicator configuration](https://canonical.com/lxd/docs/latest/reference/replicator_config/index.html.md#ref-replicator-config)
- [Cluster links](https://canonical.com/lxd/docs/latest/explanation/clusters/index.html.md#exp-cluster-links)
- [Provided metrics](https://canonical.com/lxd/docs/latest/reference/provided_metrics/index.html.md#provided-metrics)
- [Events](https://canonical.com/lxd/docs/latest/events/index.html.md#events)
