<a id="howto-replicators-dr"></a>

# How to perform disaster recovery with replicators

Active-passive disaster recovery with replicators requires advance preparation: you must [set up replicators](https://canonical.com/lxd/docs/latest/howto/replicators_create/index.html.md#howto-replicators-setup) on the primary cluster to regularly copy instances to a secondary cluster.

If the primary cluster becomes unavailable, you can promote your secondary cluster to take over workloads. Then, after the primary cluster comes back online, you can synchronize the clusters and resume the original replication setup.

#### IMPORTANT
Changing replica modes does not redirect application traffic. You must manage traffic separately from LXD.

## Fail over to the secondary cluster

If the primary cluster becomes unavailable, you can manually fail over to the secondary cluster.

On the secondary cluster, promote the replica project to `leader` mode:

CLI

```bash
lxc project promote-replica <project_name>
```

If the primary cluster is unreachable, promotion proceeds automatically without requiring validation. Use `--force` only to promote a standby project when the primary cluster is still reachable and its project is still in `leader` mode:

```bash
lxc project promote-replica <project_name> --force
```

For a planned switchover rather than a failover, do not use `--force`; instead, demote the leader first, then promote the standby, as described in [Replicators](https://canonical.com/lxd/docs/latest/explanation/disaster_recovery/index.html.md#exp-replicators).

UI

Select the project from the Project drop-down menu, then click Configuration in the navigation sidebar.

Select the Replication tab, then, under Replica mode, click Promote to leader.

If the primary cluster is unreachable, promotion proceeds automatically without requiring validation. Click Promote in the confirmation modal.

If the primary cluster is still reachable and its project has not been demoted to `standby`, checking Force before clicking Promote skips validation. For a planned switchover, demote the leader first and then promote without Force.

#### IMPORTANT
If the project is [mirrored with Ceph RBD](https://canonical.com/lxd/docs/latest/howto/replicators_create/index.html.md#howto-replicators-ceph), LXD also promotes the project’s volumes in Ceph, and different rules apply:

- If the primary cluster is unavailable, you must force the promotion.
  The primary cluster cannot demote its volumes while it is unavailable, so Ceph refuses a promotion without force.
- If the primary cluster is available, do not force the promotion.
  A forced promotion makes the volumes on the two clusters diverge.
  The volumes on the primary cluster must then be discarded and copied again, and the changes that did not reach the secondary cluster are lost.
  Instead, stop the instances, run the replicator one more time, and demote the project on the primary cluster.
  Then promote the project on the secondary cluster without force.

#### CAUTION
Forced promotion of a standby project can create a split-brain risk: if the target of replication is still in `leader` mode, then both projects are writable. Any instances created, started, or modified independently on each side during that window will diverge and cannot be automatically reconciled by the next replicator run. Only force-promote in this situation when you understand and accept that risk; for a planned switchover, demote the leader first instead.

After promotion, the project on the secondary cluster becomes writable. Start the instances to resume your workloads:

CLI

```bash
lxc start --all --project <project_name>
```

UI

Select Instances in the navigation sidebar.
Click the checkbox in the header row to select all instances, then click Start in the page header.
In the confirmation modal, click Start.

You can then verify that the instances start successfully.

CLI

Run this command to confirm that all instances have state `RUNNING`:

```bash
lxc list --project <project_name>
```

You can also run the following command to check that custom volumes have been mounted:

```bash
lxc exec <instance_name> -- df -h
```

UI

On the Instances page, verify that the status for each instance changes from Stopped to Running.

Instances may fail to boot for application-related reasons, or may require a strict boot order, which LXD does not orchestrate automatically. If a virtual machine fails to boot, you can attach its root volume to another virtual machine to investigate its contents.

Refer to the [troubleshooting guides](https://canonical.com/lxd/docs/latest/howto/troubleshoot/index.html.md#troubleshoot) if you encounter any issues.

#### IMPORTANT
Verifying that an instance is `RUNNING` does not confirm application readiness. You must manage application validation separately from LXD.

## Fail back to the primary cluster

After failover, the secondary cluster manages workloads with a project in `leader` mode. As a result, when the primary cluster comes back online, the replica projects will be out of sync. Scheduled replicator runs on the primary cluster will fail because both projects are in `leader` mode. To fail back to the primary cluster, you must synchronize the projects, restore the original replica modes, and resume the original replication direction.

### 1. Verify cluster restoration

Once the primary cluster comes back online, you can confirm the health of the cluster.

CLI

Run this command to ensure that all cluster members are back online:

```bash
lxc cluster list
```

UI

Click Clustering in the navigation sidebar, select Members from the expanded drop-down list, and confirm that all members have status Online.

### 2. Synchronize projects

How you synchronize the projects depends on whether the project is [mirrored with Ceph RBD](https://canonical.com/lxd/docs/latest/howto/replicators_create/index.html.md#howto-replicators-ceph).

#### Without Ceph RBD mirroring

You can run the replicator on the primary cluster in “restore” mode to synchronize the project on the primary cluster with the project on the secondary cluster. The replicator only runs in “restore” mode if all instances in the original source project are stopped, to prevent partial restoration of instances.

CLI

Run this command to stop running instances:

```bash
lxc stop <instance_name> [<instance_name>...] --force --project <project_name>
```

UI

Select Instances in the navigation sidebar.
Click the checkbox in the header row to select all instances, then click Stop in the page header.
In the confirmation modal, check Force stop then click Stop.

Demote the project on the primary cluster to `standby` mode:

CLI

```bash
lxc project demote-replica <project_name>
```

Though demotion does not require contact with the other project, a project can only be demoted to `standby` if [`replica.cluster`](https://canonical.com/lxd/docs/latest/reference/projects/index.html.md#project-replica:replica.cluster) is set, to specify the source of replication data. If this key is unset, use `--force` to demote the project:

```bash
lxc project demote-replica <project_name> --force
```

UI

Select the project from the Project drop-down menu, then click Configuration in the navigation sidebar.

Select the Replication tab, then, under Replica mode, click Demote to standby.

A project can only be demoted if [`replica.cluster`](https://canonical.com/lxd/docs/latest/reference/projects/index.html.md#project-replica:replica.cluster) is set, to specify the source of replication data. If this key is unset, check Force before clicking Demote.

On the primary cluster, run the replicator in “restore” mode to pull data from the secondary cluster:

CLI

```bash
lxc replicator run <replicator_name> --restore
```

UI

Click Clustering in the navigation sidebar, then select Replicators from the expanded drop-down list.

Click on the run button <span class='guilabel'><svg width='16' height='16' xmlns='http://www.w3.org/2000/svg'><path fill='currentColor' fill-rule='evenodd' clip-rule='evenodd' d='M11.3844 7.98987L5.50005 3.87809L5.50005 12.1161L11.3844 7.98987ZM12.538 9.01296C13.2485 8.51475 13.2476 7.46192 12.5363 6.96488L5.96604 2.37377C5.13747 1.7948 4.00005 2.3876 4.00005 3.3984L4.00005 12.5968C4.00005 13.6085 5.13935 14.2011 5.96773 13.6202L12.538 9.01296Z'/></svg></span> at the end of the replicator’s row.

Alternatively, click on a replicator name to view its detail page, then click on the Restore button in the header.

In the confirmation modal, check Overwrite local data, then click Restore.

Restore mode uses the instance list from the secondary cluster as the authoritative source. Any instances created on the secondary cluster during the failover period are also created on the primary cluster automatically.

The project on the primary cluster is now a standby replica of the project on the secondary cluster (the replicator remains on the primary cluster).

#### With Ceph RBD mirroring

A replicator cannot run in “restore” mode for a mirrored project.
To synchronize the projects, you must discard the volumes on the primary cluster, let Ceph copy volumes from the secondary to the primary cluster, and then replicate the instance and volume records from the secondary cluster.

When the primary cluster comes back online, its project is still in `leader` mode.
LXD might also have restarted the instances that were running when the cluster became unavailable.

1. On the primary cluster, stop all instances in the project:
   ```bash
   lxc stop --all --force --project <project_name>
   ```
2. On the primary cluster, make sure that [`replica.cluster`](https://canonical.com/lxd/docs/latest/reference/projects/index.html.md#project-replica:replica.cluster) is set to the cluster link of the secondary cluster, and demote the project:
   ```bash
   lxc project set <project_name> replica.cluster=<secondary_cluster_link_name>
   lxc project demote-replica <project_name>
   ```

   This makes the project’s volumes on the primary cluster read-only.
3. Wait until Ceph reports that the volumes have diverged from the volumes on the secondary cluster.
   To check, run the following command against the primary cluster’s Ceph cluster:
   ```bash
   rbd mirror pool status <osd_pool_name> --verbose
   ```

   Within about a minute of the demotion, the project’s volumes show the state `up+error` with the description `split-brain`.
4. On the primary cluster, demote the project again:
   ```bash
   lxc project demote-replica <project_name>
   ```

   LXD discards every volume of the project that Ceph reports as `split-brain`, and Ceph copies it again in full from the secondary cluster.
   Changes that did not reach the secondary cluster before the failover are lost.
   The second demotion is needed only after a forced promotion, because the volumes on the primary cluster then have writes that never reached the secondary cluster.
   A planned switchover demotes the project before the promotion, so its volumes do not diverge and no second demotion is needed.
5. Wait until Ceph has copied the volumes.
   To check the progress, run the `rbd mirror pool status` command again.
   The project’s volumes are ready when their state is `up+replaying`.
6. On the secondary cluster, create a replicator that targets the primary cluster, and run it:
   ```bash
   lxc replicator create <replicator_name> cluster=<primary_cluster_link_name> --project <project_name>
   lxc replicator run <replicator_name> --project <project_name>
   ```

   The instances on the secondary cluster can keep running.
   If the run fails because Ceph has not finished copying the volumes, wait and run the replicator again.
   If the run fails because it reports a volume as `split-brain`, demote the project on the primary cluster again.

The project on the primary cluster is now a standby replica of the project on the secondary cluster.

#### NOTE
If you deleted an instance or a custom volume on the secondary cluster while the primary cluster was unavailable, Ceph removes its volume from the primary cluster when it copies the volumes again, but the record remains in LXD.
Delete the record on the primary cluster yourself.

### 3. Resume original replication direction

To restore the original setup, in which the primary cluster replicates to the secondary cluster, stop any running instances in the project on the secondary cluster.

If the project is mirrored with Ceph RBD, run the replicator on the secondary cluster one more time after you have stopped the instances, so that the final state reaches the primary cluster:

```bash
lxc replicator run <replicator_name> --project <project_name>
```

The demotion of a mirrored project fails if any instance in the project is still running.

Next, demote the project on the secondary cluster back to `standby` mode:

CLI

```bash
lxc project demote-replica <project_name>
```

UI

Select the project from the Project drop-down menu, then click Configuration in the navigation sidebar.

Select the Replication tab, then, under Replica mode, click Demote to standby.

Finally, promote the project on the primary cluster back to `leader` mode. When a project is promoted from `standby` to `leader`, LXD uses the [`replica.cluster`](https://canonical.com/lxd/docs/latest/reference/projects/index.html.md#project-replica:replica.cluster) key to identify the corresponding replica project and confirm that it is no longer in `leader` mode. If this key is unset, set it to the name of the cluster link used for replication; then, promote the project:

CLI

```bash
lxc project promote-replica <project_name>
```

UI

Select the project from the Project drop-down menu, then click Configuration in the navigation sidebar.

Select the Replication tab, then, under Replica mode, click Promote to leader.

If the project is mirrored with Ceph RBD and the promotion fails because the demotion has not reached the primary cluster yet, wait a moment and try again.

Your original active-passive disaster recovery setup is now restored. You can restart your instances on the primary cluster and resume your scheduled replicator runs.

## Related topics

How-to guides:

* [How to set up replicators](https://canonical.com/lxd/docs/latest/howto/replicators_create/index.html.md#howto-replicators-setup)
* [How to manage replicators](https://canonical.com/lxd/docs/latest/howto/replicators_manage/index.html.md#howto-replicators-manage)
* [How to perform disaster recovery with storage replication](https://canonical.com/lxd/docs/latest/howto/disaster_recovery_replication/index.html.md#disaster-recovery-replication)
