Skip to content

Operations

The failure-drill runbook for the HCI cluster. Run all of these before production, from the noted node — the expected result is in brackets; anything else, stop and diagnose. The point of the whole exercise is to prove that the ZFS mirrors, 15-minute replication + HA, dual corosync rings and MLAG bonds actually survive the failures they were designed for.

Baseline first

On each node, note the numbers — they're what "normal" looks like later:

fio --name=randwrite --filename=/vmdata/fio.test --size=8G --rw=randwrite --bs=16k \
    --iodepth=32 --ioengine=libaio --direct=1 --runtime=60 --time_based --group_reporting
fio --name=randread --filename=/vmdata/fio.test --size=8G --rw=randread --bs=16k \
    --iodepth=32 --ioengine=libaio --direct=1 --runtime=60 --time_based --group_reporting
rm /vmdata/fio.test
zpool status -x ; pvecm status ; corosync-cfgtool -s ; pvesr status
# [all pools are healthy; 3 nodes quorate; both links; every replication job OK]

For the drills, create a small test VM on vm1, replicate it to vm2 (pvesr create-local-job <vmid>-0 vm2 --schedule '*/15'), and put it under HA with a strict node-affinity rule for vm1,vm2 — see Replication + HA.

1. Bond-member failure

Worth testing per bond (guest, backup, migrate):

ip link set nic10 down                # a bond-migrate member — or pull the fibre, a better test
cat /proc/net/bonding/bond-migrate    # [1 slave down, bond still up]
ping -M do -s 8972 10.128.11.12       # [no loss]
ip link set nic10 up                  # [slave rejoins aggregate]

2. Whole-switch failure (the MLAG test)

Power off / reboot VM-A (sw-vm-1):

watch cat /proc/net/bonding/bond-guest     # [degrades to 1 slave, stays up] — same for bond-backup
corosync-cfgtool -s                        # [link0 (950) down — ring0 homes to VM-A; link1 connected, quorate]
Then VM-B (bonds degrade, both rings stay up).

Power off / reboot Storage-A (sw-storage-1):

watch cat /proc/net/bonding/bond-migrate   # [degrades to 1 slave, stays up]
pvesr schedule-now <vmid>-0                # [replication still completes over the surviving leg]
corosync-cfgtool -s                        # [both links connected, unless this switch carries ring1]
Then Storage-B (sw-storage-2): bond-migrate degrades the same way, and if it's the switch carrying corosync ring1, link1 (951) drops → still quorate on 950. Power back after each, confirm slaves + ring rejoin.

3. Corosync single-ring failure

ip link set nic1 down   # ring0 (vm3: nic4)
corosync-cfgtool -s     # [link0 down, link1 connected]
pvecm status            # [still quorate, no fencing]
ip link set nic1 up

4. Disk failure + resilver

Take one side of a mirror offline while the test VM writes:

zpool offline vmdata /dev/disk/by-id/wwn-<sdc>
zpool status vmdata     # [DEGRADED, that mirror running on one disk, VM I/O continues]
zpool online vmdata /dev/disk/by-id/wwn-<sdc>
zpool status vmdata     # [resilvers only what changed while offline, back to ONLINE]

A real failure is a replacement, not an online:

zpool replace vmdata /dev/disk/by-id/wwn-<failed> /dev/disk/by-id/wwn-<new>
zpool status vmdata     # [resilvering from the surviving mirror partner; ONLINE when done]
A resilver on a 2.4 TB 10k disk takes hours and loads the partner disk. That mirror has no redundancy until it finishes, so replace failed disks promptly.

5. Replication

pvesr status                    # [test job OK, LastSync < 15 min ago]
pvesr schedule-now <vmid>-0     # force a sync
pvesr status                    # [LastSync updates, Duration small — only changes moved]
zfs list -t snapshot -r vmdata | grep __replicate_   # [one replication snapshot per disk, on both nodes]

6. Planned node reboot

With HA maintenance mode, HA-managed VMs leave the node by themselves:

ha-manager crm-command node-maintenance enable vm1
# replicated HA VMs migrate to their partner — quick, only the delta since the last sync moves
iftop -i bond-migrate           # [traffic on 10.128.11.x]
Local-only VMs have no replica, so for each one either:

  • migrate it with a full disk copy over 911 — qm migrate <vmid> vm2 --online --with-local-disks (slow for big disks: every block is read off HDDs), or
  • shut it down for the duration of the reboot.

Talos VMs are neither: drain the Kubernetes node, shut the VM down, and start it again after the reboot — see Storage on Talos.

reboot
# during: pvecm status [2/3 quorate]; HA VMs running on partners; pvesr jobs to vm1 fail until it returns (expected)
# after:
ha-manager crm-command node-maintenance disable vm1
pvesr status                    # [jobs catch up, all OK]

7. Hard node kill + HA (do last)

With the replicated HA test VM running on vm1, pull power on vm1 from the PDU (via OOB).

[~2–3 min: node fenced, VM restarts on vm2 from the last replication snapshot — anything written after that sync is gone; cluster quorate 2/3.] Check the VM's data is at the last-sync point, not corrupt.

Power vm1 back:

pvesr status    # [job now runs vm2 → vm1 — direction reversed automatically]
The old copy on vm1 is overwritten by the next sync. That's correct: the VM carried on from the replica on vm2.

Routine: planned reboots

The everyday version of drill 6: ha-manager crm-command node-maintenance enable <node> → deal with the local-only and Talos VMs → reboot → node-maintenance disable → pvesr status all OK. Patch one node at a time, never all three at once.