Operations¶
The failure-drill runbook for the HCI cluster. Run all of these before production, from the noted node — the expected result is in brackets; anything else, stop and diagnose. The point of the whole exercise is to prove that the ZFS mirrors, 15-minute replication + HA, dual corosync rings and MLAG bonds actually survive the failures they were designed for.
Baseline first¶
On each node, note the numbers — they're what "normal" looks like later:
fio --name=randwrite --filename=/vmdata/fio.test --size=8G --rw=randwrite --bs=16k \
--iodepth=32 --ioengine=libaio --direct=1 --runtime=60 --time_based --group_reporting
fio --name=randread --filename=/vmdata/fio.test --size=8G --rw=randread --bs=16k \
--iodepth=32 --ioengine=libaio --direct=1 --runtime=60 --time_based --group_reporting
rm /vmdata/fio.test
zpool status -x ; pvecm status ; corosync-cfgtool -s ; pvesr status
# [all pools are healthy; 3 nodes quorate; both links; every replication job OK]
For the drills, create a small test VM on vm1, replicate it to vm2
(pvesr create-local-job <vmid>-0 vm2 --schedule '*/15'), and put it under HA with a
strict node-affinity rule for vm1,vm2 — see
Replication + HA.
1. Bond-member failure¶
Worth testing per bond (guest, backup, migrate):
ip link set nic10 down # a bond-migrate member — or pull the fibre, a better test
cat /proc/net/bonding/bond-migrate # [1 slave down, bond still up]
ping -M do -s 8972 10.128.11.12 # [no loss]
ip link set nic10 up # [slave rejoins aggregate]
2. Whole-switch failure (the MLAG test)¶
Power off / reboot VM-A (sw-vm-1):
watch cat /proc/net/bonding/bond-guest # [degrades to 1 slave, stays up] — same for bond-backup
corosync-cfgtool -s # [link0 (950) down — ring0 homes to VM-A; link1 connected, quorate]
Power off / reboot Storage-A (sw-storage-1):
watch cat /proc/net/bonding/bond-migrate # [degrades to 1 slave, stays up]
pvesr schedule-now <vmid>-0 # [replication still completes over the surviving leg]
corosync-cfgtool -s # [both links connected, unless this switch carries ring1]
sw-storage-2): bond-migrate degrades the same way, and if it's
the switch carrying corosync ring1, link1 (951) drops → still quorate on 950. Power
back after each, confirm slaves + ring rejoin.
3. Corosync single-ring failure¶
ip link set nic1 down # ring0 (vm3: nic4)
corosync-cfgtool -s # [link0 down, link1 connected]
pvecm status # [still quorate, no fencing]
ip link set nic1 up
4. Disk failure + resilver¶
Take one side of a mirror offline while the test VM writes:
zpool offline vmdata /dev/disk/by-id/wwn-<sdc>
zpool status vmdata # [DEGRADED, that mirror running on one disk, VM I/O continues]
zpool online vmdata /dev/disk/by-id/wwn-<sdc>
zpool status vmdata # [resilvers only what changed while offline, back to ONLINE]
A real failure is a replacement, not an online:
zpool replace vmdata /dev/disk/by-id/wwn-<failed> /dev/disk/by-id/wwn-<new>
zpool status vmdata # [resilvering from the surviving mirror partner; ONLINE when done]
5. Replication¶
pvesr status # [test job OK, LastSync < 15 min ago]
pvesr schedule-now <vmid>-0 # force a sync
pvesr status # [LastSync updates, Duration small — only changes moved]
zfs list -t snapshot -r vmdata | grep __replicate_ # [one replication snapshot per disk, on both nodes]
6. Planned node reboot¶
With HA maintenance mode, HA-managed VMs leave the node by themselves:
ha-manager crm-command node-maintenance enable vm1
# replicated HA VMs migrate to their partner — quick, only the delta since the last sync moves
iftop -i bond-migrate # [traffic on 10.128.11.x]
- migrate it with a full disk copy over 911 —
qm migrate <vmid> vm2 --online --with-local-disks(slow for big disks: every block is read off HDDs), or - shut it down for the duration of the reboot.
Talos VMs are neither: drain the Kubernetes node, shut the VM down, and start it again after the reboot — see Storage on Talos.
reboot
# during: pvecm status [2/3 quorate]; HA VMs running on partners; pvesr jobs to vm1 fail until it returns (expected)
# after:
ha-manager crm-command node-maintenance disable vm1
pvesr status # [jobs catch up, all OK]
7. Hard node kill + HA (do last)¶
With the replicated HA test VM running on vm1, pull power on vm1 from the
PDU (via OOB).
[~2–3 min: node fenced, VM restarts on vm2 from the last replication snapshot —
anything written after that sync is gone; cluster quorate 2/3.] Check the VM's data is
at the last-sync point, not corrupt.
Power vm1 back:
vm1 is overwritten by the next sync. That's correct: the VM carried
on from the replica on vm2.
Routine: planned reboots¶
The everyday version of drill 6: ha-manager crm-command node-maintenance enable
<node> → deal with the local-only and Talos VMs → reboot → node-maintenance disable
→ pvesr status all OK. Patch one node at a time, never all three at once.