11. Verification & cutover¶
Prove each property behaviourally — "the page went green" is not evidence. Do the read-only checks first, then the disruptive failover tests in a maintenance window, then cut over.
11.1 Read-only checks (safe anytime)¶
- [ ] HA status (System ▸ HA ▸ Status): FW-01 MASTER on all VIPs, FW-02 BACKUP on all. Not mixed.
- [ ] Config sync: a throwaway alias created on FW-01 appears on FW-02 within seconds. Delete it after.
- [ ] advskew intact on FW-02 (
<CARP_ADVSKEW_BACKUP>, not 0) — the sync-flatten check from page 4 §4.2. - [ ] BGP asymmetry (the key invariant): on FW-01 all four neighbours Established;
on FW-02 FRR stopped, zero sessions.
vtysh -c 'show bgp summary'. - [ ] Announce discipline:
show bgp neighbor <bdr> advertised-routes= only<PUBLIC_BLOCK>. Nothing else leaks upstream. - [ ] Receive discipline: default from both borders (bdr-1 preferred by local-pref); 10/8 aggregate + venue prefixes from the CCRs.
- [ ] ECMP:
netstat -rnshows two paths for CCR-learned prefixes (net.route.multipath=1). - [ ] NAT: an internal test host egresses as
<PAT_ADDR>; a POS-net test egresses as<PAYMENTS_ADDR>(rule order correct). - [ ] OOB isolation: from an OOB client you can reach 22/443 on each node, and you cannot route through OOB to any internal host (rule 4/5 hold).
- [ ] Zabbix: both nodes reporting over OOB; backup node not throwing false "FRR down / VIP backup" criticals (CARP-aware gating works).
- [ ] Suricata: running inline on WAN/WAN2/DMZ only; EVE arriving at
<WAZUH_IP>; not assigned toSRVTRUNK.
11.2 Failover tests (maintenance window — have BMC/KVM open)¶
Take a config backup and keep BMC open before these
System ▸ Configuration ▸ Backups. If a test wedges a node, BMC/KVM is your way back.
- [ ] CARP failover: on FW-01, System ▸ HA ▸ Status → Enter Persistent CARP Maintenance Mode. FW-02 becomes MASTER, its FRR starts, all four sessions come up in BFD time, data flows resume (pfSync had the state). Exit maintenance; FW-01 reclaims (preempt).
- [ ] Border failure: shut the bdr-1 session/link. Default reconverges via bdr-2 in BFD time; no outage. Restore.
- [ ] CCR failure: shut one CCR session. Traffic continues via the other (ECMP); no isolation. Restore.
- [ ] pfSync leg pull: pull one
PFSYNCRJ45 cable. State sync continues on the other leg (active-backup bond); no split-brain (heartbeats on other interfaces). Restore. - [ ] Wazuh action: trigger a test rule; confirm the offender IP lands in the
wazuh-blocklistalias and the WAN block rule catches it.
11.3 Config backup / audit (do this once, permanently)¶
- [ ] System ▸ Configuration ▸ Backups → enable a backup target (Git or Nextcloud)
on both nodes. This is your audit trail and DR:
config.xmlis exact and complete (it even holds the interface assignment Ansible couldn't touch). - [ ] Confirm each node's own config is backed up — recall interface assignment + per-node /31s are not synced (page 2 §2.5), so a node rebuild needs its own backup, not the peer's.
11.4 Pre-cutover gate¶
- [ ] All CHOOSE values in Site values settled (public block, /31s, OOB subnet).
- [ ] RPKI ROAs + IRR for
<PUBLIC_BLOCK>published and valid (§5c) — or upstreams filter you. - [ ] Upstream borders + CCRs staged for both nodes (page 1 §1.4).
- [ ] Split-DNS records in place for published services (§6.3).
- [ ] Run old and new addressing side-by-side during migration (DESIGN §7.1), then withdraw the old ranges from routing.
11.5 Known doc debts to reconcile (net-design, not this box)¶
These are inconsistencies in the source design that this runbook parameterised around; close them before or shortly after cutover:
<PUBLIC_BLOCK>/28 vs /29 — DESIGN §5c vswan/WAN-EXISTING.md.<OOB_NET>— DESIGN §7.1 (10.128.9.0/24) vs §7.3 text (10.1.9.0/24) vs the Zabbix leg (10.201.201.0/24).- OPNSENSE.md §2/§7/§15 say "no dedicated OOB NIC" — this build uses a USB OOB NIC; amend those sections to match.