Skip to content

11. Verification & cutover

Prove each property behaviourally — "the page went green" is not evidence. Do the read-only checks first, then the disruptive failover tests in a maintenance window, then cut over.

11.1 Read-only checks (safe anytime)

  • [ ] HA status (System ▸ HA ▸ Status): FW-01 MASTER on all VIPs, FW-02 BACKUP on all. Not mixed.
  • [ ] Config sync: a throwaway alias created on FW-01 appears on FW-02 within seconds. Delete it after.
  • [ ] advskew intact on FW-02 (<CARP_ADVSKEW_BACKUP>, not 0) — the sync-flatten check from page 4 §4.2.
  • [ ] BGP asymmetry (the key invariant): on FW-01 all four neighbours Established; on FW-02 FRR stopped, zero sessions. vtysh -c 'show bgp summary'.
  • [ ] Announce discipline: show bgp neighbor <bdr> advertised-routes = only <PUBLIC_BLOCK>. Nothing else leaks upstream.
  • [ ] Receive discipline: default from both borders (bdr-1 preferred by local-pref); 10/8 aggregate + venue prefixes from the CCRs.
  • [ ] ECMP: netstat -rn shows two paths for CCR-learned prefixes (net.route.multipath=1).
  • [ ] NAT: an internal test host egresses as <PAT_ADDR>; a POS-net test egresses as <PAYMENTS_ADDR> (rule order correct).
  • [ ] OOB isolation: from an OOB client you can reach 22/443 on each node, and you cannot route through OOB to any internal host (rule 4/5 hold).
  • [ ] Zabbix: both nodes reporting over OOB; backup node not throwing false "FRR down / VIP backup" criticals (CARP-aware gating works).
  • [ ] Suricata: running inline on WAN/WAN2/DMZ only; EVE arriving at <WAZUH_IP>; not assigned to SRVTRUNK.

11.2 Failover tests (maintenance window — have BMC/KVM open)

Take a config backup and keep BMC open before these

System ▸ Configuration ▸ Backups. If a test wedges a node, BMC/KVM is your way back.

  • [ ] CARP failover: on FW-01, System ▸ HA ▸ Status → Enter Persistent CARP Maintenance Mode. FW-02 becomes MASTER, its FRR starts, all four sessions come up in BFD time, data flows resume (pfSync had the state). Exit maintenance; FW-01 reclaims (preempt).
  • [ ] Border failure: shut the bdr-1 session/link. Default reconverges via bdr-2 in BFD time; no outage. Restore.
  • [ ] CCR failure: shut one CCR session. Traffic continues via the other (ECMP); no isolation. Restore.
  • [ ] pfSync leg pull: pull one PFSYNC RJ45 cable. State sync continues on the other leg (active-backup bond); no split-brain (heartbeats on other interfaces). Restore.
  • [ ] Wazuh action: trigger a test rule; confirm the offender IP lands in the wazuh-blocklist alias and the WAN block rule catches it.

11.3 Config backup / audit (do this once, permanently)

  • [ ] System ▸ Configuration ▸ Backups → enable a backup target (Git or Nextcloud) on both nodes. This is your audit trail and DR: config.xml is exact and complete (it even holds the interface assignment Ansible couldn't touch).
  • [ ] Confirm each node's own config is backed up — recall interface assignment + per-node /31s are not synced (page 2 §2.5), so a node rebuild needs its own backup, not the peer's.

11.4 Pre-cutover gate

  • [ ] All CHOOSE values in Site values settled (public block, /31s, OOB subnet).
  • [ ] RPKI ROAs + IRR for <PUBLIC_BLOCK> published and valid (§5c) — or upstreams filter you.
  • [ ] Upstream borders + CCRs staged for both nodes (page 1 §1.4).
  • [ ] Split-DNS records in place for published services (§6.3).
  • [ ] Run old and new addressing side-by-side during migration (DESIGN §7.1), then withdraw the old ranges from routing.

11.5 Known doc debts to reconcile (net-design, not this box)

These are inconsistencies in the source design that this runbook parameterised around; close them before or shortly after cutover:

  • <PUBLIC_BLOCK> /28 vs /29 — DESIGN §5c vs wan/WAN-EXISTING.md.
  • <OOB_NET> — DESIGN §7.1 (10.128.9.0/24) vs §7.3 text (10.1.9.0/24) vs the Zabbix leg (10.201.201.0/24).
  • OPNSENSE.md §2/§7/§15 say "no dedicated OOB NIC" — this build uses a USB OOB NIC; amend those sections to match.