Skip to content

Firewall — OPNsense edge pair: setup runbook

Step-by-step build of the corporate edge HA pair fw-colo-1 / fw-colo-2. This is the how; see firewall-opnsense for the why (interface plan, zone policy, BGP design, tuning). Read that first.

Starting point (assumed done)

OPNsense ISO installed on both nodes, and the OOB management IP set on each from the console: fw-colo-1 = 172.16.201.41, fw-colo-2 = 172.16.201.42 (/24, gateway = the OOB router). The GUI is reachable over OOB.

Order of operations

Build fw-colo-1 (master) fully first, verify it standalone, then bring up fw-colo-2 with only its node-specific bits (interface IPs, OOB, CARP advskew) and let XMLRPC config-sync push everything else. Enabling HA sync before node-1 is correct will replicate mistakes. Rough sequence:

  1. Base system + plugins (both nodes)
  2. Interfaces / LAGGs / VLANs + addressing (both nodes — per-node IPs; never syncs)
  3. CARP VIPs (node-1; sync pushes to node-2)
  4. pfSync + HA config-sync (both)
  5. FRR/BGP (both nodes — per-node router-id + peer /31s; not synced)
  6. HA advertisement gating — anchor scripts (both — files on disk, not synced)
  7. ACME + LLDP + firewall/NAT (node-1; syncs)
  8. Verify

What XMLRPC sync does and does NOT replicate

Syncs (node-1 → node-2): firewall rules, NAT, aliases, CARP VIPs, certificates, ACME, DHCP, users, schedules. Per-node, set on both manually: interface assignments + IPs, hostname, FRR/BGP (own router-id + own /31 peers — the prefix-lists/route-maps are identical except node-2's CORE MED, §7), the anchor gating scripts, and the advskew that elects the master. So "build one, sync the rest" applies to the policy layer — the interface and routing layers are per-node on both.

Throughout: Save each page, then Apply changes (the amber banner). Nothing takes effect until applied.


1. Base system (both nodes)

System → Settings → General

  • Hostname fw-colo-1 / fw-colo-2, domain infra.jsm (internal/private). Note: the GUI TLS cert and the name you browse to are the public FQDN (fw-colo-<n>.infra.pubinvest.co.uk, §8) — ACME can't validate the private infra.jsm zone, and split-horizon DNS resolves the public name to the OOB IP internally.
  • DNS servers = internal resolvers 10.128.32.3, 10.128.32.4; untick "Allow DNS server list to be overridden by DHCP/PPP on WAN".
  • Timezone Europe/London.

System → Settings → Administration

  • Protocol HTTPS, Listen interfaces = OOB only (lock the GUI to the OOB net).
  • SSH: enable, listen on OOB only, key-only, root login off.
  • Create an admin user; enable TOTP 2FA (System → Access → Servers → add a TOTP server, tie it to the login). Change the root password.

System → Settings → Cron / Services → NTP — point NTP at your internal source.

2. Plugins (both nodes)

System → Firmware → Updates — update to the latest release first, reboot.

System → Firmware → Plugins — install:

Plugin Purpose
os-frr FRR — the BGP engine (WAN + CORE)
os-acme-client ACME / Let's Encrypt certificates
os-lldpd LLDP (neighbour discovery, feeds LibreNMS/Zabbix)
os-crowdsec CrowdSec agent + firewall bouncer (free; §11b)
os-zabbix-agent (optional) monitoring agent (see the arch doc)

Reboot is not required; the new menus appear immediately.

3. Interfaces, LAGGs & VLANs

This is the real port→role map we confirmed from LLDP and cabled in NetBox — not the generic "8×SFP+" plan. Driver names: igb* = i350 1G copper, ixl* = X710 10G SFP+, ue* = the USB-eth OOB NIC. Both nodes are identical except the OOB NIC (ue1 on fw-colo-1, ue0 on fw-colo-2). Build the LAGGs and VLANs before assigning interfaces.

OPNsense port Role Cabled to
ixl7 WAN A (primary) bdr-1 sfp-sfpplus11/12
ixl2 WAN B (backup) bdr-2 sfp-sfpplus11/12
ixl3 CORE A CR-COLO-1 sfp-sfpplus1/2
ixl6 CORE B CR-COLO-2 sfp-sfpplus1/2
ixl0 + ixl5 server trunk → lagg0 sw-server-1 + sw-server-2 (MLAG)
igb0 + igb1 pfSync → lagg1 the peer FW, direct
ue1 / ue0 OOB mgmt OOB-2 ether1 / ether2

Interfaces → Devices → LAGG — create two:

Device Members Proto MTU Purpose
lagg0 ixl0, ixl5 LACP 9000 server trunk → the two MLAG server switches
lagg1 igb0, igb1 Failover (active-backup) 1500 pfSync, direct to the peer

Interfaces → Devices → VLAN — on parent lagg0, tag the server VLANs. All are routed here (gateway = a CARP VIP …​.1 on OPNsense); the switch carries them L2.

VLAN Interface Subnet Gateway (CARP VIP) Purpose
910 BACKUP 10.128.10.0/24 10.128.10.1 Backup traffic
932 VM_APPS 10.128.32.0/24 10.128.32.1 AD & apps (PXE/stock)
933 VM_K8S 10.128.33.0/24 10.128.33.1 Kubernetes
934 VM_MON 10.128.34.0/24 10.128.34.1 Monitoring VMs
935 VM_DATA 10.128.35.0/24 10.128.35.1 Data / SQL
936 VM_DEV 10.128.36.0/24 10.128.36.1 Development
940 HYP_MGMT 10.128.40.0/24 10.128.40.1 Hypervisor management

L2-only VLANs — no OPNsense interface

Corosync 950 (Proxmox cluster ring) and Ceph 920/921/922 are L2-only and are not routed on OPNsense — the Proxmox/Ceph nodes talk directly on the VLAN. Do not create OPNsense VLAN interfaces, gateways or CARP VIPs for them; they're tagged on the server switch only.

Interfaces → Assignments — assign and name:

Name Device Role
WAN_A ixl7 eBGP /31 → bdr-1 (primary)
WAN_B ixl2 eBGP /31 → bdr-2 (backup)
CORE_A ixl3 eBGP /31 → CR-COLO-1
CORE_B ixl6 eBGP /31 → CR-COLO-2
BACKUP lagg0 vlan910 backup (MTU 1500)
VM_APPS lagg0 vlan932 AD & apps
VM_K8S lagg0 vlan933 Kubernetes
VM_MON lagg0 vlan934 monitoring
VM_DATA lagg0 vlan935 data / SQL
VM_DEV lagg0 vlan936 development
HYP_MGMT lagg0 vlan940 hypervisor mgmt
PFSYNC lagg1 HA sync
OOB ue1 / ue0 (already assigned) management 172.16.201.41 / .42

MTU: set 9000 on lagg0 itself (the parent); the VLAN children stay 1500 (routed). Core/WAN /31s are 1500.

Addressing (per node)

Static IPv4 on each (fw-colo-1 shown; fw-colo-2 takes the .b side of each /31 and .42/second-in-subnet):

Interface port fw-colo-1 fw-colo-2
CORE_A → CR-COLO-1 ixl3 10.254.96.5/31 10.254.96.7/31
CORE_B → CR-COLO-2 ixl6 10.254.96.9/31 10.254.96.11/31
WAN_A → bdr-1 ixl7 185.109.43.21/31 (peer .20) 185.109.43.23/31 (peer .22)
WAN_B → bdr-2 ixl2 185.109.43.25/31 (peer .24) 185.109.43.27/31 (peer .26)
PFSYNC lagg1 10.254.240.1/30 10.254.240.2/30
each server VLAN lagg0.<vid> 10.128.<x>.2/24 10.128.<x>.3/24

Transit /31s and pfSync are per-node point-to-point (no VIP) — each node runs its own BGP sessions; use the /31 values NetBox holds for ixl2/3/6/7.

CARP: a real per-node IP and the floating VIP — you need both

Each routed server-VLAN interface gets a real, unique IP on each node (.2 on fw-colo-1, .3 on fw-colo-2), plus the floating CARP VIP .1 that is the clients' gateway (§4). OPNsense requires the interface to already hold a real address in that subnet before a CARP VIP can bind, and each node needs its own IP to source traffic, sync, and stay reachable when it is BACKUP. E.g. VLAN 932 — fw-colo-1 10.128.32.2/24, fw-colo-2 10.128.32.3/24, CARP VIP 10.128.32.1/24.

4. CARP virtual IPs

Interfaces → Virtual IPs → Settings — add a CARP VIP for each floating server-VLAN gateway. These sit on directly-attached L2 segments, so they float via CARP (gratuitous ARP). The public /28 is not here — it's a routed block (see the note).

VIP Interface vhid Notes
10.128.10.1/24 BACKUP 10 backup gateway
10.128.32.1/24 VM_APPS 32 AD & apps gateway
10.128.33.1/24 VM_K8S 33 Kubernetes gateway
10.128.34.1/24 VM_MON 34 monitoring gateway
10.128.35.1/24 VM_DATA 35 data/SQL gateway
10.128.36.1/24 VM_DEV 36 development gateway
10.128.40.1/24 HYP_MGMT 40 hypervisor-mgmt gateway

The transit /31s are not CARP (per-node point-to-point).

4.1 Pick the right Mode (VIP type)

Mode What it does Use it for
CARP Floating HA address — advertises via CARP (proto 112), grat-ARPs on failover. Needs VHID + password, and a real IP already on that interface in the same subnet. Server-VLAN gateways (10.128.x.1) — attached L2, must float between the pair
IP Alias An extra address bound to an interface; no HA protocol. Can optionally be parented to a CARP VIP so it follows that VIP's state. Services can bind to it. The routed /28 (SNAT pool, inbound targets) — and anything you need to bind a service to
Proxy ARP FW answers ARP for an address it doesn't bind; can't host services. Rare — NAT-only addresses on an attached subnet you hold no IP in
Other FW neither ARPs nor binds — just "knows" the address for NAT/routing. A routed block used purely for NAT — the purist choice for our /28 if you never bind a service to it

For our /28 we use IP Alias (flexible — you can bind a service to a public IP later); Other is equally correct if it stays strictly NAT.

4.2 CARP VIP — field by field

Field Value / meaning
Mode CARP
Interface the VLAN interface the gateway lives on (e.g. VM_APPS)
Address the gateway with the VLAN's mask — e.g. 10.128.32.1/24
Virtual IP Password (+ confirm) shared secret for this VHID — must match on both nodes
VHID Group 1–255. Identical on both nodes for the same VIP, and unique within that L2 broadcast domain
Advertising Frequency → Base seconds between adverts — leave 1
Advertising Frequency → Skew priority, lower wins: 0 on fw-colo-1, 100 on fw-colo-2
Peer leave blank for normal multicast CARP. Set it to the partner's real interface IP on that segment only if you need unicast CARP (multicast filtered/not carried on that L2)
Description e.g. VM_APPS gw (VLAN 932)

Three CARP rules people trip on

  1. The interface must already hold a real per-node IP in that subnet (.2/.3) — a CARP VIP cannot bind otherwise (see §3 addressing).
  2. VHID must be identical on both nodes and unique per broadcast domain. Separate VLANs may reuse a VHID (different L2), but keep a scheme — we use the VLAN-derived numbers 10/32/33/34/35/36/40 so it stays obvious.
  3. advskew elects the master. If both nodes end up with the same skew, election is non-deterministic — always set 0 / 100 explicitly.

Do I create the VIPs on both nodes? — No, node-1 only

Virtual IPs are a synced section, so create them on fw-colo-1 and let the XMLRPC config-sync (§5) push them to fw-colo-2. Two caveats:

  • Prerequisite that does not sync: the per-node interface IPs must already exist on BOTH nodes (.2 / .3, §3). Interfaces never sync, and a synced CARP VIP cannot bind on node-2 if that interface holds no real address in the subnet — so do interfaces/IPs on both first, then VIPs on node-1.
  • Verify after the first sync: node-2 should list the VIPs and show BACKUP (Interfaces → Virtual IPs → Status). If that's what you see, leave the skew alone — your build adjusted it on the backup.

Same applies to the routed-/28 IP-Alias VIPs: created on node-1, synced to node-2, harmless there because the backup never advertises the /28 (the CARP-keyed anchor, §7) — no traffic is attracted to it.

Editing advskew on the secondary — it won't stick while VIPs are synced

You can change advskew on fw-colo-2, but the next config-sync that includes Virtual IPs pushes node-1's value straight back over it — it is not durable. If you truly need a fixed difference, untick "Virtual IPs" from the sync sections and manage the VIPs by hand on both nodes (identical address / VHID / password, skew 0 and 100).

And if both nodes claim MASTER, advskew is usually NOT the cause. Diagnose in this order:

  1. Firewall allows protocol 112 on that interface, in both directions — blocked CARP advertisements is the classic split-brain.
  2. VHID and CARP password identical on both nodes.
  3. Then look at advskew.

4.3 Global CARP behaviour (System → Settings → Tunables)

Tunable Value Why
net.inet.carp.preempt 1 all VIPs fail over together — if one CARP interface demotes, the rest follow. Without it you get split routing (some VLANs on node-1, some on node-2)
net.inet.carp.log 2 log CARP state transitions (essential when debugging flaps)
net.inet.carp.allow 1 CARP enabled (default)

4.4 Firewall + status

  • Allow CARP on every interface carrying a VIP — protocol 112 from the peer. If the nodes can't hear each other's adverts both go MASTER (split-brain) — this is the single most common CARP failure.
  • Interfaces → Virtual IPs → Status — MASTER/BACKUP per VIP, plus "Enter persistent CARP maintenance mode": use that for planned failover tests and upgrades (it demotes cleanly instead of yanking the box).

The public /28 is routed — NOT a CARP VIP

Don't create CARP VIPs for 185.109.43.176/28. The ISP routes it to you over the /31 transit and you advertise it via BGP. HA is by the CARP-keyed anchor (§7): FRR runs warm on both nodes, but only the node holding the anchor route — the CARP MASTER — announces the block (next-hop = its own /31); on failover the anchor moves and the new master's already-Established sessions send the UPDATE in ~1 s. The specific addresses you use (the /30 SNAT pool + inbound service IPs, §10) are added as IP-Alias VIPs (Type: IP Alias, not CARP) on the WAN interface or a loopback, synced to both nodes — harmless on the backup since it advertises nothing, so no traffic ever arrives there.

5. HA — pfSync + config sync

System → High Availability → Settings

  • Synchronize States (pfSync): enable; Synchronize Interface = PFSYNC; Synchronize Peer IP = the other node's PFSYNC address (10.254.240.2 on node-1, .1 on node-2).
  • Synchronize Config to backup (XMLRPC): configure on fw-colo-1 only.
  • Synchronize Config to IP = 10.254.240.2 (node-2 PFSYNC).
  • Remote user = an admin account that exists on node-2; HTTPS.
  • Tick the sections to replicate: Firewall Rules, NAT, Aliases, Virtual IPs, Schedules, Certificates, ACME, DHCP, Users/Groups — but NOT "Interfaces" or "System/General" (node-specific: IPs, hostname, advskew), and not FRR (per-node peers — configured directly on each node in step 6).

Firewall → Rules → PFSYNC — allow the pfSync protocol between the two node IPs on the PFSYNC interface (it's a direct cable, but the rule is still required).

CARP heartbeats: ensure the CARP interfaces (WAN, a server VLAN, and PFSYNC) can exchange multicast — default allow-CARP rules on those interfaces. advskew already elects node-1.

6. FRR / BGP (Routing → FRR)

Configure this on both nodes — FRR is not part of XMLRPC config-sync, and each node peers over its own /31s with its own router-id. The prefix-lists and route-maps below are identical on both with one deliberate exception: node-2's RM-CORE-OUT adds set metric 100 (§7 — the static CORE-side preference). Only the neighbour peer addresses differ otherwise. Edge AS 65510. FRR runs permanently on both nodes — HA is done by gating advertisements, not the daemon (§7).

Routing → General — Enable FRR. Router ID = a stable per-node IP (the OOB or a loopback); BFD supported.

Routing → BGP → General

  • Enable, AS number 65510, Network import/export via prefix-lists (below).
  • Advanced: bestpath as-path multipath-relax, and set net.route.multipath=1 (System → Settings → Tunables) for ECMP.

Originate the three anchored prefixes — but do NOT anchor them statically. Routing → BGP → General → Networks add:

Network Advertised to What it is
185.109.43.176/28 WAN + CORE the public block
10.128.0.0/16 CORE the whole colo block (the addressing plan's server-room /16). Connected server /24s win as more-specifics on the master; everything else in it — storage 12–17, migration 45, corosync 50/51, OOB 9, unused — hits the master's blackhole ("never routes", enforced edge-wide). OOB is mgmt-VPN-reached; if it ever needs CCR-side reachability, inject the more-specific 10.128.9.0/24 from its gateway

CCR-side PREREQUISITE: flip the CCRs' /16 blackhole to a floating backstop (distance > 20)

The colo CCRs carry 10.128.0.0/16 in their own BGP networks list, anchored by their own blackhole static at distance 1. Left like that, this design is dead on arrival: AD 1 beats eBGP (20) on an identical prefix, the FW's /16 never installs, and the CCRs blackhole all server traffic on both nodes.

The fix is one attribute — raise the CCR blackhole's distance above 20 (e.g. 200) and keep both the static and the networks entry:

  • Normal: the FW master's /16 (eBGP 20) wins, installs, and propagates metro-wide; the CCR's floating static is inactive, so its own origination goes quiet automatically.
  • Failover gap / FW pair down: the FW /16 vanishes → the floating static re-activates instantly → the CCR resumes originating and blackholing. Metro steering stays stable and colo-bound traffic gets a clean drop at the CCR instead of following a default — identical protection to the old CCR-owned design, now as a backstop.

Do the CCR change before switching the FW to the /16 (migration §7, step 0 — automated in ansible/network roles/rr: rr_aggregate_distance + a to-edge rule rejecting dst in 10.128.0.0/16, so the RRs also never send the /16 to the FW — the RR's own origination has no 65510 in path, so the FW's AS-path loop detection can't drop it and it could otherwise satisfy the backup's network import-check. FW-originated prefixes like the /28 need no such guard: any copy coming back already carries 65510 → loop-dropped.)

FRR only announces a network statement while a matching route exists in the RIB (assert bgp network import-check is on — default in current FRR; if your build has it off, the whole §7 gate is void). The blackhole routes that satisfy these are added and removed by the CARP scripts in §7 — that is the HA gate. Do not list the individual server /24s as Networks: they're connected on both nodes, so they'd advertise from the backup too and re-open the hole the aggregates close (§7.1).

Never add a persistent static/Null0 route for ANY anchored prefix

Not in System → Routes, not as an FRR static — for the /28 or the 10.128.0.0/16. A persistent anchor makes that network statement fire on both nodes permanently — the backup would advertise and attract, and the §7 gating is silently defeated. The only source of these routes is the CARP-keyed script.

Prefix-lists (Routing → BGP → Prefix Lists)

Name Rule
PL-DEFAULT-IN permit 0.0.0.0/0 exact
PL-PUBLIC-OUT permit 185.109.43.176/28 exact
PL-CORE-OUT permit 0.0.0.0/0, 185.109.43.176/28, and the anchored colo aggregate 10.128.0.0/16 (exact — not the individual connected /24s, which would leak from the backup; requires the CCR floating-backstop change, §6 danger box)

Route-maps (Routing → BGP → Route Maps)

Name Action
RM-WAN-IN-PRIMARY match PL-DEFAULT-IN, set local-preference 200
RM-WAN-IN-BACKUP match PL-DEFAULT-IN, set local-preference 100
RM-WAN-OUT match PL-PUBLIC-OUT (permit), else deny
RM-WAN-OUT-BACKUP as RM-WAN-OUT + set as-path prepend 4200000001 4200000001 (the WAN local-as)
RM-CORE-OUT match PL-CORE-OUT (permit), else deny. On fw-colo-2 ONLY: add set metric 100 (MED — the static CORE-side preference, §7)

Neighbours (Routing → BGP → Neighbors)

WAN — to the ISP borders (AS 204258), presenting our customer AS 4200000001:

Peer (fw-colo-1 / fw-colo-2) Remote AS local-as BFD In Out
bdr-1 — .20 / .22 (WAN_A ixl7) 204258 4200000001 no RM-WAN-IN-PRIMARY RM-WAN-OUT
bdr-2 — .24 / .26 (WAN_B ixl2) 204258 4200000001 no RM-WAN-IN-BACKUP RM-WAN-OUT-BACKUP

The peer IP is PER NODE — don't copy node-1's neighbours to node-2

Each FW has its own /31 to each border (see §3). fw-colo-1 peers 185.109.43.20 / .24; fw-colo-2 peers 185.109.43.22 / .26. Pointing fw-colo-2 at .24 reaches bdr-2 (it's routable) but that link's neighbour is fw-colo-1 (.25), so bdr-2 rejects the OPEN and resets — you'll see [FSM] unexpected packet received in state OpenSent followed by bgp_read_packet error: Connection reset by peer. The same two lines appear if local-as 4200000001 is missing (bdr sees the wrong AS → Bad Peer AS).

→ The FRR instance AS is 65510 (used on the CORE side); on the WAN the ISP knows us as our customer AS 4200000001, so set local-as 4200000001 no-prepend replace-as on both WAN neighbours (clean single-AS path to the ISP). BFD is OFF on the WAN — failover is driven by the tight keepalive/hold 3/9 s timers instead (BFD stays on the CORE/CCR sessions only). bdr-1 primary (local-pref 200 in), bdr-2 backup (local-pref 100 + prepend ×2 out). Announce only the /28; receive default only.

CORE — to the colo CCRs (AS 65500):

Peer Remote AS BFD In Out
CR-COLO-01 (CORE_A /31) 65500 yes (accept) RM-CORE-OUT
CR-COLO-02 (CORE_B /31) 65500 yes (accept) RM-CORE-OUT

→ enable BFD per neighbour (single-hop /31, so it just works; the CCR side keys off its bfd_enabled field). Announce 0/0 + public /28 + the anchored colo aggregate (10.128.0.0/16 — §7; needs the CCR floating-backstop prerequisite, §6); receive the server-room/internal routes.

Attaching the policy — what to select on each neighbour

On Routing → BGP → Neighbors attach the route-maps (In/Out) — not the prefix-lists. Our route-maps already match the prefix-list and set the attribute (local-pref / prepend), so the prefix-list is referenced by the map:

Neighbour Route-Map In Route-Map Out
bdr-1 (WAN primary) RM-WAN-IN-PRIMARY RM-WAN-OUT
bdr-2 (WAN backup) RM-WAN-IN-BACKUP RM-WAN-OUT-BACKUP
CR-COLO-1 / CR-COLO-2 (blank — accept) RM-CORE-OUT

The neighbour form also has Prefix-List In/Out boxes — only needed if you're filtering without setting attributes. If you attach both in the same direction, both must permit (they're ANDed). For this design: route-maps only. Leaving CORE's In blank accepts everything the CCRs send (trusted internal peers); attach a prefix-list there too if you want belt-and-braces.

Do I need a deny at the end of the prefix-list? — No

FRR prefix-lists and route-maps both have an implicit deny at the end. A prefix-list of only permit entries denies everything else; a route-map with a single permit sequence denies anything that doesn't match. A trailing explicit deny is optional documentation, not a requirement.

Does the route-map itself enforce the prefix-list? — Yes

The map's one permit clause matches the prefix-list, and anything that doesn't match falls through to the map's implicit deny. So attaching the route-map is what applies the list — you don't attach the prefix-list separately.

Two ways to silently break that:

  • A clause with no match permits everything (and applies its set) — never add a catch-all permit to these maps.
  • A typo'd / missing prefix-list name means the match can never succeed, so every route hits the implicit deny and the peer carries zero prefixes. If a session is Established but empty right after a policy edit, check the name first.

Verify what the policy actually did (vtysh):

show bgp neighbor 185.109.43.20 advertised-routes   # what we send, post out-policy
show bgp neighbor 185.109.43.20 routes              # what we accepted, post in-policy
show bgp neighbor 185.109.43.20 received-routes     # pre-policy (needs soft-reconfig in)

The real trap is le / ge, not the missing deny

A prefix-list entry matches exactly unless you add le/ge:

  • permit 0.0.0.0/0 → only the default route ✅ — this is "receive default only"
  • permit 0.0.0.0/0 le 32 → every prefix ❌ — you'd accept the full table

Same outbound: permit 185.109.43.176/28 is exact; le 32 would also permit every more-specific from it. Keep every list here exact — no le/ge.

Filters are not retroactive — soft-clear after changes

Changing a prefix-list or route-map does not re-evaluate already sent/received routes. Soft-clear the session (vtysh): clear bgp 185.109.43.20 soft in / clear bgp 185.109.43.20 soft out. Also make sure the list/map exists before you reference it on a neighbour.

7. HA advertisement gating — warm FRR + CARP-keyed anchor (both nodes)

FRR runs permanently on both nodes; failover moves one kernel route. The old stop-FRR-on-BACKUP design cost 10–20 s per failover (daemon start → BGP Idle→Established → learn → advertise, all serial). Here every session is warm and Established on both nodes at all times, so a failover is a CARP promotion plus one BGP UPDATE on an already-open session (~1 s) — and the backup holds a live default route throughout.

7.1 Design & invariants

# Invariant Enforced by
I1 Both nodes hold 0.0.0.0/0 at all times warm WAN sessions; in-policy is role-independent
I2 Every traffic-attracting prefix (185.109.43.176/28, 10.128.0.0/16) is advertised iff the node holds its anchor route FRR network + import-check semantics — by construction, not by script
I3 Exactly the CARP MASTER holds the anchors (both move together) CARP syshook + boot hook + 1-min reconcile, all calling one idempotent script whose only input is live kernel CARP state
I4 CCRs prefer node-1's default whenever it advertises one static MED 100 on node-2's RM-CORE-OUT (§6) — never touched at runtime; also the deterministic tie-break during any ≤60 s dual-advertise reconcile window

Strictness model — the backup attracts nothing, anywhere. All inbound attraction is anchor-gated: the /28 on WAN, and the whole colo block toward the CCRs via the 10.128.0.0/16 aggregate (§6). The aggregate exists precisely because the individual server /24s are connected on both nodes and therefore un-gateable (their routes always exist, so a network statement for them would always fire — including on the backup). A strict superset is gateable: only the anchor satisfies it. The FW owning the /16 requires the CCR floating-backstop prerequisite (§6 danger box) — with the CCRs' blackhole left at distance 1, the FW's /16 never installs. Consequences, all wanted:

  • Venue→server traffic can only ever enter the master — including the corner where a demoted node still has live CORE sessions (it lost its anchors, so it advertises no server space). No pfSync-dependent inbound asymmetry remains.
  • On the master, connected /24s win as more-specifics inside the /16; everything else in it — storage 12–17, migration 45, corosync 50/51, OOB 9, unused space — hits the FW's blackhole and dies at the edge instead of leaking out the default. "Storage never routes," enforced for the entire colo block.
  • Whenever no FW advertises the /16 (failover gap, both nodes down), the CCRs' floating blackhole re-activates and they resume originating it — metro steering stays stable and colo-bound traffic gets a clean drop at the CCR, not a loop.

The one deliberate exception: the default toward the CCRs. Both nodes advertise 0.0.0.0/0 at all times (node-2 at MED 100). It cannot be anchor-gated: the backup must hold a live default for itself (I1), and a route that must exist on both nodes can't act as a one-node gate. The alternatives fail the robustness bar — conditional advertisement (exist-map) is scan-timer-driven and not persistable through os-frr config regeneration (the same reason §7.7 rejects the vtysh gate). What the exception costs: a demoted node with live WAN+CORE can still carry venue→outbound traffic (better MED) — an outbound-only, pfSync-covered asymmetry in an already-degraded state. What it buys: the CCRs hold two pre-installed defaults, so venue internet survives a master death on BFD detection alone (~1 s), with no UPDATE round-trip. Every inbound flow — the ones a stateful firewall genuinely can't tolerate asymmetric — is strict by construction.

7.2 Why a kernel route is the gate (the guarantee argument)

The dynamic state — "am I the announcer?" — lives in the kernel routing table, deliberately outside FRR and outside OPNsense config:

  • Immune to the FRR/GUI lifecycle. os-frr regenerates frr.conf on every GUI edit and reload; anything injected via vtysh is wiped. A kernel route survives all of it — frr.conf stays 100 % static and identical to what the GUI manages.
  • Reboots fail closed. Kernel routes don't persist, so a booting node comes up not advertising until CARP resolves and the hook/boot-sync runs. A rebooting or crashed-and-returned node can never dual-attract.
  • Atomic + idempotent. The transition is one route add/route delete; the sync script can run any number of times, from any trigger, in any order.
  • Failure directions are asymmetric on purpose. Residual wrong states degrade to dual-advertise (upstream best-paths one node; pfSync-synced states keep the stray flows alive) rather than master-withdrawn (an outage). The ≤60 s reconcile bounds both.

7.3 The scripts (both nodes; node-local, never XMLRPC-synced)

One code path. Every trigger calls the same sync script; the CARP hook's arguments are deliberately ignored (their format varies by version — live kernel state is the single source of truth).

/usr/local/opnsense/scripts/frr-anchor-sync.sh (mode 0755):

#!/bin/sh
# Idempotent: assert ALL attraction anchors from live CARP state. Safe to run any time.
# Anchors = every prefix in §6's Networks table. All move together (I3).
ANCHORS="185.109.43.176/28 10.128.0.0/16"
VHID="10"                      # the tracked CARP vhid (SRV_T1) — set per §4

STATE=$(ifconfig | awk -v v="$VHID" '$1=="carp:" && $4==v {print $2; exit}')
RTAB=$(netstat -rn -f inet)

for A in $ANCHORS; do
    # FreeBSD netstat TRIMS TRAILING ZERO OCTETS from network routes —
    # 10.128.0.0/16 prints as "10.128/16". Match both forms, or the check
    # never sees the route (re-adds forever as MASTER; NEVER REMOVES as
    # BACKUP = fail-open).
    ABBR=$(printf '%s' "$A" | sed -E 's#(\.0)+/#/#')
    HAVE=$(printf '%s\n' "$RTAB" | awk -v a="$A" -v b="$ABBR" '$1==a || $1==b {print "yes"; exit}')
    if [ "$STATE" = "MASTER" ] && [ -z "$HAVE" ]; then
        if route -q add -net "$A" 127.0.0.1 -blackhole; then
            logger -t frr-anchor "MASTER: anchor $A added (advertising)"
        else
            # add refused with the slot empty per netstat = a foreign exact route
            # (usually a BGP-learned copy of $A -> the RR to-edge guard is not
            # deployed). Loud, not silent: this is a design-order violation.
            logger -t frr-anchor "MASTER: FAILED to add anchor $A — foreign exact route? check: vtysh -c 'show ip bgp $A' and the RR to-edge guard (§6)"
        fi
    elif [ "$STATE" != "MASTER" ] && [ -n "$HAVE" ]; then
        route -q delete -net "$A" \
          && logger -t frr-anchor "state=${STATE:-unknown}: anchor $A removed (withdrawn)"
    fi
done

/usr/local/etc/rc.syshook.d/carp/20-frr-anchor (mode 0755) — event trigger:

#!/bin/sh
# CARP transition -> re-derive from live state (args intentionally ignored)
exec /usr/local/opnsense/scripts/frr-anchor-sync.sh

/usr/local/etc/rc.syshook.d/start/99-frr-anchor (mode 0755) — boot trigger, same one-liner exec as above.

Reconcile every minute — register a cron action so it survives OPNsense config management: /usr/local/opnsense/service/conf/actions.d/actions_frranchor.conf:

[sync]
command:/usr/local/opnsense/scripts/frr-anchor-sync.sh
type:script
message:sync FRR /28 anchor to CARP state
description:FRR anchor reconcile

then service configd restart and add it in System → Settings → Cron (every minute). The reconcile is the bound on every "script didn't fire" failure mode.

7.4 Failure matrix

Event Behaviour Converges in
Planned failover (CARP maintenance) old master's hook withdraws all anchors, new master's hook advertises — both on warm sessions ~1 s
Master power loss CARP promotes (~1–3 s) → hook adds anchors → UPDATEs out. CCR-side server routes + default converge on BFD (~1 s); borders drop the dead node's stale /28 on hold expiry (9 s) — or ~1 s with WAN BFD (see 7.5) ~2–4 s (CORE) / bounded by WAN detection (inbound /28)
Node demoted, CORE sessions still live anchors move → it advertises no server space, no /28; only its MED-100 default remains (§7.1 exception) — inbound stays strict by construction ~1 s
FRR reload / GUI edit (either node) anchor is a kernel route — untouched; frr.conf is static no impact, by construction
Node reboot kernel route gone → boots withdrawn (fail-closed); CARP resolves → boot/hook sync asserts ≤ boot + seconds
Hook missing/failed on promotion new master not advertising — the bad direction ≤60 s (cron) + Zabbix alarm (7.6)
Backup wrongly holds anchor dual-advertise; borders best-path one node; pfSync covers stray flows ≤60 s (cron)
CARP split-brain (all heartbeat paths lost) both advertise — BGP cannot fix a CARP split-brain in any design; mitigated upstream by multi-segment heartbeats (WAN + server-net + pfSync bond) n/a — prevent at CARP layer

7.5 Timer budget & the WAN-BFD lever

Warm sessions make detection the entire failover budget. CORE is already ~1 s (BFD). On WAN, an unplanned master death leaves the borders best-pathing the dead node's /28 until hold (9 s) even though the new master is already advertising — only BFD (or brutal timers) shortens that. BFD-on-WAN is therefore worth reopening with the ISP: it is now the difference between ~1 s and ~9 s of inbound loss on master death, not a nice-to-have. Planned failovers don't care (the withdraw is explicit) — keep 3/9 s as the floor either way.

7.6 Monitoring the invariant (Zabbix)

Divergence must be detected, not just reconciled. UserParameter on both nodes:

# 0 = consistent; 1 = CARP state and anchor-count disagree (both present as MASTER, 0 otherwise)
# NB: netstat prints 10.128.0.0/16 as "10.128/16" (trailing-zero trimming) — match both.
UserParameter=frr.anchor.consistent,STATE=$(ifconfig | awk '$1=="carp:" && $4=="10" {print $2; exit}'); N=$(netstat -rn -f inet | grep -cE '^(185\.109\.43\.176/28|10\.128(\.0\.0)?/16) '); { [ "$STATE" = "MASTER" ] && [ "$N" -eq 2 ]; } || { [ "$STATE" != "MASTER" ] && [ "$N" -eq 0 ]; }; echo $?

Trigger: frr.anchor.consistent = 1 for >90 s (one missed reconcile) = high; the existing CARP-aware alerting (§14 pattern) stays — but note the backup's FRR being up with Established sessions is now the healthy state, so remove any "FRR down on backup is OK" muting from the old design.

7.7 Rejected alternatives (and why)

Option Why not
Stop/start FRR on CARP (old design) 10–20 s serial cold start; backup holds no default; nothing wrong when converged — just slow
vtysh prefix-list gate (incl. a top-sequence deny of the aggregates toggled on BACKUP) needs the identical hook/boot/reconcile machinery as the anchor and gates the identical scope — but the state lives in frr.conf, exactly where os-frr regeneration wipes it, and a reboot restores whatever the file says instead of failing closed. Wrong-state windows fail toward master-withdrawn = outage. Same script count, strictly weaker guarantee
FW /16 without the CCR floating-backstop change the CCRs' distance-1 blackhole beats eBGP (20) on an identical prefix — the FW's /16 never installs and the CCRs blackhole all server traffic on both nodes. The /16 design requires the CCR distance flip (§6 danger box)
Mid-size aggregates (10.128.10.0/23 + 10.128.32.0/19) instead of the /16 works without touching the CCRs (more-specifics beat their /16) — the fallback if the CCR change is off the table; costs two anchors instead of one and splits the blackhole enforcement across two layers
BGP conditional advertisement (advertise-map/exist-map) scan-timer driven (5–60 s) and not persistable through os-frr config regen — slower and weaker than the anchor. Also why the CORE default stays un-gated rather than exist-map-gated (§7.1)
Always-advertise + prepend on backup (no gating) leaves a persistent asymmetry corner: a demoted master with live WAN sessions keeps attracting inbound indefinitely

Migration from the stopped-FRR design

  1. CCR side first — automated in ansible/network: run playbooks/colo-rr.yml (rr role): rr_aggregate_distance: 200 re-anchors the CCRs' 10.128.0.0/16 blackhole as the floating backstop (§6 danger box; brief one-time withdraw as the distance-1 anchor is replaced), and the new to-edge guard rejects dst in 10.128.0.0/16 toward the FW peers — required because the RR's own /16 origination carries no 65510 in path, so FW loop detection can't drop it and it could satisfy the backup's network import-check. Until this is deployed the FW's /16 will never install.
  2. Remove /usr/local/etc/rc.syshook.d/carp/20-frr (the start/stop hook) from both.
  3. Delete any persistent Null0/static for the anchored prefixes on the FW (System → Routes and FRR statics) on both nodes — see the §6 danger box; the gate is void while one exists.
  4. Replace the per-/24 server networks with 10.128.0.0/16 in Networks and PL-CORE-OUT on both nodes. Node-2: add set metric 100 to RM-CORE-OUT. Verify bgp network import-check on both.
  5. Install the three scripts + cron action on both nodes; chmod 755.
  6. Start FRR on the backup (service frr start / enable in GUI) — it establishes all four sessions but must advertise only its MED-100 default to the CCRs and nothing on WAN: check every neighbour with vtysh -c 'show ip bgp neighbors <ip> advertised-routes'.
  7. Run frr-anchor-sync.sh by hand on both; confirm the master holds both anchors — netstat -rn -f inet | grep -E '185\.109\.43\.176|^10\.128' (the /16 shows abbreviated as 10.128/16), or vtysh -c 'show ip route 10.128.0.0/16' → Known via "kernel" … blackhole — and advertises /28 on WAN + the /16 on CORE. On a CCR, the /16 is now the eBGP route with the floating blackhole inactive.
  8. Failover test (§13) — budget is now ~1–2 s, not 10–20 s.

8. ACME / Let's Encrypt (Services → ACME Client)

Certs must use the PUBLIC name — not the private infra.jsm zone

The boxes live in infra.jsm (internal), but Let's Encrypt can only validate a publicly-resolvable name. So the GUI cert uses the public FQDNs in the pubinvest.co.uk zone — fw-colo-1.infra.pubinvest.co.uk / fw-colo-2.infra.pubinvest.co.uk — issued by DNS-01 against that zone. ACME will never work for *.infra.jsm.

  • Accounts — Let's Encrypt (production), ops email, accept ToS.
  • Challenge Types — DNS-01 against the public pubinvest.co.uk zone (add the DNS provider's API creds). DNS-01 needs no inbound reachability — ideal for an OOB-only edge FW whose public name isn't otherwise exposed. (HTTP-01 can't work here.)
  • Certificates — issue one cert on fw-colo-1 with both node names as SANs: fw-colo-1.infra.pubinvest.co.uk and fw-colo-2.infra.pubinvest.co.uk (Account + the DNS-01 challenge). A single dual-SAN cert means the XMLRPC Certificates sync gives node-2 a cert valid for its name too — no per-node ACME needed.
  • Automations — "Restart GUI" (configd) on issue/renew; enable the daily renew cron.
  • System → Settings → Administration → SSL Certificate = this cert on both nodes, and browse the GUI by the public name so it matches: split-horizon DNS on the internal resolvers (10.128.32.3/.4) resolves fw-colo-<n>.infra.pubinvest.co.uk to the OOB IP (172.16.201.41 / .42), where the GUI listens.

9. LLDP (Services → LLDPd)

Enable; select the transmit/receive interfaces (CORE_A/B, WAN_A/B, lagg0, OOB). Neighbours then show in the GUI and via SNMP → LibreNMS/Zabbix.

10. Firewall & NAT (summary — see the arch doc for policy)

  • Zones default-deny + logged: WAN, CORE, BACKUP, VM-APPS, VM-K8S, VM-MON, VM-DATA, VM-DEV, HYP-MGMT, OOB, GUEST, MGMT-VPN. East-west hairpins the trunk (firewall-on-a-stick). Split the public /28 (185.109.43.176/28) — a small pool for outbound balancing, the rest for inbound services (adjust to taste):

The /28 is routed to you — every address below is an IP-Alias VIP (not CARP), and HA rides BGP (the CARP-keyed anchor, §7), not L2. Put the IP-Aliases on the WAN interface or a loopback; config-sync replicates them to both nodes (dormant on the backup).

Range Use
.177 primary WAN service address (IP-Alias)
.178 POS / payment egress — single fixed SNAT IP (no pool)
.180 – .183 (/30) outbound SNAT pool — round-robin, 4 IPs
.184 – .190 inbound NAT — 1:1 / port-forward service addresses
.176 / .179 / .191 network / spare / broadcast

Outbound pool (Firewall → NAT → Outbound = Hybrid):

  1. Interfaces → Virtual IPs — an IP-Alias VIP for the pool 185.109.43.180/30 (standalone on WAN/lo0 — no CARP; the block fails over via BGP).
  2. Firewall → Aliases — a network alias OUT_POOL = 185.109.43.180/30 (all 4).
  3. Rule: source = the internal server nets, translation = OUT_POOL, Pool Options = Round Robin (add sticky-address if a session must keep one source IP) → SNAT round-robins across the 4 addresses.

POS / payment egress — fixed IP, ABOVE the pool rule

Higher-priority outbound rule: source = POS, destination = the payment-processor prefixes, translation = 185.109.43.178 fixed (no pool). Processors whitelist one IP; the guest firewall keeps its own distinct egress. See firewall-opnsense §NAT.

Inbound NAT (Firewall → NAT → Port Forward / 1:1), on .184 – .190:

  • Add an IP-Alias VIP for the chosen public address, then a Port Forward (or 1:1 NAT) → the internal host / service VIP. No CARP — reachability follows the BGP announcement, so it lands on whichever node is master.
  • Point split-horizon DNS at the internal address for internal clients (no hairpin); external DNS resolves to the public address.

11. Edge security (free stack)

This edge is the internet chokepoint for every venue's staff PCs, staff Wi-Fi, POS tills and the colo servers, so it carries real user traffic. The whole stack is free — no paid NGFW:

Layer Tool Notes
Exploit / malware IPS Suricata (base) inline on WAN
Reputation + brute-force CrowdSec (free) community blocklist + log scenarios
Web / category filtering separate DNS servers FW only enforces the path (§11c)
Geo / IP reputation firewall aliases free threat feeds

Zenarmor — not worth it here on free

Zenarmor's value (web categories, threat intel, user reporting) is paid, and it does netmap DPI, so it conflicts with inline Suricata on the same interface — one or the other, not both. On a free budget keep Suricata + CrowdSec and do web filtering on the dedicated DNS servers.

11a. Suricata IDS/IPS (Services → Intrusion Detection)

Base component — no plugin. Inspect the internet-facing WAN only (server VLANs run uninspected at line rate; only the CARP master sees transit).

  • Enabled, IPS mode (inline netmap), Promiscuous, Pattern matcher = Hyperscan.
  • Interfaces = WAN_A + WAN_B (ixl7/ixl2). Not lagg0 / server VLANs.
  • Runmode = workers, NUMA-local (tunables §12).
  • Rules — ET Open rulesets, daily update; run IDS (alert-only) ~a week to baseline, then flip chosen categories to drop (IPS); suppress known-good noise.
  • Home networks = the public /28 + routed server subnets.
  • Logging = EVE JSON → the monitoring collector (10.128.34.x, VLAN 934).

Disable offload + master-only

Inline netmap needs LRO/TSO/LSO off on WAN_A/WAN_B (leave offload on elsewhere). Only the MASTER forwards, so only it inspects — no CARP hook needed.

11a-i. IPS bring-up & troubleshooting (netmap on ixl/X710)

"No alerts" or "WAN drops when I turn on IPS" is almost always the netmap ↔ NIC-offload interaction, and it bites harder in IPS mode because inline Suricata is a bump-in-the-wire — if netmap won't attach cleanly you don't just lose alerts, you lose the WAN. Our edge NICs are X710 (ixl / i40e), the fussiest family for netmap, so follow this order rather than flipping IPS on blind.

1 — Prove you can see alerts in IDS first. If plain IDS gives no info, IPS won't either; the fault is upstream (interface/rules/offload), not the mode.

  • Settings: Enabled, IPS mode OFF for now, Promiscuous, Hyperscan; Interfaces = WAN_A + WAN_B (ixl7/ixl2) only; Home networks = the public /28 + routed server subnets.
  • Download tab → enable ET Open, tick categories, Download & Update. No enabled rules = no alerts — the most common "no info" cause.
  • Apply, then fire a known trigger and check Intrusion Detection → Alerts:
    curl http://testmynids.org/uid/index.html      # trips ET 2100498
    
    Shows up → pipeline works, go to step 2. Nothing → it's offload or the wrong interface.

2 — Turn off interface offload (the #1 ixl breaker). Interfaces → Settings (global):

  • Disable hardware TSO ✔, Disable hardware LRO ✔, Disable hardware checksum offload ✔.
  • VLAN Hardware Filtering → Disable — X710/i40e specifically needs this; leaving it on is the classic "netmap attaches but no packets pass" symptom.
  • Reboot — offload changes don't fully take effect live.

3 — Enable IPS. Settings → IPS mode ✔, save, apply. Keep every rule at alert for ~a week (full visibility, nothing dropped), then flip chosen categories to drop. In OPNsense a drop rule both blocks and alerts, so IPS never costs you info — going inline gives you strictly more (each event carries action: allowed vs blocked).

4 — Getting the info out. GUI Alerts tab is live-only; the durable record is EVE JSON (Settings → Logging → EVE output) shipped to the monitoring collector on VLAN 934 (10.128.34.x). Point Zabbix there, not at the GUI. alert.action: blocked = a drop rule enforced; allowed = alert-only saw + logged it.

Do the inline cutover on the BACKUP node, in a window

IPS is inline on WAN — a failed netmap attach drops internet. On the CARP pair, enable and verify IPS on fw-colo-2 while it is BACKUP (idle netmap, no live traffic at risk), fail over, verify, then repeat on the other node. Never flip both masters at once.

ixl/X710 specifics. Native i40e netmap works but is touchier than ix/igb:

  • Watch Intrusion Detection → Log File immediately after enabling IPS for a netmap attach line vs an error — an attach failure there means traffic won't pass inline.
  • If native won't attach it falls back to emulated netmap (slower, but works) — confirm with dmesg | grep -i netmap.
  • Multiqueue can trip it; if the interface flaps, cut the ixl queue count and retest.

Verify:

suricata --build-info | grep -i netmap        # NETMAP support: yes
dmesg | grep -i netmap                          # attach lines, no errors
tail -f /var/log/suricata/eve.json              # live events incl. action:blocked
curl http://testmynids.org/uid/index.html       # now BLOCKED, not just alerted

11b. CrowdSec (Services → CrowdSec — os-crowdsec, free)

Behavioural detection (brute-force, scanning) from logs + the free Community Blocklist; the firewall bouncer drops offenders at pf. Complements Suricata (signatures), no overlap.

  • Install os-crowdsec; enable the agent + the firewall bouncer.
  • Add the OPNsense/Suricata collections (log parsers + scenarios) from the Hub.
  • Enrol in the free Console (app.crowdsec.net) to receive the Community Blocklist — this shares your detection signals (reputation metadata, not payloads). Skip enrolment to stay fully local (then you only block what you detect — no community list).
  • HA/CARP: run it on both nodes; the community list is identical, locally-learned bans accrue only on the MASTER — fine, the new master keeps the list and relearns after a failover. Deploy on fw-guest-1 too (guest Wi-Fi is the best candidate of all).

11c. DNS enforcement (filtering is on the dedicated DNS servers)

Web/category filtering runs on separate DNS servers, not OPNsense — the FW only guarantees clients can't bypass them:

  • Allow client DNS (UDP/TCP 53) only to the filtering servers — NAT-redirect or drop all other outbound 53 (a device hardcoded to 8.8.8.8 is redirected or blocked).
  • Block DoT — outbound TCP 853.
  • Block DoH — a DoH-endpoints host alias (public lists exist); for user VLANs also consider blocking UDP/443 (QUIC) so browsers can't DoH-over-QUIC around it.
  • Hand out the filtering servers via DHCP at the venue routers (staff/POS/office scopes); edge enforcement is the backstop.

11d. Geo / IP-reputation aliases (Firewall → Aliases)

Cheap, high-signal drops on WAN-in (and forward):

  • URL-table aliases from free feeds — Spamhaus DROP/EDROP, abuse.ch, FireHOL level 1 — action block, logged.
  • GeoIP country blocks for regions you never transact with (needs the free MaxMind GeoLite key).
  • POS stays locked down regardless — payment prefixes + AD egress only (venue design); a till should never reach the web filter's remit in the first place.

12. Tunables & performance (System → Settings → Tunables)

Design targets from the architecture doc — set under System → Settings → Tunables (sysctl apply live; loader tunables need a reboot). Validate the exact keys against your OPNsense/FreeBSD build.

Tunable Value Why
net.isr.maxthreads -1 one netisr thread per core
net.isr.bindthreads 1 pin them (no migration)
net.isr.dispatch deferred scale RX across cores
kern.ipc.nmbclusters 1000000 mbuf headroom at 10G+
kern.ipc.nmbjumbo9 524288 9k jumbo mbufs (server trunk)
net.pf.states_hashsize 1048576 pf state hashing (loader → reboot)
net.route.multipath 1 ECMP (also §6)
dev.ixl.<n>.iflib.override_nrxds / …override_ntxds 4096 deeper X710 RX/TX rings, per unit (loader → reboot)

GUI settings (not tunables):

  • Firewall → Settings → Advanced → Firewall Maximum States = 3–5 M; Optimization = conservative (long-lived server flows).
  • Firewall → Settings → Normalization → scrub on WAN only; never scrub or set skip the server trunk (line-rate).
  • System → Settings → Miscellaneous → Power → performance (powerd off / hi-adaptive off); Cryptography = AES-NI; cap C-states at C1 for latency.
  • X710 in x8 PCIe slots, NUMA-balanced, HT on (BIOS — see arch doc §Hardware).

13. Verify

# HA
System → High Availability → Status   -> node-1 MASTER, node-2 BACKUP, all VIPs green
pfSync: Interfaces → PFSYNC            -> states counting up on both

# BGP (Routing → Diagnostics, or shell: vtysh) — run on BOTH nodes; FRR is warm on both
show ip bgp summary                    -> bdr-1, bdr-2, CR-COLO-01/02 all Established (both nodes)
show bfd peers brief                   -> CR-COLO-01/02 up
show ip bgp 0.0.0.0/0                  -> default in from WAN on BOTH nodes (bdr-1 pref 200)
# MASTER only:
netstat -rn -f inet | grep -E '185\.109\.43\.176|^10\.128'  -> both anchors
#   (the /16 prints ABBREVIATED: "10.128/16" — FreeBSD trims trailing zero octets)
show ip bgp neighbors <bdr-ip> advertised-routes -> /28 advertised
show ip bgp neighbors <ccr-ip> advertised-routes -> 0/0 + /28 + 10.128.0.0/16
# on a CCR: /16 active via eBGP from the master; its floating blackhole INACTIVE
# BACKUP: sessions Established; WAN advertised-routes EMPTY; CORE advertised-routes =
# ONLY the MED-100 default (the §7.1 exception). Anything more: run frr-anchor-sync.sh

# ACME
Services → ACME Client → Certificates  -> issued, expiry ~90d, GUI using it
# LLDP
Services → LLDPd → Neighbors           -> CR-COLO-1/2, bdr-1/2, sw-server-1/2, OOB-2
# IDS/IPS
Services → Intrusion Detection → Alerts -> events on WAN only; EVE JSON reaching 10.128.34.x
# CrowdSec
Services → CrowdSec → Decisions        -> bouncer active; community blocklist populated
# DNS enforcement (from a staff/POS test host)
dig @8.8.8.8 example.com               -> BLOCKED/redirected; :853 and DoH endpoints unreachable
# NAT
Firewall → NAT → Outbound              -> pool rule (Round Robin over 185.109.43.180/30),
                                          POS fixed-IP rule ABOVE it

Failover test: on fw-colo-1, Interfaces → Virtual IPs → Enter persistent CARP maintenance mode. Node-2 goes MASTER, its hook adds the anchor and the /28 UPDATE goes out on the already-Established sessions — total gap ~1–2 s (watch ping from outside to a /28 service IP; more than ~3 s lost means the hook didn't fire — run frr-anchor-sync.sh and check §7.6). Exit maintenance → node-1 reclaims, node-2's hook withdraws.