Firewall — OPNsense edge pair: setup runbook¶
Step-by-step build of the corporate edge HA pair fw-colo-1 / fw-colo-2.
This is the how; see firewall-opnsense for the why
(interface plan, zone policy, BGP design, tuning). Read that first.
Starting point (assumed done)
OPNsense ISO installed on both nodes, and the OOB management IP set on each
from the console: fw-colo-1 = 172.16.201.41, fw-colo-2 = 172.16.201.42
(/24, gateway = the OOB router). The GUI is reachable over OOB.
Order of operations¶
Build fw-colo-1 (master) fully first, verify it standalone, then bring up
fw-colo-2 with only its node-specific bits (interface IPs, OOB, CARP advskew) and
let XMLRPC config-sync push everything else. Enabling HA sync before node-1 is
correct will replicate mistakes. Rough sequence:
- Base system + plugins (both nodes)
- Interfaces / LAGGs / VLANs + addressing (both nodes — per-node IPs; never syncs)
- CARP VIPs (node-1; sync pushes to node-2)
- pfSync + HA config-sync (both)
- FRR/BGP (both nodes — per-node router-id + peer
/31s; not synced) - HA advertisement gating — anchor scripts (both — files on disk, not synced)
- ACME + LLDP + firewall/NAT (node-1; syncs)
- Verify
What XMLRPC sync does and does NOT replicate
Syncs (node-1 → node-2): firewall rules, NAT, aliases, CARP VIPs, certificates,
ACME, DHCP, users, schedules. Per-node, set on both manually: interface
assignments + IPs, hostname, FRR/BGP (own router-id + own /31 peers — the
prefix-lists/route-maps are identical except node-2's CORE MED, §7), the anchor
gating scripts, and the advskew that elects the master. So "build one, sync the rest" applies to the
policy layer — the interface and routing layers are per-node on both.
Throughout: Save each page, then Apply changes (the amber banner). Nothing takes effect until applied.
1. Base system (both nodes)¶
System → Settings → General
- Hostname
fw-colo-1/fw-colo-2, domaininfra.jsm(internal/private). Note: the GUI TLS cert and the name you browse to are the public FQDN (fw-colo-<n>.infra.pubinvest.co.uk, §8) — ACME can't validate the privateinfra.jsmzone, and split-horizon DNS resolves the public name to the OOB IP internally. - DNS servers = internal resolvers
10.128.32.3,10.128.32.4; untick "Allow DNS server list to be overridden by DHCP/PPP on WAN". - Timezone
Europe/London.
System → Settings → Administration
- Protocol HTTPS, Listen interfaces = OOB only (lock the GUI to the OOB net).
- SSH: enable, listen on OOB only, key-only, root login off.
- Create an admin user; enable TOTP 2FA (System → Access → Servers → add a TOTP
server, tie it to the login). Change the
rootpassword.
System → Settings → Cron / Services → NTP — point NTP at your internal source.
2. Plugins (both nodes)¶
System → Firmware → Updates — update to the latest release first, reboot.
System → Firmware → Plugins — install:
| Plugin | Purpose |
|---|---|
os-frr |
FRR — the BGP engine (WAN + CORE) |
os-acme-client |
ACME / Let's Encrypt certificates |
os-lldpd |
LLDP (neighbour discovery, feeds LibreNMS/Zabbix) |
os-crowdsec |
CrowdSec agent + firewall bouncer (free; §11b) |
os-zabbix-agent (optional) |
monitoring agent (see the arch doc) |
Reboot is not required; the new menus appear immediately.
3. Interfaces, LAGGs & VLANs¶
This is the real port→role map we confirmed from LLDP and cabled in NetBox — not the
generic "8×SFP+" plan. Driver names: igb* = i350 1G copper, ixl* = X710 10G SFP+,
ue* = the USB-eth OOB NIC. Both nodes are identical except the OOB NIC (ue1 on
fw-colo-1, ue0 on fw-colo-2). Build the LAGGs and VLANs before assigning interfaces.
| OPNsense port | Role | Cabled to |
|---|---|---|
ixl7 |
WAN A (primary) | bdr-1 sfp-sfpplus11/12 |
ixl2 |
WAN B (backup) | bdr-2 sfp-sfpplus11/12 |
ixl3 |
CORE A | CR-COLO-1 sfp-sfpplus1/2 |
ixl6 |
CORE B | CR-COLO-2 sfp-sfpplus1/2 |
ixl0 + ixl5 |
server trunk → lagg0 |
sw-server-1 + sw-server-2 (MLAG) |
igb0 + igb1 |
pfSync → lagg1 |
the peer FW, direct |
ue1 / ue0 |
OOB mgmt | OOB-2 ether1 / ether2 |
Interfaces → Devices → LAGG — create two:
| Device | Members | Proto | MTU | Purpose |
|---|---|---|---|---|
lagg0 |
ixl0, ixl5 |
LACP | 9000 | server trunk → the two MLAG server switches |
lagg1 |
igb0, igb1 |
Failover (active-backup) | 1500 | pfSync, direct to the peer |
Interfaces → Devices → VLAN — on parent lagg0, tag the server VLANs. All are
routed here (gateway = a CARP VIP ….1 on OPNsense); the switch carries them L2.
| VLAN | Interface | Subnet | Gateway (CARP VIP) | Purpose |
|---|---|---|---|---|
| 910 | BACKUP |
10.128.10.0/24 |
10.128.10.1 |
Backup traffic |
| 932 | VM_APPS |
10.128.32.0/24 |
10.128.32.1 |
AD & apps (PXE/stock) |
| 933 | VM_K8S |
10.128.33.0/24 |
10.128.33.1 |
Kubernetes |
| 934 | VM_MON |
10.128.34.0/24 |
10.128.34.1 |
Monitoring VMs |
| 935 | VM_DATA |
10.128.35.0/24 |
10.128.35.1 |
Data / SQL |
| 936 | VM_DEV |
10.128.36.0/24 |
10.128.36.1 |
Development |
| 940 | HYP_MGMT |
10.128.40.0/24 |
10.128.40.1 |
Hypervisor management |
L2-only VLANs — no OPNsense interface
Corosync 950 (Proxmox cluster ring) and Ceph 920/921/922 are L2-only and are not routed on OPNsense — the Proxmox/Ceph nodes talk directly on the VLAN. Do not create OPNsense VLAN interfaces, gateways or CARP VIPs for them; they're tagged on the server switch only.
Interfaces → Assignments — assign and name:
| Name | Device | Role |
|---|---|---|
WAN_A |
ixl7 |
eBGP /31 → bdr-1 (primary) |
WAN_B |
ixl2 |
eBGP /31 → bdr-2 (backup) |
CORE_A |
ixl3 |
eBGP /31 → CR-COLO-1 |
CORE_B |
ixl6 |
eBGP /31 → CR-COLO-2 |
BACKUP |
lagg0 vlan910 |
backup (MTU 1500) |
VM_APPS |
lagg0 vlan932 |
AD & apps |
VM_K8S |
lagg0 vlan933 |
Kubernetes |
VM_MON |
lagg0 vlan934 |
monitoring |
VM_DATA |
lagg0 vlan935 |
data / SQL |
VM_DEV |
lagg0 vlan936 |
development |
HYP_MGMT |
lagg0 vlan940 |
hypervisor mgmt |
PFSYNC |
lagg1 |
HA sync |
OOB |
ue1 / ue0 (already assigned) |
management 172.16.201.41 / .42 |
MTU: set 9000 on lagg0 itself (the parent); the VLAN children stay 1500
(routed). Core/WAN /31s are 1500.
Addressing (per node)¶
Static IPv4 on each (fw-colo-1 shown; fw-colo-2 takes the .b side of each /31 and
.42/second-in-subnet):
| Interface | port | fw-colo-1 | fw-colo-2 |
|---|---|---|---|
CORE_A → CR-COLO-1 |
ixl3 |
10.254.96.5/31 |
10.254.96.7/31 |
CORE_B → CR-COLO-2 |
ixl6 |
10.254.96.9/31 |
10.254.96.11/31 |
WAN_A → bdr-1 |
ixl7 |
185.109.43.21/31 (peer .20) |
185.109.43.23/31 (peer .22) |
WAN_B → bdr-2 |
ixl2 |
185.109.43.25/31 (peer .24) |
185.109.43.27/31 (peer .26) |
PFSYNC |
lagg1 |
10.254.240.1/30 |
10.254.240.2/30 |
| each server VLAN | lagg0.<vid> |
10.128.<x>.2/24 |
10.128.<x>.3/24 |
Transit
/31s and pfSync are per-node point-to-point (no VIP) — each node runs its own BGP sessions; use the/31values NetBox holds forixl2/3/6/7.
CARP: a real per-node IP and the floating VIP — you need both
Each routed server-VLAN interface gets a real, unique IP on each node (.2 on
fw-colo-1, .3 on fw-colo-2), plus the floating CARP VIP .1 that is the
clients' gateway (§4). OPNsense requires the interface to already hold a real address
in that subnet before a CARP VIP can bind, and each node needs its own IP to source
traffic, sync, and stay reachable when it is BACKUP. E.g. VLAN 932 — fw-colo-1
10.128.32.2/24, fw-colo-2 10.128.32.3/24, CARP VIP 10.128.32.1/24.
4. CARP virtual IPs¶
Interfaces → Virtual IPs → Settings — add a CARP VIP for each floating server-VLAN
gateway. These sit on directly-attached L2 segments, so they float via CARP
(gratuitous ARP). The public /28 is not here — it's a routed block (see the note).
| VIP | Interface | vhid | Notes |
|---|---|---|---|
10.128.10.1/24 |
BACKUP | 10 | backup gateway |
10.128.32.1/24 |
VM_APPS | 32 | AD & apps gateway |
10.128.33.1/24 |
VM_K8S | 33 | Kubernetes gateway |
10.128.34.1/24 |
VM_MON | 34 | monitoring gateway |
10.128.35.1/24 |
VM_DATA | 35 | data/SQL gateway |
10.128.36.1/24 |
VM_DEV | 36 | development gateway |
10.128.40.1/24 |
HYP_MGMT | 40 | hypervisor-mgmt gateway |
The transit /31s are not CARP (per-node point-to-point).
4.1 Pick the right Mode (VIP type)¶
| Mode | What it does | Use it for |
|---|---|---|
| CARP | Floating HA address — advertises via CARP (proto 112), grat-ARPs on failover. Needs VHID + password, and a real IP already on that interface in the same subnet. | Server-VLAN gateways (10.128.x.1) — attached L2, must float between the pair |
| IP Alias | An extra address bound to an interface; no HA protocol. Can optionally be parented to a CARP VIP so it follows that VIP's state. Services can bind to it. | The routed /28 (SNAT pool, inbound targets) — and anything you need to bind a service to |
| Proxy ARP | FW answers ARP for an address it doesn't bind; can't host services. | Rare — NAT-only addresses on an attached subnet you hold no IP in |
| Other | FW neither ARPs nor binds — just "knows" the address for NAT/routing. | A routed block used purely for NAT — the purist choice for our /28 if you never bind a service to it |
For our /28 we use IP Alias (flexible — you can bind a service to a public IP later);
Other is equally correct if it stays strictly NAT.
4.2 CARP VIP — field by field¶
| Field | Value / meaning |
|---|---|
| Mode | CARP |
| Interface | the VLAN interface the gateway lives on (e.g. VM_APPS) |
| Address | the gateway with the VLAN's mask — e.g. 10.128.32.1/24 |
| Virtual IP Password (+ confirm) | shared secret for this VHID — must match on both nodes |
| VHID Group | 1–255. Identical on both nodes for the same VIP, and unique within that L2 broadcast domain |
| Advertising Frequency → Base | seconds between adverts — leave 1 |
| Advertising Frequency → Skew | priority, lower wins: 0 on fw-colo-1, 100 on fw-colo-2 |
| Peer | leave blank for normal multicast CARP. Set it to the partner's real interface IP on that segment only if you need unicast CARP (multicast filtered/not carried on that L2) |
| Description | e.g. VM_APPS gw (VLAN 932) |
Three CARP rules people trip on
- The interface must already hold a real per-node IP in that subnet (
.2/.3) — a CARP VIP cannot bind otherwise (see §3 addressing). - VHID must be identical on both nodes and unique per broadcast domain. Separate VLANs may reuse a VHID (different L2), but keep a scheme — we use the VLAN-derived numbers 10/32/33/34/35/36/40 so it stays obvious.
- advskew elects the master. If both nodes end up with the same skew, election is
non-deterministic — always set
0/100explicitly.
Do I create the VIPs on both nodes? — No, node-1 only
Virtual IPs are a synced section, so create them on fw-colo-1 and let the
XMLRPC config-sync (§5) push them to fw-colo-2. Two caveats:
- Prerequisite that does not sync: the per-node interface IPs must already exist
on BOTH nodes (
.2/.3, §3). Interfaces never sync, and a synced CARP VIP cannot bind on node-2 if that interface holds no real address in the subnet — so do interfaces/IPs on both first, then VIPs on node-1. - Verify after the first sync: node-2 should list the VIPs and show BACKUP (Interfaces → Virtual IPs → Status). If that's what you see, leave the skew alone — your build adjusted it on the backup.
Same applies to the routed-/28 IP-Alias VIPs: created on node-1, synced to node-2,
harmless there because the backup never advertises the /28 (the CARP-keyed anchor,
§7) — no traffic is attracted to it.
Editing advskew on the secondary — it won't stick while VIPs are synced
You can change advskew on fw-colo-2, but the next config-sync that includes Virtual
IPs pushes node-1's value straight back over it — it is not durable. If you truly need a
fixed difference, untick "Virtual IPs" from the sync sections and manage the VIPs by
hand on both nodes (identical address / VHID / password, skew 0 and 100).
And if both nodes claim MASTER, advskew is usually NOT the cause. Diagnose in this order:
- Firewall allows protocol 112 on that interface, in both directions — blocked CARP advertisements is the classic split-brain.
- VHID and CARP password identical on both nodes.
- Then look at advskew.
4.3 Global CARP behaviour (System → Settings → Tunables)¶
| Tunable | Value | Why |
|---|---|---|
net.inet.carp.preempt |
1 |
all VIPs fail over together — if one CARP interface demotes, the rest follow. Without it you get split routing (some VLANs on node-1, some on node-2) |
net.inet.carp.log |
2 |
log CARP state transitions (essential when debugging flaps) |
net.inet.carp.allow |
1 |
CARP enabled (default) |
4.4 Firewall + status¶
- Allow CARP on every interface carrying a VIP — protocol 112 from the peer. If the nodes can't hear each other's adverts both go MASTER (split-brain) — this is the single most common CARP failure.
- Interfaces → Virtual IPs → Status — MASTER/BACKUP per VIP, plus "Enter persistent CARP maintenance mode": use that for planned failover tests and upgrades (it demotes cleanly instead of yanking the box).
The public /28 is routed — NOT a CARP VIP
Don't create CARP VIPs for 185.109.43.176/28. The ISP routes it to you over the
/31 transit and you advertise it via BGP. HA is by the CARP-keyed anchor (§7):
FRR runs warm on both nodes, but only the node holding the anchor route — the CARP
MASTER — announces the block (next-hop = its own /31); on failover the anchor moves
and the new master's already-Established sessions send the UPDATE in ~1 s. The specific
addresses you use (the /30 SNAT pool + inbound service IPs, §10) are added as
IP-Alias VIPs (Type: IP Alias, not CARP) on the WAN interface or a loopback,
synced to both nodes — harmless on the backup since it advertises nothing, so no
traffic ever arrives there.
5. HA — pfSync + config sync¶
System → High Availability → Settings
- Synchronize States (pfSync): enable; Synchronize Interface =
PFSYNC; Synchronize Peer IP = the other node's PFSYNC address (10.254.240.2on node-1,.1on node-2). - Synchronize Config to backup (XMLRPC): configure on
fw-colo-1only. - Synchronize Config to IP =
10.254.240.2(node-2 PFSYNC). - Remote user = an admin account that exists on node-2; HTTPS.
- Tick the sections to replicate: Firewall Rules, NAT, Aliases, Virtual IPs, Schedules, Certificates, ACME, DHCP, Users/Groups — but NOT "Interfaces" or "System/General" (node-specific: IPs, hostname, advskew), and not FRR (per-node peers — configured directly on each node in step 6).
Firewall → Rules → PFSYNC — allow the pfSync protocol between the two node IPs on the PFSYNC interface (it's a direct cable, but the rule is still required).
CARP heartbeats: ensure the CARP interfaces (WAN, a server VLAN, and PFSYNC) can
exchange multicast — default allow-CARP rules on those interfaces. advskew already
elects node-1.
6. FRR / BGP (Routing → FRR)¶
Configure this on both nodes — FRR is not part of XMLRPC config-sync, and each
node peers over its own /31s with its own router-id. The prefix-lists and
route-maps below are identical on both with one deliberate exception: node-2's
RM-CORE-OUT adds set metric 100 (§7 — the static CORE-side preference). Only the
neighbour peer addresses differ otherwise. Edge AS 65510. FRR runs permanently on
both nodes — HA is done by gating advertisements, not the daemon (§7).
Routing → General — Enable FRR. Router ID = a stable per-node IP (the OOB or a loopback); BFD supported.
Routing → BGP → General
- Enable, AS number
65510, Network import/export via prefix-lists (below). - Advanced:
bestpath as-path multipath-relax, and setnet.route.multipath=1(System → Settings → Tunables) for ECMP.
Originate the three anchored prefixes — but do NOT anchor them statically. Routing → BGP → General → Networks add:
| Network | Advertised to | What it is |
|---|---|---|
185.109.43.176/28 |
WAN + CORE | the public block |
10.128.0.0/16 |
CORE | the whole colo block (the addressing plan's server-room /16). Connected server /24s win as more-specifics on the master; everything else in it — storage 12–17, migration 45, corosync 50/51, OOB 9, unused — hits the master's blackhole ("never routes", enforced edge-wide). OOB is mgmt-VPN-reached; if it ever needs CCR-side reachability, inject the more-specific 10.128.9.0/24 from its gateway |
CCR-side PREREQUISITE: flip the CCRs' /16 blackhole to a floating backstop (distance > 20)
The colo CCRs carry 10.128.0.0/16 in their own BGP networks list, anchored by
their own blackhole static at distance 1. Left like that, this design is dead on
arrival: AD 1 beats eBGP (20) on an identical prefix, the FW's /16 never
installs, and the CCRs blackhole all server traffic on both nodes.
The fix is one attribute — raise the CCR blackhole's distance above 20 (e.g. 200) and keep both the static and the networks entry:
- Normal: the FW master's /16 (eBGP 20) wins, installs, and propagates metro-wide; the CCR's floating static is inactive, so its own origination goes quiet automatically.
- Failover gap / FW pair down: the FW /16 vanishes → the floating static re-activates instantly → the CCR resumes originating and blackholing. Metro steering stays stable and colo-bound traffic gets a clean drop at the CCR instead of following a default — identical protection to the old CCR-owned design, now as a backstop.
Do the CCR change before switching the FW to the /16 (migration §7, step 0 —
automated in ansible/network roles/rr: rr_aggregate_distance + a to-edge
rule rejecting dst in 10.128.0.0/16, so the RRs also never send the /16 to the
FW — the RR's own origination has no 65510 in path, so the FW's AS-path loop
detection can't drop it and it could otherwise satisfy the backup's network
import-check. FW-originated prefixes like the /28 need no such guard: any copy
coming back already carries 65510 → loop-dropped.)
FRR only announces a network statement while a matching route exists in the RIB
(assert bgp network import-check is on — default in current FRR; if your build has
it off, the whole §7 gate is void). The blackhole routes that satisfy these are added
and removed by the CARP scripts in §7 — that is the HA gate. Do not list the
individual server /24s as Networks: they're connected on both nodes, so they'd
advertise from the backup too and re-open the hole the aggregates close (§7.1).
Never add a persistent static/Null0 route for ANY anchored prefix
Not in System → Routes, not as an FRR static — for the /28 or the
10.128.0.0/16. A persistent anchor makes that network statement fire on
both nodes permanently — the backup would advertise and attract, and the §7
gating is silently defeated. The only source of these routes is the CARP-keyed
script.
Prefix-lists (Routing → BGP → Prefix Lists)¶
| Name | Rule |
|---|---|
PL-DEFAULT-IN |
permit 0.0.0.0/0 exact |
PL-PUBLIC-OUT |
permit 185.109.43.176/28 exact |
PL-CORE-OUT |
permit 0.0.0.0/0, 185.109.43.176/28, and the anchored colo aggregate 10.128.0.0/16 (exact — not the individual connected /24s, which would leak from the backup; requires the CCR floating-backstop change, §6 danger box) |
Route-maps (Routing → BGP → Route Maps)¶
| Name | Action |
|---|---|
RM-WAN-IN-PRIMARY |
match PL-DEFAULT-IN, set local-preference 200 |
RM-WAN-IN-BACKUP |
match PL-DEFAULT-IN, set local-preference 100 |
RM-WAN-OUT |
match PL-PUBLIC-OUT (permit), else deny |
RM-WAN-OUT-BACKUP |
as RM-WAN-OUT + set as-path prepend 4200000001 4200000001 (the WAN local-as) |
RM-CORE-OUT |
match PL-CORE-OUT (permit), else deny. On fw-colo-2 ONLY: add set metric 100 (MED — the static CORE-side preference, §7) |
Neighbours (Routing → BGP → Neighbors)¶
WAN — to the ISP borders (AS 204258), presenting our customer AS 4200000001:
| Peer (fw-colo-1 / fw-colo-2) | Remote AS | local-as | BFD | In | Out |
|---|---|---|---|---|---|
bdr-1 — .20 / .22 (WAN_A ixl7) |
204258 | 4200000001 | no | RM-WAN-IN-PRIMARY |
RM-WAN-OUT |
bdr-2 — .24 / .26 (WAN_B ixl2) |
204258 | 4200000001 | no | RM-WAN-IN-BACKUP |
RM-WAN-OUT-BACKUP |
The peer IP is PER NODE — don't copy node-1's neighbours to node-2
Each FW has its own /31 to each border (see §3). fw-colo-1 peers 185.109.43.20
/ .24; fw-colo-2 peers 185.109.43.22 / .26. Pointing fw-colo-2 at .24 reaches
bdr-2 (it's routable) but that link's neighbour is fw-colo-1 (.25), so bdr-2 rejects
the OPEN and resets — you'll see [FSM] unexpected packet received in state OpenSent
followed by bgp_read_packet error: Connection reset by peer. The same two lines
appear if local-as 4200000001 is missing (bdr sees the wrong AS → Bad Peer AS).
→ The FRR instance AS is 65510 (used on the CORE side); on the WAN the ISP knows us as
our customer AS 4200000001, so set local-as 4200000001 no-prepend replace-as
on both WAN neighbours (clean single-AS path to the ISP). BFD is OFF on the WAN —
failover is driven by the tight keepalive/hold 3/9 s timers instead (BFD stays on the
CORE/CCR sessions only). bdr-1 primary (local-pref 200 in), bdr-2 backup (local-pref
100 + prepend ×2 out). Announce only the /28; receive default only.
CORE — to the colo CCRs (AS 65500):
| Peer | Remote AS | BFD | In | Out |
|---|---|---|---|---|
CR-COLO-01 (CORE_A /31) |
65500 | yes | (accept) | RM-CORE-OUT |
CR-COLO-02 (CORE_B /31) |
65500 | yes | (accept) | RM-CORE-OUT |
→ enable BFD per neighbour (single-hop /31, so it just works; the CCR side keys
off its bfd_enabled field). Announce 0/0 + public /28 + the anchored colo
aggregate (10.128.0.0/16 — §7; needs the CCR floating-backstop prerequisite, §6);
receive the server-room/internal routes.
Attaching the policy — what to select on each neighbour¶
On Routing → BGP → Neighbors attach the route-maps (In/Out) — not the
prefix-lists. Our route-maps already match the prefix-list and set the attribute
(local-pref / prepend), so the prefix-list is referenced by the map:
| Neighbour | Route-Map In | Route-Map Out |
|---|---|---|
bdr-1 (WAN primary) |
RM-WAN-IN-PRIMARY |
RM-WAN-OUT |
bdr-2 (WAN backup) |
RM-WAN-IN-BACKUP |
RM-WAN-OUT-BACKUP |
CR-COLO-1 / CR-COLO-2 |
(blank — accept) | RM-CORE-OUT |
The neighbour form also has Prefix-List In/Out boxes — only needed if you're filtering without setting attributes. If you attach both in the same direction, both must permit (they're ANDed). For this design: route-maps only. Leaving CORE's In blank accepts everything the CCRs send (trusted internal peers); attach a prefix-list there too if you want belt-and-braces.
Do I need a deny at the end of the prefix-list? — No
FRR prefix-lists and route-maps both have an implicit deny at the end. A
prefix-list of only permit entries denies everything else; a route-map with a single
permit sequence denies anything that doesn't match. A trailing explicit deny is
optional documentation, not a requirement.
Does the route-map itself enforce the prefix-list? — Yes
The map's one permit clause matches the prefix-list, and anything that doesn't match falls through to the map's implicit deny. So attaching the route-map is what applies the list — you don't attach the prefix-list separately.
Two ways to silently break that:
- A clause with no
matchpermits everything (and applies itsset) — never add a catch-all permit to these maps. - A typo'd / missing prefix-list name means the match can never succeed, so every route hits the implicit deny and the peer carries zero prefixes. If a session is Established but empty right after a policy edit, check the name first.
Verify what the policy actually did (vtysh):
The real trap is le / ge, not the missing deny
A prefix-list entry matches exactly unless you add le/ge:
permit 0.0.0.0/0→ only the default route ✅ — this is "receive default only"permit 0.0.0.0/0 le 32→ every prefix ❌ — you'd accept the full table
Same outbound: permit 185.109.43.176/28 is exact; le 32 would also permit every
more-specific from it. Keep every list here exact — no le/ge.
Filters are not retroactive — soft-clear after changes
Changing a prefix-list or route-map does not re-evaluate already sent/received
routes. Soft-clear the session (vtysh):
clear bgp 185.109.43.20 soft in / clear bgp 185.109.43.20 soft out.
Also make sure the list/map exists before you reference it on a neighbour.
7. HA advertisement gating — warm FRR + CARP-keyed anchor (both nodes)¶
FRR runs permanently on both nodes; failover moves one kernel route. The old stop-FRR-on-BACKUP design cost 10–20 s per failover (daemon start → BGP Idle→Established → learn → advertise, all serial). Here every session is warm and Established on both nodes at all times, so a failover is a CARP promotion plus one BGP UPDATE on an already-open session (~1 s) — and the backup holds a live default route throughout.
7.1 Design & invariants¶
| # | Invariant | Enforced by |
|---|---|---|
| I1 | Both nodes hold 0.0.0.0/0 at all times |
warm WAN sessions; in-policy is role-independent |
| I2 | Every traffic-attracting prefix (185.109.43.176/28, 10.128.0.0/16) is advertised iff the node holds its anchor route |
FRR network + import-check semantics — by construction, not by script |
| I3 | Exactly the CARP MASTER holds the anchors (both move together) | CARP syshook + boot hook + 1-min reconcile, all calling one idempotent script whose only input is live kernel CARP state |
| I4 | CCRs prefer node-1's default whenever it advertises one | static MED 100 on node-2's RM-CORE-OUT (§6) — never touched at runtime; also the deterministic tie-break during any ≤60 s dual-advertise reconcile window |
Strictness model — the backup attracts nothing, anywhere. All inbound attraction
is anchor-gated: the /28 on WAN, and the whole colo block toward the CCRs via the
10.128.0.0/16 aggregate (§6). The aggregate exists precisely because the
individual server /24s are connected on both nodes and therefore un-gateable
(their routes always exist, so a network statement for them would always fire —
including on the backup). A strict superset is gateable: only the anchor satisfies
it. The FW owning the /16 requires the CCR floating-backstop prerequisite (§6
danger box) — with the CCRs' blackhole left at distance 1, the FW's /16 never
installs. Consequences, all wanted:
- Venue→server traffic can only ever enter the master — including the corner where a demoted node still has live CORE sessions (it lost its anchors, so it advertises no server space). No pfSync-dependent inbound asymmetry remains.
- On the master, connected
/24s win as more-specifics inside the/16; everything else in it — storage 12–17, migration 45, corosync 50/51, OOB 9, unused space — hits the FW's blackhole and dies at the edge instead of leaking out the default. "Storage never routes," enforced for the entire colo block. - Whenever no FW advertises the /16 (failover gap, both nodes down), the CCRs' floating blackhole re-activates and they resume originating it — metro steering stays stable and colo-bound traffic gets a clean drop at the CCR, not a loop.
The one deliberate exception: the default toward the CCRs. Both nodes advertise
0.0.0.0/0 at all times (node-2 at MED 100). It cannot be anchor-gated: the backup
must hold a live default for itself (I1), and a route that must exist on both nodes
can't act as a one-node gate. The alternatives fail the robustness bar — conditional
advertisement (exist-map) is scan-timer-driven and not persistable through os-frr
config regeneration (the same reason §7.7 rejects the vtysh gate). What the exception
costs: a demoted node with live WAN+CORE can still carry venue→outbound traffic
(better MED) — an outbound-only, pfSync-covered asymmetry in an already-degraded state.
What it buys: the CCRs hold two pre-installed defaults, so venue internet survives
a master death on BFD detection alone (~1 s), with no UPDATE round-trip. Every
inbound flow — the ones a stateful firewall genuinely can't tolerate asymmetric —
is strict by construction.
7.2 Why a kernel route is the gate (the guarantee argument)¶
The dynamic state — "am I the announcer?" — lives in the kernel routing table, deliberately outside FRR and outside OPNsense config:
- Immune to the FRR/GUI lifecycle. os-frr regenerates
frr.confon every GUI edit and reload; anything injected viavtyshis wiped. A kernel route survives all of it —frr.confstays 100 % static and identical to what the GUI manages. - Reboots fail closed. Kernel routes don't persist, so a booting node comes up not advertising until CARP resolves and the hook/boot-sync runs. A rebooting or crashed-and-returned node can never dual-attract.
- Atomic + idempotent. The transition is one
route add/route delete; the sync script can run any number of times, from any trigger, in any order. - Failure directions are asymmetric on purpose. Residual wrong states degrade to dual-advertise (upstream best-paths one node; pfSync-synced states keep the stray flows alive) rather than master-withdrawn (an outage). The ≤60 s reconcile bounds both.
7.3 The scripts (both nodes; node-local, never XMLRPC-synced)¶
One code path. Every trigger calls the same sync script; the CARP hook's arguments are deliberately ignored (their format varies by version — live kernel state is the single source of truth).
/usr/local/opnsense/scripts/frr-anchor-sync.sh (mode 0755):
#!/bin/sh
# Idempotent: assert ALL attraction anchors from live CARP state. Safe to run any time.
# Anchors = every prefix in §6's Networks table. All move together (I3).
ANCHORS="185.109.43.176/28 10.128.0.0/16"
VHID="10" # the tracked CARP vhid (SRV_T1) — set per §4
STATE=$(ifconfig | awk -v v="$VHID" '$1=="carp:" && $4==v {print $2; exit}')
RTAB=$(netstat -rn -f inet)
for A in $ANCHORS; do
# FreeBSD netstat TRIMS TRAILING ZERO OCTETS from network routes —
# 10.128.0.0/16 prints as "10.128/16". Match both forms, or the check
# never sees the route (re-adds forever as MASTER; NEVER REMOVES as
# BACKUP = fail-open).
ABBR=$(printf '%s' "$A" | sed -E 's#(\.0)+/#/#')
HAVE=$(printf '%s\n' "$RTAB" | awk -v a="$A" -v b="$ABBR" '$1==a || $1==b {print "yes"; exit}')
if [ "$STATE" = "MASTER" ] && [ -z "$HAVE" ]; then
if route -q add -net "$A" 127.0.0.1 -blackhole; then
logger -t frr-anchor "MASTER: anchor $A added (advertising)"
else
# add refused with the slot empty per netstat = a foreign exact route
# (usually a BGP-learned copy of $A -> the RR to-edge guard is not
# deployed). Loud, not silent: this is a design-order violation.
logger -t frr-anchor "MASTER: FAILED to add anchor $A — foreign exact route? check: vtysh -c 'show ip bgp $A' and the RR to-edge guard (§6)"
fi
elif [ "$STATE" != "MASTER" ] && [ -n "$HAVE" ]; then
route -q delete -net "$A" \
&& logger -t frr-anchor "state=${STATE:-unknown}: anchor $A removed (withdrawn)"
fi
done
/usr/local/etc/rc.syshook.d/carp/20-frr-anchor (mode 0755) — event trigger:
#!/bin/sh
# CARP transition -> re-derive from live state (args intentionally ignored)
exec /usr/local/opnsense/scripts/frr-anchor-sync.sh
/usr/local/etc/rc.syshook.d/start/99-frr-anchor (mode 0755) — boot trigger, same
one-liner exec as above.
Reconcile every minute — register a cron action so it survives OPNsense config
management: /usr/local/opnsense/service/conf/actions.d/actions_frranchor.conf:
[sync]
command:/usr/local/opnsense/scripts/frr-anchor-sync.sh
type:script
message:sync FRR /28 anchor to CARP state
description:FRR anchor reconcile
then service configd restart and add it in System → Settings → Cron (every
minute). The reconcile is the bound on every "script didn't fire" failure mode.
7.4 Failure matrix¶
| Event | Behaviour | Converges in |
|---|---|---|
| Planned failover (CARP maintenance) | old master's hook withdraws all anchors, new master's hook advertises — both on warm sessions | ~1 s |
| Master power loss | CARP promotes (~1–3 s) → hook adds anchors → UPDATEs out. CCR-side server routes + default converge on BFD (~1 s); borders drop the dead node's stale /28 on hold expiry (9 s) — or ~1 s with WAN BFD (see 7.5) |
~2–4 s (CORE) / bounded by WAN detection (inbound /28) |
| Node demoted, CORE sessions still live | anchors move → it advertises no server space, no /28; only its MED-100 default remains (§7.1 exception) — inbound stays strict by construction | ~1 s |
| FRR reload / GUI edit (either node) | anchor is a kernel route — untouched; frr.conf is static |
no impact, by construction |
| Node reboot | kernel route gone → boots withdrawn (fail-closed); CARP resolves → boot/hook sync asserts | ≤ boot + seconds |
| Hook missing/failed on promotion | new master not advertising — the bad direction | ≤60 s (cron) + Zabbix alarm (7.6) |
| Backup wrongly holds anchor | dual-advertise; borders best-path one node; pfSync covers stray flows | ≤60 s (cron) |
| CARP split-brain (all heartbeat paths lost) | both advertise — BGP cannot fix a CARP split-brain in any design; mitigated upstream by multi-segment heartbeats (WAN + server-net + pfSync bond) | n/a — prevent at CARP layer |
7.5 Timer budget & the WAN-BFD lever¶
Warm sessions make detection the entire failover budget. CORE is already ~1 s
(BFD). On WAN, an unplanned master death leaves the borders best-pathing the dead
node's /28 until hold (9 s) even though the new master is already advertising —
only BFD (or brutal timers) shortens that. BFD-on-WAN is therefore worth reopening
with the ISP: it is now the difference between ~1 s and ~9 s of inbound loss on
master death, not a nice-to-have. Planned failovers don't care (the withdraw is
explicit) — keep 3/9 s as the floor either way.
7.6 Monitoring the invariant (Zabbix)¶
Divergence must be detected, not just reconciled. UserParameter on both nodes:
# 0 = consistent; 1 = CARP state and anchor-count disagree (both present as MASTER, 0 otherwise)
# NB: netstat prints 10.128.0.0/16 as "10.128/16" (trailing-zero trimming) — match both.
UserParameter=frr.anchor.consistent,STATE=$(ifconfig | awk '$1=="carp:" && $4=="10" {print $2; exit}'); N=$(netstat -rn -f inet | grep -cE '^(185\.109\.43\.176/28|10\.128(\.0\.0)?/16) '); { [ "$STATE" = "MASTER" ] && [ "$N" -eq 2 ]; } || { [ "$STATE" != "MASTER" ] && [ "$N" -eq 0 ]; }; echo $?
Trigger: frr.anchor.consistent = 1 for >90 s (one missed reconcile) = high; the
existing CARP-aware alerting (§14 pattern) stays — but note the backup's FRR being up
with Established sessions is now the healthy state, so remove any "FRR down on
backup is OK" muting from the old design.
7.7 Rejected alternatives (and why)¶
| Option | Why not |
|---|---|
| Stop/start FRR on CARP (old design) | 10–20 s serial cold start; backup holds no default; nothing wrong when converged — just slow |
vtysh prefix-list gate (incl. a top-sequence deny of the aggregates toggled on BACKUP) |
needs the identical hook/boot/reconcile machinery as the anchor and gates the identical scope — but the state lives in frr.conf, exactly where os-frr regeneration wipes it, and a reboot restores whatever the file says instead of failing closed. Wrong-state windows fail toward master-withdrawn = outage. Same script count, strictly weaker guarantee |
FW /16 without the CCR floating-backstop change |
the CCRs' distance-1 blackhole beats eBGP (20) on an identical prefix — the FW's /16 never installs and the CCRs blackhole all server traffic on both nodes. The /16 design requires the CCR distance flip (§6 danger box) |
Mid-size aggregates (10.128.10.0/23 + 10.128.32.0/19) instead of the /16 |
works without touching the CCRs (more-specifics beat their /16) — the fallback if the CCR change is off the table; costs two anchors instead of one and splits the blackhole enforcement across two layers |
BGP conditional advertisement (advertise-map/exist-map) |
scan-timer driven (5–60 s) and not persistable through os-frr config regen — slower and weaker than the anchor. Also why the CORE default stays un-gated rather than exist-map-gated (§7.1) |
| Always-advertise + prepend on backup (no gating) | leaves a persistent asymmetry corner: a demoted master with live WAN sessions keeps attracting inbound indefinitely |
Migration from the stopped-FRR design¶
- CCR side first — automated in
ansible/network: runplaybooks/colo-rr.yml(rr role):rr_aggregate_distance: 200re-anchors the CCRs'10.128.0.0/16blackhole as the floating backstop (§6 danger box; brief one-time withdraw as the distance-1 anchor is replaced), and the new to-edge guard rejectsdst in 10.128.0.0/16toward the FW peers — required because the RR's own /16 origination carries no65510in path, so FW loop detection can't drop it and it could satisfy the backup'snetworkimport-check. Until this is deployed the FW's /16 will never install. - Remove
/usr/local/etc/rc.syshook.d/carp/20-frr(the start/stop hook) from both. - Delete any persistent Null0/static for the anchored prefixes on the FW (System → Routes and FRR statics) on both nodes — see the §6 danger box; the gate is void while one exists.
- Replace the per-
/24server networks with10.128.0.0/16in Networks andPL-CORE-OUTon both nodes. Node-2: addset metric 100toRM-CORE-OUT. Verifybgp network import-checkon both. - Install the three scripts + cron action on both nodes;
chmod 755. - Start FRR on the backup (
service frr start/ enable in GUI) — it establishes all four sessions but must advertise only its MED-100 default to the CCRs and nothing on WAN: check every neighbour withvtysh -c 'show ip bgp neighbors <ip> advertised-routes'. - Run
frr-anchor-sync.shby hand on both; confirm the master holds both anchors —netstat -rn -f inet | grep -E '185\.109\.43\.176|^10\.128'(the /16 shows abbreviated as10.128/16), orvtysh -c 'show ip route 10.128.0.0/16'→Known via "kernel" … blackhole— and advertises/28on WAN + the/16on CORE. On a CCR, the/16is now the eBGP route with the floating blackhole inactive. - Failover test (§13) — budget is now ~1–2 s, not 10–20 s.
8. ACME / Let's Encrypt (Services → ACME Client)¶
Certs must use the PUBLIC name — not the private infra.jsm zone
The boxes live in infra.jsm (internal), but Let's Encrypt can only validate a
publicly-resolvable name. So the GUI cert uses the public FQDNs in the
pubinvest.co.uk zone — fw-colo-1.infra.pubinvest.co.uk /
fw-colo-2.infra.pubinvest.co.uk — issued by DNS-01 against that zone. ACME
will never work for *.infra.jsm.
- Accounts — Let's Encrypt (production), ops email, accept ToS.
- Challenge Types — DNS-01 against the public
pubinvest.co.ukzone (add the DNS provider's API creds). DNS-01 needs no inbound reachability — ideal for an OOB-only edge FW whose public name isn't otherwise exposed. (HTTP-01 can't work here.) - Certificates — issue one cert on
fw-colo-1with both node names as SANs:fw-colo-1.infra.pubinvest.co.ukandfw-colo-2.infra.pubinvest.co.uk(Account + the DNS-01 challenge). A single dual-SAN cert means the XMLRPC Certificates sync gives node-2 a cert valid for its name too — no per-node ACME needed. - Automations — "Restart GUI" (
configd) on issue/renew; enable the daily renew cron. - System → Settings → Administration → SSL Certificate = this cert on both nodes,
and browse the GUI by the public name so it matches: split-horizon DNS on the
internal resolvers (
10.128.32.3/.4) resolvesfw-colo-<n>.infra.pubinvest.co.ukto the OOB IP (172.16.201.41/.42), where the GUI listens.
9. LLDP (Services → LLDPd)¶
Enable; select the transmit/receive interfaces (CORE_A/B, WAN_A/B, lagg0, OOB).
Neighbours then show in the GUI and via SNMP → LibreNMS/Zabbix.
10. Firewall & NAT (summary — see the arch doc for policy)¶
- Zones default-deny + logged: WAN, CORE, BACKUP, VM-APPS, VM-K8S, VM-MON, VM-DATA,
VM-DEV, HYP-MGMT, OOB, GUEST, MGMT-VPN. East-west hairpins the trunk (firewall-on-a-stick).
Split the public
/28(185.109.43.176/28) — a small pool for outbound balancing, the rest for inbound services (adjust to taste):
The /28 is routed to you — every address below is an IP-Alias VIP (not CARP),
and HA rides BGP (the CARP-keyed anchor, §7), not L2. Put the IP-Aliases on the WAN interface
or a loopback; config-sync replicates them to both nodes (dormant on the backup).
| Range | Use |
|---|---|
.177 |
primary WAN service address (IP-Alias) |
.178 |
POS / payment egress — single fixed SNAT IP (no pool) |
.180 – .183 (/30) |
outbound SNAT pool — round-robin, 4 IPs |
.184 – .190 |
inbound NAT — 1:1 / port-forward service addresses |
.176 / .179 / .191 |
network / spare / broadcast |
Outbound pool (Firewall → NAT → Outbound = Hybrid):
- Interfaces → Virtual IPs — an IP-Alias VIP for the pool
185.109.43.180/30(standalone on WAN/lo0 — no CARP; the block fails over via BGP). - Firewall → Aliases — a network alias
OUT_POOL=185.109.43.180/30(all 4). - Rule: source = the internal server nets, translation =
OUT_POOL, Pool Options = Round Robin (add sticky-address if a session must keep one source IP) → SNAT round-robins across the 4 addresses.
POS / payment egress — fixed IP, ABOVE the pool rule
Higher-priority outbound rule: source = POS, destination = the payment-processor
prefixes, translation = 185.109.43.178 fixed (no pool). Processors whitelist one
IP; the guest firewall keeps its own distinct egress. See
firewall-opnsense §NAT.
Inbound NAT (Firewall → NAT → Port Forward / 1:1), on .184 – .190:
- Add an IP-Alias VIP for the chosen public address, then a Port Forward (or 1:1 NAT) → the internal host / service VIP. No CARP — reachability follows the BGP announcement, so it lands on whichever node is master.
- Point split-horizon DNS at the internal address for internal clients (no hairpin); external DNS resolves to the public address.
11. Edge security (free stack)¶
This edge is the internet chokepoint for every venue's staff PCs, staff Wi-Fi, POS tills and the colo servers, so it carries real user traffic. The whole stack is free — no paid NGFW:
| Layer | Tool | Notes |
|---|---|---|
| Exploit / malware IPS | Suricata (base) | inline on WAN |
| Reputation + brute-force | CrowdSec (free) | community blocklist + log scenarios |
| Web / category filtering | separate DNS servers | FW only enforces the path (§11c) |
| Geo / IP reputation | firewall aliases | free threat feeds |
Zenarmor — not worth it here on free
Zenarmor's value (web categories, threat intel, user reporting) is paid, and it does netmap DPI, so it conflicts with inline Suricata on the same interface — one or the other, not both. On a free budget keep Suricata + CrowdSec and do web filtering on the dedicated DNS servers.
11a. Suricata IDS/IPS (Services → Intrusion Detection)¶
Base component — no plugin. Inspect the internet-facing WAN only (server VLANs run uninspected at line rate; only the CARP master sees transit).
- Enabled, IPS mode (inline netmap), Promiscuous, Pattern matcher = Hyperscan.
- Interfaces =
WAN_A+WAN_B(ixl7/ixl2). Notlagg0/ server VLANs. - Runmode = workers, NUMA-local (tunables §12).
- Rules — ET Open rulesets, daily update; run IDS (alert-only) ~a week to baseline, then flip chosen categories to drop (IPS); suppress known-good noise.
- Home networks = the public
/28+ routed server subnets. - Logging = EVE JSON → the monitoring collector (
10.128.34.x, VLAN 934).
Disable offload + master-only
Inline netmap needs LRO/TSO/LSO off on WAN_A/WAN_B (leave offload on elsewhere).
Only the MASTER forwards, so only it inspects — no CARP hook needed.
11a-i. IPS bring-up & troubleshooting (netmap on ixl/X710)¶
"No alerts" or "WAN drops when I turn on IPS" is almost always the netmap ↔ NIC-offload
interaction, and it bites harder in IPS mode because inline Suricata is a
bump-in-the-wire — if netmap won't attach cleanly you don't just lose alerts, you lose
the WAN. Our edge NICs are X710 (ixl / i40e), the fussiest family for netmap, so
follow this order rather than flipping IPS on blind.
1 — Prove you can see alerts in IDS first. If plain IDS gives no info, IPS won't either; the fault is upstream (interface/rules/offload), not the mode.
- Settings: Enabled, IPS mode OFF for now, Promiscuous, Hyperscan;
Interfaces =
WAN_A+WAN_B(ixl7/ixl2) only; Home networks = the public/28+ routed server subnets. - Download tab → enable ET Open, tick categories, Download & Update. No enabled rules = no alerts — the most common "no info" cause.
- Apply, then fire a known trigger and check Intrusion Detection → Alerts: Shows up → pipeline works, go to step 2. Nothing → it's offload or the wrong interface.
2 — Turn off interface offload (the #1 ixl breaker). Interfaces → Settings (global):
- Disable hardware TSO ✔, Disable hardware LRO ✔, Disable hardware checksum offload ✔.
- VLAN Hardware Filtering → Disable — X710/
i40especifically needs this; leaving it on is the classic "netmap attaches but no packets pass" symptom. - Reboot — offload changes don't fully take effect live.
3 — Enable IPS. Settings → IPS mode ✔, save, apply. Keep every rule at alert for
~a week (full visibility, nothing dropped), then flip chosen categories to drop. In
OPNsense a drop rule both blocks and alerts, so IPS never costs you info — going
inline gives you strictly more (each event carries action: allowed vs blocked).
4 — Getting the info out. GUI Alerts tab is live-only; the durable record is EVE
JSON (Settings → Logging → EVE output) shipped to the monitoring collector on VLAN
934 (10.128.34.x). Point Zabbix there, not at the GUI. alert.action: blocked = a drop
rule enforced; allowed = alert-only saw + logged it.
Do the inline cutover on the BACKUP node, in a window
IPS is inline on WAN — a failed netmap attach drops internet. On the CARP pair, enable
and verify IPS on fw-colo-2 while it is BACKUP (idle netmap, no live traffic at
risk), fail over, verify, then repeat on the other node. Never flip both masters at once.
ixl/X710 specifics. Native i40e netmap works but is touchier than ix/igb:
- Watch Intrusion Detection → Log File immediately after enabling IPS for a
netmapattach line vs an error — an attach failure there means traffic won't pass inline. - If native won't attach it falls back to emulated netmap (slower, but works) — confirm
with
dmesg | grep -i netmap. - Multiqueue can trip it; if the interface flaps, cut the
ixlqueue count and retest.
Verify:
suricata --build-info | grep -i netmap # NETMAP support: yes
dmesg | grep -i netmap # attach lines, no errors
tail -f /var/log/suricata/eve.json # live events incl. action:blocked
curl http://testmynids.org/uid/index.html # now BLOCKED, not just alerted
11b. CrowdSec (Services → CrowdSec — os-crowdsec, free)¶
Behavioural detection (brute-force, scanning) from logs + the free Community Blocklist; the firewall bouncer drops offenders at pf. Complements Suricata (signatures), no overlap.
- Install
os-crowdsec; enable the agent + the firewall bouncer. - Add the OPNsense/Suricata collections (log parsers + scenarios) from the Hub.
- Enrol in the free Console (
app.crowdsec.net) to receive the Community Blocklist — this shares your detection signals (reputation metadata, not payloads). Skip enrolment to stay fully local (then you only block what you detect — no community list). - HA/CARP: run it on both nodes; the community list is identical, locally-learned
bans accrue only on the MASTER — fine, the new master keeps the list and relearns after a
failover. Deploy on
fw-guest-1too (guest Wi-Fi is the best candidate of all).
11c. DNS enforcement (filtering is on the dedicated DNS servers)¶
Web/category filtering runs on separate DNS servers, not OPNsense — the FW only guarantees clients can't bypass them:
- Allow client DNS (UDP/TCP 53) only to the filtering servers — NAT-redirect or drop
all other outbound 53 (a device hardcoded to
8.8.8.8is redirected or blocked). - Block DoT — outbound TCP 853.
- Block DoH — a DoH-endpoints host alias (public lists exist); for user VLANs also consider blocking UDP/443 (QUIC) so browsers can't DoH-over-QUIC around it.
- Hand out the filtering servers via DHCP at the venue routers (staff/POS/office scopes); edge enforcement is the backstop.
11d. Geo / IP-reputation aliases (Firewall → Aliases)¶
Cheap, high-signal drops on WAN-in (and forward):
- URL-table aliases from free feeds — Spamhaus DROP/EDROP, abuse.ch, FireHOL level 1 — action block, logged.
- GeoIP country blocks for regions you never transact with (needs the free MaxMind GeoLite key).
- POS stays locked down regardless — payment prefixes + AD egress only (venue design); a till should never reach the web filter's remit in the first place.
12. Tunables & performance (System → Settings → Tunables)¶
Design targets from the architecture doc — set under System → Settings → Tunables (sysctl apply live; loader tunables need a reboot). Validate the exact keys against your OPNsense/FreeBSD build.
| Tunable | Value | Why |
|---|---|---|
net.isr.maxthreads |
-1 |
one netisr thread per core |
net.isr.bindthreads |
1 |
pin them (no migration) |
net.isr.dispatch |
deferred |
scale RX across cores |
kern.ipc.nmbclusters |
1000000 |
mbuf headroom at 10G+ |
kern.ipc.nmbjumbo9 |
524288 |
9k jumbo mbufs (server trunk) |
net.pf.states_hashsize |
1048576 |
pf state hashing (loader → reboot) |
net.route.multipath |
1 |
ECMP (also §6) |
dev.ixl.<n>.iflib.override_nrxds / …override_ntxds |
4096 |
deeper X710 RX/TX rings, per unit (loader → reboot) |
GUI settings (not tunables):
- Firewall → Settings → Advanced → Firewall Maximum States = 3–5 M; Optimization = conservative (long-lived server flows).
- Firewall → Settings → Normalization → scrub on WAN only; never scrub or
set skipthe server trunk (line-rate). - System → Settings → Miscellaneous → Power → performance (powerd off / hi-adaptive off); Cryptography = AES-NI; cap C-states at C1 for latency.
- X710 in x8 PCIe slots, NUMA-balanced, HT on (BIOS — see arch doc §Hardware).
13. Verify¶
# HA
System → High Availability → Status -> node-1 MASTER, node-2 BACKUP, all VIPs green
pfSync: Interfaces → PFSYNC -> states counting up on both
# BGP (Routing → Diagnostics, or shell: vtysh) — run on BOTH nodes; FRR is warm on both
show ip bgp summary -> bdr-1, bdr-2, CR-COLO-01/02 all Established (both nodes)
show bfd peers brief -> CR-COLO-01/02 up
show ip bgp 0.0.0.0/0 -> default in from WAN on BOTH nodes (bdr-1 pref 200)
# MASTER only:
netstat -rn -f inet | grep -E '185\.109\.43\.176|^10\.128' -> both anchors
# (the /16 prints ABBREVIATED: "10.128/16" — FreeBSD trims trailing zero octets)
show ip bgp neighbors <bdr-ip> advertised-routes -> /28 advertised
show ip bgp neighbors <ccr-ip> advertised-routes -> 0/0 + /28 + 10.128.0.0/16
# on a CCR: /16 active via eBGP from the master; its floating blackhole INACTIVE
# BACKUP: sessions Established; WAN advertised-routes EMPTY; CORE advertised-routes =
# ONLY the MED-100 default (the §7.1 exception). Anything more: run frr-anchor-sync.sh
# ACME
Services → ACME Client → Certificates -> issued, expiry ~90d, GUI using it
# LLDP
Services → LLDPd → Neighbors -> CR-COLO-1/2, bdr-1/2, sw-server-1/2, OOB-2
# IDS/IPS
Services → Intrusion Detection → Alerts -> events on WAN only; EVE JSON reaching 10.128.34.x
# CrowdSec
Services → CrowdSec → Decisions -> bouncer active; community blocklist populated
# DNS enforcement (from a staff/POS test host)
dig @8.8.8.8 example.com -> BLOCKED/redirected; :853 and DoH endpoints unreachable
# NAT
Firewall → NAT → Outbound -> pool rule (Round Robin over 185.109.43.180/30),
POS fixed-IP rule ABOVE it
Failover test: on fw-colo-1, Interfaces → Virtual IPs → Enter persistent CARP
maintenance mode. Node-2 goes MASTER, its hook adds the anchor and the /28 UPDATE
goes out on the already-Established sessions — total gap ~1–2 s (watch
ping from outside to a /28 service IP; more than ~3 s lost means the hook didn't
fire — run frr-anchor-sync.sh and check §7.6). Exit maintenance → node-1 reclaims,
node-2's hook withdraws.