Firewall — OPNsense edge pair¶
The corporate edge is an OPNsense HA pair (fw-colo-1 / fw-colo-2) at the colo.
It is the internet edge, the inter-zone firewall for the server room, and the host
of the central captive-portal VIP. The guest firewall is a separate OPNsense
instance (fw-guest-1, single-node) — see its
guest setup runbook.
Building one? See the setup runbook
This page is the design/reference. The step-by-step build (plugins, LAGGs/VLANs, CARP VIPs, pfSync, FRR/BGP, ACME, LLDP, verify) is in firewall-opnsense-setup.
Hardware (per node ×2)¶
- 2× Intel Xeon Gold 6150 (18C/36T), 128 GB RAM, BOSS 240 GB M.2 RAID1 boot.
- NICs: 1× X710+i350 combo + 3× X710 dual-SFP+ → 8× SFP+ (10G) + 2× RJ45 (1G) + dedicated IPMI/BMC per node. Each redundant pair is spread across two different cards; all X710 in x8 PCIe slots, balanced across both CPU sockets (NUMA), RSS on. See Intel X710 notes.
Interface plan (8× SFP+)¶
| # | Role | Peer | MTU |
|---|---|---|---|
| 1 | Core A (eBGP) | CR-COLO-01 /31 |
1500 |
| 2 | Core B (eBGP) | CR-COLO-02 /31 |
1500 |
| 3–4 | Server LACP a/b | VM-switch MLAG, 20G bond | 9000 |
| 5–6 | WAN → bdr-1 / bdr-2 |
upstream border /31 | 1500 |
| 7–8 | spare ×2 | → optional 40G server trunk | — |
| — | pfSync a/b | 2× 1G i350, direct, active-backup bond | 1500 |
| — | IPMI/BMC | OOB CRS326 | — |
Routing & BGP¶
- FRR (os-frr plugin), eBGP AS 65510, each node dual-homed to both CCRs.
- Warm FRR + CARP-keyed anchor — FRR runs permanently on both nodes, all
sessions Established and both holding a live default at all times. Only the CARP
MASTER holds the kernel blackhole anchor that lets the
/28networkstatement fire, so only the master attracts traffic — same strict path symmetry as gating the daemon, but failover is one BGP UPDATE on a warm session (~1–2 s, vs 10–20 s of daemon cold-start). The anchor is toggled by an idempotent CARP-hook + 1-min reconcile; CORE-side preference is a static MED on node-2. Full design, scripts and failure matrix: setup runbook §7. - BFD on the CCR sessions for fast edge failover — each is a single-hop eBGP
/31to a colo CCR, so BFD is permitted (no multihop config needed); the CCR side keys it off the per-portbfd_enabledcustom field and FRR runs BFD to match. WAN keepalive/hold 3/9 s. ECMP vianet.route.multipath=1+bestpath as-path multipath-relax. - WAN: one eBGP
/31to each border (we present customer AS4200000001to the ISP AS204258via a per-neighbourlocal-as; the FRR instance stays65510for the CORE side). No BFD on WAN — failover rides the tight keepalive/hold 3/9 s timers.bdr-1primary (local-pref 200),bdr-2backup (local-pref 100, prepend ×2 out). Announces only the routed/28(Null0-anchored, out prefix-list); receives default only. The public/28is a routed block (its addresses are IP-Alias VIPs, HA'd by the CARP-keyed anchor — not CARP-floated), decoupled from the/31transit. - To the CCRs: the master only announces the corporate public block and the
colo space — as the single anchor-gated
10.128.0.0/16aggregate (not the connected/24s, which would be un-gateable). Prerequisite: the CCRs' own /16 blackhole is a floating backstop (distance 200, networks entry kept) — the FW's eBGP /16 wins normally; whenever no FW advertises it (failover gap, pair down) the CCR static re-activates and resumes originating + blackholing, so metro steering never flaps and colo-bound traffic gets clean drops. The backup attracts nothing anywhere; all non-connected colo space (storage, migration, corosync, OOB) dies at the master's anchor blackhole. The default is the one deliberate exception: both nodes announce0.0.0.0/0(passed through — it self-withdraws if a node's WAN dies), node-2 at MED 100, so the CCRs hold two pre-installed default paths and venue internet survives a master death on BFD alone.
Zones¶
The concrete alias list and rule set is in firewall-rules.
Default-deny between zones, logged: WAN, CORE, BACKUP, VM-APPS, VM-K8S, VM-MON, VM-DATA, VM-DEV, HYP-MGMT, OOB, GUEST, MGMT-VPN. The server-room trunk (interfaces 3+4, 20G LACP → VM CRS326 MLAG) carries the routed server VLANs 910 (backup), 932 (AD/apps), 933 (k8s), 934 (monitoring), 935 (data/SQL), 936 (dev), 940 (hyp-mgmt) tagged — each VLAN's CARP VIP is its gateway. The trunk is L2 MTU 9000, routed VLAN interfaces are 1500. Corosync 950 (Proxmox) and Ceph 920/921/922 are L2-only — never routed here. Inter-zone east-west traffic hairpins the trunk (firewall-on-a-stick).
NAT¶
- The public
/28(185.109.43.176/28) is split: a/30(4 IPs) outbound SNAT pool (round-robin, IP-Alias parented to a CARP VIP so it fails over) + a fixed POS egress IP, with the remainder as inbound-NAT service VIPs. -
Never pool payment / POS egress
Card processors whitelist a fixed IP. A higher-priority rule sends POS → payment-prefixes out a single fixed/28address; the guest firewall keeps its own distinct egress IP for reputation isolation. - Hairpin/NAT-reflection: primary is split-horizon DNS (internal resolvers
10.128.32.3/.4return the internal address); fallback is Pure NAT reflection plus automatic outbound NAT for reflection.
HA¶
CARP + pfSync (2×1G active-backup direct bond) + XMLRPC config sync over the
pfSync bond. CARP heartbeats on WAN, server-net and the pfSync bond (no
split-brain); advskew makes node-1 master; all VIPs fail over together.
Edge security (free stack)¶
This is the internet edge for all venue staff/POS/office traffic + the colo servers, so it runs a full free security stack (no paid NGFW):
- Suricata — netmap inline, workers runmode, NUMA-pinned; inline IPS on WAN only
(LRO/TSO off there); server trunk + internal VLANs uninspected at line rate; only the CARP
master inspects. EVE/alerts → the monitoring VLAN (934,
10.128.34.x). - CrowdSec (free) — community blocklist + brute-force/scan scenarios, bouncer drops at pf.
- DNS-bypass enforcement — client
53allowed only to the dedicated filtering DNS servers; DoT/DoH/QUIC blocked (the category filtering itself lives on those servers). - Geo / IP-reputation firewall aliases (Spamhaus / abuse.ch / FireHOL / GeoIP).
Zenarmor is not used — its web/threat features are paid and it does netmap DPI, so it conflicts with inline Suricata (one DPI engine per interface). See the setup runbook §11 for the build steps.
Tuning highlights¶
ixl rings override_nrxds/ntxds=4096, flow-control off; net.isr maxthreads
-1, bindthreads on, deferred dispatch; nmbclusters=1000000,
nmbjumbo9=524288; pf max-states 3–5M, states_hashsize≈1048576; C-states
capped at C1, powerd off/performance, HT on, AES-NI on. Scrub on WAN only; do
not set skip on the server trunk.
Monitoring¶
Zabbix (server on the monitoring VLAN 934, 10.128.34.x), monitored over OOB/mgmt: agent-active,
SNMPv3 (source-locked), IPMI over OOB, and custom UserParameters for CARP, pfSync,
pf states, BGP/FRR and Suricata. CARP-aware alerting — the backup's BACKUP VIPs
must not alert, but its FRR must be up with all sessions Established (warm-FRR
design — a down FRR on the backup is now a fault, not the norm). Key invariant check:
CARP state vs /28 advertisement must agree on each node (mismatch >90 s = high —
setup runbook §7.6). Pair triggers: split-brain (both MASTER), no-master, redundancy
lost, config drift.