Skip to content

Firewall — OPNsense edge pair

The corporate edge is an OPNsense HA pair (fw-colo-1 / fw-colo-2) at the colo. It is the internet edge, the inter-zone firewall for the server room, and the host of the central captive-portal VIP. The guest firewall is a separate OPNsense instance (fw-guest-1, single-node) — see its guest setup runbook.

Building one? See the setup runbook

This page is the design/reference. The step-by-step build (plugins, LAGGs/VLANs, CARP VIPs, pfSync, FRR/BGP, ACME, LLDP, verify) is in firewall-opnsense-setup.

Hardware (per node ×2)

  • 2× Intel Xeon Gold 6150 (18C/36T), 128 GB RAM, BOSS 240 GB M.2 RAID1 boot.
  • NICs: 1× X710+i350 combo + 3× X710 dual-SFP+ → 8× SFP+ (10G) + 2× RJ45 (1G) + dedicated IPMI/BMC per node. Each redundant pair is spread across two different cards; all X710 in x8 PCIe slots, balanced across both CPU sockets (NUMA), RSS on. See Intel X710 notes.

Interface plan (8× SFP+)

# Role Peer MTU
1 Core A (eBGP) CR-COLO-01 /31 1500
2 Core B (eBGP) CR-COLO-02 /31 1500
3–4 Server LACP a/b VM-switch MLAG, 20G bond 9000
5–6 WAN → bdr-1 / bdr-2 upstream border /31 1500
7–8 spare ×2 → optional 40G server trunk —
— pfSync a/b 2× 1G i350, direct, active-backup bond 1500
— IPMI/BMC OOB CRS326 —

Routing & BGP

  • FRR (os-frr plugin), eBGP AS 65510, each node dual-homed to both CCRs.
  • Warm FRR + CARP-keyed anchor — FRR runs permanently on both nodes, all sessions Established and both holding a live default at all times. Only the CARP MASTER holds the kernel blackhole anchor that lets the /28 network statement fire, so only the master attracts traffic — same strict path symmetry as gating the daemon, but failover is one BGP UPDATE on a warm session (~1–2 s, vs 10–20 s of daemon cold-start). The anchor is toggled by an idempotent CARP-hook + 1-min reconcile; CORE-side preference is a static MED on node-2. Full design, scripts and failure matrix: setup runbook §7.
  • BFD on the CCR sessions for fast edge failover — each is a single-hop eBGP /31 to a colo CCR, so BFD is permitted (no multihop config needed); the CCR side keys it off the per-port bfd_enabled custom field and FRR runs BFD to match. WAN keepalive/hold 3/9 s. ECMP via net.route.multipath=1 + bestpath as-path multipath-relax.
  • WAN: one eBGP /31 to each border (we present customer AS 4200000001 to the ISP AS 204258 via a per-neighbour local-as; the FRR instance stays 65510 for the CORE side). No BFD on WAN — failover rides the tight keepalive/hold 3/9 s timers. bdr-1 primary (local-pref 200), bdr-2 backup (local-pref 100, prepend ×2 out). Announces only the routed /28 (Null0-anchored, out prefix-list); receives default only. The public /28 is a routed block (its addresses are IP-Alias VIPs, HA'd by the CARP-keyed anchor — not CARP-floated), decoupled from the /31 transit.
  • To the CCRs: the master only announces the corporate public block and the colo space — as the single anchor-gated 10.128.0.0/16 aggregate (not the connected /24s, which would be un-gateable). Prerequisite: the CCRs' own /16 blackhole is a floating backstop (distance 200, networks entry kept) — the FW's eBGP /16 wins normally; whenever no FW advertises it (failover gap, pair down) the CCR static re-activates and resumes originating + blackholing, so metro steering never flaps and colo-bound traffic gets clean drops. The backup attracts nothing anywhere; all non-connected colo space (storage, migration, corosync, OOB) dies at the master's anchor blackhole. The default is the one deliberate exception: both nodes announce 0.0.0.0/0 (passed through — it self-withdraws if a node's WAN dies), node-2 at MED 100, so the CCRs hold two pre-installed default paths and venue internet survives a master death on BFD alone.

Zones

The concrete alias list and rule set is in firewall-rules.

Default-deny between zones, logged: WAN, CORE, BACKUP, VM-APPS, VM-K8S, VM-MON, VM-DATA, VM-DEV, HYP-MGMT, OOB, GUEST, MGMT-VPN. The server-room trunk (interfaces 3+4, 20G LACP → VM CRS326 MLAG) carries the routed server VLANs 910 (backup), 932 (AD/apps), 933 (k8s), 934 (monitoring), 935 (data/SQL), 936 (dev), 940 (hyp-mgmt) tagged — each VLAN's CARP VIP is its gateway. The trunk is L2 MTU 9000, routed VLAN interfaces are 1500. Corosync 950 (Proxmox) and Ceph 920/921/922 are L2-only — never routed here. Inter-zone east-west traffic hairpins the trunk (firewall-on-a-stick).

NAT

  • The public /28 (185.109.43.176/28) is split: a /30 (4 IPs) outbound SNAT pool (round-robin, IP-Alias parented to a CARP VIP so it fails over) + a fixed POS egress IP, with the remainder as inbound-NAT service VIPs.
  • Never pool payment / POS egress

    Card processors whitelist a fixed IP. A higher-priority rule sends POS → payment-prefixes out a single fixed /28 address; the guest firewall keeps its own distinct egress IP for reputation isolation.
  • Hairpin/NAT-reflection: primary is split-horizon DNS (internal resolvers 10.128.32.3/.4 return the internal address); fallback is Pure NAT reflection plus automatic outbound NAT for reflection.

HA

CARP + pfSync (2×1G active-backup direct bond) + XMLRPC config sync over the pfSync bond. CARP heartbeats on WAN, server-net and the pfSync bond (no split-brain); advskew makes node-1 master; all VIPs fail over together.

Edge security (free stack)

This is the internet edge for all venue staff/POS/office traffic + the colo servers, so it runs a full free security stack (no paid NGFW):

  • Suricata — netmap inline, workers runmode, NUMA-pinned; inline IPS on WAN only (LRO/TSO off there); server trunk + internal VLANs uninspected at line rate; only the CARP master inspects. EVE/alerts → the monitoring VLAN (934, 10.128.34.x).
  • CrowdSec (free) — community blocklist + brute-force/scan scenarios, bouncer drops at pf.
  • DNS-bypass enforcement — client 53 allowed only to the dedicated filtering DNS servers; DoT/DoH/QUIC blocked (the category filtering itself lives on those servers).
  • Geo / IP-reputation firewall aliases (Spamhaus / abuse.ch / FireHOL / GeoIP).

Zenarmor is not used — its web/threat features are paid and it does netmap DPI, so it conflicts with inline Suricata (one DPI engine per interface). See the setup runbook §11 for the build steps.

Tuning highlights

ixl rings override_nrxds/ntxds=4096, flow-control off; net.isr maxthreads -1, bindthreads on, deferred dispatch; nmbclusters=1000000, nmbjumbo9=524288; pf max-states 3–5M, states_hashsize≈1048576; C-states capped at C1, powerd off/performance, HT on, AES-NI on. Scrub on WAN only; do not set skip on the server trunk.

Monitoring

Zabbix (server on the monitoring VLAN 934, 10.128.34.x), monitored over OOB/mgmt: agent-active, SNMPv3 (source-locked), IPMI over OOB, and custom UserParameters for CARP, pfSync, pf states, BGP/FRR and Suricata. CARP-aware alerting — the backup's BACKUP VIPs must not alert, but its FRR must be up with all sessions Established (warm-FRR design — a down FRR on the backup is now a fault, not the norm). Key invariant check: CARP state vs /28 advertisement must agree on each node (mismatch >90 s = high — setup runbook §7.6). Pair triggers: split-brain (both MASTER), no-master, redundancy lost, config drift.