OPNsense Edge Pair — Hardware & Cabling¶
Related pages
Design-phase hardware & cabling doc. Operational reference: Firewall — OPNsense; GUI build: OPNsense build runbook.
Version: 1.0 — 2026-07-06 Physical build and interface plan for the corporate edge HA pair (DESIGN §6). Guest is a separate OPNsense (DESIGN §8) — not this pair.
1. Hardware (per node, ×2 nodes)¶
| Item | Spec |
|---|---|
| CPU | 2× Intel Xeon Gold 6150 (2×18C/36T) |
| RAM | 8× 16 GB = 128 GB |
| Boot | BOSS card 240 GB (M.2 RAID1) |
| NICs | 1× X710+i350 combo (2× SFP+ + 2× RJ45) + 3× X710 dual-SFP+ (6× SFP+) = 4 cards |
| Total ports/node | 8× SFP+ (10G) + 2× RJ45 (1G) + dedicated IPMI/BMC |
pfSync moves to 1G copper, HA over both ports. State sync is low-bandwidth (~40 Mbps even at 50k new-conn/s; a 5M-state bulk sync on node-rejoin is ~4 s over 1G), so both i350 1G ports become a redundant pfSync — two direct cables to the peer, active-backup bond (redundancy, not bandwidth), so a NIC/cable failure never interrupts state sync. This frees two SFP+. No dedicated OOB NIC: the separate BMC on the OOB switch is OPNsense's out-of-band path (console/KVM/power), and OS web/SSH management is in-band via a mgmt VLAN on the server trunk. Trade-off: a total server-trunk failure drops in-band OS access → fall back to BMC KVM (acceptable break-glass). Result: 6 SFP+ essential + 2 SFP+ spare.
NIC role discipline: spread each redundant pair across two different cards so one card failure never kills both halves. Suggested: Core-A + WAN-bdr1 on the combo card; Core-B + WAN-bdr2 on X710 #1; Server-a on X710 #2; Server-b on X710 #3. A lost card = degraded-but-up (one core, one WAN, LACP to one link), not an outage.
2. Interface plan (per node) — 8× SFP+, 2× 1G RJ45, separate BMC¶
| # | Role | Port | Peer | MTU |
|---|---|---|---|---|
| 1 | Core A (eBGP) | SFP+ (combo card) | CR-COLO-01 (/31) | 1500 |
| 2 | Core B (eBGP) | SFP+ (X710 #1) | CR-COLO-02 (/31) | 1500 |
| 3 | Server LACP a | SFP+ (X710 #2) | VM-switch-A (MLAG) | 9000 |
| 4 | Server LACP b | SFP+ (X710 #3) | VM-switch-B (MLAG) | 9000 |
| 5 | WAN → bdr-1 | SFP+ (combo card) | upstream border-1 (/31) | 1500 |
| 6 | WAN → bdr-2 | SFP+ (X710 #1) | upstream border-2 (/31) | 1500 |
| 7–8 | spare ×2 | SFP+ (X710 #2, #3) | — → or server links c/d for 40G trunk | — |
| — | pfSync a | RJ45 #1 (1G i350) | direct cable to peer node ⎫ active-backup bond | 1500 |
| — | pfSync b | RJ45 #2 (1G i350) | direct cable to peer node ⎭ (redundant pfSync) | 1500 |
| — | IPMI / BMC | dedicated BMC port (separate) | OOB CRS326 switch | 1500 |
Used: 6 SFP+ + 2 RJ45 (HA pfSync) + BMC. 2 SFP+ spare. OS management is in-band (mgmt VLAN on the server trunk); OOB/break-glass is the BMC on the OOB switch. CARP heartbeats on all CARP interfaces (server + WAN + pfSync bond) — no dedicated heartbeat port. The 2 spare SFP+ can upgrade the server trunk to 40G (4×10G) if east-west/backup is heavy.
3. Core uplinks — BGP, dual-homed to both CCRs, CARP-gated FRR¶
Each node connects to both CR-COLO-01 and CR-COLO-02 (not one-each) — a single CCR failure never isolates a node. FRR runs eBGP AS 65510 to each CCR over a routed /31 (4 physical core links across the pair).
Only the CARP master speaks BGP — use os-frr's CARP failover mode. The FRR plugin (Routing ▸ General) has a CARP-tracking option: pick a CARP VIP to track, and FRR is stopped whenever that VIP is in BACKUP state and started on MASTER. This solves the stateful-firewall symmetry problem structurally — the backup node announces nothing, so the CCRs can only ever route via the master, and no prepend/local-pref trickery is needed.
- Track a purpose-made "routing" CARP VIP (e.g. on the server-net, or one of the /28 public VIPs) so all of FRR (core + both WAN border sessions) follows one master election. All CARP VIPs share the same advskew, so master status moves as one. (There is no WAN-transit VIP — the WAN links are /31 p2p, §5.)
- Failover behaviour: BGP sessions do not survive failover (FRR TCP state isn't synced — pfSync only syncs firewall state). On CARP transition the new master starts FRR, re-establishes core + WAN sessions and re-announces; expect a few seconds of routing convergence. Mitigate with BFD on the CCR sessions and modest keepalive/hold timers (e.g. 3/9 s) on WAN. Data-plane flows resume seamlessly after convergence because pfSync has the state table on the new master.
- Backup node still needs its own routes while FRR is stopped: give each node a low-priority static default (gateway = upstream hand-off, or via the core /31) marked "far gateway / do not use for policy" so updates, monitoring and pfSync-adjacent services on the backup keep working. It carries no transit traffic (no VIPs, no announcements), so this is management-plane only.
- What OPNsense announces to the CCRs (master only):
0.0.0.0/0(learned from upstream, passed through — §5a), the corporate public routed block, and the server-room zone subnets. Receives the 10/8 aggregate + venue prefixes (DESIGN §6). Enablebgp bestpath as-path multipath-relax-style ECMP + FreeBSDROUTE_MPATH(net.route.multipath=1) so the two CCR sessions load-share. - Alternative (not chosen): both nodes peering permanently with prepend on the backup gives pre-established sessions (faster failover) at the cost of asymmetry risk and doubled session count. CARP-gated FRR is simpler and self-consistent — take the few seconds of reconvergence.
4. Server-room uplink — dedicated LACP to the VM MLAG pair¶
Interfaces #3+#4 are a dedicated LACP bond to the VM CRS326 MLAG pair (one link to each member) — the "direct uplink to the server switch network with dedicated interfaces." 20G, survives either switch or either link.
- This trunk carries the routed server-room VLANs (Tier-1 930, Tier-2 931, DMZ 932, backup 910, hyp-mgmt 940) as tagged VLANs; each VLAN's CARP VIP is its gateway (DESIGN §7.2 — every routed server VLAN gateways on OPNsense). The Ceph VLANs (920/921/922) are NOT here — L2-only, never routed.
- Trunk sizing (20G, or 40G using the 2 spare SFP+): OPNsense gateways every routed server VLAN, so inter-zone east-west hairpins this trunk (in and back out — counts double), plus backups. Baseline is 2×10G = 20G; with pfSync on 1G there are 2 spare SFP+, so if east-west/backup is heavy, add server links c/d (2 to each VM switch) for 40G. If it ever still saturates, keep the heaviest server-to-server flows within a zone (switched, same VLAN) rather than routing them. Firewall-on-a-stick trade-off, accepted for zero-trust.
- Trunk L2 MTU 9000 (headroom); the routed VLAN interfaces stay 1500. Storage VLANs (Ceph/iSCSI) do NOT traverse OPNsense — they're L2-only on the storage fabric, never routed (DESIGN §7.2).
- Physical per-zone isolation (a separate bond for Tier-1) isn't possible with 8 SFP+ — zone separation is by VLAN on the shared trunk, enforced by the OPNsense zone firewall.
5. WAN — dual upstream border routers, one eBGP /31 each (no VRRP)¶
Two upstream border routers (bdr-1 / bdr-2 — we hold a /28 within their block). No VRRP/VIP on the upstream pair: stacking a first-hop-redundancy protocol under BGP is redundant and interaction-prone — BGP already does failover via session state + BFD. Instead: a separate eBGP session over a point-to-point /31 to each border router, and let BGP pick primary/backup.
Interfaces #5 → bdr-1, #6 → bdr-2 (per node — both used). Each node reaches both borders, so the master role works whichever node holds it and whichever border is up. Four transit links across the pair (node×border) — the port budget easily covers it.
5a. Sessions, HA, and why there is no WAN VIP here¶
- A
/31has no room for a VIP — so the BGP source is each node's real /31 address, not a CARP VIP. "Which node announces" is handled entirely by CARP-gated FRR (§3): the backup's FRR is stopped, so only the master's sessions are ever Established. - The borders configure neighbours for both nodes' /31 addresses on each link; at any moment only the master's two sessions are up. On failover the new master's FRR starts and its two sessions come up — the borders already have the config waiting. No VIP, no shared address, no VRRP.
- Public addressing is decoupled from transit. The /31s carry only the BGP session + packets. All public service/NAT addresses come from the
/28routed block, which OPNsense announces via BGP and floats as CARP VIPs (on a loopback/opt interface). NAT egress masquerades to a /28 VIP; inbound services bind /28 VIPs; both follow the master via CARP — independent of which /31 transit link is up. - Prefer one border, fail to the other: local-pref higher on bdr-1's default (both nodes agree via iBGP/FRR); prepend our announcement to bdr-2 for inbound preference. BFD on both sessions.
FRR (master only; two neighbours — os-frr GUI maps 1:1):
router bgp 65510
neighbor <BDR1-/31> remote-as <UPSTREAM-AS>
neighbor <BDR2-/31> remote-as <UPSTREAM-AS>
neighbor <BDR1-/31> bfd
neighbor <BDR2-/31> bfd
address-family ipv4 unicast
network <ROUTED-BLOCK> ! e.g. 185.109.40.8/28
neighbor <BDR1-/31> route-map WAN-IN-PRIMARY in ! set local-pref 200
neighbor <BDR2-/31> route-map WAN-IN-BACKUP in ! set local-pref 100
neighbor <BDR1-/31> prefix-list WAN-OUT out
neighbor <BDR2-/31> route-map WAN-OUT-PREPEND out ! WAN-OUT + prepend x2
!
ip route <ROUTED-BLOCK> Null0 ! RIB anchor so `network` fires
ip prefix-list WAN-OUT permit <ROUTED-BLOCK> ! announce ONLY our block
ip prefix-list WAN-OUT deny 0.0.0.0/0 le 32
ip prefix-list WAN-IN permit 0.0.0.0/0 ! accept default only from each
ip prefix-list WAN-IN deny 0.0.0.0/0 le 32
- Announce: exactly the corporate routed block on both sessions — the out prefix-list guarantees a fat-fingered redistribute can never leak 10/8, server zones, or the core's routes upstream. Null0 anchor keeps it announced while the box is up.
- Receive: default only from each border (no full tables at a stub edge). bdr-1 preferred via local-pref; bdr-2 hot standby.
- Failure modes: bdr-1 link/session down → default via bdr-2 in BFD time; master node down → new master brings both sessions up. Any single element (a border, a link, or a node) can fail with automatic recovery.
5b. The default-route chain (why nothing here is static)¶
Upstream ──0/0──▶ OPNsense (master) ──0/0──▶ CR-COLO-01/02 ──default-originate if-installed──▶ RRs ──▶ whole fleet
The learned default is simply passed through to the CCRs (the to-core route-map permits 0.0.0.0/0 + the public block + zone subnets). No default-originate on OPNsense, no static default anywhere in the chain. Consequence: upstream withdraws or the WAN dies → OPNsense loses 0/0 → stops announcing it to the CCRs → the RRs stop originating it fleet-wide → the payment prefixes' only remaining route is via Starlink, and bulk traffic (guest, staff, backups) correctly blackholes rather than crushing the backup uplink (CONFIG-GUIDE §4.4). The whole failover is emergent from route withdrawal — there is nothing to "switch."
- NAT vs routed block: the public routed block hosts (DMZ/public services) are routed with real public IPs behind the firewall — not NAT'd. General corporate/venue outbound egress is source-NAT'd to an address from the routed
/28(a CARP VIP on the /28) — never a WAN /31 transit address (those aren't in your announced space). Keep them separate: routed block = inbound services + the outbound NAT address(es); WAN /31 = transit only. - /28 sub-allocation (16 addresses — carve deliberately): reserve some for inbound service VIPs / 1:1 NAT; use one (or a small pool) for outbound PAT — one PAT address ≈ 64k concurrent sessions, so 1–2 covers the whole corporate + venue estate, spread across a small pool only if session scale ever demands it. The guest firewall's separate egress IP (DESIGN §8) also draws from routed space — keep it a distinct address for reputation isolation.
- RPKI/IRR: publish ROAs for the routed block (origin = the announcing AS) + matching IRR objects before cutover, or upstreams filter you (same discipline as WAN-DESIGN §10).
5c. Outbound source-NAT pool (variable source IP)¶
You can spread outbound NAT across a pool of /28 addresses instead of a single PAT IP — raises session headroom and distributes reputation risk.
Mechanics (OPNsense): 1. VIPs (HA-aware): create a CARP VIP as the carrier, then an IP Alias VIP for the pool range parented to that CARP VIP — so the whole pool inherits CARP state and floats to the master as one. (A bare IP Alias doesn't fail over; parenting it to the CARP VIP is what makes the pool HA.) 2. Firewall ▸ NAT ▸ Outbound (Hybrid or Manual): rule matching corporate/venue source nets → Translation target = the pool alias, and set Pool Options.
Pool type — pick by whether the far end cares about your source IP: | Pool type | Behaviour | Use when | |---|---|---| | Round-robin + sticky-address | spreads new sources across the pool, but a given internal host stays on one egress IP for its session lifetime | default — load spread + per-host consistency | | Source-hash | internal IP deterministically maps to a fixed egress IP | you want a stable, reproducible client→egress mapping | | Round-robin (no sticky) | every connection may use a different egress IP | pure scale, far end doesn't care about source |
pf under the hood: nat on $wan from <src> to any -> { 185.109.40.a - .b } round-robin sticky-address (or source-hash).
⚠ Do NOT pool payment/POS egress. Card processors whitelist a specific, stable source IP — round-robin would rotate you off the whitelist and break transactions. Put a higher-priority outbound NAT rule above the pool: POS nets → payment-prefixes translating to a single fixed /28 address, then the general pool rule catches everything else. (This dovetails with the Starlink payment-failover design — payments want a known, stable egress either way.)
Reputation trade-off: pooling means one blacklisted address only affects a fraction of traffic — but it also "dirties" more of your scarce /28 with general egress. With ~14 usable addresses, a pool of 2–4 for general egress + 1 fixed for payments + inbound VIPs is a sensible carve. Guest stays on its own separate firewall/address regardless (DESIGN §8).
6. HA — CARP + pfSync + config sync¶
- pfSync (HA, 2× 1G bond): two direct cables between the nodes, i350 1G ports, active-backup — no switch, redundant, can't be disrupted by a single NIC/cable fault. Carries firewall state so failover is stateful. (1G is ample — §2.)
- CARP heartbeat: advertises on the WAN, server-net (LAN), and the pfSync bond — multiple independent paths, so a single link failure can't cause split-brain (both nodes master). No dedicated heartbeat port needed.
- Config sync (XMLRPC): over the pfSync bond — master pushes config to backup automatically.
- Master/backup: advertise-skew makes node-1 master by default; all CARP VIPs (the /28 public VIPs, each server VLAN gateway) fail over together. (No VIP on the WAN /31 transit links.)
7. OOB & IPMI¶
- No dedicated OOB NIC — the two 1G ports are used for HA pfSync (§6). OPNsense OS management is in-band: a mgmt VLAN on the server trunk (per-node IP + CARP VIP), reachable via the routed path / mgmt VPN.
- IPMI/BMC (the server's dedicated management port, separate from OS NICs) → the OOB CRS326 switch. This is the box's out-of-band path: power-cycle / serial console / KVM / virtual-media even when OPNsense is down. Both nodes' BMCs on OOB, reached via the OOB WireGuard entry (DESIGN §7.3).
- Break-glass: if the server trunk is fully down (in-band mgmt gone), reach the OS via BMC KVM. Accepted trade-off for using both 1G ports on pfSync.
8. Physical / resilience notes¶
- Diverse power: node-1 and node-2 on different PDUs/feeds; each NIC-pair's two links ideally across two cards and the cabling routed diversely to the two CCRs / two MLAG members.
- Core diversity: node-1↔CR-COLO-01 and node-2↔CR-COLO-01 should not share a single break point; same for CR-COLO-02.
- Cabling colour code (suggest): core = yellow, server-trunk = blue, WAN = red, sync = green (crossover between nodes), OOB = grey. Label every DAC/patch both ends with role + peer.
- DACs vs optics: in-rack (sync, core, server-trunk to co-located CRS) use SFP+ DAC; anything leaving the rack uses optics.
9. Cabling diagram (per node shown once; mirror for node-2)¶
graph LR
subgraph N1["OPNsense FW-01"]
P1[SFP+ Core A]
P2[SFP+ Core B]
P34[SFP+×2 Server LACP]
P5[SFP+ WAN]
P7[SFP+ pfSync]
P8[RJ45 OOB]
BMC[IPMI]
end
P1 --- CR1[CR-COLO-01]
P2 --- CR2[CR-COLO-02]
P34 === VM[VM CRS326 MLAG pair]
P5 --- WAN[ISP handoff]
P7 -. direct DAC .- FW2[OPNsense FW-02]
P8 --- OOB[OOB switch]
BMC --- OOB
10. PCIe lanes, NUMA & throughput (sinking max traffic)¶
The Xeon Gold 6150 has 48 PCIe 3.0 lanes/socket = 96 dual-socket. Four NIC cards at x8 each = 32 lanes + BOSS + chipset fits with room to spare — but placement still decides whether you get line rate or a bottleneck:
- Every X710 must sit in a physically-x8 (electrical) PCIe 3.0 slot. A dual-10G X710 needs x8 for full 2×10G under load; an x4 slot caps it (~32 Gbps gen3 — OK for 20G, but no headroom and a trap if you don't check). Verify slot electrical width in the manual, not just physical size.
- Balance the 4 cards across both CPUs' lanes (NUMA). Two cards on CPU1's root complex, two on CPU2's. Traffic whose NIC, RSS queues, and handling cores are all on the same NUMA node goes fast; cross-NUMA (NIC on CPU1, cores on CPU2) pays a UPI hop and kills pps. So: put the server-trunk cards on one CPU's slots, the core+WAN cards on the other, and pin each interface's queues to local-NUMA cores.
- RSS on, queues = local-NUMA core count. X710 (FreeBSD
ixl) has solid multiqueue/RSS — spread flows across queues so no single core is the ceiling. Keep hyperthreading on (helps Suricata worker threads). - Memory channels: 8× DIMMs on a 6-channel/socket platform = 4+4, leaving 2 channels/socket empty. Fine for firewalling, but high-pps + IDS is memory-bandwidth sensitive — if you ever push this box hard, populating all 12 channels (12× DIMMs) lifts the pps ceiling more than more RAM capacity would. Flag for the parts order.
- Offloads: checksum/TSO/LRO on for max raw throughput on non-inspected paths — but see §11, IDS needs some of these off on inspected interfaces.
- Realistic ceiling: as a pure L3 firewall/router this hardware pushes line-rate across multiple 10G (tens of Gbps aggregate) with the tuning above — the limit is PCIe/NUMA/memory, not CPU. With IDS/IPS inline the bottleneck moves to Suricata (§11), which is why inspection placement matters.
FreeBSD/OPNsense tuning checklist for throughput: RSS enabled; hw.ix/ixl queues per NUMA; disable flow-control; bump net.isr, mbuf clusters, and kern.ipc.nmbclusters; AES-NI on (WireGuard/IPsec); pin Suricata to a NUMA node.
11. IDS / IPS — placement is the whole game¶
You want IDS on and max throughput — those pull against each other, so place inspection deliberately rather than "everywhere."
- IDS (detection, passive) vs IPS (inline, blocking): IPS sits in the forwarding path and its throughput is your throughput; IDS on a mirror inspects a copy and doesn't cap forwarding but can miss packets under load. Decide per interface.
- Do NOT inline-inspect the server-room trunk. East-west + backup traffic on the 20G bond would be crushed by Suricata, and it's already behind the zone firewall. Leave the server-trunk uninspected (offloads on, line rate).
- Inline IPS on WAN (and the guest/DMZ edges) — that's where threats enter, and real WAN utilisation is far below 10G line rate, so Suricata keeps up. This is the right place to spend the inspection budget.
- Suricata mode: netmap inline on the inspected interfaces (WAN/DMZ). Netmap requires LRO/TSO off on those NICs (Suricata must see real, un-coalesced packets) — so the offload note in §10 applies per-interface: off where inspected, on elsewhere.
- Workers mode, pinned: run Suricata in
workersrunmode, thread count matched to a NUMA node's cores,cpu-affinitypinned so inspection doesn't cross NUMA. 36 threads of headroom means a rich ruleset (ET Open/Pro) runs comfortably at WAN rates. - Realistic inline number: with a full ruleset, plan low-single-digit to ~5–10 Gbps sustained inline per box depending on ruleset and tuning — ample for a 10G WAN at real utilisation, not for inspecting the 20G server trunk (hence: don't).
- HA interaction: Suricata runs on both nodes; only the CARP master actively inspects live traffic. Keep rulesets identical (config sync) so a failover doesn't change the security posture. IPS block state is per-node — a mid-session failover re-evaluates on the new master (acceptable).
- Logging off-box: ship Suricata EVE/alerts to the Tier-2 monitoring/syslog (
10.128.36.x) — don't let alert logging chew the BOSS card or local IO.
Net rule: line-rate everywhere by default; inline IPS only on WAN/DMZ; server-trunk and storage never inline. That gives you "IDS on" where it matters without capping the high-bandwidth internal paths.
12. Performance tunables (baseline, then tune from counters)¶
Set these as a starting point, then raise from observed drops, not blindly — monitor the counters in §12g and only push a knob when its counter shows pressure. GUI locations noted (System ▸ Settings ▸ Tunables = sysctl/loader; Interfaces ▸ [iface] = offloads; Firewall ▸ Settings ▸ Advanced = pf).
12a. NIC / driver (X710 ixl / iflib)¶
| Tunable | Value | Why |
|---|---|---|
dev.ixl.<n>.iflib.override_nrxds / override_ntxds |
4096 / 4096 |
Deeper rings absorb bursts at 10G+ (default 1024 drops under microbursts). Loader tunable. |
| RX/TX queues | auto (= cores, capped 8/port on X710) | iflib pins one queue/core; leave auto unless NUMA-tuning by hand |
dev.ixl.<n>.fc |
0 |
Disable flow control — Ethernet PAUSE causes head-of-line blocking; let TCP/queues handle it |
| Hardware offloads (checksum, TSO, LRO) | ON on uninspected ifaces (uncheck the three "disable offload" boxes); OFF on WAN/DMZ inspected ifaces | Offload = throughput; but netmap/Suricata must see raw packets (§11) |
| Interrupt moderation (AIM) | default on | X710 adaptive moderation balances latency/pps well |
12b. Kernel network stack (sysctl)¶
| Tunable | Value | Why |
|---|---|---|
net.isr.maxthreads |
-1 (all cores) |
Spread netisr across all cores for high pps |
net.isr.bindthreads |
1 |
Pin netisr threads to cores (NUMA locality) |
net.isr.dispatch |
deferred |
Queue to isr threads — best for multi-core forwarding (vs direct) |
kern.ipc.nmbclusters |
1000000 |
mbuf pool — many queues × deep rings need headroom or RX stalls |
kern.ipc.nmbjumbo9 |
524288 |
9K jumbo mbufs for the jumbo server-trunk |
net.inet.ip.intr_queue_maxlen |
2048 |
Deeper IP input queue |
net.route.multipath |
1 |
ECMP — the two CCR sessions actually load-share |
net.inet.ip.redirects / net.inet.icmp.drop_redirect |
0 / 1 |
Hygiene on a router |
net.inet.ip.forwarding |
1 |
(OPNsense sets this when routing) |
12c. Firewall (pf)¶
| Setting | Value | Why |
|---|---|---|
| Firewall Maximum States | 3–5 million | A busy edge with many venues/flows blows past the default; sized to RAM (128 GB is ample) |
net.pf.states_hashsize |
scale with state count (e.g. 1048576) |
Hash big enough to keep state lookup O(1) |
| Firewall Optimization | normal (→ aggressive if state churn high) |
Aggressive expires idle states faster, trims the table |
| Scrub / reassembly | on WAN, off internal fast paths if not needed | Scrub costs CPU; keep it at the untrusted edge |
| State-policy | if-bound only where needed |
Default floating is fine; if-bound helps asymmetric edge cases (not needed here with CARP-gating) |
Do not set skip on the server-trunk |
— | You want stateful zero-trust between venues and server zones (DESIGN §10) |
12d. CPU / NUMA / power¶
| Tunable | Value | Why |
|---|---|---|
hw.acpi.cpu.cx_lowest |
C1 |
Avoid deep C-states — deep sleep adds latency/jitter to interrupt handling |
Power profile (powerd) |
off / performance | No frequency scaling on a forwarding box — keep cores at full clock |
| HyperThreading | on | More Suricata worker threads; forwarding also benefits |
| NIC↔CPU placement | balance cards across both sockets, pin queues local-NUMA (§10) | Cross-NUMA is the silent pps killer |
12e. Suricata / IDS (suricata.yaml, mostly GUI-exposed)¶
| Setting | Value | Why |
|---|---|---|
| Runmode | workers |
Best multi-core scaling |
Threads / cpu-affinity |
= one NUMA node's cores, pinned | Inspection stays on-socket; doesn't fight forwarding |
| Capture | netmap, inspected ifaces only (WAN/DMZ) | Never assign Suricata to the server-trunk |
detect.profile |
high |
128 GB RAM affords big detection contexts |
| Flow/stream/reassembly memcaps | generous (GB-scale) | Prevents inspection drops under load (stats.log → *.memcap_drops) |
max-pending-packets |
4096+ |
Deeper pipeline for burst tolerance |
| Ruleset | ET Open/Pro, prune with rule-profiling | Fewer, relevant rules = higher inline throughput |
| EVE/alert logging | remote (Tier-2 syslog 10.128.36.x) |
Don't let logging chew the BOSS card / local IO |
12f. WireGuard / crypto (if road-warrior VPN lands here)¶
- AES-NI on (BIOS +
aesnimodule) — hardware crypto. Xeon Gold 6150 has it; ensures VPN isn't CPU-bound. - WireGuard on OPNsense is kernel-mode (ROS7-comparable) — multi-core, scales with the cores available.
12g. Monitor these — the tunables are demand-driven¶
netstat -m→ mbuf denied/delayed ≠ 0 → raisenmbclusters/nmbjumbo9.pfctl -si→ current vs max states near limit → raise Maximum States.sysctl dev.ixl.<n>.mac.rx_discards/.rx_no_buffers→ RX drops → deeper rings / more mbufs / check NUMA.vmstat -i→ interrupt distribution across cores → confirms RSS/NUMA spread.- Suricata
stats.log→capture.kernel_drops/*.memcap_drops→ inspection can't keep up → prune rules, raise memcaps, or move that iface off inline. top -HP→ per-core load → spot a single hot core (RSS not spreading) or NUMA imbalance.
Guiding principle: the ceiling is PCIe/NUMA/memory and (with IPS) Suricata — not raw CPU. Set the §12a–d baseline, place IDS per §11, then let the §12g counters tell you what to raise. Don't max every knob preemptively; oversized mbuf/state pools waste RAM and can hide real issues.
13. Hairpin NAT / NAT reflection (published services)¶
Services are published as port forwards on the WAN public IP (CARP VIP) → internal targets — most to the server network (DMZ/Tier zones, 10.128.x), occasionally one to a venue LAN (10.V.x). The hairpin problem: an internal client (a venue till, a staff PC) that reaches the service by its public IP sends to an address the firewall owns; without help, the reply returns from the internal server's real IP, the client (which expected a reply from the public IP) drops it, and the connection hangs.
13a. Primary fix — split-horizon DNS (avoid the hairpin entirely)¶
Internal clients resolve service names to internal IPs, so they never touch the public IP or hairpin at all:
- On the internal resolvers (
10.128.32.3/.4, Tier-1), publish an internal view of each service zone:service.pubinvest.co.uk → 10.128.x(the real internal target), while public DNS keeps the public IP. - Result: venue/staff client → internal IP → normal east-west path (venue → core → OPNsense zone firewall → server). Client source IP is preserved (good for logging/zero-trust), no reflection latency, no extra firewall load, no hairpin to the colo and back.
- This is the recommended default for every published service where you control the client's DNS — which here is all corporate/venue clients (they use internal resolvers by design).
13b. Fallback — NAT reflection (Pure NAT) for hardcoded-IP cases¶
For anything that can't use split-DNS (a device with a hardcoded public IP, or a client on external DNS), enable reflection so the firewall hairpins correctly:
Firewall ▸ Settings ▸ Advanced ▸ Network Address Translation:- Reflection for port forwards = Pure NAT (not "NAT + Proxy" — proxy is TCP-only and doesn't scale).
- Automatic outbound NAT for Reflection = on — this is the essential half: it source-NATs the reflected traffic to OPNsense's internal interface IP so the server's reply returns through OPNsense to be un-NAT'd. Without it, replies go direct and the client drops them.
- Reflection for 1:1 = on (if you use 1:1 NAT for any server needing a whole public IP).
- Or set reflection per-port-forward rule (override the global) so you only reflect where needed.
- Caveat: reflection SNATs the source, so the server sees OPNsense's IP, not the real client — you lose client-IP visibility and per-client firewalling on reflected flows. That's the trade-off, and exactly why 13a (split-DNS) is preferred wherever possible.
13c. The common case — forwards to the server network¶
- Port forward: /28 public VIP : port → server target (
10.128.xDMZ/Tier). Target a /28 CARP VIP, not a node or WAN-transit IP, so forwards fail over with the master. - Auto-create the associated pass rule (WAN → target : port). For a DMZ web service, prefer a reverse proxy (Caddy/HAProxy on OPNsense or a DMZ host) over raw forwards — TLS termination, one public port, WAF-friendly.
- Add the split-DNS internal record (13a). Reflection (13b) covers stragglers.
13d. The odd case — a forward to a venue LAN¶
A forward whose target is a venue device (10.V.x) has more moving parts than a server-network forward, because the venue sits behind its own router/zone firewall across the metro:
- OPNsense: port forward WAN VIP : port →
10.V.x, + pass rule. - Return path: the venue device replies toward the client — for an internal hairpin client this needs reflection outbound-NAT (13b) or split-DNS (13a) so symmetry holds; for a genuine internet client the reply routes back via the core to OPNsense normally.
- Venue router (the extra step): the venue's zone firewall must permit the inbound flow to
10.V.x:port — add the source (OPNsense's server/WAN reach, ortrusted-inbound) to that venue's allow rules (CONFIG-GUIDE §3.2). A server-network forward doesn't need this because OPNsense is the server VLANs' gateway; a venue forward crosses a second firewall you must open. - Prefer relocating the service to the DMZ if it's genuinely internet-facing — a public port pointing into a venue LAN widens that venue's exposure. Keep venue-LAN forwards rare and tightly scoped (single host, single port), exactly as you framed it ("maybe the odd one").
13e. CARP & security notes¶
- All port forwards and 1:1 NAT target the CARP VIP → they follow the master on failover; pfSync keeps the state so existing sessions survive.
- Keep the forward list minimal and documented — every forward is public attack surface. Reverse-proxy web services; restrict source where the audience is known (partner IPs → an address-list on the pass rule).
- Reflection only works on the master (it's stateful NAT) — fine, since only the master forwards anyway.
14. Zabbix monitoring¶
Monitor both nodes over the OOB/management interface (not the data path — a data-path or CARP event must never read as "host down"). Zabbix server/proxy lives on the Tier-2 monitoring VLAN (10.128.36.x). This replaces the earlier LibreNMS mention for the corporate side.
14a. Collection methods (layer them)¶
| Layer | Method | Covers |
|---|---|---|
| OS/base | Zabbix agent (active) via os-zabbix-agent plugin + template FreeBSD by Zabbix agent |
CPU, per-core, memory, swap, disk, processes, uptime, base interfaces |
| Interfaces | SNMPv3 (os-net-snmp), source-locked to the Zabbix IP |
High-rate per-port in/out/errors/discards on the X710s |
| Hardware | IPMI polled from the BMC over OOB (Zabbix IPMI monitoring) | PSU redundancy, fans, temps, hardware event log — independent of the OS |
| Firewall/routing | Custom UserParameters (agent) | CARP, pfSync, pf states, BGP (FRR), Suricata — nothing standard covers these |
Prefer agent-active (the firewall pushes; no inbound poll holes in the OOB firewall). Don't apply SNMP and agent templates to the same metric — pick one source per item or you double every graph and trigger.
14b. Trim / turn OFF from the stock templates (they over-monitor a firewall)¶
- Interface LLD is the big one. OPNsense presents dozens of interfaces — VLANs,
carp*,pfsync0,enc0,pflog0,lo0,gif/gre,usbus*, and your ~8 intentionally-down spare SFP+. Left unfiltered you get item bloat + a wall of "link down" alerts. Filter the LLD (macro{$NET.IF.IFNAME.MATCHES}/NOT_MATCHES) to physical + key logical ports; exclude^(lo|pflog|enc|pfsync|usbus|gif|gre). Monitorpfsync0andcarp*via the custom items in 14c instead of the generic interface template. - Link-down triggers on spare ports — the ~8 unused SFP+ are meant to be down. Either don't discover them, or gate the link trigger on
ifAdminStatus=up(only alert when a port that's supposed to be up goes down). - Duplicate CPU/mem if both SNMP and agent templates land — disable one side.
- Aggressive interface-error triggers — default "any error delta" is noisy near 60 GHz-adjacent paths. Change to a rate over time (e.g. discards/errors > N/s sustained 5 m).
- Disk-IO / low-value FS items on the BOSS card — the box isn't IO-bound; trim if noisy. Keep disk space + SMART.
- ICMP host-availability — point it at the OOB address, not a data/VIP address, so failover isn't "host down."
14c. ADD — custom items (the firewall-specific metrics no template has)¶
Ship these as agent UserParameters (in the OPNsense zabbix-agent config / a dropin). All the "is it down" ones must be CARP-aware (see 14d).
| Item | Source (UserParameter) | Alert on |
|---|---|---|
| CARP master/backup per VIP | ifconfig \| grep -c 'carp:.*MASTER' (and BACKUP) |
designated master unexpectedly BACKUP; BOTH nodes MASTER = split-brain (critical); neither MASTER = no gateway (critical) |
| pfSync health | pfsync0 up + state-count delta vs peer | sync iface down; large state divergence between nodes |
| pf state table | pfctl -si current vs pfctl -sm max |
> 80% of Maximum States (ties to §12c); any "state limit" drops |
| BGP sessions (FRR) | vtysh -c 'show bgp summary json' parsed per neighbour (CCR1, CCR2, upstream) |
any neighbour ≠ Established while this node is MASTER |
| Suricata | EVE stats / stats.log: capture.kernel_drops, *.memcap_drops, alert rate |
drop rate > 0 sustained (inspection failing); alert-rate spike (possible attack) |
| NIC HW counters | sysctl dev.ixl.<n>.mac.rx_discards / rx_no_buffers / crc_errors |
discards/no-buffers rising → §12g tuning needed |
| Per-core CPU | kern.cp_times (or agent per-CPU) |
one hot core = RSS/NUMA imbalance (§10) |
| Temp / PSU / fan | IPMI (14a) | over-temp, PSU redundancy lost, fan fail |
| Cert expiry / config revision | web-UI cert; OPNsense config change | cert < 21 days; config changed without a ticket (audit) |
| WireGuard peer handshakes (if VPN here) | wg show last-handshake age |
staff peer stale (optional) |
14d. CARP-aware alerting — the recurring gotcha¶
On the BACKUP node, by design: FRR is stopped (§3), all VIPs are in BACKUP, and it carries no traffic. So "FRR down / BGP sessions down / VIPs not master / ~0 pps" are normal on the backup and must not alert. Two ways to handle it, pick one and apply consistently:
- Gate triggers on role: every "service down" trigger includes and last(/…/carp.master)=1 (only fires when this node is master), or
- Trigger dependencies: a "node is BACKUP" trigger that suppresses the dependent service-down triggers.
Otherwise a perfectly healthy standby generates a constant stream of false criticals — the single most common mistake monitoring a CARP+FRR pair.
14e. Pair-level triggers (the ones that actually matter for HA)¶
Beyond per-node health, alert on pair state: - Split-brain: both nodes report CARP MASTER → critical (sync-link failure / misconfig). - No master: neither node MASTER → critical (total edge down). - Redundancy lost: backup node unreachable/failed while master is fine → you are now a SPOF — warning even though service is up. - Config drift: master's config revision ≠ backup's → HA config sync broke. - Both RADIUS/upstream down, both transits down (if applicable) — service-affecting.
14f. Dashboards & routing¶
- One "Edge HA" dashboard: master indicator per node, BGP session matrix (CCR1/CCR2/upstream × node), pf state gauge, Suricata drop/alert rate, WAN throughput, temps.
- Route critical (split-brain, no-master, both-transits) to on-call immediately; warning (redundancy lost, cert expiry, NIC discards) to the daily digest.
- Suricata high-severity alerts are better handled in the log/SIEM stack (EVE → syslog
10.128.36.x), with Zabbix watching only the health metrics (drops, running) — don't turn Zabbix into an IDS console.
15. Summary of the answer to each requirement¶
- Direct uplink to server switch net, dedicated interfaces → #3+#4 dedicated LACP bond to the VM MLAG pair (§4).
- Direct core with BGP, redundant to the two core routers → #1+#2, each node dual-homed to both CCRs, eBGP AS65510 + BFD (§3).
- WAN + upstream BGP → one eBGP /31 to each border router (bdr-1, bdr-2) — no VRRP; BGP does redundancy (bdr-1 primary via local-pref, BFD both). /31s carry transit only; the public /28 is announced and floats as CARP VIPs for NAT/services. Announce only the routed block, take default only, pass default inward. "Which node" = CARP-gated FRR, not a VIP (§5).
- FRR CARP failover → os-frr CARP-tracking stops FRR on the BACKUP node — only the master peers/announces, which makes routing symmetric with no prepend hacks (§3).
- HA CARP → CARP VIPs on WAN+server VLANs, pfSync over a dedicated direct DAC, config sync, dual heartbeat paths (§6).
- OOB → #8 mgmt to OOB switch + BMC to OOB, independent of production (§7).
- Ports → 8× SFP+ + 2× 1G RJ45 + separate BMC. 2 core + 2 server + 2 WAN = 6 SFP+; HA pfSync bonded over the 2× 1G copper (active-backup, direct to peer); 2 SFP+ spare (→ optional 40G server trunk). OS mgmt in-band; OOB via BMC.
- PCIe / throughput → all X710s in x8 gen3 slots, NICs balanced across both CPUs, RSS + NUMA-pinned; line-rate across multiple 10G as a router (§10).
- IDS on → inline IPS on WAN/DMZ only (netmap, workers, NUMA-pinned), server-trunk uninspected — inspection where threats enter without capping internal bandwidth (§11).
- Tunables → NIC rings/offloads, netisr, mbuf pools, ECMP, pf state table, C-states/power, Suricata affinity — baseline in §12, then tune from the §12g counters.
- Outbound NAT pool → variable source across a /28 pool (IP-Alias parented to a CARP VIP; round-robin+sticky or source-hash), but payments on a fixed whitelisted address, not pooled (§5c).
- Hairpin NAT → split-DNS primary (preserves client IP), Pure-NAT reflection fallback; most forwards to server DMZ, the odd one to a venue LAN needs the venue router opened too (§13).
- Zabbix monitoring → agent + SNMP + IPMI over OOB; trim interface-LLD/spare-port noise; add CARP/pfSync/pf-states/BGP/Suricata custom items; CARP-aware alerting so the healthy backup stays quiet; pair-level split-brain/redundancy triggers (§14).