10. Performance tunables¶
Do this on both nodes — tunables are per-node (not synced). Set the baseline, then
raise from observed counters, not blindly (§12). GUI locations: System ▸ Settings
▸ Tunables (sysctl/loader), Interfaces ▸ [iface] (offloads), Firewall ▸ Settings ▸
Advanced (pf).
10.1 NIC / driver (X710 ixl)¶
| Tunable | Value | Why |
|---|---|---|
dev.ixl.<n>.iflib.override_nrxds / override_ntxds |
4096 / 4096 |
deep rings absorb 10G+ microbursts (default 1024 drops) |
dev.ixl.<n>.fc |
0 |
disable flow-control (PAUSE causes head-of-line blocking) |
| Offloads (checksum/TSO/LRO) | ON uninspected ifaces; OFF on WAN/WAN2/VL_DMZ |
offload = throughput, but netmap/Suricata needs raw packets (§11) |
Offloads OFF on the Suricata interfaces
WAN, WAN2, VL_DMZ run inline netmap IPS — untick "hardware CRC/TSO/LRO" (i.e.
disable offload) on those three. Leave offloads on for SRVTRUNK and the rest.
10.2 Kernel network stack (sysctl)¶
| Tunable | Value | Why |
|---|---|---|
net.isr.maxthreads |
-1 |
spread netisr across all cores |
net.isr.bindthreads |
1 |
pin to cores (NUMA locality) |
net.isr.dispatch |
deferred |
best for multi-core forwarding |
kern.ipc.nmbclusters |
1000000 |
mbuf pool for many deep-ring queues |
kern.ipc.nmbjumbo9 |
524288 |
9K jumbo mbufs for the jumbo trunk |
net.inet.ip.intr_queue_maxlen |
2048 |
deeper IP input queue |
net.route.multipath |
1 |
ECMP — the two CCR sessions load-share |
net.inet.icmp.drop_redirect |
1 |
router hygiene |
10.3 Firewall (pf)¶
| Setting | Value |
|---|---|
| Firewall Maximum States | 3–5 million (sized to 128 GB) |
net.pf.states_hashsize |
1048576 (scale with state count) |
| Optimization | normal → aggressive if state churn is high |
| Scrub | on WAN, off internal fast paths |
10.4 CPU / NUMA / power¶
| Tunable | Value | Why |
|---|---|---|
hw.acpi.cpu.cx_lowest |
C1 |
avoid deep C-state latency/jitter |
Power profile (powerd) |
off / performance | no freq scaling on a forwarding box |
| HyperThreading | on | more Suricata worker threads |
| NIC↔CPU placement | balance cards across sockets, pin queues local-NUMA | cross-NUMA is the silent pps killer (§10) |
10.5 Suricata (mostly GUI-exposed)¶
| Setting | Value |
|---|---|
| Runmode | workers |
| Threads / cpu-affinity | = one NUMA node's cores, pinned |
detect.profile |
high |
| flow/stream/reassembly memcaps | generous (GB-scale) |
max-pending-packets |
4096+ |
10.6 Monitor these — tunables are demand-driven (§12g)¶
Don't max every knob preemptively. Watch, then raise the one under pressure:
| Counter | Meaning | Action |
|---|---|---|
netstat -m denied/delayed ≠ 0 |
mbuf starvation | raise nmbclusters/nmbjumbo9 |
pfctl -si near max |
state table filling | raise Max States |
dev.ixl.N.mac.rx_discards/rx_no_buffers |
RX drops | deeper rings / more mbufs / check NUMA |
vmstat -i |
IRQ distribution | confirm RSS/NUMA spread |
Suricata capture.kernel_drops/*.memcap_drops |
inspection can't keep up | prune rules, raise memcaps, or move iface off inline |
top -HP |
per-core load | spot a hot core (RSS not spreading) |
Guiding principle
The ceiling is PCIe/NUMA/memory and (with IPS) Suricata — not raw CPU. Set this baseline, place IDS per page 9, then let the counters tell you what to raise. Oversized mbuf/state pools waste RAM and hide real issues.