Skip to content

10. Performance tunables

Do this on both nodes — tunables are per-node (not synced). Set the baseline, then raise from observed counters, not blindly (§12). GUI locations: System ▸ Settings ▸ Tunables (sysctl/loader), Interfaces ▸ [iface] (offloads), Firewall ▸ Settings ▸ Advanced (pf).

10.1 NIC / driver (X710 ixl)

Tunable Value Why
dev.ixl.<n>.iflib.override_nrxds / override_ntxds 4096 / 4096 deep rings absorb 10G+ microbursts (default 1024 drops)
dev.ixl.<n>.fc 0 disable flow-control (PAUSE causes head-of-line blocking)
Offloads (checksum/TSO/LRO) ON uninspected ifaces; OFF on WAN/WAN2/VL_DMZ offload = throughput, but netmap/Suricata needs raw packets (§11)

Offloads OFF on the Suricata interfaces

WAN, WAN2, VL_DMZ run inline netmap IPS — untick "hardware CRC/TSO/LRO" (i.e. disable offload) on those three. Leave offloads on for SRVTRUNK and the rest.

10.2 Kernel network stack (sysctl)

Tunable Value Why
net.isr.maxthreads -1 spread netisr across all cores
net.isr.bindthreads 1 pin to cores (NUMA locality)
net.isr.dispatch deferred best for multi-core forwarding
kern.ipc.nmbclusters 1000000 mbuf pool for many deep-ring queues
kern.ipc.nmbjumbo9 524288 9K jumbo mbufs for the jumbo trunk
net.inet.ip.intr_queue_maxlen 2048 deeper IP input queue
net.route.multipath 1 ECMP — the two CCR sessions load-share
net.inet.icmp.drop_redirect 1 router hygiene

10.3 Firewall (pf)

Setting Value
Firewall Maximum States 3–5 million (sized to 128 GB)
net.pf.states_hashsize 1048576 (scale with state count)
Optimization normal → aggressive if state churn is high
Scrub on WAN, off internal fast paths

10.4 CPU / NUMA / power

Tunable Value Why
hw.acpi.cpu.cx_lowest C1 avoid deep C-state latency/jitter
Power profile (powerd) off / performance no freq scaling on a forwarding box
HyperThreading on more Suricata worker threads
NIC↔CPU placement balance cards across sockets, pin queues local-NUMA cross-NUMA is the silent pps killer (§10)

10.5 Suricata (mostly GUI-exposed)

Setting Value
Runmode workers
Threads / cpu-affinity = one NUMA node's cores, pinned
detect.profile high
flow/stream/reassembly memcaps generous (GB-scale)
max-pending-packets 4096+

10.6 Monitor these — tunables are demand-driven (§12g)

Don't max every knob preemptively. Watch, then raise the one under pressure:

Counter Meaning Action
netstat -m denied/delayed ≠ 0 mbuf starvation raise nmbclusters/nmbjumbo9
pfctl -si near max state table filling raise Max States
dev.ixl.N.mac.rx_discards/rx_no_buffers RX drops deeper rings / more mbufs / check NUMA
vmstat -i IRQ distribution confirm RSS/NUMA spread
Suricata capture.kernel_drops/*.memcap_drops inspection can't keep up prune rules, raise memcaps, or move iface off inline
top -HP per-core load spot a hot core (RSS not spreading)

Guiding principle

The ceiling is PCIe/NUMA/memory and (with IPS) Suricata — not raw CPU. Set this baseline, place IDS per page 9, then let the counters tell you what to raise. Oversized mbuf/state pools waste RAM and hide real issues.