Skip to content

WAF network and BGP

The WAF is reached through a virtual IP (VIP) that the WAF nodes announce to the edge firewall over BGP. The firewall routes the VIP to whichever node is announcing it, so a node can be taken out of service by withdrawing its route, with no DNS or NAT change.

Traffic path

  1. DNS for each protected site points at the WAF's public address.
  2. The edge firewall NATs that address (ports 80 and 443) to the VIP.
  3. The firewall routes the VIP to a WAF node, learned over BGP.
  4. The node applies the WAF checks and forwards allowed requests to the app's origin.

How the VIP is announced

  • FRR runs BGP (and BFD, for sub-second failure detection) on each WAF node.
  • The VIP is configured on each node's loopback as a /32 and is never answered by ARP on the LAN; traffic reaches it only via the BGP route.
  • A small watcher (waf-bgp-health) checks the node's own HAProxy health endpoint every second. After 3 good checks it announces the VIP; after 3 failures it withdraws it. A node that is unhealthy or drained never attracts traffic.
  • The node firewall only accepts BGP (tcp/179) and BFD (udp/3784–3785) from the configured BGP peers.
  • The BGP session password is optional and controlled by a setting; when enabled, a deploy refuses to run without it.

Modes

Set once for the pair in the Ansible inventory, then deploy:

Mode Behaviour
Active/passive The primary announces the VIP normally. The backup announces it with its AS path prepended, so the firewall only uses the backup when the primary's route is gone. All traffic goes to one node at a time.
Active/active Both nodes announce the VIP equally and the firewall spreads traffic across them (ECMP). Rate limits can be up to twice as generous for a client whose flows hash to both nodes.

In both modes the nodes keep rate-limit stick tables (HAProxy peers) and CrowdSec decisions (cross-reported) in sync, so a switchover loses no state.

Failback is immediate: when the primary comes back (health restored, or undrained), all traffic moves back to it at once, and long-lived connections such as WebSockets reconnect.

Checking BGP

On a WAF node:

sudo vtysh -c 'show bgp ipv4 unicast summary'                     # session to each peer: Established
sudo vtysh -c 'show ip bgp neighbors <peer> advertised-routes'     # should list the VIP /32
sudo /usr/local/sbin/waf-bgp status --json                         # {"announced": true, "drained": false, "sessions_established": true}

On the firewall's BGP diagnostics page, the VIP /32 should appear with a WAF node as next hop. In active/passive with two nodes there are two paths: the primary's selected as best, the backup's with a longer AS path. In active/active both are marked multipath.

Draining a node (maintenance)

sudo /usr/local/sbin/waf-bgp drain     # withdraw the VIP; stays drained across reboots
sudo /usr/local/sbin/waf-bgp undrain   # announce again once healthy

Deploys drain a node automatically before disruptive work (package upgrades, a Coraza restart), but only when another node is announcing, so a deploy can never take the pair to zero nodes. An operator's manual drain is left in place by deploys.

With a single node, draining it takes the WAF offline.

Deploy safety checks for BGP

  • Before touching any node, a deploy checks the BGP settings (mode, exactly one primary in active/passive, peers configured) and that at least one node is announcing the VIP.
  • After updating each node, it waits for that node to announce again and for its BGP session to be Established (up to 150 seconds) before moving to the next node.
  • The very first BGP deploy of an environment needs a one-off first-deploy flag, because nothing is announcing yet. Remove it straight afterwards.

Adding a node

  1. Build the VM like the existing nodes (same OS, deploy user, key and sudo setup).
  2. On the firewall, add a BGP neighbour for the new node with the WAF nodes' AS number, and accept the VIP /32 from it.
  3. Add the node to the inventory (with its role, in active/passive), then Preview and Apply.
  4. Check the firewall now has a path for the VIP from each node, with the primary preferred.