Skip to content

WAF troubleshooting

All commands run on a WAF node unless stated otherwise.

A site returns 403

Find the request in the log and read waf_action / policy_src:

sudo tail -500 /var/log/haproxy.log | grep '"app": "<app>"' | sed 's/^[^{]*//' \
  | jq -c '{src,country,status,waf_action,policy_src,waf_rule_ids,bot_id,bot_verified,path}' | tail
waf_action Cause Fix
block, policy_src: geo Country policy. Internal clients have no country and fail allow lists Test from outside, or set the geo action to log
block, waf_rule_ids set OWASP rules in block mode matched Check it's not a false positive; add an exclusion by rule ID and path
ip_ban / crowdsec_ban Manual ban or CrowdSec decision IP lists page, or sudo cscli decisions list / cscli decisions delete --ip <ip>
bot_block Bot policy (spoofed claim, AI category, unverified) The app's Bots section
ratelimit (429) Per-IP rate limit for the app Raise the app's rate limit

A site returns 503

  • Origin down: HAProxy's health check on the origin is failing. Check the origin answers on its address and port with the app's hostname, and for TLS mode verify that its certificate matches the SNI name and is trusted (or switch the origin to plain).
  • Coraza down, app set to fail closed: waf_action: spoa_error. systemctl status coraza-spoa.

A site returns 421 "unknown host"

The hostname isn't in any app's hostname list, or DNS points at the WAF before the app exists.

Nothing answers on the VIP

sudo /usr/local/sbin/waf-bgp status --json
sudo vtysh -c 'show bgp ipv4 unicast summary'
systemctl is-active frr waf-bgp-health haproxy
ip -4 addr show lo            # the VIP /32 should be there
  • "drained": true: someone (or a failed deploy) drained the node. sudo waf-bgp undrain.
  • "announced": false, not drained: HAProxy's health check is failing, so the watcher withdrew the VIP. Check systemctl status haproxy and the node's /.waf/health endpoint.
  • Session not Established: check the neighbour on the firewall, and that BGP from the peer is allowed.
  • waf-bgp: vtysh failed … bgpd is not running: FRR is up without its BGP daemon. sudo systemctl restart frr, then check again.

A deploy failed

Open the failed deploy in the GUI or the CI pipeline. The play stops on the first failing node, restores its previous config and never touches the next node, so the site keeps serving.

Symptom Cause
All jobs fail waiting for a runner (e.g. Insufficient cpu) CI runner capacity, not the code. Retry when runners are free
No WAF node is announcing the VIP Nothing is announcing: fix BGP first, or for the very first BGP deploy set the one-off first-deploy flag
BGP session to … not Established after 150s The node's BGP session didn't come up after the change; check the firewall side
Stuck in "pending", then timeout The pipeline was queued behind another job; check CI before retrying

Instant change shows "pending"

A node didn't accept it (unreachable or restarting). It's already saved in the database and is retried automatically, and every node is re-synced every minute. It clears by itself once the node answers.

Bot lists look stale

cat /etc/haproxy/live/maps/bot_lists.built   # last sync time
wc -l /etc/haproxy/live/maps/bot_ranges.map

Check the nightly bot-list job ran and went green. A red run keeps the previous data.