WAF troubleshooting¶
All commands run on a WAF node unless stated otherwise.
A site returns 403¶
Find the request in the log and read waf_action / policy_src:
sudo tail -500 /var/log/haproxy.log | grep '"app": "<app>"' | sed 's/^[^{]*//' \
| jq -c '{src,country,status,waf_action,policy_src,waf_rule_ids,bot_id,bot_verified,path}' | tail
waf_action |
Cause | Fix |
|---|---|---|
block, policy_src: geo |
Country policy. Internal clients have no country and fail allow lists | Test from outside, or set the geo action to log |
block, waf_rule_ids set |
OWASP rules in block mode matched | Check it's not a false positive; add an exclusion by rule ID and path |
ip_ban / crowdsec_ban |
Manual ban or CrowdSec decision | IP lists page, or sudo cscli decisions list / cscli decisions delete --ip <ip> |
bot_block |
Bot policy (spoofed claim, AI category, unverified) | The app's Bots section |
ratelimit (429) |
Per-IP rate limit for the app | Raise the app's rate limit |
A site returns 503¶
- Origin down: HAProxy's health check on the origin is failing. Check the origin answers on its
address and port with the app's hostname, and for TLS mode
verifythat its certificate matches the SNI name and is trusted (or switch the origin toplain). - Coraza down, app set to fail closed:
waf_action: spoa_error.systemctl status coraza-spoa.
A site returns 421 "unknown host"¶
The hostname isn't in any app's hostname list, or DNS points at the WAF before the app exists.
Nothing answers on the VIP¶
sudo /usr/local/sbin/waf-bgp status --json
sudo vtysh -c 'show bgp ipv4 unicast summary'
systemctl is-active frr waf-bgp-health haproxy
ip -4 addr show lo # the VIP /32 should be there
"drained": true: someone (or a failed deploy) drained the node.sudo waf-bgp undrain."announced": false, not drained: HAProxy's health check is failing, so the watcher withdrew the VIP. Checksystemctl status haproxyand the node's/.waf/healthendpoint.- Session not Established: check the neighbour on the firewall, and that BGP from the peer is allowed.
waf-bgp: vtysh failed … bgpd is not running: FRR is up without its BGP daemon.sudo systemctl restart frr, then check again.
A deploy failed¶
Open the failed deploy in the GUI or the CI pipeline. The play stops on the first failing node, restores its previous config and never touches the next node, so the site keeps serving.
| Symptom | Cause |
|---|---|
All jobs fail waiting for a runner (e.g. Insufficient cpu) |
CI runner capacity, not the code. Retry when runners are free |
No WAF node is announcing the VIP |
Nothing is announcing: fix BGP first, or for the very first BGP deploy set the one-off first-deploy flag |
BGP session to … not Established after 150s |
The node's BGP session didn't come up after the change; check the firewall side |
Stuck in "pending", then timeout |
The pipeline was queued behind another job; check CI before retrying |
Instant change shows "pending"¶
A node didn't accept it (unreachable or restarting). It's already saved in the database and is retried automatically, and every node is re-synced every minute. It clears by itself once the node answers.
Bot lists look stale¶
cat /etc/haproxy/live/maps/bot_lists.built # last sync time
wc -l /etc/haproxy/live/maps/bot_ranges.map
Check the nightly bot-list job ran and went green. A red run keeps the previous data.