Guide 05 — Alerting Design: Dependencies, Escalations, Maintenance¶
The reporting requirement is alerting-first: the system's job is to page once, about the right thing, with venue context. This guide is what makes 700 endpoints livable for one operator.
1. Trigger dependency tree (the keystone)¶
Dependencies mirror physical topology so root cause pages and symptoms stay quiet:
OPNsense: venue VPN tunnel down ← pages: "Venue X unreachable"
└── venue MikroTik unreachable (suppressed)
├── venue switch unreachable (suppressed)
│ └── each AP (suppressed)
└── each till (suppressed)
metro hub device down ← pages once
└── all venue tunnels via that hub (suppressed)
Implementation notes:
- Venue-device reachability triggers depend on the venue router's ICMP trigger; the router's trigger depends on the tunnel item on OPNsense.
- Because hosts are sync-generated, dependencies must be applied systematically,
not by hand: either template-level dependencies where possible, or a small
scripted pass (Zabbix API) that runs after the sync and wires venue-device
triggers to their venue router using the
site:tag. Treat that script as part of the sync tooling, in git. - Ceph/Proxmox triggers depend on the relevant node's reachability similarly.
2. Severity policy¶
| Severity | Meaning | Delivery |
|---|---|---|
| Disaster | Colo-level outage (Ceph ERR, cluster quorum, SQL down) | Page immediately, repeat |
| High | Venue fully unreachable, metro device down, Ceph WARN | Page immediately |
| Average | Single device down (one AP, one till), disk >90% | Notify, no repeat page |
| Warning | Trends: capacity forecast, sustained high util | Daily digest / dashboard only |
| Info | Everything else | Dashboard only |
A single till or AP down is deliberately not a page — it shows on the NOC grid and per-venue dashboard, and pages only if you choose to escalate (e.g. ≥2 tills down in one venue = High, via a calculated item/trigger per venue).
3. Actions and escalations¶
- One trigger action with escalation steps: notify → repeat after 30 min if unacked → (optional) second contact. Keep it to one action with conditions on severity/tags rather than many overlapping actions.
- Recovery notifications on for High+, off for the noise tiers.
- Route by tag when others come aboard later (
site:androle:tags are already on every host courtesy of the sync).
4. Test the tree before trusting it¶
Before parallel-run sign-off, run controlled failure drills:
- Drop one venue's tunnel (or ACL the venue) → expect exactly one High alert.
- Power off one till → expect one Average notification, no page.
- Pull one AP → same.
- Simulate metro path loss out of hours → expect one page, all venue alerts suppressed.
5. Maintenance windows¶
- Zabbix maintenance objects per host group: venue refits (
Venue/<name>group, data collection continues, alerts suppressed), colo patch nights, RouterOS upgrade windows. - Make creating a maintenance window part of the change habit — it's the difference between clean alerting and learned alert-blindness.
6. Noise budget¶
Working rule: if it paged and you took no action, change the trigger. Review weekly during parallel run: raise thresholds, add dependencies, or demote severity. Target steady state is a page a week, not a page a day.