Monitoring Architecture Design — Zabbix 7 + NetBox + Grafana¶
Date: 2026-07-08 Status: Draft for review Owner: Lewis
1. Goal and drivers¶
Replace the current split monitoring estate (Zabbix for servers/network + LibreNMS for core network) with a single, NetBox-driven Zabbix platform that covers the whole infrastructure with no gaps and produces clean, low-noise alerting.
Primary driver: coverage gaps — Ceph, Proxmox, tills, and UniFi are under-monitored or unmonitored today. Secondary drivers: one inventory, one alert pipeline, one pane of glass, low upkeep for a single operator.
2. Estate¶
| Location | Kit |
|---|---|
| Colo facility | ~60 servers: Proxmox hosts with Ceph, physical SQL servers, OPNsense firewalls, plus all VMs |
| Metro hubs | ~20 core network devices |
| 40 venues | Each: 1 MikroTik router, 1 UniFi switch, ~5 UniFi APs, ~5 Windows till PCs |
Total ~600–700 monitored endpoints. All venues are on always-on VPN/WAN; every device is directly routable from the colo.
Operator: one person. Reporting requirement: alerting-first with a NOC overview; formal SLA reports are secondary and can come later from the same data.
3. Decision summary¶
| Decision | Choice | Rejected alternatives |
|---|---|---|
| Monitoring platform | Zabbix 7.0 LTS, single central server at colo | Full Prometheus/VictoriaMetrics rebuild (metrics-first, alerting/SNMP/inventory all DIY — too much ongoing tax for one person); hybrid Zabbix+VM (recreates today's split-brain) |
| LibreNMS | Retired after parallel run | Keeping it alongside (two alert pipelines) |
| Database | PostgreSQL 16 with TimescaleDB extension (natively supported by Zabbix) | Plain Postgres (housekeeper pain at scale over time); separate TSDB product (unnecessary moving part) |
| Source of truth | NetBox; hosts sync to Zabbix via netbox-zabbix-sync |
Hand-managing hosts in Zabbix UI |
| Dashboards | Grafana with Zabbix datasource plugin | Zabbix native dashboards only (weaker NOC/drilldown experience) |
| Zabbix proxies | None — central polling; tills use active agents | Per-site proxies (no venue hardware to run them, and full routability makes them unnecessary at this scale) |
4. Architecture¶
NetBox (source of truth: sites, devices, VMs, IPs, roles)
│ netbox-zabbix-sync (webhook-triggered + hourly reconcile)
▼
Zabbix 7.0 LTS ── PostgreSQL 16 + TimescaleDB
│ (single server at colo; polls everything over the VPN/WAN)
▼
Grafana (Zabbix datasource) — NOC overview + per-venue drilldown
Placement¶
All monitoring services run as VMs on Proxmox at the colo:
- Zabbix server VM
- PostgreSQL 16 + TimescaleDB VM (separate from the Zabbix server — the DB is what hurts at upgrade/restore time)
- Grafana VM (or co-hosted with Zabbix frontend)
- NetBox (already deployed)
Because monitoring lives inside the infrastructure it monitors, two safeguards:
- Dead man's switch — Zabbix server pings an external service (e.g. Healthchecks.io) every few minutes; the external service alerts if pings stop. This catches "monitoring itself is down".
- VPN tunnels are first-class monitored objects — tunnel state and throughput monitored from the OPNsense/MikroTik side, with trigger dependencies hanging off them.
5. Coverage plan¶
| Layer | Method | Template |
|---|---|---|
| Proxmox cluster | API token, HTTP agentless checks | Proxmox VE by HTTP — auto-discovers nodes and VMs (LLD) |
| Ceph | Zabbix agent 2 on monitor node(s) | Ceph by Zabbix agent 2 — health, OSDs, pools, capacity |
| VMs (guest level) | Zabbix agent 2, active mode | Linux by Zabbix agent / Windows by Zabbix agent |
| Physical SQL | Agent 2 + native SQL template | MSSQL by ODBC or PostgreSQL by agent 2 — open question: flavour? |
| OPNsense | os-zabbixagent plugin + SNMP |
Agent template + community OPNsense template |
| MikroTik (venues + metro) | SNMPv3 | Official MikroTik by SNMP (per-model auto-discovery) |
| Metro/core kit | SNMPv3 | Vendor SNMP templates — this is what replaces LibreNMS |
| UniFi switches/APs | SNMP on devices + controller API | SNMP templates + HTTP checks against UniFi controller — open question: controller self-hosted or cloud? |
| Tills | Zabbix agent 2 (Windows), active mode | Trimmed Windows by Zabbix agent: reachability, disk, CPU/mem, reboot detection |
Tills use active agents: they phone home over the VPN, polling load is nil, and deployment is one scripted MSI install parameterised with server + hostname (hostname must match the NetBox device name).
6. Host structure and alert hygiene¶
- Host groups are generated from NetBox sites:
Colo,Metro/<hub>,Venue/<name>. - Tags are generated from NetBox device roles:
role:till,role:ap,role:router, etc. - Templates are assigned by the sync based on NetBox device role + platform.
The alert-quality keystone: trigger dependencies mirror physical topology.
A venue VPN drop pages once ("Venue X unreachable"), not 12 times. A metro hub failure pages once, not 12 venues × 12 devices. Escalation chains and maintenance windows configured in Zabbix actions.
7. Database and retention¶
- PostgreSQL 16 with TimescaleDB extension (Zabbix-supported combination).
- History: 31 days raw. Trends: 2 years.
- Native TimescaleDB compression on chunks older than 7 days (~90% reduction).
- Retention enforcement via chunk drops (instant) instead of housekeeper DELETEs.
At this scale (~300–500 new values/sec) the DB is not a bottleneck; TimescaleDB is chosen for operational hygiene (compression + retention), not raw performance need.
8. NetBox as source of truth¶
netbox-zabbix-sync creates/updates/disables Zabbix hosts from NetBox. Operating
rule: nobody creates hosts in the Zabbix UI. A device that is not in NetBox does
not exist. NetBox data-model requirements are specified in
01-netbox-data-model.md — in brief: every monitored device needs a site,
a device role from the canonical list, a platform, a unique name following the naming
convention, a primary IPv4, and status discipline (active = monitored).
Sync runs hourly by cron for reconciliation, plus NetBox webhooks for near-instant adds/changes.
9. Grafana¶
Zabbix datasource plugin. Initial dashboards:
- NOC overview — all-venue status grid, colo health (Proxmox/Ceph/SQL), metro links.
- Per-venue drilldown — driven by a site variable; router, switch, APs, tills.
SLA/uptime reporting can be added later from the same data; venue membership is authoritative NetBox data, not naming convention.
10. Migration plan (parallel run, no big bang)¶
- Build new Zabbix 7 + TimescaleDB + Grafana at colo.
- Finish populating NetBox to the data-model spec; stand up
netbox-zabbix-sync. - Onboard colo first (Proxmox, Ceph, SQL, OPNsense — the gap list), then metro SNMP, then venues (scripted till agent rollout, one pilot venue first).
- Parallel-run against old Zabbix + LibreNMS for 2–4 weeks; new system alerts to operator only.
- Cut alert routing over; retire old Zabbix and LibreNMS.
11. Open questions¶
- SQL flavour on the physical database servers (MSSQL vs PostgreSQL vs MySQL) — determines template and agent plugin.
- UniFi controller — self-hosted or Ubiquiti cloud? Determines API check approach.
- Whether Proxmox VMs should be auto-populated into NetBox (e.g. via a netbox-proxmox sync) or maintained by hand — hypervisor-level VM monitoring works either way via the Proxmox template's LLD; this only affects guest-agent hosts.
12. How-to guides¶
| Guide | Covers |
|---|---|
01-netbox-data-model.md |
NetBox changes needed to drive Zabbix cleanly |
02-zabbix-platform-build.md |
Zabbix 7 + TimescaleDB + Grafana build |
03-netbox-zabbix-sync.md |
Installing and configuring the sync |
04-coverage-rollout.md |
Per-layer onboarding (Proxmox, Ceph, SQL, OPNsense, MikroTik, UniFi, tills) |
05-alerting-design.md |
Dependencies, escalations, maintenance windows |
06-migration-cutover.md |
Parallel run and decommissioning |