Skip to content

Monitoring Architecture Design — Zabbix 7 + NetBox + Grafana

Date: 2026-07-08 Status: Draft for review Owner: Lewis

1. Goal and drivers

Replace the current split monitoring estate (Zabbix for servers/network + LibreNMS for core network) with a single, NetBox-driven Zabbix platform that covers the whole infrastructure with no gaps and produces clean, low-noise alerting.

Primary driver: coverage gaps — Ceph, Proxmox, tills, and UniFi are under-monitored or unmonitored today. Secondary drivers: one inventory, one alert pipeline, one pane of glass, low upkeep for a single operator.

2. Estate

Location Kit
Colo facility ~60 servers: Proxmox hosts with Ceph, physical SQL servers, OPNsense firewalls, plus all VMs
Metro hubs ~20 core network devices
40 venues Each: 1 MikroTik router, 1 UniFi switch, ~5 UniFi APs, ~5 Windows till PCs

Total ~600–700 monitored endpoints. All venues are on always-on VPN/WAN; every device is directly routable from the colo.

Operator: one person. Reporting requirement: alerting-first with a NOC overview; formal SLA reports are secondary and can come later from the same data.

3. Decision summary

Decision Choice Rejected alternatives
Monitoring platform Zabbix 7.0 LTS, single central server at colo Full Prometheus/VictoriaMetrics rebuild (metrics-first, alerting/SNMP/inventory all DIY — too much ongoing tax for one person); hybrid Zabbix+VM (recreates today's split-brain)
LibreNMS Retired after parallel run Keeping it alongside (two alert pipelines)
Database PostgreSQL 16 with TimescaleDB extension (natively supported by Zabbix) Plain Postgres (housekeeper pain at scale over time); separate TSDB product (unnecessary moving part)
Source of truth NetBox; hosts sync to Zabbix via netbox-zabbix-sync Hand-managing hosts in Zabbix UI
Dashboards Grafana with Zabbix datasource plugin Zabbix native dashboards only (weaker NOC/drilldown experience)
Zabbix proxies None — central polling; tills use active agents Per-site proxies (no venue hardware to run them, and full routability makes them unnecessary at this scale)

4. Architecture

NetBox (source of truth: sites, devices, VMs, IPs, roles)
   │  netbox-zabbix-sync (webhook-triggered + hourly reconcile)
   ▼
Zabbix 7.0 LTS ── PostgreSQL 16 + TimescaleDB
   │  (single server at colo; polls everything over the VPN/WAN)
   ▼
Grafana (Zabbix datasource) — NOC overview + per-venue drilldown

Placement

All monitoring services run as VMs on Proxmox at the colo:

  • Zabbix server VM
  • PostgreSQL 16 + TimescaleDB VM (separate from the Zabbix server — the DB is what hurts at upgrade/restore time)
  • Grafana VM (or co-hosted with Zabbix frontend)
  • NetBox (already deployed)

Because monitoring lives inside the infrastructure it monitors, two safeguards:

  1. Dead man's switch — Zabbix server pings an external service (e.g. Healthchecks.io) every few minutes; the external service alerts if pings stop. This catches "monitoring itself is down".
  2. VPN tunnels are first-class monitored objects — tunnel state and throughput monitored from the OPNsense/MikroTik side, with trigger dependencies hanging off them.

5. Coverage plan

Layer Method Template
Proxmox cluster API token, HTTP agentless checks Proxmox VE by HTTP — auto-discovers nodes and VMs (LLD)
Ceph Zabbix agent 2 on monitor node(s) Ceph by Zabbix agent 2 — health, OSDs, pools, capacity
VMs (guest level) Zabbix agent 2, active mode Linux by Zabbix agent / Windows by Zabbix agent
Physical SQL Agent 2 + native SQL template MSSQL by ODBC or PostgreSQL by agent 2 — open question: flavour?
OPNsense os-zabbixagent plugin + SNMP Agent template + community OPNsense template
MikroTik (venues + metro) SNMPv3 Official MikroTik by SNMP (per-model auto-discovery)
Metro/core kit SNMPv3 Vendor SNMP templates — this is what replaces LibreNMS
UniFi switches/APs SNMP on devices + controller API SNMP templates + HTTP checks against UniFi controller — open question: controller self-hosted or cloud?
Tills Zabbix agent 2 (Windows), active mode Trimmed Windows by Zabbix agent: reachability, disk, CPU/mem, reboot detection

Tills use active agents: they phone home over the VPN, polling load is nil, and deployment is one scripted MSI install parameterised with server + hostname (hostname must match the NetBox device name).

6. Host structure and alert hygiene

  • Host groups are generated from NetBox sites: Colo, Metro/<hub>, Venue/<name>.
  • Tags are generated from NetBox device roles: role:till, role:ap, role:router, etc.
  • Templates are assigned by the sync based on NetBox device role + platform.

The alert-quality keystone: trigger dependencies mirror physical topology.

metro hub uplink
   └── venue MikroTik ICMP
         ├── venue UniFi switch
         │     └── venue APs
         └── venue tills

A venue VPN drop pages once ("Venue X unreachable"), not 12 times. A metro hub failure pages once, not 12 venues × 12 devices. Escalation chains and maintenance windows configured in Zabbix actions.

7. Database and retention

  • PostgreSQL 16 with TimescaleDB extension (Zabbix-supported combination).
  • History: 31 days raw. Trends: 2 years.
  • Native TimescaleDB compression on chunks older than 7 days (~90% reduction).
  • Retention enforcement via chunk drops (instant) instead of housekeeper DELETEs.

At this scale (~300–500 new values/sec) the DB is not a bottleneck; TimescaleDB is chosen for operational hygiene (compression + retention), not raw performance need.

8. NetBox as source of truth

netbox-zabbix-sync creates/updates/disables Zabbix hosts from NetBox. Operating rule: nobody creates hosts in the Zabbix UI. A device that is not in NetBox does not exist. NetBox data-model requirements are specified in 01-netbox-data-model.md — in brief: every monitored device needs a site, a device role from the canonical list, a platform, a unique name following the naming convention, a primary IPv4, and status discipline (active = monitored).

Sync runs hourly by cron for reconciliation, plus NetBox webhooks for near-instant adds/changes.

9. Grafana

Zabbix datasource plugin. Initial dashboards:

  1. NOC overview — all-venue status grid, colo health (Proxmox/Ceph/SQL), metro links.
  2. Per-venue drilldown — driven by a site variable; router, switch, APs, tills.

SLA/uptime reporting can be added later from the same data; venue membership is authoritative NetBox data, not naming convention.

10. Migration plan (parallel run, no big bang)

  1. Build new Zabbix 7 + TimescaleDB + Grafana at colo.
  2. Finish populating NetBox to the data-model spec; stand up netbox-zabbix-sync.
  3. Onboard colo first (Proxmox, Ceph, SQL, OPNsense — the gap list), then metro SNMP, then venues (scripted till agent rollout, one pilot venue first).
  4. Parallel-run against old Zabbix + LibreNMS for 2–4 weeks; new system alerts to operator only.
  5. Cut alert routing over; retire old Zabbix and LibreNMS.

11. Open questions

  1. SQL flavour on the physical database servers (MSSQL vs PostgreSQL vs MySQL) — determines template and agent plugin.
  2. UniFi controller — self-hosted or Ubiquiti cloud? Determines API check approach.
  3. Whether Proxmox VMs should be auto-populated into NetBox (e.g. via a netbox-proxmox sync) or maintained by hand — hypervisor-level VM monitoring works either way via the Proxmox template's LLD; this only affects guest-agent hosts.

12. How-to guides

Guide Covers
01-netbox-data-model.md NetBox changes needed to drive Zabbix cleanly
02-zabbix-platform-build.md Zabbix 7 + TimescaleDB + Grafana build
03-netbox-zabbix-sync.md Installing and configuring the sync
04-coverage-rollout.md Per-layer onboarding (Proxmox, Ceph, SQL, OPNsense, MikroTik, UniFi, tills)
05-alerting-design.md Dependencies, escalations, maintenance windows
06-migration-cutover.md Parallel run and decommissioning