Guide 02 — Zabbix 7 Platform Build (Zabbix + TimescaleDB + Grafana)¶
High-level build guide for the central monitoring stack at the colo. Target: three VMs on Proxmox, all covered by the monitoring they host, plus an external dead man's switch.
1. VM layout¶
| VM | Sizing (start) | Runs |
|---|---|---|
colo-mon-db-01 |
4 vCPU / 8 GB / 200 GB (SSD-backed Ceph) | PostgreSQL 16 + TimescaleDB |
colo-mon-zbx-01 |
4 vCPU / 8 GB / 50 GB | Zabbix server 7.0 LTS + frontend |
colo-mon-gfx-01 |
2 vCPU / 4 GB / 20 GB | Grafana |
At ~700 hosts / 300–500 values per second this is generous; grow later if needed. NetBox stays on its existing VM.
2. Database build¶
- Install PostgreSQL 16 + the TimescaleDB 2.x package (use the Zabbix docs' supported-version matrix for the exact pairing).
CREATE EXTENSION timescaledb;in thezabbixdatabase.- Load the Zabbix server schema, then run the Zabbix-provided
timescaledb.sqlscript — this converts history/trends tables to hypertables. - In Zabbix frontend (Administration → Housekeeping): enable override item history/trend period, history 31d, trends 731d, and enable native TimescaleDB compression for chunks older than 7 days.
- Tune Postgres basics:
shared_buffers~25% RAM,max_connectionssized to Zabbix pollers + frontend.
Backups: nightly pg_dump (config tables matter most; history is replaceable) plus
Proxmox VM backup. Test a restore once before go-live.
3. Zabbix server build¶
- Install Zabbix server 7.0 LTS + frontend (nginx) from official repos, pointed at the DB VM.
- Set the frontend behind HTTPS (internal CA or Let's Encrypt via the VPN).
- Global secret macros for shared credentials: SNMPv3 user/auth/priv, Proxmox API token, UniFi controller creds, SQL monitoring account.
- Enable and size pollers modestly (defaults are close at this scale); watch
zabbix[process,*,avg,busy]internal items after onboarding and tune once. - Configure alerting media (email/Telegram/Slack/SMS as preferred) but leave routing to operator-only during parallel run (see Guide 06).
4. Grafana¶
- Install Grafana OSS + the Alexander Zobnin Zabbix datasource plugin
(
alexanderzobnin-zabbix-app); connect via the Zabbix API with a read-only user. - Build two dashboards first:
- NOC overview — venue status grid (one cell per venue driven by the venue router trigger), colo health row (Proxmox cluster, Ceph health, SQL, OPNsense), metro link states.
- Venue drilldown — templated on site/host group; router WAN + tunnel, switch, per-AP status/clients, per-till status.
- Anonymous read-only viewer access on the NOC dashboard if it goes on a wall screen.
5. Monitoring the monitoring¶
- Dead man's switch: cron on the Zabbix server VM curls a Healthchecks.io (or similar external) URL every 5 minutes; the external service emails/pushes if pings stop. This is the only alert path that doesn't depend on the colo.
- Zabbix monitors its own DB VM, Grafana VM, and NetBox VM via agent (they're just hosts in NetBox like everything else).
- Watch Zabbix internal health items (queue, cache usage, poller busy rates) with the built-in Zabbix server health template.
6. Order of operations¶
- DB VM → schema → TimescaleDB conversion
- Zabbix server + frontend → secret macros → media types
- Dead man's switch
- Grafana + datasource
- Then proceed to Guide 03 (sync) — do not hand-create hosts in the meantime