Guide 06 — Migration and Cutover¶
Parallel-run migration from old Zabbix + LibreNMS to the new platform. No big bang; the old systems keep alerting until the new one has proven itself.
Phase 0 — Build (Guides 02, 03)¶
- New Zabbix 7 + TimescaleDB + Grafana built at colo; dead man's switch live.
- NetBox pre-flight checklist (Guide 01) passes for colo + metro at minimum.
- Sync running, restricted to colo scope.
Phase 1 — Colo (the gap list first)¶
- Onboard Proxmox, Ceph, SQL, OPNsense, colo VMs per Guide 04.
- This phase delivers the actual driver of the project (coverage gaps) even before anything migrates — new coverage, no risk to existing alerting.
- Exit: colo dashboards populated, triggers drilled, one week of clean data.
Phase 2 — Metro (replaces LibreNMS's job)¶
- NetBox population for metro kit → sync → SNMP templates.
- Compare against LibreNMS for 1–2 weeks: same devices visible, interface counters agree, alerts fire on both for any real event.
- Exit: no metro condition LibreNMS caught that new Zabbix missed.
Phase 3 — Venues (pilot, then batches)¶
- One pilot venue: NetBox pattern applied, sync, SNMP on MikroTik/UniFi, till agents installed on-site. Run the Guide 05 §4 failure drills there.
- Fix what the pilot teaches (it will teach something — usually naming or the till install).
- Roll remaining 39 venues in batches of ~5/week: NetBox script → sync → RouterOS SNMP script → till agent deploy. ~2 months at a comfortable pace, faster if till deployment is remote.
Phase 4 — Parallel run and cutover¶
- 2–4 weeks with both systems alerting; new system routes to operator only.
- Weekly noise review (Guide 05 §6). Cutover criteria:
- Every alert the old systems raised was also raised (or consciously demoted) by the new one.
- No unsupported items / dead hosts in new Zabbix.
- Dependency drills pass.
- Cut alert routing to the new system. Old systems to notify-nobody mode for a final 2 weeks (insurance), then:
Phase 5 — Decommission¶
- Final config/DB backup of old Zabbix and LibreNMS (archived, not kept running).
- VMs stopped for 2 weeks, then deleted; monitoring of those VMs removed via NetBox status change (which is the workflow working as designed).
- Update any runbooks/bookmarks pointing at old UIs.
Rollback posture¶
At every phase the old systems are still alerting, so rollback is simply "don't cut over". The only irreversible step is Phase 5, gated behind two clean weeks of new-system-only operation.