Terraform / Terragrunt Adjustment Plan¶
Version: 1.0 — 2026-07-06
How ../terraform-mikrotik (Terragrunt, per-device) and ../terraform-modules (mikrotik/core, mikrotik/venue) evolve to deliver DESIGN.md / CONFIG-GUIDE.md.
1. What the current modules encode (and what must change)¶
| Current code | Where | Target |
|---|---|---|
iBGP full mesh — every device lists every peer in bgp_peers |
core/bgp.tf + every terragrunt.hcl |
RR model: clients declare nothing but the two RR loopbacks (a _common constant); RRs generate client connections from the shared registry |
redistribute = "connected,static" on template and connections |
core/bgp.tf |
Removed. Explicit origination via output.network address-list only |
OSPF redistribute = ["connected"] |
core/ospf.tf |
Removed. Interface-templates only (loopback passive + p2p infra links) |
Per-venue static routes (venue_interfaces[].routes, check_gateway=ping) |
core/static_routes.tf |
Deleted entirely — venue /16s arrive via per-venue eBGP |
| Venue: static default(s) to WAN gateway, no routing protocol | venue/routes.tf |
eBGP CE: blackhole anchor + 1–2 eBGP sessions w/ BFD, default learned |
Venue VLAN 20 (lan_public) routed locally + optional local hotspot |
venue/interfaces.tf, hotspot.tf |
VLAN 20 becomes L2-only, bridged to a tagged uplink; hotspot module retained but default-off (fallback only) |
Defaults pointing at 10.1.x services (dns_servers, ad_servers, RADIUS) |
both variables.tf |
Single source in shared data; flip to 10.128.x at Phase 7 by changing one file |
| OSPF on every device incl. venues | both modules | OSPF/LDP only in the transport module; venue module has neither |
Also: only ~5 core + ~16 venue devices are under Terraform. The design assumes the whole fleet is — un-managed routers (e.g. CR-MATHEWSTREET-01) need importing as part of Phase 0.
2. Module restructure — composable roles instead of two monoliths¶
Split terraform-modules/modules/mikrotik/ by function, compose per device class. Reuse existing code where it's already right (identity, snmp, ntp, dns, services, input firewall, venue VLAN/DHCP/firewall files move mostly unchanged).
modules/mikrotik/
baseline/ identity, users/aaa(radius), services, dns, ntp, snmp, syslog,
romon, input-firewall skeleton, address-lists (from shared data)
transport/ OSPF (no redistribute) + BFD + LDP + iBGP client (ip+l2vpn) to RRs
rr/ adds RR duty to a transport device: client connections generated
from registry, default-originate, eBGP to edge appliances
(OPNsense/Starlink/guest-FW) with routing-filter chains
venue-handoff/ used BY transport devices (metro PEs): per-venue /30 address,
eBGP session + from-venue filter, guest VLAN on the venue port,
BGP-VPLS spoke stanza ← replaces venue_interfaces+static routes
venue/ venue router: VLANs/DHCP/zone-firewall (existing files) +
eBGP CE (new) + guest VLAN bridged to tagged uplink; no OSPF/LDP
vpls/ BGP-VPLS instance definitions (spoke + hub variants) — isolated
because of provider coverage risk (§5). Covers guest AND the ISP
POP circuits (L2 carriage for the separate WAN business)
No bng module here. The retail-ISP BNG/borders are a separate business (AS 204258, wan/WAN-DESIGN.md) with its own Terraform scope if automated. The metro only carries its POP circuits as L2 VPLS (the vpls module handles that, handing off to tagged ports toward the WAN borders — not to any metro device).
Device composition:
| Device class | Modules |
|---|---|
| Colo CCR | baseline + transport + rr + venue-handoff (colo-fed venues) + vpls (hub) |
| Metro/roof CCR | baseline + transport + venue-handoff + vpls (spokes) |
| Venue router | baseline + venue |
| Starlink RB4011 | baseline + thin bespoke module (eBGP filters + payment-prefixes/starlink-announce lists — must stay in lockstep with the CCRs' from-starlink contract; same registry) |
Scope boundary — what is deliberately NOT (fully) in Terraform:
| Kit | Treatment | Why |
|---|---|---|
| Ceph/VM CRS326 pairs | baseline module only; L2 fabric (MLAG/bridges/bonds/VLANs) = versioned .rsc at build, then hands-off |
Change rate ≈ 0 after build; blast radius = the whole storage fabric; provider modify-as-destroy is scariest on bridge-port resources. Oxidized export + git keeps "config is code" without plan/apply risk. Promote to a switch module only if VLAN churn appears |
| OOB switch + OOB gateway | Fully manual — versioned .rsc + nightly export |
Never automate the rescue path with the thing it rescues — OOB is how you fix a bad apply, so it must sit outside Terraform's blast domain |
| OPNsense pair | Not RouterOS — its own config discipline (config.xml backups; Ansible later if wanted) | Different provider entirely |
3. Shared data — kill the hand-typed peer lists and /30s¶
The biggest current fragility: both ends of each /30 and every mesh peer are typed by hand per device. Introduce one registry consumed everywhere, e.g. terraform-mikrotik/_common/network.yaml:
rr_loopbacks: ["10.255.255.255", "10.255.255.254"]
bgp_core_asn: 65500
ospf_costs: { fibre10g: 10, wired1g: 100, ghz60: 500, lastresort: 5000 }
services: # flip these at Phase 7, every device follows
dns: ["10.1.84.3", "10.1.6.1"] # -> 10.128.32.x
radius: ["10.1.88.11"] # -> 10.128.36.11
syslog: "10.1.88.20" # -> 10.128.36.20
management: ["10.201.201.0/24"]
venues:
rocking-horse:
venue_id: 33 # => 10.33.0.0/16, AS 64545
uplinks:
- { hub: cr-charlottestreet-002, port: sfp-sfpplus2,
p2p: 10.254.31.0/30, class: fibre10g, primary: true }
guest_vpls: true
_common/mikrotik.hcl loads it (yamldecode(file(...))) and exposes derived values. Then:
- Venue terragrunt.hcl shrinks to: name/venue-id/credentials + which uplink ports face the LAN — the /30, AS, hub port, and eBGP peers are all derived from the registry.
- Metro PE terragrunt.hcl derives its
venue-handofflist from the same records (filtered byhub == this device) — the two ends of a /30 can no longer disagree. - RR client connections on the colo CCRs are generated from the registry's transport-device list; adding a metro router = one registry entry.
- AS numbering is a function, not data:
64512 + venue_id.
4. Key new resources (sketches)¶
transport/bgp.tf — RR client, both address families, no redistribute:
resource "routeros_routing_bgp_template" "core" {
name = "pubinvest-core"
as = var.bgp_core_asn
router_id = var.loopback_address
local { address = var.loopback_address, role = "ibgp" }
output { network = "bgp-networks" } # explicit origination only
# address-families = "ip,l2vpn" # verify provider attr (§5)
}
resource "routeros_routing_bgp_connection" "rr" {
for_each = toset(var.rr_loopbacks)
name = "RR-${each.value}"
templates = [routeros_routing_bgp_template.core.name]
remote { address = each.value, as = var.bgp_core_asn }
}
venue/bgp.tf — CE side:
resource "routeros_ip_route" "aggregate_anchor" {
dst_address = "10.${var.venue_id}.0.0/16"
blackhole = true
}
resource "routeros_routing_bgp_connection" "uplink" {
for_each = { for u in var.uplinks : u.p2p => u }
name = "uplink-${each.value.primary ? "primary" : "backup"}"
as = 64512 + var.venue_id
local { role = "ebgp", address = cidrhost(each.value.p2p, 2) }
remote { address = cidrhost(each.value.p2p, 1), as = var.bgp_core_asn }
output { network = "bgp-networks" }
input { filter = "default-only" }
use_bfd = true # verify provider attr
}
venue-handoff (metro PE side) — per venue: /30 address on port, eBGP connection with input.filter = "from-venue-${id}", a routeros_routing_filter_rule chain accepting only 10.V.0.0/16 (+ community 65500:100, local-pref 50 on backup hubs), guest VLAN interface on the venue port, VPLS spoke (§5).
Deletions: core/static_routes.tf (all three resource blocks), venue/routes.tf statics, OSPF redistribute, BGP redistribute, no_client_to_client_reflection (RRs must reflect).
5. Provider coverage — verify before building (risk register)¶
Built on terraform-routeros. Known-good from existing code: OSPF, BGP template/connection, routes, firewall, VLANs, DHCP. Verify against the provider version you pin (check registry docs, then a lab CHR):
| Feature | Needed by | Risk |
|---|---|---|
routeros_routing_filter_rule (ROS7 filter syntax) |
venue filters, edge/Starlink policy | Medium — exists in recent versions; confirm rule-string handling |
BGP address-families incl. l2vpn on template/connection |
VPLS signalling | Medium |
MPLS LDP (/mpls ldp instance|interface) resources |
transport module | Medium — MPLS support in provider is newer |
BGP VPLS (/routing bgp vpls) |
vpls module | High — likely missing |
| BFD configuration resource | all fast failover | Medium |
Mitigation for gaps: the vpls/ (and if needed BFD/LDP) modules wrap a routeros_system_script + one-shot scheduler that idempotently applies the CLI stanza (script content templated from the same variables), so the interface to the rest of the code is already final; swap internals to native resources when the provider catches up (and consider contributing the resource upstream — the provider accepts PRs readily). Do not let a provider gap push VPLS config into unmanaged/manual state.
6. State & cutover strategy (maps to CONFIG-GUIDE §9 phases)¶
Renaming/splitting modules re-addresses resources — on this provider, a destroy is a live config deletion on a production router. Rules:
- Never
applya restructure blind:moved {}blocks (orterragrunt state mv) for every resource that survives the split; CI must show a plan with only expected changes before any apply. - Coexistence is the mechanism, matching the network migration: Phase 3 runs old mesh + new RR sessions simultaneously — in TF terms, add the
transportmodule alongside the legacybgp_peersinput, and only after table comparison remove the mesh entries from the device's terragrunt inputs (a pure-delete plan you can read and approve). - Per-device cutover = per-directory
terragrunt apply, in the phase order: transport routers first (Phase 2–4), venues one at a time (Phase 3), VPLS spokes — guest then ISP POP circuits (Phase 5), service-IP flip innetwork.yaml(Phase 7). - Import the unmanaged fleet first (Phase 0): create terragrunt dirs for Mathew St, Fenwick, Temple Court, Holmes, remaining venues;
terraform importtheir existing resources so plans start from truth, not from scratch. The example-configs exports are the checklist of what to import. - Pin provider + module versions per device dir (already the pattern via
source.hcl); upgrade the provider once, in a lab, when adopting the MPLS/VPLS resources.
7. CI additions (GitLab)¶
- MR pipeline:
terragrunt planfor every changed device dir (+ every dir when_common/network.yamlchanges — that file fans out). - Plans posted to MR; apply = manual job per directory, in phase order.
- Nightly drift job:
terragrunt plan -detailed-exitcodeacross the fleet; non-empty plan without an open MR ⇒ alert (someone changed a router by hand — the design's "config is code" rule, enforced). - Keep RouterOS
/exportbackups (Oxidized) as the independent second record; Terraform state is not a backup.
8. Suggested build order¶
- Registry (
network.yaml) +baselinemodule; adopt on the 5 already-managed core devices (no behaviour change, pure refactor withmovedblocks). - Import remaining fleet onto
baseline(+ legacy inputs as-is). transport+rrmodules; Phase 2–3 rollout; delete mesh/static inputs per device.venueeBGP +venue-handoff; convert venues one at a time.vplsmodule (after provider verification/lab); Phase 5 guest pilot, then ISP POP circuits (L2 carriage for the separate WAN business).- Flip
services:block to 10.128.x; Phase 7.