Skip to content

Terraform / Terragrunt Adjustment Plan

Version: 1.0 — 2026-07-06 How ../terraform-mikrotik (Terragrunt, per-device) and ../terraform-modules (mikrotik/core, mikrotik/venue) evolve to deliver DESIGN.md / CONFIG-GUIDE.md.


1. What the current modules encode (and what must change)

Current code Where Target
iBGP full mesh — every device lists every peer in bgp_peers core/bgp.tf + every terragrunt.hcl RR model: clients declare nothing but the two RR loopbacks (a _common constant); RRs generate client connections from the shared registry
redistribute = "connected,static" on template and connections core/bgp.tf Removed. Explicit origination via output.network address-list only
OSPF redistribute = ["connected"] core/ospf.tf Removed. Interface-templates only (loopback passive + p2p infra links)
Per-venue static routes (venue_interfaces[].routes, check_gateway=ping) core/static_routes.tf Deleted entirely — venue /16s arrive via per-venue eBGP
Venue: static default(s) to WAN gateway, no routing protocol venue/routes.tf eBGP CE: blackhole anchor + 1–2 eBGP sessions w/ BFD, default learned
Venue VLAN 20 (lan_public) routed locally + optional local hotspot venue/interfaces.tf, hotspot.tf VLAN 20 becomes L2-only, bridged to a tagged uplink; hotspot module retained but default-off (fallback only)
Defaults pointing at 10.1.x services (dns_servers, ad_servers, RADIUS) both variables.tf Single source in shared data; flip to 10.128.x at Phase 7 by changing one file
OSPF on every device incl. venues both modules OSPF/LDP only in the transport module; venue module has neither

Also: only ~5 core + ~16 venue devices are under Terraform. The design assumes the whole fleet is — un-managed routers (e.g. CR-MATHEWSTREET-01) need importing as part of Phase 0.

2. Module restructure — composable roles instead of two monoliths

Split terraform-modules/modules/mikrotik/ by function, compose per device class. Reuse existing code where it's already right (identity, snmp, ntp, dns, services, input firewall, venue VLAN/DHCP/firewall files move mostly unchanged).

modules/mikrotik/
  baseline/        identity, users/aaa(radius), services, dns, ntp, snmp, syslog,
                   romon, input-firewall skeleton, address-lists (from shared data)
  transport/       OSPF (no redistribute) + BFD + LDP + iBGP client (ip+l2vpn) to RRs
  rr/              adds RR duty to a transport device: client connections generated
                   from registry, default-originate, eBGP to edge appliances
                   (OPNsense/Starlink/guest-FW) with routing-filter chains
  venue-handoff/   used BY transport devices (metro PEs): per-venue /30 address,
                   eBGP session + from-venue filter, guest VLAN on the venue port,
                   BGP-VPLS spoke stanza     ← replaces venue_interfaces+static routes
  venue/           venue router: VLANs/DHCP/zone-firewall (existing files) +
                   eBGP CE (new) + guest VLAN bridged to tagged uplink; no OSPF/LDP
  vpls/            BGP-VPLS instance definitions (spoke + hub variants) — isolated
                   because of provider coverage risk (§5). Covers guest AND the ISP
                   POP circuits (L2 carriage for the separate WAN business)

No bng module here. The retail-ISP BNG/borders are a separate business (AS 204258, wan/WAN-DESIGN.md) with its own Terraform scope if automated. The metro only carries its POP circuits as L2 VPLS (the vpls module handles that, handing off to tagged ports toward the WAN borders — not to any metro device).

Device composition:

Device class Modules
Colo CCR baseline + transport + rr + venue-handoff (colo-fed venues) + vpls (hub)
Metro/roof CCR baseline + transport + venue-handoff + vpls (spokes)
Venue router baseline + venue
Starlink RB4011 baseline + thin bespoke module (eBGP filters + payment-prefixes/starlink-announce lists — must stay in lockstep with the CCRs' from-starlink contract; same registry)

Scope boundary — what is deliberately NOT (fully) in Terraform:

Kit Treatment Why
Ceph/VM CRS326 pairs baseline module only; L2 fabric (MLAG/bridges/bonds/VLANs) = versioned .rsc at build, then hands-off Change rate ≈ 0 after build; blast radius = the whole storage fabric; provider modify-as-destroy is scariest on bridge-port resources. Oxidized export + git keeps "config is code" without plan/apply risk. Promote to a switch module only if VLAN churn appears
OOB switch + OOB gateway Fully manual — versioned .rsc + nightly export Never automate the rescue path with the thing it rescues — OOB is how you fix a bad apply, so it must sit outside Terraform's blast domain
OPNsense pair Not RouterOS — its own config discipline (config.xml backups; Ansible later if wanted) Different provider entirely

3. Shared data — kill the hand-typed peer lists and /30s

The biggest current fragility: both ends of each /30 and every mesh peer are typed by hand per device. Introduce one registry consumed everywhere, e.g. terraform-mikrotik/_common/network.yaml:

rr_loopbacks: ["10.255.255.255", "10.255.255.254"]
bgp_core_asn: 65500
ospf_costs: { fibre10g: 10, wired1g: 100, ghz60: 500, lastresort: 5000 }

services:            # flip these at Phase 7, every device follows
  dns:    ["10.1.84.3", "10.1.6.1"]        # -> 10.128.32.x
  radius: ["10.1.88.11"]                   # -> 10.128.36.11
  syslog: "10.1.88.20"                     # -> 10.128.36.20
  management: ["10.201.201.0/24"]

venues:
  rocking-horse:
    venue_id: 33                            # => 10.33.0.0/16, AS 64545
    uplinks:
      - { hub: cr-charlottestreet-002, port: sfp-sfpplus2,
          p2p: 10.254.31.0/30, class: fibre10g, primary: true }
    guest_vpls: true

_common/mikrotik.hcl loads it (yamldecode(file(...))) and exposes derived values. Then:

  • Venue terragrunt.hcl shrinks to: name/venue-id/credentials + which uplink ports face the LAN — the /30, AS, hub port, and eBGP peers are all derived from the registry.
  • Metro PE terragrunt.hcl derives its venue-handoff list from the same records (filtered by hub == this device) — the two ends of a /30 can no longer disagree.
  • RR client connections on the colo CCRs are generated from the registry's transport-device list; adding a metro router = one registry entry.
  • AS numbering is a function, not data: 64512 + venue_id.

4. Key new resources (sketches)

transport/bgp.tf — RR client, both address families, no redistribute:

resource "routeros_routing_bgp_template" "core" {
  name      = "pubinvest-core"
  as        = var.bgp_core_asn
  router_id = var.loopback_address
  local  { address = var.loopback_address, role = "ibgp" }
  output { network = "bgp-networks" }          # explicit origination only
  # address-families = "ip,l2vpn"              # verify provider attr (§5)
}

resource "routeros_routing_bgp_connection" "rr" {
  for_each = toset(var.rr_loopbacks)
  name      = "RR-${each.value}"
  templates = [routeros_routing_bgp_template.core.name]
  remote { address = each.value, as = var.bgp_core_asn }
}

venue/bgp.tf — CE side:

resource "routeros_ip_route" "aggregate_anchor" {
  dst_address = "10.${var.venue_id}.0.0/16"
  blackhole   = true
}

resource "routeros_routing_bgp_connection" "uplink" {
  for_each = { for u in var.uplinks : u.p2p => u }
  name = "uplink-${each.value.primary ? "primary" : "backup"}"
  as   = 64512 + var.venue_id
  local  { role = "ebgp", address = cidrhost(each.value.p2p, 2) }
  remote { address = cidrhost(each.value.p2p, 1), as = var.bgp_core_asn }
  output { network = "bgp-networks" }
  input  { filter = "default-only" }
  use_bfd = true                                # verify provider attr
}

venue-handoff (metro PE side) — per venue: /30 address on port, eBGP connection with input.filter = "from-venue-${id}", a routeros_routing_filter_rule chain accepting only 10.V.0.0/16 (+ community 65500:100, local-pref 50 on backup hubs), guest VLAN interface on the venue port, VPLS spoke (§5).

Deletions: core/static_routes.tf (all three resource blocks), venue/routes.tf statics, OSPF redistribute, BGP redistribute, no_client_to_client_reflection (RRs must reflect).

5. Provider coverage — verify before building (risk register)

Built on terraform-routeros. Known-good from existing code: OSPF, BGP template/connection, routes, firewall, VLANs, DHCP. Verify against the provider version you pin (check registry docs, then a lab CHR):

Feature Needed by Risk
routeros_routing_filter_rule (ROS7 filter syntax) venue filters, edge/Starlink policy Medium — exists in recent versions; confirm rule-string handling
BGP address-families incl. l2vpn on template/connection VPLS signalling Medium
MPLS LDP (/mpls ldp instance|interface) resources transport module Medium — MPLS support in provider is newer
BGP VPLS (/routing bgp vpls) vpls module High — likely missing
BFD configuration resource all fast failover Medium

Mitigation for gaps: the vpls/ (and if needed BFD/LDP) modules wrap a routeros_system_script + one-shot scheduler that idempotently applies the CLI stanza (script content templated from the same variables), so the interface to the rest of the code is already final; swap internals to native resources when the provider catches up (and consider contributing the resource upstream — the provider accepts PRs readily). Do not let a provider gap push VPLS config into unmanaged/manual state.

6. State & cutover strategy (maps to CONFIG-GUIDE §9 phases)

Renaming/splitting modules re-addresses resources — on this provider, a destroy is a live config deletion on a production router. Rules:

  1. Never apply a restructure blind: moved {} blocks (or terragrunt state mv) for every resource that survives the split; CI must show a plan with only expected changes before any apply.
  2. Coexistence is the mechanism, matching the network migration: Phase 3 runs old mesh + new RR sessions simultaneously — in TF terms, add the transport module alongside the legacy bgp_peers input, and only after table comparison remove the mesh entries from the device's terragrunt inputs (a pure-delete plan you can read and approve).
  3. Per-device cutover = per-directory terragrunt apply, in the phase order: transport routers first (Phase 2–4), venues one at a time (Phase 3), VPLS spokes — guest then ISP POP circuits (Phase 5), service-IP flip in network.yaml (Phase 7).
  4. Import the unmanaged fleet first (Phase 0): create terragrunt dirs for Mathew St, Fenwick, Temple Court, Holmes, remaining venues; terraform import their existing resources so plans start from truth, not from scratch. The example-configs exports are the checklist of what to import.
  5. Pin provider + module versions per device dir (already the pattern via source.hcl); upgrade the provider once, in a lab, when adopting the MPLS/VPLS resources.

7. CI additions (GitLab)

  • MR pipeline: terragrunt plan for every changed device dir (+ every dir when _common/network.yaml changes — that file fans out).
  • Plans posted to MR; apply = manual job per directory, in phase order.
  • Nightly drift job: terragrunt plan -detailed-exitcode across the fleet; non-empty plan without an open MR ⇒ alert (someone changed a router by hand — the design's "config is code" rule, enforced).
  • Keep RouterOS /export backups (Oxidized) as the independent second record; Terraform state is not a backup.

8. Suggested build order

  1. Registry (network.yaml) + baseline module; adopt on the 5 already-managed core devices (no behaviour change, pure refactor with moved blocks).
  2. Import remaining fleet onto baseline (+ legacy inputs as-is).
  3. transport + rr modules; Phase 2–3 rollout; delete mesh/static inputs per device.
  4. venue eBGP + venue-handoff; convert venues one at a time.
  5. vpls module (after provider verification/lab); Phase 5 guest pilot, then ISP POP circuits (L2 carriage for the separate WAN business).
  6. Flip services: block to 10.128.x; Phase 7.