5. Routing — FRR / BGP¶
Configure on the primary only (config sync replicates FRR). This is the core of the edge: dual eBGP to the CCRs, dual eBGP to the WAN borders, and CARP-gated FRR so only the master ever speaks BGP.
5.1 Install and enable FRR¶
- System ▸ Firmware ▸ Plugins → install
os-frr. - Routing ▸ General → Enable the routing service.
- Set the Router ID to a stable address (the routing VIP
<PUBLIC_ROUTING_VIP>or a loopback).
5.2 CARP-gated failover¶
Routing ▸ General, the CARP failover option:
| Field | Value |
|---|---|
| Enable CARP Failover / CARP Demotion | on |
| CARP VIP to track | the routing VIP (VHID 50) from page 3 §3.4 |
This stops FRR whenever that VIP is in BACKUP and starts it on MASTER. The backup therefore announces nothing, the CCRs/borders can only route via the master, and routing is symmetric with no prepend tricks (§3). This is the mechanism that makes the whole dual-node design consistent — don't skip it.
Failover behaviour
BGP sessions do not survive failover (FRR's TCP state isn't synced — pfSync only syncs firewall state). On a CARP transition the new master starts FRR, re-establishes all four sessions and re-announces; expect a few seconds of convergence. BFD (§5.5) and modest timers keep it short. Data-plane flows resume seamlessly because pfSync already has the state table on the new master.
5.3 BGP general¶
Routing ▸ BGP ▸ General:
| Field | Value |
|---|---|
| Enable | on |
| AS Number | <OPN_AS> (65510) |
| Router ID | <PUBLIC_ROUTING_VIP> |
| Network (announce) | <PUBLIC_BLOCK> + the server-room zone subnets + 0.0.0.0/0 (passed through — §5b) |
Enable ECMP so the two CCR sessions load-share: set bestpath as-path multipath-relax
(BGP ▸ General advanced) and the sysctl net.route.multipath=1 (page 10).
Announce via network statements + a Null0 anchor, never redistribute
Add <PUBLIC_BLOCK> as a static route to Null0 (Routing ▸ Static, or an FRR
ip route <PUBLIC_BLOCK> Null0) so the network statement has a RIB anchor and
stays announced while the box is up. Do not redistribute connected — the
out-filter (§5.4) is the guarantee that 10/8, server zones, and core routes can
never leak upstream.
5.4 Prefix-lists and route-maps¶
Routing ▸ BGP ▸ Prefix Lists and Route Maps. Implements §5a verbatim.
# Announce ONLY our block upstream
prefix-list WAN-OUT permit <PUBLIC_BLOCK>
prefix-list WAN-OUT deny 0.0.0.0/0 le 32
# Accept default ONLY from each border
prefix-list WAN-IN permit 0.0.0.0/0
prefix-list WAN-IN deny 0.0.0.0/0 le 32
Route-maps:
| Map | Applied | Action |
|---|---|---|
WAN-IN-PRIMARY |
in, from bdr-1 | set local-pref 200 |
WAN-IN-BACKUP |
in, from bdr-2 | set local-pref 100 |
WAN-OUT |
out, to bdr-1 | permit WAN-OUT prefix-list |
WAN-OUT-PREPEND |
out, to bdr-2 | WAN-OUT + prepend our AS ×2 |
Result: bdr-1 preferred inbound (higher local-pref) and outbound (no prepend);
bdr-2 hot standby. Announce exactly <PUBLIC_BLOCK>, take default only.
For the core (CCR) sessions, the to-core out policy permits 0.0.0.0/0 +
<PUBLIC_BLOCK> + the server-zone subnets, and accepts the 10/8 aggregate + venue
prefixes inbound (§3, DESIGN §6).
5.5 Neighbours¶
Routing ▸ BGP ▸ Neighbors. Four neighbours, all from the master's real interface addresses (the /31s — no VIP source, §5a). Enable BFD on all four.
WAN borders (AS <UPSTREAM_AS>)¶
| Neighbour | Remote AS | In map | Out map | BFD |
|---|---|---|---|---|
<BDR1_FW01> (bdr-1) |
<UPSTREAM_AS> |
WAN-IN-PRIMARY |
WAN-OUT |
on |
<BDR2_FW01> (bdr-2) |
<UPSTREAM_AS> |
WAN-IN-BACKUP |
WAN-OUT-PREPEND |
on |
Core CCRs (AS <CORE_AS>)¶
| Neighbour | Remote AS | Policy | BFD |
|---|---|---|---|
<CR1_FW01> (CR-COLO-01) |
<CORE_AS> |
to-core in/out (§5.4) | on |
<CR2_FW01> (CR-COLO-02) |
<CORE_AS> |
to-core in/out | on |
The far ends must already be configured for BOTH nodes
Because of CARP-gating, the borders/CCRs carry neighbour config for FW-01 and FW-02's addresses; only the master's are ever Established. Confirm the upstream staging from page 1 §1.4 before expecting sessions to come up.
5.6 Backup-node static default¶
While the backup's FRR is stopped it still needs management-plane reachability (updates, monitoring, pfSync-adjacent services). Give each node a low-priority static default (System ▸ Gateways + Routes) marked "far gateway / do not use for policy", via the upstream hand-off or the core /31. It carries no transit (no VIPs, no announcements on the backup) — management-plane only (§3).
This is the only static route on the box
Everything else is learned. The default-route chain (upstream → OPNsense → CCRs → RRs → fleet) is entirely dynamic — nothing to switch on failure, it's emergent from route withdrawal (§5b).
5.7 Verify (master only)¶
- Routing ▸ Diagnostics ▸ BGP (or
vtysh -c 'show bgp summary'): on FW-01 all four neighbours Established; on FW-02 FRR is stopped and there are no sessions. That asymmetry is the design working, not a fault. - Confirm you receive default from both borders (bdr-1 preferred) and the 10/8 aggregate + venue prefixes from the CCRs.
- Confirm you announce only
<PUBLIC_BLOCK>upstream (show bgp neighbor <bdr> advertised-routes) — nothing else.