Industrial networking-Design Review: VRRP + PRP Hybrid Architecture

Hi Experts,

I'm validating a hybrid design combining L3 gateway redundancy (VRRP router-on-a-stick) with IEC 62439-3 PRP. Would appreciate your critique on standards alignment and latent failure modes.

### Proposed Architecture
* **Core Gateway:** Dual routers running router-on-a-stick with per-vlan VRRP (unique Virtual IP/VMAC per substation VLAN).
* **PRP Integration:** R909 RedBoxes interfacing core routers into parallel LAN-A / LAN-B infrastructure.
* **Topology Path:** Core Routers ↔ RedBoxes ↔ Aggregation Switches (A & B) ↔ 10 Substation Downstreams.
* **VLAN Strategy:**
* Field/Client VLANs: Unique per substation (VLAN 1–10).
* Management VLAN: Common `VLAN 555` stretched across all 10 substations.
* Trunk Policy: All trunks unpruned end-to-end.

### Key Questions for Review
1. Does terminating VRRP gateway virtual IPs on a router-on-a-stick fed via RedBox interlinks comply cleanly with standard PRP integration patterns?
2. Do you see stability or security risks with unpruned trunks across all downstream aggregation links?
3. Are there regulatory/security (IEC 62443) or operational red flags with stretching a common management `VLAN 555` across 10 discrete substations?]

Zoo365_0-1789651194746.png
 
A few points from the security / operations side:

1. Unpruned trunks: I would prune. With every VLAN on every trunk, a loop or broadcast storm in one substation VLAN can reach all 10 sites, and a single misconfigured or compromised port has layer-2 reach everywhere. Allow only the VLANs each substation actually needs (its own field VLAN + its management VLAN), and set the native VLAN to an unused ID.

2. Stretched VLAN 555: this is the one I would flag in an IEC 62443 review. One L2 management domain across 10 substations puts all sites in a single zone, so a compromise at one site gives L2 adjacency (ARP spoofing, discovery protocols) to the managed devices at every other site. A management VLAN per substation, routed at the core, with an ACL/firewall conduit that only allows management traffic from the jump host / NMS subnet, gives you per-site isolation and a zone/conduit model you can actually document.

3. VRRP behind RedBoxes: workable in principle, but remember PRP protects the LAN paths, not the gateway. If each router hangs off a single RedBox interlink, that RedBox and its interlink are still single points of failure, so make sure each router has its own RedBox (or its own interlink) and that the VRRP priority/preemption timers are tuned so a RedBox failure triggers a clean failover.

Before go-live I would run a simple fault-injection test (pull LAN-A, pull LAN-B, fail the master router, fail a RedBox) and record the recovery time for your SCADA polling at each step.
 
Top