Ring Redundancy: The Number That Decides Whether the Line Stops
TL;DR: Ring redundancy is not a yes or no property. It is a number, and the number that matters is not on the switch datasheet, it is in your controller. If the network takes longer to heal than the PLC’s watchdog allows, the line stops anyway. RSTP heals in seconds. MRP guarantees hundreds of milliseconds. ERPS rings on the switches we supply publish under 20 ms. PRP and HSR do not heal at all, because nothing ever broke.
Last updated: 15 September 2026
On this page
Key takeaways
- The specification you are designing to is your controller’s tolerance for lost frames, not the switch’s marketing number.
- RSTP is a spanning tree protocol built for office networks. It converges in seconds, and it is the default on almost every managed switch.
- MRP, standardised as IEC 62439-2, is a ring protocol designed to “react deterministically on a single failure”, under a dedicated media redundancy manager.
- The ATOP EHG9508 and EHG9512 we supply publish a ring recovery time of under 20 ms across 40 switches using ERPS and Compatible Ring.
- PRP and HSR, standardised as IEC 62439-3, send every frame twice and provide “seamless switchover with zero recovery time”.
- Unmanaged PoE switches cannot participate in any ring redundancy scheme at all.

What is the number you are actually designing to?
The controller’s tolerance, not the switch’s. A PLC talking to remote IO over Ethernet expects an update every cycle, and it faults when a configured number of cycles pass without one. That timeout is the specification. Any ring redundancy scheme that heals faster than it is invisible to the process, and any ring redundancy scheme that heals slower produces exactly the same outcome as none at all: a stopped line and an alarm.
This reframes the whole purchasing decision. People ask us whether a switch “supports ring redundancy”, which is nearly always yes and nearly always useless. The question that decides the design is what update rate the controller runs at, how many missed updates it will accept, and therefore how many milliseconds of network outage the plant can absorb without noticing.
Work in that direction and the protocol selects itself. Tens of milliseconds of tolerance rules out spanning tree immediately and points at a deterministic ring. Zero tolerance, which is the normal situation for protection signalling in a substation, rules out rings entirely and points at parallel paths. A control loop with a second of slack can honestly use RSTP and save money.
The failure we see most often is a network designed to a switch datasheet and a process designed to a controller manual, with nobody comparing the two numbers until commissioning.
Why RSTP is the wrong default on a plant floor
Because it was designed for a different problem. Spanning tree exists to stop loops in arbitrary, unplanned office topologies where nobody knows in advance which link will fail or what the network will look like next year. It achieves that beautifully, and it achieves it by negotiating, which takes time.
Rapid Spanning Tree improved matters enormously over the original protocol, whose convergence was measured in tens of seconds. It is still a negotiation between peers rather than a pre-agreed plan, and its recovery is measured in seconds rather than milliseconds. For a file server that is invisible. For a servo drive it is a fault.
The trap is that RSTP is on by default. Buy a managed industrial switch, wire it in a ring, and it will not loop, because spanning tree has quietly blocked a port for you. The ring looks redundant, and it is. It is simply ring redundancy at the wrong timescale, and nobody discovers that until the day a fibre is cut and the plant finds out how long “a few seconds” feels.
There is one genuinely good reason to keep RSTP running: it is the safety net that stops a broadcast storm if somebody patches two access ports together. Run it alongside the fast protocol rather than instead of it.
MRP, ERPS and deterministic ring redundancy
Deterministic ring redundancy replaces negotiation with a plan. One switch is appointed manager, it knows the ring topology in advance, it deliberately blocks one port to break the loop, and it watches for a break. When the break happens, it does not discover the new topology, it already knows it, and simply unblocks the port it was holding.
The Media Redundancy Protocol is the standardised version of that idea. IEC 62439-2:2016, edition 2.0, “Industrial communication networks. High availability automation networks. Part 2: Media Redundancy Protocol (MRP)” specifies, in its own words, “a recovery protocol based on a ring topology, designed to react deterministically on a single failure of an inter-switch link or switch in the network, under the control of a dedicated media redundancy manager node”.
ITU-T G.8032, usually called ERPS, solves the same problem from the telecoms side, and it is the protocol most industrial vendors publish their headline number against. On the ATOP managed switches we supply, the published figure is concrete: the EHG9508 and EHG9512 specify a ring recovery time of under 20 ms across 40 switches using ERPS and Compatible Ring. The same families list ITU-T G.8032 ERPS, STP, RSTP, MSTP, MRP in manager and client roles, Compatible Ring, Chain and U-Ring.
Note what that specification is careful to include: a switch count. Ring redundancy recovery time scales with ring size, because the healing message has to travel round the ring. A number quoted without a node count is not a specification, it is a hope.


PRP and HSR: ring redundancy with no recovery at all
Where zero frames may be lost, ring redundancy stops trying to heal quickly and starts sending everything twice. IEC 62439-3:2021, edition 4.0 specifies the Parallel Redundancy Protocol and High-availability Seamless Redundancy, both of which provide seamless switchover with zero recovery time.
PRP attaches each device to two completely independent networks and sends a copy of every frame down both. The receiver takes whichever copy arrives first and discards the duplicate. If an entire network fails, nothing recovers, because the other copy was already arriving. HSR does the same thing around a ring, with each node forwarding in both directions and discarding duplicates as they return.
The cost is honest and substantial: two networks worth of cable and ports, or nodes that natively speak HSR, and roughly double the traffic. That is why it lives where the consequence of a lost frame is a protection operation rather than a late data point, which in practice means IEC 61850 substations and comparable critical infrastructure. Our substation and smart grid work is where this conversation usually starts.
The switches that cannot do ring redundancy at all
This is worth saying plainly before somebody orders the wrong thing. An unmanaged industrial PoE switch has no management interface, no spanning tree, no ring protocol, and no ability to participate in ring redundancy of any kind. Wire unmanaged switches in a ring and you get a broadcast storm, not resilience.
They are excellent at the job they are for. An unmanaged PoE switch is a camera or sensor aggregation switch at the edge of a star, and for that job managed features are money you do not need to spend.
- Star topology, cameras and sensors, no resilience requirement: an unmanaged PoE switch. Watch the power budget rather than the protocol list.
- Ring redundancy with a millisecond requirement: ATOP managed, quoted per project. The redundancy protocol list and the published recovery time are on the product page, and we will tell you the node count the figure assumes.
- Zero frame loss: PRP or HSR capable hardware, and a design conversation before anything is quoted.
One caution on the port speeds, because it catches people out. Many unmanaged PoE switches carry Gigabit uplinks and 10/100 access ports. That is a sensible design for PoE cameras and sensors, and it is not a Gigabit switch. Read the port line, not the title.


The physical mistake that undoes all of it
Running both halves of the ring through the same duct. Ring redundancy protects against a single failure of a link or a switch. It does not protect against a single failure that takes out two links at once, and an excavator does not care that your two fibres were logically independent.
The same logic applies inside the building. Two paths through the same riser, the same tray, the same panel, or the same UPS are one path wearing a disguise. A ring redundancy design is only as good as the diversity of the physical routes it runs over, and that diversity is decided by civils and containment, not by the switch configuration.
Power is the version people forget most often. A beautifully engineered ring of eight switches, all fed from the same distribution board, has a single point of failure with a handle on it. Dual power inputs on industrial switches exist for this reason, and a relay output that tells you when one of them has gone is worth wiring up on day one rather than discovering silently on day two.
Test it before you need it. Pull a fibre during commissioning, with the process running in a safe state, and measure what actually happens. A recovery time you have observed on your ring, with your node count, is worth more than any datasheet figure including the ones quoted above.
Frequently asked questions
Is RSTP good enough for ring redundancy on a plant?
Sometimes, and it depends entirely on the controller. For SCADA polling, HMI traffic and general plant IT, seconds of recovery is usually tolerable. For remote IO, motion or protection, it is not. Keep RSTP enabled as a loop-prevention backstop even when a faster protocol is doing the real work.
Can I mix vendors in one ring redundancy scheme?
With a standardised ring redundancy protocol like MRP or G.8032 ERPS, in principle yes. With a vendor’s proprietary ring, no. This is the practical reason to specify a standard rather than a brand, even when you intend to buy one brand, because it keeps the second phase of the project open.
How many switches can a ring redundancy loop have?
The protocol sets a limit and the recovery time is quoted against a node count, so the two questions are the same question. The ATOP figure above is under 20 ms at 40 switches. Smaller rings recover faster, which is one good reason to split a large plant into several rings rather than one heroic loop.
Does ring redundancy protect against a switch failing?
Ring redundancy protects the ring, not the devices hanging off that switch. Everything connected to the failed switch is gone until it is replaced. If a particular device must survive its own switch failing, it needs two connections to two switches, which is what PRP is for.
Do we need managed switches everywhere?
No, and we will not quote it that way. Managed switches belong on the ring redundancy loop itself. The unmanaged switches hanging off the ring, aggregating cameras or sensors in a cabinet, do not need to speak the ring protocol and are considerably cheaper for it.
Where this leaves you
Ring redundancy is one of the few areas of network design where the right answer is a number rather than an opinion. Find the controller’s tolerance first, then pick the protocol that clears it with margin, then check that your two physical paths are genuinely two paths. Everything else is detail.
Send us the node count, the controller and its update rate, and whether the routes are physically diverse, and we will come back with a topology, the protocol it needs and a price inside one working day. Start with our industrial networking solution, and check the PoE power budget before you size anything with cameras on it.
Next step
Get a priced kit list for your site
Answer three quick questions and tell us where to send it. An engineer replies within one working day with the parts, the prices and the lead time.
Rather talk it through? Call 023 9223 3611
Received. An engineer has it.
You will hear back within one working day with the parts, the prices and the lead time. A confirmation is on its way to your inbox.
If it is urgent, call 023 9223 3611.