Skip the Quick Fixes That Don't Hold
The failure looks like this. A Layer 2 loop forms on the plant network, often because Spanning Tree Protocol (STP) drops out after a switch software update. Traffic starts circulating around the fiber ring. Every P3-550 on that segment reboots itself and comes back in STOP, with E01701 and E01005 logged. Nobody touched power and nobody moved the RUN/STOP switch. DL06 and DL205 CPUs on the same segment, with the same Modbus TCP and programming connections, keep running.
These are the usual night-shift moves and why none of them hold:
| Quick fix | Why it fails |
|---|---|
| Put the CPU back in RUN and walk away | The loop is still active. The CPU crashes again as soon as the flood hits the external Ethernet port. |
| Power-cycle the CPU | Same result. The fault is on the wire, not in the CPU's retained state. |
| Re-download the project | The logic did nothing wrong. The CPU went down while processing incoming frames. |
Roll back or step forward within 1.1.15.99, 1.1.15.101, 1.1.15.104
|
All three builds crash the same way. The fault does not depend on which of these versions is loaded. |
| Replace the CPU | A new CPU on the same firmware and the same looped network fails the same way. |
Production comes back in this order: stop the storm, recover the CPU, load firmware that ignores the flood, then configure the switches so a storm cannot reach the PLCs again. Get it running, then fix it properly.
Kill the Loop Before You Touch the CPU
At Layer 2, frames have no hop limit. When STP is not blocking a redundant path, a broadcast frame goes around the ring and every switch floods it out every port on every pass. Within seconds each host on the segment receives the entire circulating load. Most endpoints drop frames they have no use for. On pre-1.1.15.111 firmware, the P3000 external Ethernet port tried to process that traffic instead and crashed the processor.
- Find the switch or switches that were updated most recently. On a plant network, a recent switch software change is the first suspect.
- Check the STP status on every switch in the ring. Confirm that STP is enabled, that a single root bridge is elected, and that exactly one ring port is in blocking state.
- If STP will not converge, break the ring by hand. Shut down or unplug one fiber uplink between two ring nodes. The ring becomes a line, which cannot loop.
- Restore the STP configuration on the updated switch. Compare it against a switch that was not updated, or against your saved config backup.
- Put the redundant link back only after STP shows the expected blocking port.
Check: Watch broadcast and multicast counters on the switch ports that feed the PLCs. The counters should drop back to their normal background rate, and link and activity LEDs should stop showing solid, continuous activity. Stop here if the counters are still climbing. Any CPU you put back in RUN will go down again.
Read the Log, Then Return the CPU to RUN
- Connect with the programming software and open the CPU's error and event history before you clear anything.
- Record the codes, their timestamps, and the loaded firmware version. The key fact is
E01701logged together withE01005after an unprompted restart. Look up the exact text for both codes in the programming software help for your firmware build. - Match the timestamps against the switch logs. The CPU restart should line up with the moment STP went down or the loop formed. If it does, you have confirmed the cause.
- Clear the fault and switch the CPU to RUN.
Check: The CPU stays in RUN. Modbus TCP clients reconnect, and the programming software stays online without dropouts. If the CPU restarts again on a quiet network, stop. You have a different problem. Save the log and escalate.
Load Firmware 1.1.15.111
AutomationDirect released 1.1.15.111 with changes to the external Ethernet port handling. On this build, the CPU ignores unexpected traffic such as network storms instead of trying to process it and crashing. With this change, the P3000 behaves like the DL-series processors on the same segment. Treat it as protection against crashes. It does not replace a working network.
The update needs the physical RUN/STOP switch in the STOP position. You cannot bypass this, and you cannot run the update remotely with only a program stop. This is deliberate. If communications drop partway through a remote flash, the partially written firmware can leave the CPU unrecoverable. You then have to travel to the site and replace the CPU anyway, with far longer downtime than a planned visit. Having someone on site during the flash cuts the outage if something goes wrong.
- Download
1.1.15.111from AutomationDirect's official site. Read the release notes for any project conversion or software version requirements. - Save a backup of the project from the CPU before flashing.
- Get a qualified person to the cabinet. Take the process to a safe state and move the switch to STOP.
- Run the firmware update from a laptop on a stable connection. Do not run it across the plant backbone that just stormed. Use a short, direct connection or a local switch.
- Do not remove power or the cable until the update reports that it is complete.
- Let the CPU restart. Download the project again if the software asks for it, move the switch back to RUN, and confirm I/O.
Check: Read the firmware version back from the CPU. It must show 1.1.15.111 or a later official release. Check that retained values and recipes survived. Compare them against your backup.
Harden the Switch Ports Feeding PLCs
New firmware keeps the CPU alive through a storm. The storm still takes out Modbus TCP polling and programming access while it lasts. Contain the storm at the switch.
- Storm control / broadcast rate limiting: Enable it on every access port that connects to a PLC. Set the threshold from the traffic you measure during normal operation, with headroom. Use the switch's own documentation for units and limits.
- Edge port protection: Set PLC ports as edge ports and enable BPDU guard or the switch's equivalent, so a stray loop at the panel disables the port instead of flooding the ring.
- Loop detection: Enable it if your switches support it. It is a second layer of protection if STP fails the way it did here.
- Segmentation: Move the more than 30 PLCs off the general plant segment and onto their own VLAN or physical network, routed or firewalled from the office network. A smaller broadcast domain limits how many machines one loop can reach.
- Change control: After any switch software update, check STP state and port roles before you leave the switch. Keep configuration backups of each switch before and after the change.
Check: Review each PLC-facing port. Storm control should be enabled, the port should show as an edge port, and the STP topology should match your drawing. Save the running config so it survives a switch reboot.
Plan the Rollout Across the Plant
With more than 200 PLCs, many in remote or restricted locations and some processes allowing only a few minutes of downtime, flashing every CPU in one night is not practical. Put the network work first and the firmware on a schedule.
| Priority | Which CPUs | Reason |
|---|---|---|
| 1 | P3000 CPUs on the segment that stormed | Proven exposure. They will crash again if the loop comes back. |
| 2 | P3000 CPUs running continuous processes with short outage windows | An unplanned STOP costs the most here. Flash during a planned stop. |
| 3 | Remaining P3000 CPUs | Flash during routine maintenance visits. |
| — | DL06 / DL205 | Not affected by this failure. No action needed for this issue. |
Bundle each firmware visit with other on-site work, so the trip, the permit, and the special access personnel count toward more than one job. Carry the project backup and a spare CPU on visits to the most remote or critical sites. A failed flash then costs one swap, not a second trip.
Check: Keep a register of every P3000 with its location, current firmware version, and update date. Close it out when every row shows 1.1.15.111 or later.
Verify End to End
Confirm each layer independently. Do not take the first quiet night as proof.
- Network: STP shows one root bridge and the expected blocking port. Broadcast counters on PLC ports sit at their normal background rate.
- Switch protection: On a bench switch or an isolated test segment, create a deliberate loop with a patch cable and confirm that storm control or BPDU guard trips as configured. Never create a loop on the live production ring.
-
CPU: On that same bench setup, connect a spare or out-of-service P3000 running
1.1.15.111and let it see the storm. It should stay in RUN and log no newE01701/E01005restart. - Application: In production, confirm that Modbus TCP clients and programming connections recover on their own after any network disturbance, without anyone touching the CPU.
- Next switch change: Watch the PLCs during the next planned switch update. Check the CPU logs afterward for unprompted restarts.
FAQ
What happens if a P3000 on old firmware sees another broadcast storm?
On 1.1.15.99, 1.1.15.101, or 1.1.15.104, the CPU tries to process the flood on its external Ethernet port and crashes. It restarts into STOP with E01701 and E01005 logged, even though nobody cycled power or moved the RUN/STOP switch. Update to 1.1.15.111 and enable storm control on the switch port that feeds it.
What happens if a remote P3000 firmware update loses communications?
A partial firmware write can leave the CPU unrecoverable. You then have to travel to the site and replace it. That is why the update requires the physical switch in STOP and a person at the cabinet.
What happens if STP is down but the ring stays connected?
The redundant fiber path forms a Layer 2 loop. Broadcast frames circulate with no hop limit and flood every host on the segment. Break the ring by shutting down one uplink until STP converges again with one blocking port.
When should I stop and call AutomationDirect support?
Call AutomationDirect support if a P3000 on 1.1.15.111 or later still restarts into STOP during a storm, or if it logs E01701/E01005 on a network with normal broadcast counters. Also call them if a firmware update fails to complete and the CPU will not come back online. Have the CPU event log, the firmware version, and the matching switch log timestamps ready before you call.