Write Down the Failure Before You Power Anything Down
This failure happens in two stages, and the order tells you where to look. First the running CPU degrades. It stays online, but Modbus TCP traffic to some slaves stops completely while the others answer very slowly. A Stop-to-Run transition changes nothing. Then a normal power cycle leaves the CPU unable to boot. A controller that degrades in Run and then will not boot has a problem with its own internal state. The ladder program and the I/O are not the cause.
The configuration here was a Productivity3000 5-slot base with a P3-550 CPU polling six Modbus TCP slaves over a weekend. The hang looked like this:
| Symptom | What it means | What it rules in or out |
|---|---|---|
| CPU online, some Modbus slaves dead, others very slow | The client is stalled on timeouts, or the comm stack is short on resources | Fault sits below the user program |
| Stop then Run, no change | Stop/Run restarts the program scan. It does not restart the Ethernet stack | Program logic alone is not the cause |
| "automationdirect.com booting" splash, then the text clears and the spinner keeps rotating | Boot started but never handed over to normal operation | CPU firmware or hardware state, not the project |
| Front buttons beep but do nothing | The keypad and display handler is alive. The main boot is not servicing it | CPU is partly alive, so this is not a dead supply |
| Red CPU LED flashes intermittently | CPU-level fault or boot-state indication | Read the LED table in the CPU hardware manual for the exact meaning |
| No connection over USB or Ethernet | Boot never reaches the point where the ports serve the programming software | You cannot recover it through software |
| All I/O modules removed, no change | No I/O module is loading or corrupting the backplane | Focus on the CPU, its base connection, and power |
Check before moving on: you have recorded the LED pattern, the display state, and the approximate time comms started to degrade. Pull the time from slave-side logs, an HMI alarm history, or a historian if you have one. Support will ask for all three.
Skip the Quick Fixes That Don't Work
Night shift reaches for these first. On this failure none of them restores production.
- Stop/Run toggle. This resets the scan, not the TCP stack or the Modbus client connections. If comms are stuck below the program, Stop/Run cannot reach them.
- A quick power cycle. The power supply and the CPU both have bulk capacitance. A short off-time may not bring every rail and peripheral controller to zero, so the CPU can come back up in a bad state. In this case the quick cycle turned a degraded CPU into one that would not boot at all.
- Pulling the I/O modules. This was already tried and changed nothing. It is a useful elimination step, but it is not a fix.
- Guessing a front-button "factory reset". Some other platforms have a keypad memory-clear combination. None was used or needed in this recovery. The buttons were not being serviced during the hang anyway: they beeped and did nothing. Use only a clear procedure documented in the P3-550 hardware user manual. Guessing combinations on a CPU stuck in boot will either do nothing or clear something you wanted to keep.
- Boxing it for return immediately. Try the drain-and-reseat first. It costs about an hour and it recovered this unit.
Get it running, then fix it properly. The next section gets it running.
Pull the CPU and Let It Drain Out of the Base
What recovered this CPU was taking it out of the base and leaving it on a bench for about an hour. After it was reinserted, it booted and ran normally.
- Switch off power to the base at the supply feed, not only at the module.
- Remove the CPU from the base. Leave the I/O modules out for now.
- Set the CPU on an ESD-safe surface and leave it out of the base. The unit that recovered sat out for about an hour. Do not shorten this to a few seconds, because that is just the quick power cycle again.
- While it drains, measure the incoming supply voltage at the power supply terminals. Compare it with the rated input range printed on the supply module or in its datasheet.
- Reinsert the CPU, apply power, and watch the complete boot.
Why this works when a power cycle does not: taking the module out removes the backplane path and any stored charge on the base side. Given enough time, all the CPU's on-board rails and volatile state fall fully to zero. The Ethernet controller, flash interface, and main processor then come out of a true cold start instead of a partial reset. Pulling the module also wipes the backplane contacts. The next two sections separate those two effects.
Check: the splash screen clears to normal status text instead of a lone spinner. The front buttons navigate the display menus. The red CPU LED stops its intermittent flashing.
Stop here if the CPU still hangs at the spinner after a full drain and reseat, with no I/O installed and a verified supply voltage. If a spare base or power supply is on the shelf, try the CPU in it once. If it still hangs, go straight to the escalation at the end of this page.
Reseat the CPU and Prove the Base Connection
The recovery involved both a long drain and a physical reseat, so a marginal backplane contact is still a suspect. Rule it out now, while the rack is empty.
- Power down. Remove the CPU and inspect the base connector and the CPU edge connector for debris, bent pins, or discoloration.
- Reinsert the CPU and seat it fully. Engage any locking hardware the module uses.
- Power-cycle three times using a normal off-time and no extended drain. Watch every boot to completion.
- Between cycles, press lightly on the CPU housing while it runs. The display and LED must not flicker.
Check: three clean boots in a row with normal-length power cycles. If a normal cycle now hangs it again, the fault is inside the CPU and is not a contact problem. Log it and escalate.
Reconnect Productivity Suite Over USB, Then Ethernet
Connect over USB first. It is a point-to-point link with no switch, no IP settings, and no other traffic, so it tells you whether the CPU's programming interface is healthy on its own.
- Connect over USB and go online with Productivity Suite.
- Read the CPU firmware revision and hardware information. Write them down.
- Open the CPU's system status and any error or event history the software exposes. Record every entry from the failure window.
- Compare the project on the CPU with your saved master copy. Confirm the retentive values are sane and not zeroed or garbage.
- Check the firmware revision against the current release on AutomationDirect's site. Read the release notes for any fixes to boot, the Ethernet stack, or Modbus TCP client behavior. If a newer release covers either, update before you resume testing.
- Disconnect USB and go online over Ethernet using the CPU's configured IP address.
Check: both links connect, the project matches your master, and the firmware revision is recorded along with the before and after revisions of any update.
Bring the Modbus TCP Slaves Back One at a Time
The "some dead, some slow" pattern has a specific mechanism. A Modbus TCP client that polls several servers waits the full timeout on every request to a server that has stopped answering. If requests are serviced in order, every healthy slave queued behind the dead one slows down too. A second mechanism looks the same: the CPU's TCP stack runs short of connection resources, for example sockets that are never closed after repeated timeouts or reconnect storms. In that case new connections fail and the existing ones crawl. Stop/Run will not clear either condition.
Rebuild the network so that each slave proves itself before you add the next.
- From a PC on the same switch, ping each of the six slaves and poll it with a Modbus TCP test tool. Any slave that fails from the PC is a slave or network fault, not a CPU fault.
- Confirm there are no duplicate IP addresses on the segment. A slave or PC that picks up a duplicate address partway through a run produces exactly this kind of intermittent dead-or-slow behavior.
- In the program, wire the success, error, and timeout status outputs of every Modbus read and write instruction into counters. Also latch the timestamp of the first error.
- Set the per-slave timeout and poll interval to what the process actually needs. Do not poll all six slaves at the fastest rate at once.
- Enable one slave. Let it run while you watch its counters.
- Add the next slave only when the previous one shows a rising success count and flat error and timeout counts.
Check: with all six slaves enabled, every success counter climbs and every error and timeout counter stays flat. When you unplug one slave on purpose, the other five keep their update rate within your timeout budget. If they all slow down, add per-slave enable logic that skips a slave after repeated timeouts and retries it on a slower interval.
Reinstall the I/O and Run the Overnight Soak
The original failure appeared after unattended running over a weekend. A test that passes in ten minutes proves nothing. Run the soak for at least as long as the original failure took to appear.
- Reinstall the I/O modules and confirm each one is recognized.
- Log the Modbus success, error, and timeout counters, a free-running heartbeat, and timestamps. Log to the CPU's own data logging or to an external PC polling the CPU.
- If the switch is managed, keep its port error counters and link-state log for the same window.
- Leave it running overnight and do not touch it.
- In the morning, read the counters first, then perform a deliberate power cycle with a normal off-time.
| Soak result | Action |
|---|---|
| Counters clean, CPU reboots normally | Keep the status counters in the production program permanently and proceed with the job |
| Comms degrade again, CPU still boots | Save the logs and error history. Check for a firmware update, then contact AutomationDirect support with the data |
| Comms degrade and the boot hang returns | CPU firmware or hardware fault. Escalate for return |
| Boot hang with clean comms counters | Suspect the CPU hardware, base contact, or supply. Swap the base or supply if spares exist, then escalate |
End-to-end check: after the soak and the deliberate power cycle, the CPU reaches normal display text. USB and Ethernet both connect. All six slaves resume polling with no errors. Every counter from the night is saved with the firmware revision you recorded earlier.
FAQ
Can I factory reset a P3-550 CPU with the front buttons?
Do not guess at button combinations. During this hang the buttons beeped but were not serviced, and the recovery that worked was powering off, leaving the CPU out of the base for about an hour, and reseating it. Use only a memory-clear procedure documented in the P3-550 hardware user manual.
Does removing the I/O modules fix a Productivity3000 CPU stuck at boot?
No. Removing every I/O module made no difference here. It rules out the modules, so your next suspects are the CPU itself, its base connection, and the power supply.
Does a Stop/Run cycle reset Modbus TCP comms on the P3-550?
No. Stop/Run restarts the program scan but leaves the Ethernet stack and client connections as they were, so dead or slow slaves stay that way. Restart the CPU only after you have captured the comm counters and error history.
Can I keep running a P3-550 that hung at boot once?
For testing, yes, once it has passed an overnight soak with clean Modbus counters and several deliberate power cycles. Stop and contact AutomationDirect technical support if the boot hang returns after a full drain and reseat, or if comms degrade again on current firmware. Send them the firmware revision, the LED and display behavior, the failure timeline, and the logged counters.