Resolving P3-550 Modbus TCP Stalls with Terminator I/O Slaves

Brian Holt9 min read
AutomationDirectModbusTroubleshooting
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

A Productivity3000 P3-550 polls several Modbus TCP slaves with MRX/MWX instructions. In this setup there were two Terminator I/O (T1H) racks, used for high-speed counter inputs, and four Koyo DL06 PLCs. It runs cleanly while every slave is online. Pull one slave off the network and every slave stops communicating. The CPU status bit for Ethernet buffer at 95% comes on. Traffic only resumes after a STOP-to-RUN cycle. After the dead-slave problem is fixed, a second fault appears: with four or more slaves enabled, the Terminator racks start timing out intermittently, even though the DL06s never fail.

Four separate logic faults cause this. None of them is a network fault. Get it running, then fix it properly.

Skip the Quick Fixes That Don't Hold

These are the usual first moves. Each one fails for a specific reason.

  • Cycling the CPU STOP to RUN. This clears the Ethernet transmit queue and resets every instruction's status bits, so traffic resumes. The sequencing logic is unchanged, so the stall returns the next time a slave drops or a fast response lands at the wrong moment. Use it to restore production tonight, not as the fix.
  • Switching from manual to automatic polling. Automatic polling still locked up in this installation. It also takes away control of when each transaction fires, which you need to isolate a dead slave.
  • Adding status-bit interlocking with one shared sequencer. Waiting for Success or Error before sending the next message is correct. Using one counter to step through every transaction on every slave is the problem. When one slave is down, every healthy slave waits for the dead slave's full timeout on each pass.
  • Slowing down polling to the Terminator racks. Adding delays to the T1H read/write polling did not stop the timeouts. The fault is a missed edge, not an overloaded slave, so slower polling only changes how often the race happens.
  • Blaming the T1H TCP stack or the buffer. The rising Ethernet buffer is a symptom. Instructions keep queuing because the logic never sees the transactions finish.

Find Where the Scan Actually Stalls

A Modbus TCP slave should have only one outstanding request from a given master at a time. Serialize transactions per device. Ethernet lets you run different devices in parallel. The logic breaks this rule in four ways.

1. One sequencer across all devices

A single counter steps MB01 read, MB01 write, MB02 read, MB02 write, and so on. When MB01 goes offline, its step waits the full timeout before the counter moves on, and this happens every cycle. Throughput to every healthy slave falls to roughly one pass per dead-slave timeout. Meanwhile, anything that re-triggers the instruction adds requests to the Ethernet queue.

2. The In Progress bit timing conflicts with your watchdog

The instruction's In Progress bit follows the timeout set in the CPU hardware configuration. If your own timeout timer is enabled by In Progress, the two timeouts overlap. The code can occasionally stall when a device drops out. Enable the watchdog timer from the same sequencer permissive that triggers the MRX/MWX.

3. A level contact where an edge is required

The sub-rung that advances the counter to the next instruction used a normally open contact on the permissive. That contact stays true for several scans, so the counter can advance more than once. With several counters, it causes a race condition between them. Use an edge contact so the counter advances exactly once per completion.

4. The Complete bit never drops when the slave answers too fast

This fault hides behind the other three. An instruction's Complete bit stays on until that instruction runs again. When the instruction re-fires, Complete clears and then sets again when the response arrives.

A PLC-based slave such as a DL06 or an H2-Ecom100 answers after its own scan. The P3-550 therefore sees Complete go false for at least one scan, and the one-shot re-arms.

A T1H is not limited by a scan and answers much faster. The bit can go off and back on between two evaluations of your one-shot. The one-shot never sees a false-to-true change, so it never fires again. That sequencer step waits forever, and the 1-second "Complete stuck" watchdog reports a comms failure.

This fault depends on timing. Adding MB04 and above changes scan and Ethernet service timing enough to make the race common, which is why the slave count seemed to matter. When the T1H racks were replaced with H2-Ecom100 units on the same project, the problem did not appear. It only appeared with the T1H hardware.

Match the Symptom to the Cause

Symptom Cause Fix
One slave offline, all slaves stop; Ethernet buffer 95% status bit on Single sequencer waits on the dead slave's timeout; instructions keep queuing One sequencer per slave
Occasional stall when a device drops, even with per-device interlocks Watchdog timer enabled by the In Progress bit, which follows the CPU hardware-config timeout Enable the watchdog from the sequencer permissive that triggers the instruction
Counters skip steps or interfere with each other Normally open contact on the advance permissive Replace it with an edge contact
T1H racks time out intermittently when 4+ slaves are active; DL06s never fail Complete bit re-sets faster than the one-shot can detect RST the Complete bit on the same rung that latches it
Delaying T1H polling does not help The fault is a missed edge, not slave overload Complete-bit reset, not a slower poll rate
STOP-to-RUN restores comms temporarily Mode change clears the queue and status bits; the logic fault remains All of the above

Split the Sequencer: One Interlock per Slave

Give each Modbus TCP slave its own called task and its own step counter. Each counter interlocks only that slave's transactions. Do not interlock between devices. Healthy slaves then keep polling at full speed while a dead slave waits out its timeouts on its own.

  1. Create one called task per slave (for example MB01 through MB06). Call each task from the main task through an enable bit, so you can disable any slave from the HMI or from a fault bit.
  2. Inside each task, create one step counter covering that slave's transactions only. Example: step 1 is the MRX read and step 2 is the MWX write.
  3. Trigger each instruction from its step permissive, and only while the previous transaction has finished with Success or Error.
  4. Advance the counter on an edge of Complete or Error, never on a level contact.
  5. Reset the counter to step 1 after the last transaction so the slave polls in a continuous loop.
  6. Enable the per-step watchdog timer from the same step permissive that fires the instruction, not from In Progress.

In this installation, the per-slave structure produced transaction rates of around 10 ms per slave. This was measured with counters written to the DL06s. It proved faster and more reliable than automatic polling.

Reset the Complete Bit the Moment You See It

Latch Complete into your own SET bit for downstream logic. On the same rung, and in parallel, reset the instruction's Complete bit. The instruction then always starts its next transaction from a known false state. The one-shot sees a clean rising edge no matter how fast the slave answers.

Change the completion rung in each task. Example for the MB01 read (rung 6 of MB01_Comms):

One Shot "MB01_R01_COMPLETE" ----+----(SET) MB01_R01_COMPLETE_SET
                                 |
                                 +----(RST) MB01_R01_COMPLETE
  • Apply this to every MRX and MWX instruction in the project, not just the T1H tasks. It is only strictly required for the fast responders, but it removes the race for any slave, including a PLC slave that is later replaced with faster hardware.
  • Use MB01_R01_COMPLETE_SET, or your equivalent, for the counter advance and for any data-valid logic. Never use the raw Complete bit for these after this rung.
  • Clear the SET bit when the sequencer leaves that step. Otherwise the next pass sees a stale completion.

Build Offline Detection and Auto-Recovery

Once each slave has its own sequencer, handling a dead slave is local to that slave's task.

  1. Detect. Run a watchdog timer per step, enabled by the step permissive. The comms-failure alarm in this installation used 1 second of Complete not arriving. Set your own value above the slave's normal response time and above the timeout configured in the CPU hardware configuration, so the instruction reports Error first.
  2. Count. Advance on Error as well as Complete. A timed-out read must not freeze the sequencer, or that slave never recovers.
  3. Flag. After a set number of consecutive Errors, set a slave-offline bit for the HMI and hold that slave's process data at a safe value.
  4. Throttle, don't stop. While a slave is offline, keep its sequencer cycling on a slow retry interval instead of disabling the task permanently. Retries keep testing the link without holding up other slaves, because nothing outside the task waits on it.
  5. Recover. On the first Success after an offline period, clear the offline bit, reset the error counter and return to the normal poll rate. No STOP-to-RUN cycle is needed.

Stop here if the Ethernet buffer 95% status bit still comes on after these changes. Look for an instruction that fires outside its sequencer permissive. Common causes are a duplicated rung, a leftover test rung or an automatic-poll entry still configured for the same slave.

Verify Under Load Before You Walk Away

  1. Enable all slaves (MB01 through MB06). Watch each task's step counter. Each one should cycle continuously with no watchdog alarms.
  2. Put a heartbeat counter in each slave task and write it to a DL06. Confirm the transaction rate per slave is steady, in the range seen here (around 10 ms).
  3. Unplug one Terminator rack. Only that slave's offline bit should set. Every other slave's heartbeat must keep counting at the same rate, and the Ethernet buffer 95% bit must stay off.
  4. Reconnect the rack. Confirm the offline bit clears and data updates resume without a mode change.
  5. Repeat steps 3 and 4 with a DL06, then with two slaves offline at the same time.
  6. Leave the full configuration running for several days and log watchdog trips per slave. Any trip on a T1H task with the network healthy means a Complete bit reset is missing on that task.

FAQ

Why does one offline Modbus TCP slave stop all P3-550 communications?

A single sequencer shared across all slaves waits the full timeout on the dead slave before moving on. Instructions keep queuing until the CPU's Ethernet buffer 95% status bit sets. Give each slave its own step counter and task so only that slave waits.

Why does Terminator I/O time out when DL06 slaves on the same P3-550 don't?

T1H racks answer without a scan delay, so the instruction's Complete bit can clear and re-set before the one-shot evaluates it. The one-shot never fires, and the step hangs until the watchdog trips. Add a parallel RST of the Complete bit on the rung that latches it.

Why does STOP to RUN fix Productivity3000 Modbus comms temporarily?

The mode change clears the Ethernet queue and resets every instruction status bit, so sequencers restart from a clean state. The interlock, edge-contact and Complete-bit faults are still in the logic, so the stall returns on the next dropout or fast response.

When should I call AutomationDirect support about MRX/MWX lockups?

Call support if the buffer bit still sets or a task still stalls after you have made every change: per-slave sequencers, edge-triggered advance, a watchdog enabled from the step permissive and a Complete bit RST on every instruction. Send them the project file, the slave models and a record of which step counter freezes.

Back to blog