Troubleshooting Modbus TCP Connection Churn in PLCs

Daniel Price7 min read
ModbusOther ManufacturerTechnical Reference
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

Modbus TCP connection churn occurs when a client opens a TCP connection, sends one query, receives one reply, and immediately closes the connection. The completed request can be valid, but repeating the TCP handshake and teardown for every transaction wastes socket, CPU, memory, and network resources. Follow the packet from the PLC client through the network to the server, then separate physical-link faults, connection failures, transaction failures, and teardown behavior.

Where does the request path stop?

Map one complete exchange before changing either endpoint. The PLC initiates the connection from a local IP address and temporary source port. The request crosses the physical link and any switches or routed hops, reaches the server's configured Modbus TCP listening port, receives a response, and then enters TCP teardown. A failure at one stage can resemble a Modbus application timeout at the client.

Path point Reading to take Outcome Next check
PLC interface Link state, error counters, negotiated speed and duplex Errors or link transitions indicate a layer-one problem Correct media or negotiation before inspecting Modbus
Network path Client and server IP addresses, route, packet loss No bidirectional IP traffic indicates addressing or path failure Trace the first hop that stops forwarding
Server listener Configured address and listening port A rejected or unanswered open indicates a listener, filtering, or resource problem Inspect connection attempts and active socket count
Modbus exchange Request and matching response A response proves that transaction completed before teardown Determine who closes and how often
TCP teardown Close initiator and resulting socket states One close per response identifies connection-per-request behavior Measure reconnect rate and resource pressure

Does layer one remain healthy during the reconnect cycle?

Layer one first. Watch both endpoint interfaces while the PLC repeatedly polls. Rising receive errors, discarded frames, link transitions, or mismatched negotiation can interrupt handshakes and responses independently of the client's close policy. Check intermediate switch ports as well as the two endpoints because the packet can disappear between healthy interfaces.

If counters remain stable and each request produces a reply, move to TCP analysis. If counters rise, correct the cable, connector, interface, or negotiation fault and repeat the test. Changing connection behavior cannot repair damaged frames or intermittent carrier loss.

Does the failure occur before or after TCP opens?

Classify the first missing event. If the server never accepts the connection, inspect the destination address, configured listening port, access rules, listener backlog, and available sockets. If TCP opens but no Modbus response returns, inspect the request, server processing, and transaction matching. If the response returns and the PLC immediately closes, the data exchange succeeded; the remaining issue is connection policy and its cumulative load.

Observed sequence Meaning Decision
Open attempt with no established connection Failure is below the Modbus transaction layer Check reachability, filtering, listener state, and exhausted resources
Connection opens; request has no reply Transport exists, but transaction processing failed or stalled Check request validity, server diagnostics, and response timeout
Request and reply complete; PLC closes Connection-per-request policy is active Measure the cost of repeated opens and closes
Several queries share one connection Persistent-session policy is active Test recovery after an intentional disconnect

A Modbus reply and TCP close are separate events. Closing after the reply does not invalidate the completed transaction. It does discard the established transport context, so the next query must wait for another handshake. A valid single exchange therefore does not prove that rapid repetition will interoperate reliably with every server.

What do socket states and resource readings show?

Capture traffic at the client or server and correlate it with server diagnostics. Count connection attempts, established connections, completed request-response pairs, closes initiated by each endpoint, rejected opens, and retry intervals. On the server, trend active sockets, queued opens, task or thread count, CPU load, and memory use.

TCP implementations retain state after closure to prevent delayed packets from an old connection being mistaken for data in a new one. Repeated reuse can therefore accumulate sockets in states such as , consume temporary source ports, fill an accept queue, or reach a server's socket limit. A server that creates a task and allocates buffers for every connection can also suffer CPU loading and memory fragmentation even when every Modbus transaction succeeds.

Reading Healthy interpretation Churn warning
Connection attempts per second Tracks the intended polling demand Approaches the transaction rate because every query reconnects
Established socket count Stable for the configured clients Oscillates rapidly or reaches the server limit
Closed-state socket count Clears without affecting new opens Grows while reconnects begin to fail
CPU, memory, tasks, or threads Remain bounded during polling Rise with connection rate rather than request workload
Open latency Stable across the test Increases as queues or throttling engage

Aggressive clients have been observed attempting 50 to 100 opens and closes per second. Some servers deliberately penalize this behavior; a 1-second delay before accepting a replacement connection is one implementation example. Treat any such delay as a server-specific protective control, measure its effect on the PLC timeout and scan demand, and never insert it blindly.

Which connection policy resolves the active branch?

If physical transport, routing, and transactions are sound but resources rise with reconnect frequency, keep the TCP connection open across multiple Modbus queries. Reconnect only after an explicit communication error, peer closure, application shutdown, or an idle policy chosen from the endpoint documentation. Persistent operation removes repeated handshake and teardown work while preserving normal Modbus request-response processing.

If end users cannot control open and close behavior, the correction belongs in the PLC client's firmware or communication component. Expose a persistent-session option if compatibility requires retaining the existing mode. The server still needs bounded, graceful handling of clients that reconnect per query because deployed clients may not be configurable.

Setting choice Transport effect Use decision
Close after every reply Handshake, allocation, and teardown repeat for every transaction Retain only where an endpoint requirement makes it necessary
Keep connection open Multiple transactions reuse established transport state Preferred when server capacity and client recovery logic support it
Server-side open throttle Limits resource consumption from reconnecting clients Use as protection after validating client timeouts
Bounded server resources Caps sockets, workers, and buffers Apply regardless of client behavior

How should the correction be applied and verified?

  1. Record a baseline packet capture and server resource trend during representative polling. Mark each open, query, reply, close, retry, and rejected connection.
  2. Correct any physical errors, address mismatch, route failure, listener problem, or filtering issue found earlier in the decision tree.
  3. Change the PLC client to retain the established TCP connection for successive requests. If that behavior is fixed internally, update the firmware or communication component rather than asking operators to alter unavailable controls.
  4. Implement reconnect handling for server closure and communication failure. Prevent simultaneous retry loops from creating an uncontrolled connection storm.
  5. On the server, bound connection queues, socket allocation, worker creation, and per-client resource use. Add throttling only after measuring how its delay interacts with the client's timeout and retry behavior.
  6. Repeat the same workload. Confirm that multiple query-response pairs now traverse one connection, request latency remains acceptable, socket-state counts stay bounded, and CPU and memory return to stable levels.
  7. Force one controlled disconnect. Verify that the PLC detects the loss, opens one replacement connection, resumes matched Modbus request-response traffic, and then keeps that connection open.

FAQ

What happens if a Modbus TCP client closes after every reply?

The completed reply remains valid, but the next query requires a new TCP handshake. At high polling rates, socket states, queues, CPU work, and memory allocation can become the limiting factors.

What happens if the server runs out of sockets?

New connection attempts may be queued, delayed, rejected, or time out. Compare rejected opens and active socket counts with the server's documented limits while capturing the failed handshake.

preserves closed-connection state so delayed packets cannot contaminate a later connection. Rapid reconnects can accumulate that state and delay or prevent reuse of the same endpoint combination.

What happens if the server adds a 1-second reconnect delay?

The throttle can reduce CPU and allocation pressure, but it can also exceed the PLC's connection timeout or slow polling. Test the 1-second delay against the configured timeout and retry sequence before deployment.

How do I verify the Modbus TCP connection fix?

Capture the final test and verify that several matched requests and replies use one TCP connection, server resources remain bounded, and one forced disconnect produces exactly one successful reconnect followed by resumed traffic.

Back to blog