When an RTU slave goes offline, queued requests can each consume the configured timeout, delaying other nodes and eventually filling the client’s transmit queue. Use the token to correlate asynchronous callbacks, and control how many requests you place in the queue for a slave that is not responding.
What does the token identify in the callback path?
The client application calls MBclient.addRequest(...), which places a request in the client’s queue for transmission to the selected remote ID. The RTU bus carries the request to the addressed slave; a valid reply or an error is then delivered asynchronously to the client’s handler. The token provides the application a way to identify which request produced that callback.
Use a distinct token for each request whose result must be distinguished from other outstanding requests. If multiple requests use the same token, the protocol exchange may still occur, but the application cannot reliably tell those results apart using the token in onDataHandler or onErrorHandler. A token is an application-level correlation value; it does not make an offline slave reply or change RTU bus timing.
| Value | Role in this path | Engineering implication |
|---|---|---|
token |
Correlates asynchronous result with a request | Use distinct values for requests requiring separate identification |
modbus_RemoteId |
Selects the addressed remote slave | Check it against the slave’s configured unit ID |
WRITE_HOLD_REGISTER |
Selects the Modbus operation | The callback still reports asynchronously |
10 - 1 |
Register argument in the shown call | Confirm the library’s addressing convention and the slave’s register map |
Where does an unanswered request stop the data path?
A master sends a request over the serial link and waits for the addressed slave. Modbus RTU serial traffic is serialized: a client cannot usefully advance another transaction on the same bus while the current exchange is unresolved. A slow slave may need time to respond, so a client must wait for either a response or the timeout decision before treating the request as failed.
For a missing response, the client reports an E0 TIMEOUT error through the error handler. A malformed response produces an error as well. The error callback is therefore part of normal transaction handling, not an optional diagnostic hook. It lets the application record failure and decide whether to retry, suppress future polling, or mark the remote unavailable.
Trace failures from the physical link upward. First check that the serial wiring and interface are intact and that the slave is powered and connected. Then verify the remote ID, serial configuration, function/register arguments, and response integrity against the slave configuration and client diagnostics. A timeout indicates no valid response reached the client before its timer expired; it does not by itself distinguish a disconnected device from a configuration or link fault.
Why does one offline slave delay other nodes?
Each request already in the client queue can incur its own wait for a response or timeout. If the application continues adding requests for an offline slave, the queue accumulates work that cannot complete successfully. The client reports timeouts, but requests already queued still have to be processed; the failure of one transaction does not automatically remove every later request for that slave.
In the described setup, the timeout was configured as DEFAULTTIMEOUT 2000 milliseconds, requests were added every 5 seconds, and 10 requests had accumulated for an offline slave. If those 10 queued requests each consume the 2,000 ms timeout in sequence, the delay is approximately 20 seconds before that backlog is cleared. The queue’s default maximum was reported as 100 requests. Once full, requests for reachable nodes can also be blocked from entering the queue.
| Observed condition | Mechanism | Check or response |
|---|---|---|
| Offline slave produces timeout callbacks | No valid reply arrives before the request timeout | Inspect the link and slave configuration; handle the error callback |
| Repeated polling continues during outage | New requests accumulate behind requests awaiting timeout | Stop or suppress polling for that remote while it is unavailable |
| Other nodes stop being serviced | Shared queue reaches its maximum capacity | Reduce queue growth and monitor queue use and callback results |
| Several seconds of delay after reconnection or recovery | Earlier timed-out requests still occupy processing time | Limit the number of requests outstanding per slave |
Which queue-control approaches fit this case?
The practical choices are to reduce how long the client waits, stop adding work for a failed remote, or continue queueing and accept the resulting delay. A shorter timeout can improve recovery time, but it also increases the risk of classifying a slow but healthy slave as failed. Suppressing polling avoids unbounded accumulation but requires the application to decide when to resume it.
| Approach | Benefit | Cost or limit |
|---|---|---|
| Keep adding requests at the normal interval | No application state change | Offline node requests consume timeout time and can fill the shared queue |
| Set a shorter timeout | Reduces time spent waiting on each unanswered request | Must still allow the slowest expected valid server response |
| Mark the slave offline and stop adding its requests | Prevents further backlog for that node | Application needs a recovery or re-poll policy |
| Remove a server’s queued requests with a dedicated service call | Would clear that node’s backlog directly |
dropRequestsForServer() was discussed as a possible future service call, not an available capability to rely on |
For this pattern, prefer application-level suppression: issue a request, process its callback, and stop scheduling new requests for a remote that has timed out until a defined recovery check permits polling again. Tune the timeout only after measuring the slowest expected response; do not shorten it merely to mask a growing queue.
How should request scheduling and error handling work?
- Assign a distinct token to each request that needs to be identified independently in a callback.
- Add the request for the intended remote ID and operation, then wait for either the data callback or error callback before deciding what to do next for that remote.
- In
onDataHandler, associate the returned result with its token and mark the remote responsive. - In
onErrorHandler, associate the error with its token. ForE0 TIMEOUT, mark the remote unavailable and stop enqueuing its normal polling requests. - Keep requests for other nodes from being starved by avoiding repeated additions for the failed remote. Define a controlled recovery check rather than continually adding requests during the outage.
- Set the client timeout to a value that permits expected slave response latency. Read the configured timeout and queue limit from the actual client/library configuration rather than relying on defaults across deployments.
Do not treat “one request per polling interval” as a guarantee of one outstanding request. If the interval is shorter than the time needed to finish or timeout a transaction, or if the application queues work independently of callback completion, multiple requests can accumulate. Track in-flight work per remote or schedule its next request only after handling the previous result.
How can you verify the fix on the bus?
Test with one reachable slave and one unavailable slave while observing callback tokens, timeout events, queue occupancy, and service to the reachable node. Confirm that each successful response maps to the expected token and that a missing or malformed response is handled through the error callback. If repeated identical requests are queued, verify the application can distinguish their callbacks or prevent the overlap.
With the unavailable slave, confirm the first timeout changes its application state and prevents normal polling requests from continuing to accumulate. The reachable slave should continue receiving service, and queue usage should remain below its configured maximum. After restoring the offline slave, run the recovery check and confirm valid data callbacks resume before normal polling is re-enabled. Finally, compare measured response time for the slowest valid slave with the configured timeout and verify that the timeout still leaves enough response margin.
FAQ
Can I reuse the same token for multiple Modbus RTU requests?
Only if the application does not need to distinguish their asynchronous results. Use different tokens for requests that must be correlated individually in onDataHandler or onErrorHandler.
Does the client retry one unanswered request forever?
The described behavior reports E0 TIMEOUT for a missing response. Continued queue growth occurs when requests keep being added; handle the timeout and decide whether to stop scheduling or retry.
Can a 2-second timeout delay other Modbus RTU nodes?
Yes. Ten sequential queued requests each consuming the described DEFAULTTIMEOUT 2000 can take about 20 seconds; verify the fix by confirming other nodes remain serviced and queue occupancy stays below its configured maximum.