Troubleshooting MQTT Engine Connect/Disconnect Loops

Daniel Price7 min read
Industrial NetworkingOther ManufacturerTroubleshooting
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

The request path starts at MQTT Engine, leaves the server through its network interface, crosses the routed network, reaches the broker listener, establishes an MQTT session, and publishes the host state. A connect/disconnect loop means one of those stages succeeds briefly and the next stage fails, or another client or broker policy tears down the session. Follow the packet and prove each hop in order.

Where does the MQTT path stop?

Map the configured path before changing software. Record values directly from MQTT Engine and the broker rather than relying on assumed defaults.

Path item Value to record Passing check
Source server Hostname and active interface address The expected interface carries broker traffic
Broker destination Configured hostname or IP address Name resolution returns the intended broker address
Broker port Port from the connection settings A TCP connection reaches that listener
Security setting Plain or encrypted transport, certificate settings, and credentials The broker accepts the transport and authentication exchange
MQTT identity Configured client identity No other connection replaces the session
Publish operation Host-state destination identified by configuration and logs The broker accepts the host-state publication
Failure timing Connect time, disconnect time, and interval between retries The session remains established beyond the former failure interval

The reported installation had failures on two of eight servers, and those two were the only servers running the 8.3 branch. That correlation makes a side-by-side configuration and traffic comparison useful, but it does not identify the failing layer. The issue began on 7/26, so correlate that boundary with network, broker-policy, certificate, credential, and server changes. The first checkpoint passes when the destination address, listener port, transport mode, and client identity are known for both an affected and an unaffected connection.

Does layer one and TCP remain stable?

Check the physical and operating-system path before interpreting MQTT messages. Inspect link state, interface errors, dropped packets, route selection, DNS results, firewall counters, and broker reachability from the affected server. Also check server disk space and resource exhaustion because a full volume or constrained process can interrupt logging, persistence, and module operation even when the network remains reachable.

  1. Confirm that the configured broker name resolves to the intended address from the affected server.
  2. Test TCP establishment to the configured broker port from that same server and interface.
  3. Capture traffic on the interface while MQTT Engine connects and disconnects.
  4. Check whether the server sends a TCP reset, the broker sends one, or the flow ends after retransmissions and timeout.
  5. Repeat the same capture on an unaffected server using its actual configuration, then compare handshake direction, retransmissions, and connection lifetime.

If no TCP handshake completes, stay below MQTT: correct routing, DNS, firewall, listener binding, or link faults. If TCP connects and then closes cleanly or resets after application data, continue upward. This checkpoint passes when the affected server can maintain the broker TCP connection without loss or unexplained resets.

Does the broker accept the MQTT session?

Once TCP is stable, determine whether the failure occurs during transport security, MQTT session establishment, or the first publication. With encrypted transport, a packet capture normally exposes TCP behavior and the security handshake but not MQTT packet contents; pair it with broker and MQTT Engine timestamps. With readable MQTT traffic, follow the connect request, broker response, publication, and any acknowledgement required by the configured delivery level.

Observed stopping point Likely fault domain Next check
Before MQTT session establishment Transport security, certificate validation, listener mode, or credentials Compare client and broker security logs at the same timestamp
Immediately after session establishment Duplicate client identity, broker session policy, or initial publish rejection Search broker logs for the client identity and disconnect reason
When publishing host state Publish authorization, invalid destination policy, or broker-side processing Check the exact destination against the broker access-control rule
After a repeatable interval Keepalive path, idle timeout, intermediary timeout, or resource pressure Compare the interval with configured and logged timers

The key symptom is trouble publishing the host state. A successful network connection therefore does not prove that the application path works. The checkpoint passes only when the broker accepts the MQTT session and the first host-state publication.

Are identity and publish permissions correct?

A broker can accept credentials yet deny publication to a particular destination. Review the broker authorization rule using the exact authenticated principal and host-state destination shown by the configuration or diagnostic logs. Check case, separators, prefixes, and wildcard scope exactly; access-control matching is commonly literal.

Also verify that the two affected servers do not share an MQTT client identity with another active connection. Brokers commonly permit one live session per client identity, so a second connection can displace the first and create a repeating connect/disconnect pattern. Broker logs decide this case by showing the same identity arriving from another source address near each disconnect.

  1. Locate the authenticated principal, client identity, and rejected publish destination in the engine and broker records.
  2. Compare them with a working server without copying identities between servers.
  3. Correct the broker rule if it denies the required host-state publication.
  4. Assign a distinct client identity if another live client uses the same value.
  5. Reconnect once and watch the broker record through the first state publication.

This checkpoint passes when one unique client session remains active and the broker records an accepted host-state publish.

How should the fix be applied?

Disabling and enabling MQTT Engine briefly restored the connection early in the incident, then stopped restoring it. Treat that action as a session reset, not a repair: it can clear transient state but cannot correct a denied publish, duplicate identity, broken route, invalid security configuration, or broker policy.

Apply only the change indicated by the stopping point. Correct the network path when TCP fails; correct transport-security settings when the handshake fails; correct credentials when authentication fails; correct identity collisions when the established session is displaced; correct the destination permission when the host-state publish is rejected. Change one fault domain at a time so the next capture proves causality.

Upgrading the server software from 8.3.6 to 8.3.8 and CirrusLink MQTT Engine from 5.0.3 to 5.0.4 did not stop the loop in this installation. Repeating those upgrades is therefore not the primary diagnostic step. Preserve the module, server, broker, and packet timestamps for escalation through official support channels if the broker accepts the publication but MQTT Engine still disconnects. This checkpoint passes when the targeted change removes the previously observed failure at the same protocol stage.

How is the repair verified end to end?

  1. Start a synchronized packet capture and collect MQTT Engine and broker logs.
  2. Enable the connection once.
  3. Confirm DNS resolution, TCP establishment, and any configured security handshake.
  4. Confirm that the broker accepts the MQTT session under the intended client identity.
  5. Confirm that the host-state publication is accepted rather than immediately followed by rejection or disconnect.
  6. Observe the connection beyond its former repeatable failure interval and check for retransmissions, resets, identity replacement, and new publish failures.
  7. Confirm downstream data flow through the same broker session.

The repair passes only when the capture, broker record, and MQTT Engine log agree that one session stayed active and completed the host-state publication.

FAQ

How do I find where an MQTT Engine connection loop stops?

Capture the path in order: name resolution, TCP handshake, security handshake, MQTT session establishment, and host-state publication. Match packet timestamps with MQTT Engine and broker logs.

How do I tell whether the broker is rejecting the host state?

Find the host-state destination and authenticated principal in the diagnostic records, then compare them with the broker access-control rule. A completed MQTT connection followed by failure at the first publish points to authorization or destination policy.

How do I detect a duplicate MQTT client identity?

Search broker logs for the same client identity connecting from another source near each disconnect. Configure a distinct identity for every simultaneously active client.

How do I use a packet capture when MQTT traffic is encrypted?

Use the capture to identify TCP establishment, retransmissions, resets, and security-handshake failure. Correlate those timestamps with broker and MQTT Engine logs to locate the MQTT-level rejection.

How do I verify the MQTT Engine reconnect-loop fix?

Reconnect once, confirm broker acceptance of the session and host-state publication, then observe beyond the former failure interval. Finish by confirming downstream data flow through the same continuously active session.

Back to blog