Resolving Ignition OPC UA Faults When a Tag Value Changes

Daniel Price7 min read
OPC / OPC UAOther ManufacturerTroubleshooting
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

This Ignition OPC UA connection shows Connected and serves live values, but it faults for about 2 seconds each time the device publishes a new value. It recovers by itself, and the tag shows the new value only after the fault. Roughly 1 in 10 times it also faults between value changes. Raising Keep Alive Failures Allowed made no difference. That result points to the cause. A keep-alive counter only tolerates missed keep-alives. It does not tolerate a failed service call, a rejected message, or a closed secure channel. Each of those faults the connection on the first occurrence. The new value appears only after recovery, so the notification that should carry it never reached the tag. The value arrived through the initial read or subscription that Ignition rebuilds on reconnect.

Which hops does the value cross before it fails?

In OPC UA the client pulls data. The Ignition gateway queues Publish requests, and the device server answers one when a monitored item changes or when a keep-alive is due. Whichever hop drops or rejects that Publish response is where the fault begins.

Hop Carries Failure that produces a ~2 s fault
Device application to its OPC UA server stack Value update into the address space Server stalls while it processes the update, so the Publish response is late
Server stack to TCP Publish response (data change notification) Message is malformed, too large, or the server closes the channel
Switches, firewalls, routing TCP segments Drops, retransmissions, resets
Gateway OPC UA client Decode and apply the notification Client rejects the message (encoding limit or decode error) or the request times out
Ignition tag Final value Tag gets the value only after reconnect

Other devices from the same vendor never fault. That shifts suspicion toward what differs on this unit: firmware, server configuration, or the data type and size of the changing tag. A string or array that grows is a common cause. Collect the facts below before you change any settings.

Is the network path clean when the fault occurs?

Check the physical layer first. It takes minutes and removes a whole class of causes.

  1. Read the error counters on the switch port that serves the device: CRC errors, drops, and duplex. A duplex mismatch shows up as late collisions or FCS errors, and it gets worse under bursty traffic.
  2. Run a continuous timestamped ping from the gateway host to the device. Force or wait for a tag change and watch the ping during the fault window.
  3. List every firewall or inspection device between gateway and device. Check its logs for session drops at the fault timestamps.

Check: if ping keeps answering at normal latency through the fault and the port counters stay flat, the network is not the cause. The fault is in the OPC UA session layer on one of the two endpoints.

What does the gateway log record at the fault timestamp?

The connection status only tells you it faulted. The gateway log tells you why.

  1. In the gateway log configuration, set the OPC UA client loggers to DEBUG. Use TRACE only for a short capture window.
  2. Trigger a value change on the device and note the exact time.
  3. Find the first error or warning at that time. Record the exact OPC UA status code or exception text.
Log content at fault time Points to
Request or service timeout Server stalled and did not answer within the request timeout
Decoding error or encoding-limit exceeded Notification payload is malformed or exceeds the client's message, string, or array limits
Secure channel closed or error message received from server Server-side fault or security/channel renewal problem
Connection reset / socket closed TCP-level teardown; confirm which side sent it in a capture

Check: give the vendor the exact status code and timestamp. A server can log nothing and still send a response the client rejects, so an empty server log does not clear the device.

What does an unsecured Wireshark capture show?

An encrypted session hides the payload. To see the actual messages, reconfigure the connection to use the device's no-security endpoint, with security policy None and message mode None, for the duration of the test.

  1. Confirm the device exposes a None endpoint. If it does not, ask the vendor to enable one temporarily.
  2. Point the Ignition connection at that endpoint and confirm it connects and the fault still reproduces.
  3. Capture on the gateway host, filtered to the device IP and the port in the endpoint URL. 4840 is the registered OPC UA default, but read the actual port from the URL. Apply the opcua display filter.
  4. Trigger a value change. Stop the capture after the connection recovers.
  5. Find the last Publish response before the teardown and note which side sends the close, error, or TCP reset first.
Capture pattern Meaning Owner
Publish response arrives, then the gateway closes the channel Client rejected the notification (size limit or decode failure) Raise client limits, or have the vendor fix the encoding or payload size
Publish requests outstanding, no response until the gateway gives up Server stalled during the value update Vendor; widen the request timeout as a workaround
Server sends an OPC UA error message or TCP reset Server-side fault Vendor, with the capture attached
TCP retransmissions or zero-window from the device Device CPU or buffer starvation, or network loss Vendor or network

Check: if the faults stop completely with security set to None, the problem is in the secured channel. Suspect secure channel token renewal or security processing on the device, and compare the fault times against the channel token lifetime. This also fits the faults that occur between value changes.

Which Ignition settings actually move the fault?

Match the setting to the failure you captured. Changing settings at random only hides the evidence.

Captured cause Ignition-side change Effect
Server responds late Increase the connection's request timeout above the observed stall Connection rides through the stall; the value arrives late but without a fault
Client rejects a large notification Raise the maximum message, string, or array size limits in the connection's advanced settings Notification decodes and applies without a reconnect
Subscription path on the device is unstable Move the affected tags to a separate tag group and switch its OPC data mode from subscription to polled read Removes the device's Publish path from the loop, at the cost of polling load
Missed keep-alives only Keep-alive interval and failure count Helps only when the log shows keep-alive timeouts, which does not match this failure

Change one setting at a time. Keep each change only if the soak test below passes. Otherwise revert it and send the capture to the vendor. Compare firmware and server configuration against the sibling devices that never fault, because that difference is usually what the vendor has to fix.

How do you prove the connection is stable end to end?

  1. Restore the production security policy if you ran the test on None, unless the None test was the fix itself.
  2. Leave gateway logging at DEBUG and run a soak test covering at least ten value changes. Ten matters because the between-change faults showed up at about 1 in 10.
  3. Confirm the connection status never leaves Connected and the log shows no channel close or timeout.
  4. For each change, confirm the tag updates when the device publishes, not after a reconnect. Compare the tag timestamp against the device-side change time.
  5. Take one final capture during a value change. It should show the Publish response carrying the new value followed by the next Publish request on the same channel, with no close or reset.

FAQ

Does raising Keep Alive Failures Allowed stop OPC UA faults on data change?

No. That setting only tolerates missed keep-alives. A rejected notification, a request timeout, or a closed secure channel faults the connection on the first occurrence regardless of the keep-alive count.

Can I run Wireshark on an encrypted Ignition OPC UA connection?

You can capture it, but the payload is unreadable. Switch the connection to the device's None security endpoint for the test, then filter with opcua on the endpoint port.

Why does the tag update only after the OPC UA fault clears?

The data change notification carrying the new value failed, either because the client rejected it or the server stalled. The value reaches the tag through the initial read or subscription that Ignition rebuilds on reconnect.

Can I switch an Ignition tag group from subscription to polled reads to avoid the fault?

Yes, as a workaround. Put the affected tags in their own tag group with polled/read data mode, which bypasses the device's Publish path. Verify the added polling load on the device before you leave it that way.

Does a clean device error log rule out the OPC UA server?

No. A server can send a late, oversized, or malformed response without logging anything. The gateway log status code and a packet capture show which side closed the channel.

Back to blog