Here is what the panel shows. Every ControlLogix value on the Perspective screens is populated and looks normal, but none of the values are changing. Alarms tied to process limits stay quiet. The ControlLogix device connection on the active gateway intermittently goes to Idle. Reconnecting the device does not recover it. Disabling and re-enabling the device does not recover it either. The only action that restores live data is failing over to the redundant gateway.
In one mission-critical installation this state ran for two days before anyone noticed, because the screens never looked broken. The natural next request is a Perspective button that lets shift operators force failover without opening the gateway web interface.
You can build something close to that button. First confirm that a button is the right fix. Work the checks below in order.
Check 1: Prove the Data Is Stale, Not the Screen
Start here. A value that is displayed is not a value that is updating.
- Pick a tag that always moves, such as a running counter, an analog that is never perfectly flat, or a PLC clock value.
- Open the tag in the Designer Tag Browser or on a diagnostics view. Read its value, quality, and timestamp.
- Compare the timestamp to the current gateway time.
| Reading | Meaning | Next check |
|---|---|---|
| Timestamp is current and the value moves | Data is live. The screen binding or view is the problem. | Stop here. Debug the view, not the gateway. |
| Timestamp is frozen and quality is Good | Subscription is stalled with no quality change. This is the dangerous case and matches the two-day miss. | Check 2 |
| Timestamp is frozen and quality is Bad or Uncertain | The driver has flagged the loss. Alarms on quality should already fire. | Check 2 |
If values freeze while quality stays Good, a quality-based alarm will never catch it. That is why you need a heartbeat, which is covered in the verification section.
Check 2: Read the Device Connection Status on the Active Gateway
Open the gateway web interface on the active node. Go to the device connections status page and find the ControlLogix device.
- Connected while tags are frozen: the driver believes its session is healthy. The stall is in the request and response path, not in connection establishment. Go to Check 3.
- Idle or cycling between states: the driver has no working session to the controller. Note the time. Pull the gateway logs filtered on the device and driver loggers for that window. Then go to Check 3.
- Disabled or Faulted: read the fault text in the device diagnostics. A configuration or addressing fault is a different problem from an intermittent stall.
Record the exact status string, the timestamps, and the log entries every time the fault occurs. You need this record for Check 4 and for any support case.
Check 3: Locate the Stall on the Gateway Side or the PLC Side
Failover fixes the problem, and a device disable/enable does not. That points at state held inside the active gateway's driver instance that a device restart does not clear. When the backup gateway takes over, it opens brand-new connections to the controller from a fresh process. Anything stuck in the old process is left behind.
The PLC side can still be the trigger. It can refuse or drop connections that the driver then fails to recover from cleanly. Separate the two cases:
| Symptom | Likely cause | Test that decides it |
|---|---|---|
| Device shows Idle, and disable/enable does not recover it, but failover does | Driver state inside the active gateway process is not fully released on device restart | On the next event, check whether the backup gateway could connect to the same PLC at that moment. If it can, the PLC is accepting connections and the stuck state is on the gateway. |
| Idle events line up with PLC-side activity such as downloads, online edits, or other HMIs and historians connecting | Controller or Ethernet module running out of connection resources, or communications overhead starving the requests | Open the Ethernet module web diagnostics. Read the active connection count and error counters during an event. Compare them against the module datasheet limits. |
| Values frozen, quality Good, device shows Connected | Subscription or scan stall with no connection loss | Watch the device diagnostics request and response counters. If they stop incrementing, the driver is not polling. |
| Events cluster at specific times of day | Network maintenance, switch or firewall session timeouts, backups | Correlate the gateway log timestamps with network device logs. |
Do not keep re-trying disable/enable on the device. It has already failed, and it burns time during an event.
Check 4: Pull the Ignition Version and the PLC Firmware
Get these two facts before you build anything:
- The Ignition version on both redundant nodes, from the gateway status page.
- The ControlLogix controller firmware and the Ethernet module firmware, from the module web page or from Studio 5000.
The Logix driver has received many fixes across recent Ignition releases. An intermittent connection stall on an older gateway may already be fixed. Check the release notes for your version against driver fixes. If you are behind, upgrading is the real fix, and a failover button is a patch over a known defect.
Third-party EtherNet/IP driver modules for Ignition also exist. Swapping the driver is a legitimate way to test whether the stall follows the driver or follows the PLC. Treat it as a controlled test with a rollback plan, not an emergency change.
Decide Whether to Build a Failover Button at All
Ignition has no documented scripting call that triggers a redundancy failover. No built-in function exists that you can call from a button event to say "hand off to the backup now."
Two facts shape any workaround:
- Perspective scripts already run on the gateway. You do not need to send a message from the button to the gateway. The component event script executes in the gateway process that serves the session, which is the active node.
-
The only lever is to take the active gateway down.
system.util.executeruns OS commands in a separate process. The Gateway Command-line Utility,gwcmd, provides basic gateway commands, including restart and shutdown. Restart or stop the active node, and the backup takes over when it detects that the master is gone.
Be clear about what this does. It is a sledgehammer used to swat a fly:
- It kills every Perspective session, every gateway script, and every other connection on that node, not just the stuck device.
- The script that issues the command dies with the gateway. You get no confirmation back on the button.
- Whether the master takes control back when it returns depends on your redundancy recovery configuration. Read that setting before you test.
- It depends on the OS allowing the gateway service account to execute
gwcmd. Hardened hosts may block it.
If root cause work is active and stale values are unacceptable in the meantime, a guarded restart button is a defensible stopgap. If you can upgrade or fix the connection issue within days, skip the button.
Build the Guarded Restart Button
Use this procedure only as an interim measure on a mission-critical system.
-
Confirm the path. On each redundant node, find where
gwcmdlives in the Ignition install directory. Run it by hand from an OS shell with its help option. Read the documented option for restart and for shutdown. Use the exact option from your version's help output. -
Confirm permissions. Run the same command as the account the Ignition service runs under. If it fails there, it will fail from
system.util.executeas well. - Restrict the button. Gate the component with Perspective security levels or roles so that only authorized operators can see and fire it.
- Add a confirmation step. Open a popup that states which gateway will restart, and require a second press. Log the user, time, and reason to a database or audit table before the command fires, because nothing runs after it.
- Write the event script. Keep it minimal:
# Perspective button onActionPerformed (runs on the active gateway)
# Replace the path and option with the values confirmed in step 1.
import system
logger = system.util.getLogger("ManualFailover")
user = self.session.props.auth.user.userName
logger.warn("Manual failover requested by %s" % user)
# Write the audit record here, before the restart.
system.util.execute(["<path-to-gwcmd>", "<restart-option-from-gwcmd-help>"])
- Test during a maintenance window. Fire the button with the backup healthy. Time how long sessions take to reconnect on the backup and how long the backup takes to reach its active state.
- Document the operator action. Write down when to press the button (a stale-data alarm is active and confirmed) and what to check afterward.
Do not wire this to an automatic trigger. A heartbeat glitch that restarts the active gateway unattended can cause an outage bigger than the one you are trying to prevent.
Verify Detection, Failover, and Recovery
The two-day miss was a detection failure first. Fix detection whether or not you build the button.
- Heartbeat tag: Add a free-running counter or a toggling bit in the ControlLogix program. Read it through the same device connection as the process tags. Alarm when its value has not changed within a window longer than your scan class allows. Quality-based alarms will not catch a frozen value that still reads Good.
- Per-PLC stale alarm: Keep the stale-data alarms you already added. Confirm that each one fires by disabling the device on a test node and watching the alarm trip.
- After a failover: On the new active node, confirm that the device status reads Connected, the heartbeat is incrementing, and tag timestamps are current. Then check the redundancy status page and confirm the old master is back as backup, or still down, as your recovery configuration intends.
- Capture evidence: Save the gateway logs from the failed node, covering the minutes before the Idle state. That log is what closes the root cause.
FAQ
How do I force Ignition gateway failover from a Perspective button?
No documented failover call exists. The only lever is to restart or stop the active gateway with system.util.execute running gwcmd, which lets the backup take over. Lock the button behind security roles and a confirmation popup, and write an audit record before the command runs.
How do I detect stale ControlLogix values in Ignition when quality still shows Good?
Add a heartbeat counter in the PLC program, read it over the same device connection, and alarm when it stops changing. Also compare tag timestamps against gateway time, because a frozen value with Good quality will not trip a quality alarm.
How do I tell whether the stall is in the Ignition driver or in the PLC?
During an event, check whether the backup gateway can connect to the same PLC, and read the Ethernet module's connection count and error counters. If the backup connects cleanly while the active node stays Idle, the stuck state is in the active gateway process. Compare your Ignition version against driver fixes in the release notes.
How do I know when to escalate to Inductive Automation support?
Escalate once you have the Ignition version on both nodes, the controller and Ethernet module firmware, and gateway logs covering at least one Idle event, and a device disable/enable still does not recover it. Open the case through Inductive Automation's official support channel with that data attached. If module diagnostics show connection exhaustion or faults, bring in Rockwell Automation support in parallel.