Testing ICS/SCADA Redundancy with Controlled Failures

Karen Mitchell6 min read
Best PracticesHMI / SCADAOther Manufacturer
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

The SCADA screen freezes, values turn stale, or alarms appear during a network failover even though controller logic continues to run. Treat that display as the first observation, then follow the data path through the tag, driver, routing layer, and redundant site. The objective is not merely to prove that backup equipment powers up; it is to prove that the complete system changes paths without creating an operationally significant interruption or conflicting active states.

What is the screen telling you?

Record the operator-visible symptom before changing anything. Capture the first stale value or communication alarm, the duration, the affected displays, and the recovery sequence. A brief loss across every display points toward a shared driver, server, gateway, or site transition. A problem limited to one display or tag group starts the investigation higher in the application stack.

Reading Location Meaning Next check
Display value and update time Operator workstation Shows when presentation stopped receiving current data Read tag quality and timestamp
Tag quality and source timestamp SCADA tag database Separates a stale source from a display binding problem Inspect the driver session
Connection and subscription state SCADA communication driver Shows whether the server retained or rebuilt its controller connection Check the active network path
Controller operating state Independent controller diagnostic path Distinguishes loss of supervision from loss of control Compare network and server event times

If the source tag remains current while the display freezes, the tag is right; the binding or presentation path is wrong. If tag quality drops while the controller remains operational, continue through the driver and network. In a train-control architecture, distinguish SCADA supervision from the signalling controllers: an interruption to SCADA does not by itself prove that controller execution stopped.

Did the failure remain inside one layer?

Correlate timestamps from the operator display, SCADA servers, communication drivers, switches, routers, and firewalls. Use one time basis across the records. Event order matters more than the presence of isolated alarms: the sequence shows whether the route disappeared first, the driver disconnected first, or a server changed role before connectivity was ready.

  1. Confirm whether the process controllers remained in their normal operating state.
  2. Compare the tag timestamp with the screen timestamp. If the tag continued updating, inspect display bindings and the client-to-server session.
  3. If the tag stopped, check whether the driver retained its connection or entered a reconnect cycle.
  4. If multiple drivers or servers lost connectivity together, inspect gateway ownership and route availability.
  5. If one site remained reachable but both sites claimed the active role, investigate redundancy arbitration and path convergence as one event.

This sequence prevents a network transition from being misdiagnosed as a server defect and prevents a server-role problem from being hidden behind a generic communication alarm.

Was the redundant gateway ready to forward traffic?

A gateway becoming active and a routed path becoming usable are separate events. In the documented legacy failure, a core switch activated VRRP before OSPF had finished establishing its routing table. The gateway address was present, but the required routes were not. Main and disaster-recovery sites then operated in a split-brain condition for 15 seconds.

Network reading Outcome Decision
VRRP ownership changes after required routes are installed The new gateway can forward traffic when it becomes active Continue to driver and SCADA role checks
VRRP ownership changes before OSPF convergence The virtual gateway is reachable without a complete forwarding path Coordinate gateway eligibility with route readiness
Routes converge, but application recovery remains slow The network is available while sessions or server roles are still recovering Measure driver reconnect and server arbitration

Two correction patterns can work. Route-aware gateway tracking ties eligibility to actual reachability and is preferred because it responds to state. Timer coordination delays gateway activation until measured routing convergence has completed; use it only when convergence is bounded and repeatable. Do not select a delay by guesswork—derive it from timestamped failure tests and retain margin for the slowest accepted transition.

Did the SCADA platform survive the network event?

The final design contains two sites with five servers per site, so network recovery alone is not the acceptance criterion. Record each server's role, peer visibility, client ownership, driver state, and data freshness before, during, and after every injected failure. A hot-standby active-active platform must preserve its intended role model while the network changes beneath it.

Check for a server that remains nominally active after losing the resources needed to serve current data. Also check whether both sites become active without the arbitration path required to prevent conflicting ownership. When connectivity returns, watch for repeated role changes, duplicate alarm generation, delayed history transfer, or clients remaining attached to the recovered but non-preferred endpoint.

Firewall behavior belongs in the same timeline. Traditional clustered firewalls may require matching firmware, while independent active devices that synchronize sessions separate the forwarding roles more cleanly. For the installed architecture, document the actual upgrade and failover constraints; the cited clustered arrangement could make an upgrade messy or cause approximately 15 minutes of outage at that site.

Which controlled failures expose the hidden dependency?

Build the test matrix from the dependency boundary inward. The commissioned test sequence covers enterprise or upstream C-network links and devices at each site, simultaneous failure of field A and B networks at both Main and DRS sites, and a hard power loss to the entire network at each site. These cases test different assumptions and should not be collapsed into one generic failover result.

Injected condition Primary observation Failure exposed
Enterprise or upstream C link/device failure Route, gateway, client, and external-service continuity Unexpected carrier properties or third-party routing dependencies
Field A and B network failure at both sites Driver quality, controller reachability, and server role stability Shared dependencies hidden by nominally separate field paths
Complete network power loss at Main Transfer to DRS and operator-session recovery Site-level arbitration, routing, and session defects
Complete network power loss at DRS Main-site continuity and clean DRS reintegration Standby isolation or uncontrolled role changes on restoration

Run disruptive tests only inside an approved operating window with explicit stop criteria and a recovery owner. A lab representing about 80% of the network can validate logic and expected transitions, but it cannot reproduce every production load, managed-service behavior, or third-party boundary. One third-party routing export produced a two-minute failover instead of the intended seconds; only an end-to-end test exposed the operational delay.

How do you correct and verify the resolving branch?

  1. Preserve the commissioned baseline, including routing, gateway, firewall, driver, and SCADA role settings.
  2. Inject one failure and correlate display, tag, driver, server, gateway, and routing timestamps.
  3. If gateway ownership precedes route readiness, apply route-aware tracking or a delay derived from measured convergence.
  4. If a third-party route or managed service controls recovery, reconcile the exported configuration and retest across that service boundary.
  5. If networking recovers first, tune the application recovery mechanism identified by the driver or server logs rather than masking it with a longer network timer.
  6. Repeat the same failure several times, then restore the failed component and test reintegration separately.
  7. Accept the case only when controller operation, tag freshness, operator displays, server roles, alarms, and histories all return in the required sequence without split brain.

FAQ

Why does SCADA freeze when the controllers are still running?

The supervisory path can lose its driver session, route, gateway, or server connection while controller logic continues locally. Compare controller state with tag quality and source timestamps to separate loss of supervision from loss of control.

Why does VRRP failover cause a split brain?

VRRP can activate a gateway before OSPF has installed the required routes. The observed sequence created a 15-second split brain, so gateway eligibility must follow verified route readiness.

Why does a redundant link take two minutes to fail over?

Check the actual service type and the third-party routing configuration at the boundary. A misaligned exported routing configuration caused a two-minute transition where recovery was expected in seconds.

Why does redundancy pass in the lab but fail in production?

The lab may omit production load, carrier behavior, and third-party routing boundaries even when it represents most of the network. Repeat the controlled production failure and verify the final operator-screen timestamp against tag freshness, driver recovery, route readiness, and the intended server roles.

Back to blog