The operator sees an analog value disappear, freeze, or jump when the primary controller fails. With two 1769-L18ER-BB1B controllers, the practical target is a controlled standby scheme that preserves selected data and redirects SCADA after a detected failure. It is not controller redundancy: the process may stop, an update can be incomplete at failure, and the switchover can create a bump.
What must remain visible after the primary fails?
Define “analog values” before building the failover. Live analog inputs, operator setpoints, recipe values, calculated states, and totals require different handling. A live input is a measurement owned by I/O hardware; copying its last value preserves a snapshot but does not keep measuring the process. A setpoint or recipe value is controller data and can be replicated to the standby controller or recorded outside both controllers.
| Value class | What failure removes | Suitable retention path | Commissioning check |
|---|---|---|---|
| Live analog input | The primary controller's current I/O update | Independent standby access to the field signal, or hold the last replicated sample | Disconnect the primary and confirm whether the standby receives a new measurement or only retains the last one |
| Operator setpoint | The primary controller's current memory value | Produced/consumed tag transfer to standby memory | Change the setpoint and confirm the standby copy updates before failure testing |
| Recipe value | The active recipe stored only in one controller | Peer replication, historian, or data-recording software | Load a recipe, remove the primary, and compare every required field |
| Calculated value or accumulator | The latest execution state | Replicate the state needed for restart | Compare primary and standby values over the operating range |
Create a retention list containing only the data required after a failure. For each item, record its owner, update direction, acceptable age, and whether the standby must continue acquiring it. The check is complete when every requested value is classified as a live measurement or retained controller data.
How should the two controllers exchange retained data?
Use produced/consumed tags to transfer the selected data between the controllers. Group related values into a defined memory structure so the standby receives a coherent set of setpoints, recipe fields, and state values. Do not treat the communication link as proof that both controllers contain the same scan-level state; a failure can occur while an update is in progress.
Assign one writer for each replicated value. Under normal operation, controller 1 owns the process data and controller 2 copies the received data into a standby memory area. Avoid logic that permits both controllers to write the same command or setpoint, because a restored link can otherwise create an ownership conflict.
- Create a source data group in controller 1 containing the values from the retention list.
- Produce that group for controller 2.
- Create the matching consumed data in controller 2.
- Copy or move valid received values into a separate standby memory group. Keeping received and accepted values separate prevents a communication transition from immediately overwriting the retained copy.
- If controller 2 can later become the owner, provide an equivalent transfer path back to controller 1 for recovery. Gate that path with the active-controller state.
Track a monotonically changing heartbeat or sequence field with the data group. The standby can then distinguish a healthy connection from a value that remains numerically unchanged. The check passes when the standby data and sequence follow the primary during normal operation, then hold their last accepted state when the link is removed.
How does each controller decide that communication has failed?
Build an explicit health state in each controller. Use a GSV-based diagnostic or the controller's communication status information to evaluate the peer connection, and combine that result with the changing heartbeat. Connection status alone can report a configured path while the application data is no longer advancing; heartbeat logic alone can mistake a paused task or communications interruption for a failed controller.
| Observed condition | Likely interpretation | Required response |
|---|---|---|
| Connection healthy and sequence changing | Peer and data exchange are active | Primary retains ownership; standby refreshes its memory copy |
| Connection healthy but sequence stopped | Application execution or update path has stopped | Declare data stale using the project-approved detection interval |
| Connection faulted | Controller, network path, switch, cable, or power may have failed | Hold the last valid data and evaluate standby activation |
| Both controllers report themselves active | Ownership arbitration has failed | Block outputs and require a defined recovery action |
Select the failure-detection interval from the process response requirement and measured network behavior; no interval is established for this installation. Read the connection diagnostic and trend its transitions during commissioning rather than guessing. A short interval can cause nuisance transfers during network disturbances, while a long interval extends process downtime.
The check passes when stopping the heartbeat, breaking the controller path, and removing primary power each create the intended health state without falsely declaring both controllers active.
What makes the standby active without creating two masters?
A controlled standby needs ownership arbitration, not just duplicated logic. The standby must remain unable to command the process while the primary is healthy. When the primary fails, the standby can enter an active state only after peer health is lost and the external ownership conditions permit it.
A controller-to-controller message cannot by itself resolve every failure. Loss of communication leaves each controller unable to tell whether the peer CPU failed or only the network path failed. If both retain access to outputs, each can conclude that it should run. Use an external ownership signal, switched communication path, or other hard interlock appropriate to the process so only one controller can command outputs.
- Define explicit standby, active, transfer-pending, and faulted states.
- Permit controller 1 to be active only while its ownership input or equivalent arbitration condition is present.
- Permit controller 2 to become active only after primary health is lost and controller 2 receives ownership.
- Block all output commands when ownership is absent or contradictory.
- Require a deliberate recovery sequence before transferring control back to a restored controller.
A proposed hardware arrangement used an “online” output to operate a Black Box SW1040A, held the secondary logic paused until the primary output failed, and reported approximately 300 ms of downtime. Treat that as one installation-specific design result, not a guaranteed transfer time. Its behavior depends on how the switch, ownership output, controller logic, I/O, and SCADA paths interact.
The check passes only when simulated CPU loss, network loss, power loss, and restoration cannot produce two output owners.
How should field analog signals reach the standby?
Replicated tags preserve the last known number; they do not create a redundant measurement path. If the standby must continue controlling from current analog inputs, it needs a valid electrical and I/O path to those signals. The proposed arrangement of wiring I/O to both controllers requires an engineering review of input impedance, signal loading, isolation, commons, fault current, and the field device's permitted load. Do not parallel outputs from two controllers.
| Configuration | What it provides | Main limitation |
|---|---|---|
| Produced/consumed value only | Last accepted analog sample in standby memory | The value becomes stale after primary or communication failure |
| Signal connected to both input paths | Both controllers may observe the field measurement | Electrical loading and isolation must be validated for the actual transmitter and input modules |
| External historian or recorder | Independent history and recovery of values | Does not make the standby controller capable of controlling the process |
Both peer replication and an external historian work for data preservation, but they solve different problems. Use peer replication when the standby needs the latest controller state for restart. Use a historian or data-recording system when the requirement is to retain and review values even if both controllers or their common network are unavailable.
The check passes when each live signal is electrically measured at both intended input paths without changing the field reading, and retained-only values are clearly identified as stale after loss of updates.
How should SCADA address both controllers?
The tag is right; the binding is wrong if the screen points only to controller 1. SCADA cannot use one controller-specific tag path and automatically infer that controller 2 now owns the process. Configure separate controller connections and separate tag references, then select which reference the display and commands use according to a verified active-controller status.
- Create one SCADA communication path for controller 1 and another for controller 2.
- Expose the same logical data set from both controllers, using distinct SCADA references.
- Expose active, standby, communication-health, and data-valid status from each controller.
- For display-only objects, use a global object or equivalent visibility logic to show the active controller's value and hide the inactive value.
- For commands, gate writes so the object writes only to the controller that has confirmed ownership. Visibility alone is not command arbitration.
- Display stale or unavailable quality when neither controller has valid ownership and current data.
Two SCADA configurations can work. Separate visible objects are direct and easy to diagnose but require more engineering. A reusable global object reduces repeated screen work but must pass both controller references and the active-state selection correctly. In either case, the screen must never show an old standby value as a current process measurement.
The check passes when the displayed value, quality indication, and command destination all move together during a controlled ownership transfer.
What infrastructure remains a single point of failure?
Two CPUs do not make a redundant system when they share one power source, one switch, one cable route, one I/O path, or one SCADA communication path. Map every shared dependency before presenting the design as improved availability.
| Dependency | Failure effect | Commissioning decision |
|---|---|---|
| Power feed | Both controllers can stop together | Provide independent feeds if common power loss is in scope |
| Controller network | Peer synchronization and SCADA access can disappear together | Separate paths where the availability requirement demands it |
| Physical routing | One incident can damage both paths | Separate controller placement and cable routing where practical |
| I/O hardware | The standby can run but cannot observe or command the process | Define which I/O failures the design must tolerate |
| SCADA server or driver | Both controller paths can remain healthy while the operator loses control | Test the complete display and command chain |
If the requirement is full, bumpless redundancy, migrate to a controller platform and architecture designed for redundancy rather than presenting custom logic around the 1769-L18ER-BB1B as an equivalent. The custom standby design can preserve selected data and automate a controlled restart, but the stated process may stop and the transfer may cause an upset.
The check passes when the failure-mode list identifies every common dependency and states whether that failure is tolerated, detected only, or outside the project scope.
How is the complete failover verified?
Test from the operator screen backward through SCADA, controller ownership, peer data, and field I/O. Record the last valid primary sequence, the first valid standby sequence, the displayed value, command destination, output ownership, and process response for each case.
- Run controller 1 as active and confirm controller 2 is receiving current replicated data while its outputs remain blocked.
- Change each retained setpoint or recipe field and compare both controller copies.
- Apply known changes to live analog inputs and confirm whether controller 2 receives independent measurements or only replicated samples.
- Interrupt only the peer communication path. Confirm that the data becomes stale, ownership arbitration prevents two masters, and SCADA reports the correct quality.
- Remove primary controller power. Confirm that controller 2 activates only through the defined ownership path, uses the last accepted retained data, and routes SCADA commands to the active controller.
- Fail each shared infrastructure component identified in the dependency table and verify the documented response.
- Restore controller 1. Confirm that it does not seize control automatically or overwrite newer active data.
- Execute the approved transfer-back sequence and verify one output owner, current analog data, correct SCADA binding, and a changing heartbeat end to end.
FAQ
Why does the analog value freeze when the primary CompactLogix fails?
Produced/consumed transfer preserves the last accepted sample, but it cannot refresh a live measurement after the primary or its communication path fails. Continued measurement requires a validated field-signal and I/O path to the standby.
Why does SCADA not switch automatically to the second controller?
SCADA tags resolve through configured controller paths. Create separate paths and references for both controllers, then select the display and command destination from confirmed active-controller and data-valid states.
Why does communication loss risk making both controllers active?
Each controller can interpret a broken network as failure of its peer. Use external ownership arbitration or a hard interlock so loss of communication cannot grant output authority to both controllers.
Why does a two-controller failover still cause a process bump?
The 1769-L18ER-BB1B arrangement is a controlled standby design, not true controller redundancy. Detection, ownership transfer, SCADA rebinding, I/O state, and partially updated data can interrupt execution; verify the final state by confirming one output owner, current analog data, correct SCADA binding, and a changing heartbeat.