Automated containment that is acceptable on an IT network becomes a process-safety event when it acts on a controller. Revoking a compromised session costs one user about an hour of work. Isolating a controller that was flagged as compromised can black out a substation. This walkthrough builds a containment design in which automation handles detection and decision preparation, and a person who knows the plant state executes every action that can change a physical process. Each section adds one element and ends with a check that proves it before the next.
Classify each containment action by reversibility
Blast radius is the set of physical and business outcomes that follow if an action fires on a false positive. Reversibility is whether the pre-action state can be restored with no loss beyond time. "Contain" covers actions with very different values of both, so classify the verbs individually instead of the domain as a whole.
| Action | Domain | Wrong-call consequence | Execution |
|---|---|---|---|
| Revoke a user session | IT | One user loses about an hour of work | Automate |
| Re-image a host | IT | Recoverable rebuild | Automate |
| Cut the OT segment from upper layers | OT boundary | Loss of telemetry and work-order communication; the running operation continues if control is local | Human decision, lowest OT impact |
| Isolate a controller or I/O module | OT | Can de-energize the process it drives | Human executes |
Rule: automate where the blast radius is recoverable (IT). In OT, automation prepares the decision and a human executes it.
- List every action in the response playbook and write the worst-case wrong-call outcome next to it.
- Tag each action
AUTOorHUMAN.
Check 1: every action that can change a controller, module, or process state carries HUMAN. Any OT action tagged AUTO fails the review.
Separate detection and decision preparation from execution
A model trained on traffic patterns learns communication behavior. It has no representation of the physical process behind a node, so its confidence score says nothing about consequence. Keep the model in the detection role and remove its ability to act.
- Route OT alerts to an operations-visible queue, not to a firewall or switch API.
- Give the automation service account read access and ticket creation only. Remove write credentials for OT enforcement points (firewalls, switches, controllers).
- Generate a decision package per alert: affected asset, connected physical process, proposed action, scope, expected loss (telemetry, control, interlocks), rollback path, current operating state.
- Name an operations owner as approver. Security proposes; operations executes or authorizes execution.
Map this rule to your sector or national standard and read the wording. For Swiss utilities that document is the IKT-Minimalstandard für Schweizer EVU, published in German; verify what it states before citing it as requiring this split.
Check 2: replay a synthetic false-positive alert against a test enforcement device. Expected: a decision package is created, the device configuration is unchanged, and the audit log shows zero write operations from the automation account.
Document which physical process hangs off each module
The first context gap is static: topology and process dependency. It is a property of the plant, it can be documented, and it rarely is. Without it the decision package cannot state what isolation costs.
- Create one record per controller and remote I/O module: process served, upstream and downstream dependencies, safety implications, and the shutdown or restart path.
- Assign operations as owner of the records. Security consumes them read-only.
- Attach the record to the decision package automatically by asset ID.
Check 3: pick a controller at random and read its decision package. Expected: an engineer who has never seen the alert can state the affected process, the consequence of isolation, and the restart path without phoning anyone. If not, the record is incomplete.
Add operating state as a second context input
The second gap is temporal. The same action on the same controller is a nuisance in one month and a total loss in another. Take a perishable-goods line with a seasonal peak: off-peak, a stopped line costs a couple of hours. At peak, product arrives by truck on a schedule that cannot be paused and degrades while it waits, so an isolation on a false positive keeps costing every minute and cannot be undone. The OT equivalents are a cold start that takes hours, or an isolation during peak load when the recovery path runs through people who are already busy.
Recoverability is therefore a function of time, and the IT/OT line alone does not capture it. Add these fields to the decision package:
| Field | Purpose |
|---|---|
| Production calendar state (peak / off-peak) | Scales cost of any stop |
| Throughput or load state | Shows what is queued behind the process |
| Restart cost (for example cold-start duration) | Sets cost of recovery after isolation |
| Staff availability for recovery | Tests whether the recovery path is actually staffed |
Check 4: replay the same alert under two operating states. Expected: the package rating and recommended action differ, and the peak-state package flags the isolation cost explicitly.
Prefer boundary containment that leaves the running process alone
Cutting the OT segment from upper layers stops telemetry and work-order communication. It halts the next work order but not the current operation, so only a person can stop the operation itself. This holds only when the control loop executes locally in the controller. If setpoints, recipes, or permissives arrive from above, the cut is a process change, so trace those dependencies in the Section 3 record first.
For cases where containment escalates, pre-write escalation templates per process class and let the human choose the rung. Process size, speed, interdependencies, and safety implications decide which rung applies:
- Shut down gracefully, then isolate.
- Isolate at a wider scope, shut down gracefully, then isolate at a narrower scope.
- Last resort: shut everything down, then isolate.
These are candidate sequences to validate with operations for each process class. None is a default.
Check 5: in a maintenance window with four-eyes approval, cut the segment from upper layers. Expected: the current operation completes, local HMI and control stay functional, a telemetry-loss alarm raises at the upper layer, and the work-order queue stops advancing. Any process trip means the loop depends on an upstream signal; correct the dependency record.
Turn existing human gates into machine-readable authorization
Maintenance windows, four-eyes approval on interventions, and keyswitch position (RUN, as opposed to a program or remote-enabled position) already encode which actions are allowed in which state. Humans parse them today. A machine-readable version changes the question from "human or machine" to "which actions are pre-authorized in which operating state". Build it only after Checks 1-5 pass, and treat the state feed as a controlled asset.
| Design question | Required answer |
|---|---|
| Who owns the context feed? | A named operations role, with a maintenance cycle |
| What happens when the feed is stale? | Default to the most conservative state: human-only execution |
| What if an attacker knows containment goes conservative at peak? | Timing advantage exists; treat the calendar and state inputs as attack targets with authenticated writes and change monitoring |
| What if the calendar is unmaintained? | Worse than none, because it produces confident wrong decisions; retire it |
Check 6: stop updating the feed past its allowed age. Expected: the system falls back to human-only execution, and an alert reaches the feed owner. Then change a calendar entry without authorization. Expected: a change alert fires.
Set the evidence bar for relaxing the human gate
"Most incidents are caused by humans" is a base rate with no comparison group, because humans run nearly every operational decision today. It does not show humans are the weaker operator. The argument for the gate rests on blast radius and reversibility, not on error rates. A human operator acts inside a chain of friction: four-eyes, interlocks, procedures, a colleague saying "hold on". An automated response to a false positive acts at machine speed with none of that friction, and the failure mode is physical. Error probability can be equal while the consequence profile differs.
- Log each response with time to decision, time to execute, and wrong-call outcome.
- Review the log after every real event and drill.
- Relax the gate for a specific action only when data shows that automated containment against live control systems yields fewer and less severe incidents than the human-gated path.
Check 7: the log shows time to decision for every OT event. Expected: no OT action lacks a recorded human approver and operating state.
End-to-end containment drill
- Inject a synthetic compromise alert on a controller in a test or maintenance-window environment. Expected: detection fires, and a decision package is created with process record and operating state attached.
- Confirm the automation account made no write to any OT enforcement device. Expected: audit log shows read and ticket operations only.
- Repeat the alert under peak and off-peak states. Expected: the rating and recommended action differ.
- Approve a boundary cut as operations owner with a second approver. Expected: the current operation continues, telemetry drops, and the work-order queue stalls.
- Let the context feed go stale. Expected: fallback to human-only execution with an alert to the feed owner.
- Restore the connection and measure recovery. Expected: restart cost matches the figure in the decision package.
FAQ
Why does automated containment carry more risk in OT than in IT?
The wrong-call consequence is physical. A revoked session costs one user about an hour, while an isolated controller can de-energize a process or black out a substation. Automate where the blast radius is recoverable and keep a human as the final gate on OT actions.
Why does a traffic-learning model misjudge isolating an OT device?
It learns communication patterns, not the physical process behind a module or what is queued behind that process. The consequence depends on plant topology (static) and operating state such as peak season or cold-start cost (temporal), and neither is visible on the network.
Why does cutting a controller from upper layers not stop the running operation?
When the control loop runs locally in the controller, the cut removes telemetry and work-order communication but not the current operation, so only a person can stop it. If setpoints or permissives come from above, the cut is a process change; verify the dependency record and test in a maintenance window.
Why does a context-aware containment feed create a new attack target?
Once containment goes conservative at peak, an attacker who knows the production calendar gains a timing advantage, so the calendar and state inputs become worth attacking. Authenticate writes, monitor changes, name an owner, and default to human-only execution when the feed is stale.