Structuring IT-OT Incident Response for Manufacturing Plants

Patricia Callen8 min read
Industrial NetworkingOther ManufacturerTechnical Reference
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

A suspicious network event near PLC networks or a historian forces a containment call that neither the SOC nor the OT lead can make alone. Security sees a threat path. Operations sees a line that stops when a controller loses its network. The checks below are ordered so each reading either clears you to act or sends you to the next check. Run them in sequence, and take the measurement at each step before you decide.

Which side of the process boundary did the event touch?

Read the asset role and the traffic direction first. Assets that only read from the process (historian collectors, DMZ mirrors, reporting servers) sit on the read side. Assets that can write to the process (controllers, engineering workstations, HMIs with write access, remote-access paths into the control network) sit on the write side.

  • Read side only: run the standard IT playbook and notify the OT lead. A compromised historian is a pivot risk. Cutting it costs trend data, not process control. Go to check 4 and pick a tier.
  • Write side touched or possibly touched: the event needs joint handling. An engineering workstation can download logic, so a compromise there puts controller program integrity in question, not just the workstation. Go to check 2.
  • Cannot tell which side: treat it as write side. Misclassifying downward is the expensive error.

What does the operator see if this path is cut?

List every flow a proposed block would break, then have an operator state the physical consequence of each. Typical flows are HMI to controller, controller to controller, cyclic I/O connections, historian collection, and remote access. The mechanism matters. Cyclic connections run on a watchdog, so a firewall block does not fail quietly. The connection times out, and the device then applies its configured fault action, which may be hold last state, go to a defined safe state, or fault the machine. Read that fault action from the project configuration of the affected devices. Do not rely on memory or on what the network diagram implies.

  • Loss is tolerable (remote access, historian collection, an idle engineering workstation): go to check 4.
  • Loss stops or degrades the line: the decision is a production decision as well as a security decision. Go to check 3 and keep the block off until both approvers agree.
  • Nobody can state the consequence: treat the loss as intolerable. Do not block. Use monitoring-only measures until the OT engineer can answer.

Who signs the containment, and is it written down?

Check whether a written rule names the approvers before the event. The rule that removes most of the IT-OT friction is joint authority: any containment that can touch the control network requires sign-off from both the security lead and the senior OT engineer, and neither approves alone. Security cannot unilaterally block something that stops a line, and operations cannot dismiss a real threat, because the decision is structurally shared. Writing it down in advance matters because nobody should be discovering who may stop production while the incident clock runs.

  • Rule exists in writing: go to check 4.
  • Rule does not exist: get the security lead and the senior OT engineer on one call, record both approvals with a timestamp, and write the rule after the incident closes.

The written rule should contain named roles plus deputies, the definition of an OT containment action, the documentation required for each action (time, action, approvers, expected process effect, rollback), and the trigger for bringing in the owner of any safety-related function. Joint authority between security and OT does not replace the safety function owner when a path touches a safety-related system.

Which containment tier fits the readings so far?

Choose the lowest tier that closes the threat path. This ladder is a suggested structure. Adapt the tiers to your architecture and record them in your own plan.

  • Tier 0, observe: raise monitoring on the affected segment, capture traffic and logs, take no network action.
  • Tier 1, remove access paths: disable remote access, revoke or rotate credentials, isolate the IT-side or DMZ asset. No control-network flow is cut.
  • Tier 2, restrict specific flows: block named flows at the boundary of the affected segment, with the operator-stated consequence for each flow on record.
  • Tier 3, isolate a cell or segment: production owners must confirm the cell can run or be stopped safely in isolation.
  • Tier 4, stop the process: a production and safety decision that goes above the two-person security and OT approval.

Tiers 0 and 1 on the read side can be pre-authorized in writing if security and OT agree they carry no process effect. Everything from Tier 2 up on the write side goes through joint sign-off.

Which signals feed the decision, and how does each one mislead?

The containment decision is a control loop with people in it. It fails the same way a control loop fails: a bad measurement produces a confident wrong action. Verify each input before it drives a block.

Signal Source Wrong-value symptom
Alert severity and affected asset SIEM, OT network monitoring A stale asset inventory maps the alert to the wrong device. The block lands on a running controller instead of a spare workstation.
Process state and criticality Operators An answer given before a batch transition or startup is out of date when the block executes. Ask again at approval time.
Connection and fault status of controllers and HMIs Controller diagnostics, HMI comms status Watchdog faults caused by your own block get read as attacker activity, which drives further blocks. Timestamp the containment so its faults are separable.
Controller program integrity Running project compared against a known-good offline copy With no baseline, you cannot clear a controller. The team defaults to isolating it, which costs production.
Historian data continuity Historian collection status A gap is read as compromise, or a block-induced gap hides an earlier real one. Record when each gap started.
Approval state Written containment record A verbal approval from one party is contested later, and the joint rule is undermined.

Does the framework hold without a daily working relationship?

No. Predefined authority and role definitions work far better on top of an existing relationship, and daily communication is what makes the formal framework function under pressure. A perfect runbook between teams that only talk during incidents fumbles the real event. A rougher plan between people who talk every morning holds.

Build the rhythm from OT monitoring. The OT engineers receive a daily morning brief from the monitoring system. Any critical finding becomes a direct conversation between the security lead and the senior OT engineer about the course of action, not a ticket. That conversation builds the shared vocabulary and the picture of normal traffic that make an anomaly recognizable to both sides.

On unified war room versus separate OT plan, use both. Keep the OT annex to your incident response plan as the document that defines authority, and run the war room as the place where those roles execute with cyber, safety, and operations seats filled. The annex sets who decides. The daily contact determines whether the decision is fast.

How do you run the joint containment call for a write-side event?

  1. Classify the event as read side or write side using check 1. Default to write side if unclear.
  2. Open one call with the security lead, the senior OT engineer, and an operator for the affected area. Add the safety function owner if a safety-related system is in the path.
  3. List every flow the candidate action breaks. The operator states the physical effect of each. The OT engineer states the configured fault action of the affected devices.
  4. Select the lowest tier from check 4 that closes the threat path.
  5. Obtain both approvals and record time, action, approvers, expected process effect, and rollback trigger.
  6. Preserve logs and volatile evidence before any reboot or isolation that would destroy them.
  7. Execute. The operator watches HMI and process behavior while the OT engineer watches controller and connection diagnostics.
  8. Compare observed behavior against the recorded expected effect. If an unexpected fault appears, apply the rollback trigger, then re-decide with a fresh operator reading.

How do you prove the structure works before a live event?

Rehearse it with the people who bear the consequences. OT tabletops should bring operators and OT engineers into active roles alongside security, not security running a drill at OT. Operators know what a given containment action costs on the floor, and they surface the "you cannot block that, here is what happens physically" objection during rehearsal instead of during a live decision. If your exercises are IT-led with OT as an audience, move operators into decision roles first. CISA publishes tabletop exercise packages and cybersecurity scenarios that give you ready scenario material.

Score each exercise with measurements, not impressions:

  • Time from detection to a two-party approval decision.
  • Whether every proposed block had an operator-stated physical effect before approval.
  • Whether the decision record contained all five required fields.
  • Whether the team separated its own containment-induced faults from attacker activity.
  • Whether a controller could be cleared against a known-good copy, or the team defaulted to isolation.

A rehearsed authority structure stays unproven until a live event runs it. Review the first real containment decision against the written rule and amend the rule where it broke.

FAQ

Can the SOC block traffic to a PLC network on its own?

Not under a joint-authority rule. Any containment that can touch the control network needs approval from both the security lead and the senior OT engineer. Pre-authorize in writing only those actions, such as disabling a remote-access account, that both sides agree carry no process effect.

Does a separate OT incident response plan replace a unified war room?

No. The OT plan defines who approves containment and what gets documented, and the war room is where cyber, safety, and operations execute those roles together. Daily contact between security and OT decides how well either one performs.

Can I keep handling an OT incident internally if a containment action causes unexplained faults?

Stop further blocks, apply the rollback trigger from the decision record, and escalate when the affected system is a safety-related function, when controller logic cannot be cleared against a known-good copy, or when faults persist after rollback. Take those cases to the controller or drive manufacturer's official support channel and to your national CERT or regulator per your plan. Do not test additional blocks on a live line while waiting for their response.

Back to blog