Configuring OT Network Monitoring Without Production Risk

David Krause9 min read
Best PracticesIndustrial NetworkingOther Manufacturer
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

OT network monitoring must add visibility without transferring uncontrolled change authority into production systems. Security controls that are routine on office endpoints—remote management, domain policy, automated patching, active discovery, antivirus scans, reboots, dynamic addressing, and removable-media blocking—can interrupt deterministic communications, invalidate application compatibility, or remove the only recovery path for an isolated machine. Commission each control as a plant change with an owner, test boundary, rollback method, and measurable acceptance result.

Control and change boundaries

The term OT here means the controllers, operator interfaces, supervisory applications, industrial computers, networks, and support tools that operate or recover production. IT protects enterprise services and information; OT adds process continuity, equipment state, and controlled recovery to that security objective. Neither group can manage the boundary alone.

Assign decision rights before connecting remote monitoring and management (RMM) tooling. IT can own the monitoring platform and corporate security requirements while OT owns production impact assessment, compatibility testing, maintenance timing, and functional acceptance. Plant management resolves risk decisions that neither technical group can accept independently.

Change class Required OT decision Required IT decision Acceptance evidence
Firewall rule Identify the production service and operating consequence Review exposure and restrict the permitted path Required session passes; unrelated traffic remains blocked
Host policy or RMM agent Validate application, licensing, recovery, and reboot behavior Configure policy, logging, and administrative access Production application completes its functional test
Update or scan Approve the test case and maintenance window Define package, scope, and rollback control No unexpected restart, service loss, or data gap
Incident response Place the process in a safe operating state Contain the cyber event without uncontrolled production changes Named responders can execute the documented sequence

Check 1: expect every proposed IT control to have an IT owner, an OT approver, a rollback owner, and a production acceptance test before deployment.

Firewall and DMZ interconnection

Connect the enterprise and OT networks through one controlled logical boundary. A demilitarized zone (DMZ) is an intermediate network that terminates or relays authorized services so enterprise clients do not receive unrestricted routes to controllers and production hosts. The firewall boundary protects enterprise assets from OT-originated faults and protects OT assets from enterprise-originated changes.

OPC UA data transfer does not justify general enterprise access to the control network. Document each required conversation by source, destination, protocol, direction, business purpose, application owner, and recovery consequence. Allow only those conversations. If a service requires traffic in both directions, record that behavior explicitly instead of describing the connection as read-only.

  1. Draw the enterprise, DMZ, and OT security zones, including every route between them.
  2. Place externally consumed services at the controlled boundary or behind an approved relay architecture.
  3. Build firewall rules from observed application requirements and configuration records, not broad subnet access.
  4. Remove alternate paths such as unmanaged wireless bridging, secondary network adapters, or temporary remote-access links.
  5. Record the firewall configuration and a tested rollback before activation.

Check 2: expect each approved OPC UA or management session to establish through its documented path, while a connection attempt outside the approved source, destination, or direction fails and appears in the boundary log.

Passive monitoring connection

Start network visibility with passive collection. A switch mirror port copies selected traffic to a sensor; the sensor analyzes the copies without polling controllers or opening sessions to production devices. An npcap-based or equivalent capture sensor can forward observations to an external monitoring platform.

Passive and active monitoring are different controls. A passive sensor connected to a correctly configured receive-only path does not need authority to write traffic into OT. An RMM agent, credentialed scanner, discovery probe, or remote shell is active because it can execute code, initiate traffic, change state, or restart a host. Treat each active function as a separate production change.

  1. Select the switch interfaces or network segments whose traffic must be observed.
  2. Configure the mirror source and dedicated sensor destination according to the switch design.
  3. Disable unneeded services and inbound management paths on the sensor-facing production interface.
  4. Send monitoring results outward through the approved firewall or DMZ path.
  5. Baseline switch load, communication errors, and application response before adding further monitoring functions.

Check 3: expect the sensor to receive copied OT frames and export telemetry, with no routable session from its capture interface back to a controller and no new production communication errors.

Address and service dependency register

Build an OT dependency register before applying address, firewall, domain, or endpoint policy. Production applications often reference fixed addresses, device names, local accounts, firewall ports, hardware licenses, mapped paths, and startup order. Changing one dependency can make a healthy controller appear offline.

Retain static addressing wherever the installed configuration depends on a stable endpoint. A move to dynamic addressing is valid only after every consumer has been reconfigured and tested against the replacement naming or reservation method. Closing an industrial application port at the firewall produces the same symptom as a failed PLC connection, so test network policy before replacing hardware.

Dependency Failure symptom Commissioning record
PLC or server address Timeout, offline indication, or stale data Configured address, consumers, and approved address method
Application traffic through a firewall One service fails while basic reachability may remain Source, destination, direction, and required service
OPC UA path Downstream information stops or quality degrades Server, client, boundary path, and certificate ownership
Hardware licensing device Engineering or supervisory software will not start Physical interface, driver, custody, and recovery procedure
Offline-machine backup Vision or machine recovery cannot load the saved image Backup location, approved transfer method, and restore test

Check 4: expect every production endpoint to retain its approved identity and every listed consumer to read current data after the boundary rules are applied.

Host policy and removable-media controls

Least privilege remains the target, but a corporate policy must account for OT application behavior. Domain membership, local-rights removal, device-control policy, and automatic configuration can affect legacy services, engineering software, drivers, scheduled tasks, and unattended startup. Validate the complete application workflow under the proposed account, not merely the Windows login.

A blanket mass-storage block can also disable required recovery hardware. An isolated vision system may need removable media or a USB-to-SATA adapter to load its backup. Moving files through SharePoint does not solve that path when the target machine is intentionally not networked. Hardware licensing can likewise depend on an enabled USB interface.

  1. Classify each USB dependency as licensing, diagnostics, backup, restore, or ordinary file transfer.
  2. Replace ordinary transfer with the approved network workflow where the machine supports it.
  3. Create a controlled exception for indispensable offline recovery and licensing hardware.
  4. Limit custody and use of approved media, scan it through the site’s controlled process, and record each transfer.
  5. Keep an offline copy of the recovery procedure so a domain or network outage does not remove access to it.

Check 5: expect the production account to launch all required services, the hardware license to remain visible, and an authorized technician to read the stored backup through the documented offline recovery path.

Updates, scans, and reboot control

An OT software update is a compatibility project, not a desktop maintenance task. It can change application behavior, drivers, communications components, licensing, or restart requirements. A plant-wide move to Studio5000 v38, for example, requires a controller-by-controller compatibility assessment; one evaluated scope included a line item for 52 1756-L83 units. The version request alone does not define the required hardware work.

Use a sandbox test rig that represents the production application, communications path, licensing method, and startup sequence. Internal validation remains necessary for niche software when no accepted third-party validation covers the installed combination. Budget the test environment as part of the change.

  1. Inventory the installed application, operating environment, drivers, controllers, and licensing dependencies.
  2. Read the product compatibility and lifecycle information supplied for each installed component.
  3. Apply the exact proposed update, endpoint agent, scan policy, and reboot behavior to the test rig.
  4. Run application startup, communications, data recording, backup, restore, and restart tests.
  5. Deploy to one low-consequence production scope during an approved window.
  6. Hold expansion until OT accepts the results and the rollback remains available.

Do not permit RMM to install updates or reboot production hosts on an enterprise schedule. Do not introduce active virus scans or network discovery into OT until resource use and device interaction have passed the same staged test.

Check 6: expect the updated test system and first production scope to restart only when commanded, reconnect to every dependency, preserve production data, and complete the operating sequence without intervention.

Incremental rollout and symptom isolation

Change one controlled element at a time. Bundling a firewall revision, domain policy, RMM installation, address change, and antivirus rollout destroys fault isolation. Record the before-state, exact change, start and finish, affected assets, operator observation, and rollback result for each increment.

Observed symptom First suspected change Discriminating check
PLC appears offline after a boundary change Required application traffic blocked Compare firewall denies with the dependency register
SCADA server stops or loses data Update, reboot, service-account, or scan policy Correlate service state and data timestamps with the change log
Device changes address after restart Dynamic addressing introduced Compare the active address with the approved endpoint record
Engineering software loses its license USB or driver policy Check device recognition under the production account
Offline vision recovery cannot start Mass-storage policy removed the transfer path Execute the approved backup-media access test
Intermittent slowdown follows monitoring deployment Active polling, discovery, scanning, or host resource contention Disable only the new function under change control and compare the baseline

Check 7: expect each rollout increment to have one attributable configuration change, stable baseline readings, operator acceptance, and a proven rollback before the next increment starts.

End-to-end production verification

  1. Check 8.1: expect controllers, operator interfaces, supervisory servers, and isolated support systems to report their approved addresses and normal communication state.
  2. Check 8.2: expect OPC UA clients to receive current values through the documented firewall and DMZ path, with no unexplained quality or history gap.
  3. Check 8.3: expect the passive sensor to observe the selected traffic and export monitoring data without initiating a production session.
  4. Check 8.4: expect unauthorized cross-boundary and remote-management attempts to fail and produce a reviewable log entry.
  5. Check 8.5: expect production applications to survive the approved restart sequence, recover communications, recognize required licensing hardware, and resume data recording.
  6. Check 8.6: expect the backup and restore path for each non-networked machine to work under the active removable-media policy.
  7. Check 8.7: expect operations, OT, and IT to record acceptance against the same change identifier before expanding the deployment.

FAQ

Can I connect IT monitoring directly to OT switches?

Yes, when the connection is a controlled mirror destination feeding a passive sensor. Route exported telemetry through the approved firewall or DMZ path and verify that the capture interface cannot initiate sessions into OT.

Does OT monitoring need write access?

Passive traffic collection does not need write access. RMM, discovery, polling, remote shell, update, and reboot functions are active controls and require separate testing and approval.

Can I move PLCs and SCADA servers to dynamic IP addresses?

Only after reconfiguring and testing every consumer that depends on a fixed endpoint. If the installed applications reference static addresses, retain them until the replacement identity method passes the full communication test.

Does blocking USB affect OT recovery?

It can block hardware licenses, authorized backup media, and USB-to-SATA recovery for isolated equipment. Keep a controlled exception and prove the restore path with the active device policy.

Does commissioning end when the monitoring dashboard is green?

No. The final verification is an operating test: run the production sequence, confirm current OPC UA data, verify passive telemetry, reject an unauthorized boundary attempt, restart through the approved sequence, and prove the offline restore path.

Back to blog