When this RapidSCADA installation stopped updating, ScadaCommService repeatedly failed to connect to localhost:10000 because ScadaServerService had stopped; the later ThreadAbortException entries appeared during manual shutdown and restart, not as a record of the original failure.
Which message appears first when the server stops updating?
Use the earliest timestamp in the server, communicator, and automatic-control logs, not the last exception visible after a restart. In the reported event, automatic-control triggers issued commands at 05:15, but the server log had no captured crash entry at that time. By 10:00, the communicator was trying to reconnect and receiving connection-refused errors. The operator restarted the server around 10:05, and the server log then recorded shutdown-related exceptions.
This order matters: the messages at restart describe the state of a service being stopped. They do not show what first caused the server to become unavailable hours earlier. The same distinction applies to the later restarts: the server started, accepted the local client, processed automatic-control commands, and stopped again. That sequence makes the AutoControl/current-data path a priority for investigation, but a time correlation alone does not identify a defective trigger or prove a specific software defect.
| First useful observation | What it indicates | Next check |
|---|---|---|
| SCADA values stop changing, but no service state is recorded | The operator symptom alone cannot separate a server stop from a data-collection or display problem. | Compare server and communicator service state and log timestamps. |
CommSvc reports connection refused on 127.0.0.1:10000
|
No process is accepting connections on that local endpoint at that moment. | Check whether ServerSvc is stopped and whether its connection listener is running. |
| CommSvc reports a transport timeout | The connection attempt did not receive a timely response; on loopback, inspect server responsiveness and listener state before investigating field wiring. | Check the ServerSvc process, listener, and event timing. |
ThreadAbortException appears after a manual stop or restart |
The service is aborting worker threads during shutdown; this is not the initiating crash signature by itself. | Recover earlier logs or reproduce with additional capture enabled. |
Does the refusal at localhost:10000 point to a tag or a stopped listener?
A refused connection is below the channel/tag layer. CommSvc was configured to connect to the local SCADA server, and its log explicitly reports refusal at 127.0.0.1:10000. That means the server-side TCP listener was not accepting the connection then. A bad input-channel value, alarm predicate, or Modbus register mapping cannot by itself explain a local TCP refusal.
A timeout is a different observation. The reported CommSvc log alternated between connection timeouts and refusals after a server restart. A timeout means the attempt did not complete with a response in the allotted time; it can occur while a process or listener is unresponsive, and it does not identify why. Check the service and listener at the same timestamp. Do not begin by changing PLC addresses or tag bindings unless CommSvc has re-established its server connection and a separate field-data symptom remains.
Use service state and connection state to choose the branch:
- If ServerSvc is stopped and CommSvc receives refusal, restore the server service only after preserving logs and automatic-control state.
- If ServerSvc is running but CommSvc times out, inspect server responsiveness, listener state, and application logs before restarting either service.
- If CommSvc connects and receives commands but field values or device actions are missing, move downstream to the communication line, device driver, protocol response, and PLC/device feedback.
Did ServerSvc stop after AutoControl processed current data?
Read ServerSvc and ModAutoControl logs side by side. In the first captured restart, the server recorded Connection listener is stopped, an exception in current-data processing, an exception while executing actions calculated in ModAutoControl, then Server is aborted and ScadaServerService is stopped. The stack includes RaiseOnCurDataProcessing and ProcCurData. This places the shutdown interruption in the current-data processing path; it does not reveal the initiating condition because the entries were emitted as the service stopped.
The later sequence is more useful for correlation. ServerSvc started successfully, loaded its channel configuration and formulas, started the connection listener, and accepted the local client. ModAutoControl then logged trigger state changes and sent commands. Around 10:17:51, the server listener stopped and the service stopped; a further restart showed the same general sequence. Such recurrence after active trigger processing is a reason to test the module path before changing unrelated network settings.
Keep the wording precise in incident notes: “ServerSvc stopped after AutoControl activity” is supported by the log sequence; “a particular trigger crashed the server” is not yet demonstrated. An exception mentioning an aborted thread during shutdown is evidence of interruption, not proof of a deadlock, race, overload, or particular trigger defect.
Does disabling ModAutoControl change the failure?
Run a controlled isolation test when the process can safely operate without automatic commands. The reported test project used simulator-generated data with a one-second update interval and AutoControl triggers that issued commands to a Modbus relay. The problem reproduced within roughly one to two minutes in the described test. A later test reported no crash during a 1.5-hour run with AutoControl off. The project author also reported that the failure occurred whether the relay was connected or disconnected.
Those results make AutoControl and its interaction with current-data processing the leading test branch. They argue against requiring a live relay failure to trigger the symptom, but they do not establish that all installations or all crashes share the same cause. Repeat the comparison under equivalent data generation, request cadence, and observation period. Record whether the module is enabled, whether a trigger fires, what command is issued, and the elapsed time to any service stop.
- Capture a baseline with the normal configuration: service state, channel update cadence, trigger events, command events, and CommSvc connection status.
- With the process in an approved condition that does not depend on AutoControl commands, disable the module and repeat the observation period.
- Compare equivalent runs. If the failure appears only with AutoControl active, focus on trigger evaluation, action dispatch, and load/cadence interaction. If it also occurs with the module inactive, widen the investigation to other server modules, service events, host resource pressure, and operating-system logs.
- Re-enable the module only for a controlled test, and monitor it long enough to include the normal trigger bursts that preceded the event.
Disabling a control module is a diagnostic isolation step, not a production fix. Before using it, account for any process action the module normally commands.
Do channel conditions or command dispatch explain the trigger activity?
Separate trigger evaluation from command delivery. A data trigger depends on its configured channel, comparison condition, deadband, delay, repeat behavior, and channel-status checks. An event trigger can depend on an event’s channel and new value. If a trigger does not fire when expected, inspect the live/history value and status for that channel and compare them with the trigger’s actual configured predicates. If it fires unexpectedly or repeatedly, compare the value transitions with the threshold, delay, deadband, and repeat settings rather than assuming a driver binding fault.
For the captured activity, the ModAutoControl log names synchronization triggers such as F14\F14 MCnt Sync and F14\CCuMix AccWt Sync; the server log records output channels 1402 and 9001. These entries establish that the module generated commands on those output channels. They do not, by themselves, prove that a PLC or relay executed the command.
| Setting or observation | Where to read it | Diagnostic effect |
|---|---|---|
| Trigger active state, type, source channel, condition, deadband, delay, repeat count | AutoControl trigger configuration and module log | Shows whether channel data should cause a trigger and whether repeat firing is configured. |
| Current channel value and status at trigger time | SCADA current-data view and archive, if available at sufficient resolution | Distinguishes a real condition transition from stale, invalid, or unexpected channel data. |
| Output channel command event | Server and ModAutoControl logs | Shows that the server-side trigger path issued or enqueued the command. |
| Command received by CommSvc | CommSvc log | Confirms server-to-communicator delivery, not physical actuation or PLC acceptance. |
| Protocol response and resulting device feedback | Communication-line/device diagnostics and corresponding feedback channel | Determines whether the command traversed the driver/protocol/device path and changed the intended state. |
This staged reading prevents two common misdiagnoses: treating a channel predicate problem as a server-to-communicator binding problem, and treating a server-side “command sent” entry as proof of a field-device action. In the reported outage, CommSvc lost its connection to ServerSvc, so fix that server-service branch before interpreting a lack of downstream device response.
Does command or simulator pacing change the crash rate?
Test pacing as a controlled variable, not as proof of root cause. One suggested workaround was to set command delay to one second. The report says this reduced the crash probability or frequency but did not eliminate the failure. In the test project, commands already had a one-second delay; the failure still occurred. Adding one second after the simulator request cycle or changing the simulator device interval from one second to two seconds also reduced the reported frequency in limited testing.
The mechanism is relevant: a command delay can place work in a queue rather than issuing it directly at trigger evaluation time, changing how trigger activity overlaps with data requests and command processing. A slower request cycle also changes event arrival cadence. Either change can reduce contention or timing sensitivity, but neither demonstrates a corrected server defect. The logs show bursts of triggers and commands, including repeated synchronization commands, so record both trigger rate and request cadence when comparing runs.
Use a small test matrix rather than changing multiple variables at once:
- Hold trigger conditions and simulator update interval constant; compare command delay disabled versus one second.
- Hold command delay constant; compare the original request interval with the tested slower interval.
- Record trigger count, command count, server uptime, and whether CommSvc remains connected for each run.
Do not infer that an offline device is the cause just because commands target it. The reported reproduction occurred with the relay disconnected as well as connected. Conversely, do not conclude that the device is irrelevant to every failure without checking the driver logs and protocol response in the installation being diagnosed.
What evidence must be saved before restarting the services?
In the reported incident, the server log had no useful entry at the time of the silent stop, and older CommSvc entries had rolled out. Restarting then produced the ThreadAbortException messages that could be mistaken for the cause. Before cycling services, preserve the diagnostic state available at that moment.
- Copy the current ServerSvc, CommSvc, and ModAutoControl logs, including rotated files, and note the host’s clock/time zone and the last known-good update time.
- Record whether each service is running, whether the server listener is active, and the exact connection error and timestamp. Capture the process/service state before restart if operational procedures allow.
- Save the AutoControl configuration and module state. Identify active trigger groups, trigger types, source channels, conditions, delays/repeats, command channels, and any recent configuration changes.
- Capture current values and statuses for trigger source channels, relevant output channels, and feedback channels. If the normal archive interval is 30 seconds, recognize that short trigger bursts or value changes may fall between archived samples; a targeted higher-resolution snapshot can preserve the needed state if the product configuration supports it.
- Record host CPU and memory pressure and operating-system service/application events around the failure. These readings help distinguish an application-path fault from a host-wide stall.
- Only after capture, perform the planned restart and note which service was started first, whether CommSvc authenticated/connected, and whether current values resumed.
A command log alone is not enough to reconstruct the triggering data. The incident report specifically lacked a snapshot of the channel values at failure time; preserving the source-channel values and status alongside trigger state can show whether a trigger condition was normal, stale, or repeatedly oscillating.
How should the resolving branch be tested and verified?
Use the check results to choose a bounded remediation. If CommSvc reports refusal because ServerSvc is stopped, restore ServerSvc after preserving evidence; restarting CommSvc alone cannot create a missing server listener. If ServerSvc remains running but unresponsive, capture its state and logs before terminating it. If disabling AutoControl prevents recurrence under equivalent test conditions, keep the module isolated while reviewing trigger predicates and command pacing, then reintroduce it under controlled observation.
- Confirm process conditions permit a controlled test; disable or isolate AutoControl only when the resulting loss of automatic actions is acceptable.
- Save the existing module configuration and current trigger/channel state so a test change can be reversed.
- Change one variable at a time: first isolate AutoControl, then test command delay or request cadence independently. Keep the original setting and result for comparison.
- Start ServerSvc and verify its log reports startup completion and a started connection listener. Then start or reconnect CommSvc and verify it connects to the server rather than retrying
localhost:10000. - Verify current input channels update, the expected trigger transitions occur, and each intended command appears in server and CommSvc logs. Check protocol response or device feedback separately; do not count a server-side command event as proof of actuation.
- Observe through the trigger/request pattern associated with the failure and record uptime and service state. If the issue returns, preserve the new logs and state before another restart.
The installation’s one-second command delay and slower simulator cadence reduced the reported failure frequency but did not eliminate it; a 1.5-hour no-crash test with AutoControl disabled is useful isolation evidence, not a validated permanent fix. Apply a change as a fix only after the server listener remains available, CommSvc stays connected, expected channel updates continue, and the required command path behaves correctly through representative trigger activity.
RapidSCADA AutoControl crash troubleshooting FAQs
Why does CommSvc say connection refused on 127.0.0.1:10000?
The local ServerSvc listener is not accepting connections at that time, commonly because the server service has stopped. Check ServerSvc state and listener status before changing channel or Modbus settings.
Why does ThreadAbortException appear after I restart RapidSCADA?
Stopping a service aborts its worker threads, so this exception can be a shutdown consequence. It does not identify the original silent-stop cause; compare earlier ServerSvc, CommSvc, and AutoControl timestamps.
Why does a one-second AutoControl command delay not prevent the crash?
It changed command scheduling and reduced the reported frequency, but failures still occurred with a one-second delay configured. For final verification, confirm ServerSvc remains started, CommSvc stays connected, input values continue updating, and expected commands reach the communicator through representative trigger activity.