A 5-minute profile of an OPC UA server with 200 subscribed tags puts ClearChangeMasks at 36.59%, OnMonitoredNodeChange at 36.53% and QueueValue at 36.49% of CPU. Those three numbers are one call chain, not three separate costs. The server application writes a value, marks the node changed, and the stack pushes a copy into every monitored item queue watching that node. The load scales with writes per second, not with tag count alone. Scaling to 2000+ tags multiplies it by ten.
Where does the CPU go between the application write and the client queue?
The data path on the server side is:
- Application code assigns a new value to a variable node.
- The application calls
ClearChangeMaskson that node, which raises a change notification. -
OnMonitoredNodeChangeruns once per monitored item attached to the node. -
QueueValuecopies the value, status and timestamps into that item's notification queue. - The publish cycle drains the queues into a
Publishresponse toward the client (UaExpert here).
The three profiler percentages are nearly identical because each function is inclusive of the one it calls. The leaf, QueueValue, carries almost all the cost. Time is spent on queue allocation, value copies and locking. Nothing in this chain checks whether the value actually changed. Every ClearChangeMasks call becomes a queued notification.
Estimate the call rate with this formula (assumption: one monitored item per tag, every tag written every cycle):
QueueValue calls/s = tags x (1000 / write_cycle_ms) x monitoring_clients
200 tags, 100 ms cycle, 1 client = 2,000 calls/s
2000 tags, 100 ms cycle, 1 client = 20,000 calls/s
Check before moving on: confirm the profiler shows the same three frames in one stack, with QueueValue as the deepest. If a different frame sits under QueueValue, the cost is in that frame (allocation, lock contention), and the fixes below still apply because they cut the call count.
How do you measure write rate against real change rate?
Add two counters in the node manager write routine: writesAttempted and valuesChanged.
| Result | Meaning | Action |
|---|---|---|
valuesChanged ≈ writesAttempted
|
Process values truly change every cycle (noisy analogs) | Apply deadband and slower publishing (sections below) |
valuesChanged far below writesAttempted
|
The application re-publishes unchanged values every cycle | Gate ClearChangeMasks on change |
Check: the ratio is known per tag group before any configuration edit. Without it, config changes are guesses.
How do you stop unchanged values from reaching the monitored item queue?
Push only changes into the stack. Compare the new value with the current node value and call ClearChangeMasks only when they differ. Apply a source-side deadband to floating-point tags so instrument noise does not register as a change.
if (!newValue.Equals(variable.Value)) // arrays: element-wise compare
{
variable.Value = newValue;
variable.Timestamp = DateTime.UtcNow;
variable.ClearChangeMasks(SystemContext, false);
}
Update the whole tag batch inside one lock scope, then raise the change masks. Do not take and release the node manager lock per tag. Set the status code and source timestamp from the field data, not from the write moment, so a client can tell a stale value from a fresh one.
Check: re-run the counters. QueueValue call rate must now track valuesChanged, not writesAttempted.
Which client and server settings cut notification volume?
The client requests the publishing interval, sampling interval, queue size and data change filter. The server clamps those requests to the limits in the config file. Read the supplied config against the request path:
| Setting | Current value | Effect on load | Change |
|---|---|---|---|
MinPublishingInterval |
100 | Floor for how often a subscription publishes | Raise to the fastest interval the process needs |
MaxNotificationQueueSize |
100 | Cap on queue depth per monitored item; worst case 2000 items x 100 = 200,000 queued values (derived) | Use a client queue size of 1 to a few entries unless data loss between publishes is unacceptable |
MaxNotificationsPerPublish |
1000 | 2000 changed items need at least 2 notification messages per cycle | Raise only if publish response count is the bottleneck |
MaxMessageQueueSize |
10 | Retransmission buffer per subscription | Leave unless the client loses messages |
On the client side, set the monitoring mode to report data changes with a deadband on analog items, a publishing interval no faster than the operator needs, and a small queue with discard-oldest.
Check: read back the revised values the server returns in CreateMonitoredItems and CreateSubscription responses. The revised publishing and sampling intervals must match what you intended; if not, the config clamp is overriding the request.
How do client-to-server writes fit into the 2000+ tag scenario?
Client writes enter through the Write service, not the change-notification path. Send one Write request carrying an array of WriteValue entries for a tag group instead of one request per tag. Each request costs a secure-channel decrypt/verify, a session lookup and a response encode, so batching removes per-request overhead. Apply the same change gate to values coming from the client: an unchanged write that still calls ClearChangeMasks fans out to every other subscriber of that node.
Check: with the client writing a 2000-item batch, confirm from the counters that only the changed nodes raise change notifications on the server.
What proves the fix under the 2000+ tag load?
- Record the baseline: 200 tags, 5-minute profile,
ClearChangeMasks36.59%,OnMonitoredNodeChange36.53%,QueueValue36.49%. - Apply the change gate, deadband, revised sampling groups and client queue settings, then repeat the identical 5-minute profile at 200 tags. The three frames must drop well below baseline.
- Compare
QueueValuecalls per second againstvaluesChangedper second from the counters. They must match. - Scale to 2000+ tags with realistic change behavior and profile again. Record process CPU and the share held by the three frames.
- Run a worst-case step where every tag changes each cycle. This gives the real ceiling of the server; if CPU is still unacceptable there, cut the fast subscription's item count or lengthen its publishing interval.
- At the client, confirm no
BadTooManyPublishRequests-style backlog or queue-overflow status bits appear on items, and that item timestamps advance at the expected rate.
What happens if the application calls ClearChangeMasks every cycle on unchanged values?
Every call runs OnMonitoredNodeChange and QueueValue for each monitored item on the node, so CPU scales with write rate x subscribed items even when nothing changed. Compare against the current value first and call it only on a real change.
The first AvailableSamplingRatesRemove that group, or raise the client's requested interval, when the process does not need sub-100 ms data.
What happens if MaxNotificationQueueSize stays at 100 with 2000 monitored items?
The worst-case buffered depth is 2000 x 100 = 200,000 queued values when the client publishes slower than values change. Request a small client queue size to hold memory and per-item queue work down.