Overview
A running (or rolling) average is the unweighted mean of the most recent N samples of a process signal. In discrete control, it acts as a single-pole digital low-pass filter that suppresses high-frequency measurement noise without the phase lag that a first-order analog RC filter introduces at very low cutoff frequencies. The canonical implementation keeps an array of length N and, on every scan, subtracts the oldest sample from a running sum, replaces it with the newest sample, and adds the new value to the sum. The average is then Sum / N.
Two storage strategies are common on a PLC:
- Shift register — copy the entire array on every sample. Conceptually simple, but the cost grows linearly with window length and creates a 1-scan race condition if the array is read elsewhere mid-shift.
- Circular buffer with index modulo — overwrite the oldest slot in place using a write index that wraps at N. The data does not move; only the index does. This is the recommended pattern on SIMATIC S7-1200 (firmware V4.0 and later) and S7-1500/1500T controllers where SCL compilation targets the optimized machine code path.
The Mathematics of the Running Average
For a window size N and samples xi:
yk = (1 / N) · Σi=k-N+1k xi
Recomputing the sum from scratch every scan costs N additions. The incremental form maintains the sum:
Sk = Sk-1 − xk-N + xk
Two additions and one indexed load — independent of N. This is the form the circular buffer exploits. Average is then yk = Sk / N.
The transfer function of an N-point running average is H(z) = (1/N) · (1 + z−1 + … + z−(N−1)). It has unity gain at DC, a −3 dB bandwidth of roughly 0.443 / N in normalized frequency, and a first null at fs / N. For a 100 ms cycle and N = 16 the effective cutoff is ≈ 0.28 Hz, which is the right ballpark for filtering pulse-flow, weight, and temperature signals.
Circular Buffer vs Shift Register Implementations
| Property | Shift register (array copy) | Circular buffer (index modulo) |
|---|---|---|
| Memory writes per sample | N | 1 |
| Scan-time scaling | Linear in N | Constant |
| Atomicity | Not atomic — mid-shift reads see partial buffer | Atomic — only the slot being written is transient |
| Index management | None | Wrap at 0..N−1 |
| Cold start | Fill counter to 0..N−1 | Same; prefill with first sample or zero |
| Recommended window size | N ≤ 16 typical | Any N; tested up to 4096 |
The shift-register version on a Siemens CPU uses the BLKMOV (block move) instruction in STL, or a FOR loop in SCL. Both generate one MOVE per element when the optimizer cannot prove aliasing safety, which on S7-1500 typically hits the SIMATIC load memory bandwidth before the cyclic OB timing does. The circular buffer avoids the move entirely.
Reference Implementation in SCL
The following SCL code is the canonical incremental form. It is intended to live inside a function block (FB) with Tijd declared as ARRAY[0..9] OF DINT (the original problem window) and is straight-line code — no FOR loop, no BLKMOV.
// === Running average FB, SCL (S7-1200/1500) ===================
// Inputs : GemetenTijd : DINT – newest sample
// Outputs: Gemiddelde : INT – mean of the populated window
// Statics: Tijd : ARRAY[0..IndexGrootte-1] OF DINT
// Som : DINT
// Index : DINT
// Teller : DINT – number of valid samples in buffer
// IndexGrootte: DINT – window length, e.g. 10
// ================================================================
// 1) Subtract the value about to be overwritten
#Som := #Som - #Tijd[#Index];
// 2) Write the new sample into the slot
#Tijd[#Index] := #GemetenTijd;
// 3) Add the new sample to the running sum
#Som := #Som + #Tijd[#Index];
// 4) Advance the write index with modulo wrap
#Index := #Index + 1;
IF #Index = #IndexGrootte THEN
#Index := 0;
END_IF;
// 5) Track how many slots hold a valid sample (used during fill-up)
IF #Teller < #IndexGrootte THEN
#Teller := #Teller + 1;
END_IF;
// 6) Compute the average using the count, not the window size
IF #Teller > 0 THEN
#Gemiddelde := DINT_TO_INT(#Som / #Teller);
END_IF;
- The original code wrote after advancing the index, which caused the value to land in the wrong slot on the wrap-around scan. The fixed sequence subtracts the slot that is about to be overwritten, then writes, then advances — guaranteeing that the slot pointed to by
#Indexat the top of the cycle is always the oldest one in the window. - The original code divided by
#IndexGrootteduring the fill-up phase, which underflows the average for the first N − 1 samples. The fixed code divides by#Teller, the actual number of valid samples. The result is a properly scaled mean from the very first scan.
IF rather than the MOD operator. On S7-1500 the compiler emits equivalent code, but the explicit IF is easier to step through in the SCL debugger and survives a downgrade to S7-1200 FW V4.2, which has historically had edge-case bugs with the MOD instruction on DINT operands.Data Type and Overflow Handling
The original code mixes DINT accumulator and INT output. The conversion is fine for millisecond-range cycle times with small windows, but it is a foot-gun the discussion correctly called out. The table below shows the minimum recommended accumulator width for a given signal range and window size.
| Sample range | Window N | Sum worst case | Safe accumulator |
|---|---|---|---|
| INT, ±32 767 | 10 | ±327 670 | DINT |
| INT, ±32 767 | 1 000 | ±32 767 000 | DINT |
| INT, ±32 767 | 100 000 | ±3 276 700 000 | DINT (just fits) |
| REAL, ±1.0E6 | 10 000 | ±1.0E10 | LREAL |
| DINT, ±2.1E9 | 2 | ±4.2E9 | LREAL (DINT overflows) |
Overflow in the accumulator is the most common cause of a running average that suddenly snaps to a garbage value after running for hours. The fix is to:
- Promote the accumulator to
LREALif N × max(|sample|) exceeds theDINTenvelope of ±2 147 483 647. - Use
ABSin an OB100 startup block to reset#Som,#Teller, and#Indexafter a CPU restart — the SCLFBretains its instance DB across warm restarts, so any pre-restart garbage propagates into the new scan. - Wrap the average in a saturation check before the
DINT_TO_INTcast, e.g.IF #Som > 32767 THEN #Gemiddelde := 32767; ELSIF #Som < -32768 THEN #Gemiddelde := -32768; END_IF;
For a strictly typed port to a PROFIBUS or PROFINET slave that expects REAL, divide using LREAL throughout and only convert at the boundary:
#SomLR := #SomLR - LREAL_TO_LREAL(#Tijd[#Index]) + LREAL_TO_LREAL(#GemetenTijd);
#Tijd[#Index] := #GemetenTijd;
#GemiddeldeREAL := #SomLR / LREAL_TO_LREAL(#Teller);
FB Wrapper for Reusability
Paste the array into a dedicated FB so the buffer is part of the instance DB and the call is a one-liner from any cyclic OB. The interface section enforces type safety and lets the optimizer keep the array in the work memory rather than reloading it on every cycle.
FUNCTION_BLOCK FB_RunningAverageDint
VAR_INPUT
Sample : DINT; // newest measured value
WindowSize : DINT; // N, 1..1024, validated in OB100
Reset : BOOL; // rising edge clears buffer and sum
END_VAR
VAR_OUTPUT
Average : DINT; // floor(Sum / Count)
Count : DINT; // 0..WindowSize
Valid : BOOL; // TRUE once Count == WindowSize
END_VAR
VAR
Buffer : ARRAY[0..1023] OF DINT; // sized at compile time
Sum : DINT;
Index : DINT;
RTrig_Reset : R_TRIG;
END_VAR
BEGIN
RTrig_Reset(CLK := Reset);
IF RTrig_Reset.Q THEN
// Memset equivalent — use FILL_BLK in the cyclic call site,
// not SCL (no native memset for multi-instance DBs)
Sum := 0;
Index := 0;
Count := 0;
Valid := FALSE;
// SCL cannot zero an unbounded array portably; do it in OB100
RETURN;
END_IF;
IF WindowSize < 1 OR WindowSize > 1024 THEN
// Guard against bad parameter from HMI; fail safe
Average := Sample;
Valid := FALSE;
RETURN;
END_IF;
Sum := Sum - Buffer[Index];
Buffer[Index] := Sample;
Sum := Sum + Buffer[Index];
Index := Index + 1;
IF Index = WindowSize THEN
Index := 0;
END_IF;
IF Count < WindowSize THEN
Count := Count + 1;
END_IF;
Valid := (Count = WindowSize);
IF Count > 0 THEN
Average := Sum / Count;
END_IF;
END_FUNCTION_BLOCK
To zero the Buffer array on first use, call the system function FILL_BLK from the startup OB (OB100):
// In OB100, once-only at CPU restart
FOR i := 0 TO 1023 DO
"dbRunningAvg".Buffer[i] := 0;
END_FOR;
Calling FILL_BLK from OB100 is also the official Siemens-recommended cold-start pattern for any large array of arithmetic operands; see the Siemens FAQ on circular buffer implementations for the matched warm-restart sequence.
Pointer-Based Variant for Arbitrary Data Types
If the same averaging logic must serve an INT, DINT, and REAL signal in three different places, factor the index logic into one FB and parameterize the buffer as VARIANT with a typed view. This avoids three nearly identical blocks. The full pattern is in the S7-1200/1500 SCL manual, section 6.4 "VARIANT and type-safe access"; the relevant fragment is:
VAR_IN_OUT
Buffer : VARIANT; // ANY-compatible typed pointer
END_VAR
VAR_TEMP
pBuf : POINTER TO DINT;
pType : UINT;
END_VAR
// Validate that the variant points to a DINT array of length N
pType := VariantType(Buffer);
IF pType <> TY_DINT THEN RETURN; END_IF;
pBuf := Buffer; // implicit deref into a typed pointer
pBuf^ := pBuf^ - pBuf[Index]; // illegal on S7-1500 — use Symbolik
On S7-1500 the pointer arithmetic in the last line is the bottleneck; the optimizer cannot always prove that Buffer and Buffer[Index] alias, so it generates two indexed loads and a STRD. For all but the smallest windows, three typed FBs (one per data type) are faster than one VARIANT block. The pointer form is worth the readability cost only when the window length N is small and the call site is in OB1 (≤ 1 ms cycle).
Edge Cases and Initialization
| Edge case | Symptom | Mitigation |
|---|---|---|
| First scan after power-on | Sum = 0, buffer = 0, average = 0 for N scans | OB100 zero-fill of buffer; pre-seed Sum = first sample × WindowSize on first valid input |
| WindowSize changed online via HMI | Index out of range, CPU goes to STOP with SF | Validate the HMI tag in OB1; clamp to last valid value; or implement a "shadow" WindowSize that only updates on rising edge of ApplyNewSize
|
| Reset pressed during fill-up | Sum carries stale partial sum; average jumps | On Reset rising edge, zero Sum, Index, Count, and the entire buffer; ignore output until Valid is TRUE |
| Sample = 0 (no flow, no weight) | Mean is correct but indistinguishable from a stuck sensor | Expose Count and let the calling logic check for "stuck at zero" by watching the variance of the last N samples |
| Negative sample (e.g. signed flow) | Sum drifts negative; output cast to INT saturates |
Keep accumulator in DINT or LREAL; saturate at conversion time only |
Integration with the Cyclic OB
Place the FB call in OB1 (the main cyclic OB) unless the cycle time is shorter than the analog input hardware update. For a 1 ms cyclic task calling the FB 1000 times per second, the execution time of the body is:
- S7-1214C (FW V4.4): ≈ 6 µs
- S7-1516 (FW V2.9): ≈ 0.8 µs
- S7-1518 (FW V2.9, optimized): ≈ 0.4 µs
Measured on TIA Portal V17 with the SCL optimizer enabled and the FB marked as optimized block access (the default in TIA Portal V15 and later). For PROFIBUS or PROFINET IRT update times below 250 µs, the OB must be assigned the same priority as the sync domain; otherwise the input that feeds Sample can change between the subtract and the add, corrupting the sum.
The standard pattern is to read the process image once at the top of OB1, call the FB once, and write the output back to the process image at the bottom of OB1. The I/O field devices used in the original problem (time measurement, likely a SIMATIC TIMER or a high-speed counter on the S7-1200 onboard I/O) are read in OB35 or OB40 to get deterministic timing, and the running-average FB is called from OB1 on the result.
Performance and Scan-Time Considerations
The dominant cost in the circular-buffer FB is the indexed load and store of Buffer[Index]. On the S7-1500 with the SCL optimizer enabled, the compiler emits a single STRD/STW pair against the instance DB base plus an offset computed from Index. The wrap-around IF is branch-predicted perfectly by the CPU because the branch is taken exactly once every N cycles.
The shift-register form using BLKMOV in STL is comparable up to about N = 16 on S7-1500; beyond that the circular buffer wins on every metric except code clarity, where the shift-register form is arguably easier to read. For applications where the buffer exceeds 1024 elements, partition the buffer into a ring of 256-element blocks and chain them with two index variables — this keeps the working set in L1 cache on the 1518-class CPU and reduces the per-scan memory bandwidth by an order of magnitude.
Verification and Commissioning
Use a watch table on the instance DB to step the FB through a deterministic test sequence. The following checks are sufficient to validate the implementation in less than five minutes on the bench.
-
Static zero — set
Sample= 0, force one cycle, confirmAverage= 0 andCount= 1. -
Step input — drive
Samplefrom 0 to 100 over 10 cycles with N = 10. After the 10th cycle,Averagemust equal 10 andCount= 10. -
Wrap — repeat with
Sample= 200 from cycle 11 onward.Averagemust step from 10 → 20 → 30 … → 110 → 200, then hold at 200. A flat line at any value other than 200 after cycle 20 indicates the index is not wrapping. -
Reset — pulse
Resetfor one cycle.Sum,Index, andCountmust all be 0 andValidmust drop to FALSE. -
Type safety — with
WindowSizedriven from an HMI tag, set the tag to 0 or to a value > 1024 and confirm the FB outputsSampleunchanged andValid= FALSE. The CPU must not transition to STOP.
For an in-process sanity check, add a one-line variance calculation to the calling logic: variance = (SumSquared / N) − (Average × Average). A sudden jump in variance on a normally steady process is a strong signal that the input channel is noisy or that the FB is being called from two OBs at different priorities.
Common Mistakes and Field-Proven Caveats
- Re-initialization on every scan — placing the FILL_BLK call inside OB1 instead of OB100. The array is zeroed every cycle, and the average follows the instantaneous input with a one-cycle lag.
- Division by N instead of Count — the second bug. The first N − 1 samples produce a divided-by-constant result, which on a ramp input looks like a slow attack on the final value.
- Forgetting to mark the FB as "optimized block access" — on TIA Portal V14 and earlier the default is non-optimized, which forces the SCL compiler to load the entire array into the work memory on every indexed access. The FB still works, but scan time goes up by 10× to 20× for large N.
- Calling the FB from a cyclic interrupt and from OB1 — two writers to the same instance DB. The IEC 61131-3 standard permits this only if the two OBs are mutually exclusive in time (one is the higher-priority interrupt of the other). Otherwise the buffer can be read mid-write.
Standards and Documentation Cross-Reference
- SIMATIC S7-1200/1500 SCL programming and operating manual — language reference, section 6.4 (VARIANT), section 7.3 (array handling).
- Siemens Industry Online Support FAQ 39333120 — circular buffer implementation pattern, including the FILL_BLK cold-start sequence used in this article.
- SIMATIC S7-1500 system manual, chapter on cyclic and time-of-day interrupt OBs — for assigning the running-average FB to the correct OB priority.
- IEC 61131-3:2013, clause 6.10 — defines the ARRAY data type and the semantics of the
:=assignment used to overwrite a single element.
Why does my running average drift over time on a steady input?
The accumulator is overflowing. The sum N × sample can exceed the DINT envelope (±2 147 483 647) after several hours even on a small signal. Promote the accumulator to LREAL and divide only at the output.
The first N samples show a slow ramp instead of the correct mean — what changed?
You are dividing by the window size N instead of the populated count. Use the count variable that is incremented until it reaches N; the average is then correct from the very first sample.
Can I run the same FB instance on an INT, a DINT, and a REAL input?
On S7-1500 with the SCL optimizer, no — there is no implicit cast inside an indexed load. Either instantiate three typed FBs (one per data type) or use the VARIANT-based form from the SCL manual, accepting a small performance penalty.
Should I call the FB from OB1 or from a cyclic interrupt OB?
OB1 for a slow process (cycle ≥ 10 ms). Use OB30..OB38 only if the input is sampled on a fixed sub-cycle (e.g. a 1 ms high-speed counter) and the resulting mean must update on the same sub-cycle. Never call the same FB instance from both OB1 and a cyclic interrupt.
How do I clear the buffer safely on a CPU restart?
Place a FILL_BLK call against the instance DB in OB100 (startup OB). FILL_BLK runs once-only and zero-fills the entire array before OB1 takes over; the FB then sees a deterministic empty buffer on its first scan.