Overview
Indirect addressing on SIMATIC controllers is, by design, more expensive than direct symbolic addressing. The CPU must first read the index or pointer operand, resolve the effective address, and only then execute the read or write at the resolved target. On a S7-300/S7-400 the cost is visible in scan time. On a S7-1200/S7-1500 the cost is still real, but modern TIA Portal toolchain and optimized data block architecture let you remove most of the reasons to use raw pointer arithmetic in the first place.
This reference explains the instruction-level performance delta between a legacy S7-417 (S7-400 family) and a S7-1518 (S7-1500 family), then moves into the engineering patterns TIA Portal provides so that the same reusability that used to require POINTER, ANY, or P# in STL is now expressed with arrays, UDTs, and symbolic tag access. The intent is to give a controls engineer a defensible answer to the question "does indirect addressing slow the CPU?" with measured numbers, the official Siemens guideline that drives the recommendation, and the code patterns that deliver the best cycle time on current hardware.
Direct vs Indirect Addressing Mechanics
A SIMATIC instruction in STL/SCL can resolve its operand through three addressing modes:
- Direct, symbolic: The compiler resolves the tag name to an absolute address at compile time. The MC7/SCL runtime issues a single load/store against that absolute address.
-
Direct, absolute: The engineer writes the absolute address in the instruction (e.g.
L MW 20). Still one load, but no symbolic check. - Indirect: The operand is computed at runtime from another tag (index, pointer, ANY, VARIANT). The CPU must first read the index/pointer word, add an offset, then access the target. This is at minimum two memory operations before the actual data move.
In STL on a S7-400 this is literally what you see in the status word: an indirect access such as L IW [MD 10] forces the CPU to fetch MD10, mask the byte/word part, and then issue the peripheral load. The instruction list reference in the STEP 7 classic manual tags this with the "indirect" descriptor for a reason.
On the S7-1500 the underlying mechanism is the same, but the operand lookup is performed by the optimized block compiler. The compiler can often constant-fold a loop index when the loop bounds are known, eliminating what would have been an indirect access in STL. This is why a hand-written SCL FOR loop over an array frequently benchmarks faster than a hand-written STL LOOP with P# arithmetic on the same 1518.
Instruction Cost Model
The minimum cost of a fully-indirect read can be modeled as:
T_indirect = T_index_fetch + T_address_resolve + T_target_access
For a fully-symbolic direct read the cost collapses to T_direct = T_target_access because the address is baked into the compiled code. On any Siemens PLC the ratio T_indirect / T_direct is roughly 1.5 to 3.0 for simple bit/word operations, and 2.0 to 4.0 for floating-point loads through a pointer. The absolute time is small, but at high scan rates on a small CPU it can dominate the OB1 budget.
Performance Comparison: S7-417 vs S7-1518
The numbers below are taken directly from the Siemens performance data sheets and the Programming Guideline for S7-1200/S7-1500 (entry ID 81318674). They represent execution time per instruction at default priority class 1 (OB1).
| Operation | S7-417 (S7-400) | S7-1518-4 PN/DP (S7-1500) | Speedup Factor |
|---|---|---|---|
| Bit operation (U, O, =, S, R) | 7.5 ns | 1 ns | 7.5x |
| Bit-string operation (UW, OW, XOW) | 7.5 ns | 2 ns | 3.75x |
| Fixed-point math (+I, -I, *I, /I) | 7.5 ns | 2 ns | 3.75x |
| Floating-point math (+R, *R, /R) | 15 ns | 6 ns | 2.5x |
The S7-1518-4 PN/DP (6ES7518-4AP00-0AB0) is the current high-end of the S7-1500 line; the S7-417 (6ES7417-4HL04-0AB0) was the high-end of the S7-400 family and is now in product phase "out of production" per the Siemens product portfolio. Across the four instruction groups the 1518 is between 2.5x and 7.5x faster per instruction. The S7-319 (mid-range S7-300) is closer to the 1518 on some fixed-point operations but is still measurably slower on floating point and bit-string work.
L IW [MD10] instead of L IW 20) add the index/pointer fetch overhead discussed above. On the 1518 that overhead is small (single-digit ns per indirection) but on a S7-314 with a busy OB1 it is often the difference between meeting and missing a 10 ms scan time budget.What This Means for Indirect Overhead
Assume the same indirect penalty on both families. A S7-417 with a 7.5 ns base bit operation plus a 7.5 ns indirect fetch settles around 15 ns per bit access through a pointer. The 1518 with a 1 ns base plus the same constant indirect overhead is closer to 4 to 5 ns. The penalty as a percentage of the instruction shrinks as the hardware generation increases, which is one reason the Siemens guideline explicitly tells S7-1200/S7-1500 programmers to prefer optimized data blocks and arrays instead of pointer arithmetic: the headroom gained by the new CPU is so large that you can use the symbol-friendly access pattern without paying a meaningful cycle-time tax.
TIA Portal Optimization: Optimized Blocks and Arrays
The single most important performance setting in TIA Portal for any S7-1200/S7-1500 program is the "Optimized block access" property. The Siemens guideline (entry ID 81318674) calls this out as a foundation for any performance-sensitive program. With an optimized block the compiler is allowed to:
- Reorder struct members to align on natural word boundaries.
- Store the symbolic name and the absolute offset in a side table so symbolic debug does not break when the layout is reordered.
- Generate code that uses the symbolic identifier directly, eliminating the legacy
OPN DB/DBX/D BWindirection.
You enable optimization on a data block from its properties dialog in TIA Portal, or by adding the attribute directly in the source:
DATA_BLOCK "dbRecipe"
{ S7_Optimized_Access := 'TRUE' }
VERSION : 0.1
STRUCT
temperature : REAL;
pressure : REAL;
valveOpen : BOOL;
END_STRUCT;
END_DATA_BLOCK
Combined with optimized access, arrays of structured types are the modern replacement for pointer-based indexing. The SCL compiler lowers arr[i].field to a single load with a scaled index, which is faster than any P#DB.dbx0.0 arithmetic in classic STL.
Variable Index Without Pointer
Indexed access into an array in SCL is not considered "indirect addressing" in the performance sense used by the guideline. The S7-1500 backend has a dedicated addressing mode for ARR[i] that uses a base register plus a scaled offset, which is one CPU cycle plus a memory access. Compare this to a fully-qualified pointer dereference in STL, which is at least two cycles plus two memory accesses.
// Tight loop over an array, variable index - the modern pattern
FOR #i := 0 TO 99 DO
#arrFiltered[#i] := #arrRaw[#i] * #gain;
END_FOR;
If i is bounded at compile time, the compiler may even unroll the loop. If i is read from an HMI tag or another block output, the index is still resolved in a single instruction on the 1518.
SCL Indexed Access Patterns
SCL on the S7-1500 supports three indexed access patterns. The choice matters for both performance and maintainability.
1. Direct Symbolic Array Access (Fastest, Recommended)
// Symbolic access with literal index - constant-folded at compile time
#temp := "dbRecipe".recipe[5].temperature;
// Symbolic access with variable index - single-instruction on 1518
#temp := "dbRecipe".recipe[#idx].temperature;
2. AT Construct (View Overlay, No Copy)
Use the AT keyword to overlay a structured view onto a byte array. This is the right pattern for parsing a Profinet/Modbus buffer that you received as ARRAY OF BYTE.
FUNCTION_BLOCK "fbParseFrame"
VAR
rawBuffer : ARRAY[0..31] OF BYTE;
frameView : AT "rawBuffer" : STRUCT
cmd : BYTE;
length : UINT;
value : REAL;
END_STRUCT;
END_VAR
BEGIN
#value := #frameView.value;
END_FUNCTION_BLOCK
The AT view does not copy data. The compiler emits a single load against the resolved offset. This is the recommended pattern for protocol parsing on the S7-1200/S7-1500 and is documented in the TIA Portal help under "AT view on a data block".
3. VARIANT Parameter (Slowest, Most Flexible)
For block parameters that must accept any data type, use VARIANT. The receiving FB uses VariantGet / VariantPut to dereference, which is significantly slower than a typed reference parameter. Reserve it for libraries that need type-agnostic behavior, and benchmark if used in a hot path.
FUNCTION_BLOCK "fbScaleGeneric"
VAR_INPUT
pSource : VARIANT;
END_VAR
VAR_TEMP
rValue : REAL;
bOK : BOOL;
END_VAR
BEGIN
#bOK := VariantGet(SRC := #pSource, DST => #rValue);
IF #bOK THEN
// ...
END_IF;
END_FUNCTION_BLOCK
UDT-Based Data Modeling for Remote I/O
When you read or write data from a Profinet/Profibus/Ethernet/RS232/RS485/Modbus device, the historical pattern was to declare a single shared DB and walk it with P# pointers to extract fields. On a S7-1500 that is unnecessary in nearly every case. The modern pattern is to model the device payload as a UDT and then store an ARRAY of that UDT in the receive DB.
TYPE "UDT_AnalogModule"
VERSION : 1.0
STRUCT
channel0 : INT;
channel1 : INT;
channel2 : INT;
channel3 : INT;
status : WORD;
END_STRUCT;
END_TYPE
DATA_BLOCK "dbAnalog"
{ S7_Optimized_Access := 'TRUE' }
VERSION : 0.1
STRUCT
modules : ARRAY[1..16] OF "UDT_AnalogModule";
END_STRUCT;
END_DATA_BLOCK
Now a single BLK_MOV or MOVE_BLK instruction copies the entire Profinet frame into the UDT array. The fields are then accessed by symbolic name and index, with no pointer anywhere in the code. This is the pattern Siemens recommends in the S7-1500 communication manual and in the programming guideline entry ID 81318674.
Symmetrically, an FB that sends UDP or Modbus data takes the DB number as an input, performs a single bulk copy into the send buffer, and returns. No P# arithmetic in the parameter list is required:
// Pseudocode for a UDP send FB that takes a UDT DB
FUNCTION_BLOCK "fbUdpSendGeneric"
VAR_INPUT
iDbNumber : INT; // ANY DB number, not a pointer
iOffset : DINT; // optional start offset
iLength : INT;
END_VAR
// Internally: POKE_BLK from #iDbNumber, #iOffset, #iLength into the socket buffer
The interface still has indirect elements (the DB number), but they are simple integer parameters, not pointer arithmetic. The runtime cost of looking up a DB by number is a single instruction on the 1518, and the bulk copy is a DMA-class block move.
Pointer, ANY, and P# in Legacy Code
On the S7-1500, STL is still available, and so are POINTER, ANY, and P#. The Siemens guideline does not prohibit them. The guideline is clear that you cannot use them in optimized blocks - this is the reason a direct port of a S7-400 STL block that uses P# parameters will fail to compile or will throw an "invalid pointer" error at runtime when the destination block is set to optimized access. You must either:
- Convert the block to non-optimized access (Properties → Attributes → uncheck "Optimized block access"), or
- Refactor the block to use UDTs, arrays, and symbolic access.
Option 1 is appropriate for legacy communication blocks that have to interop with classic S7-300/400 stations. Option 2 is the long-term path and the one Siemens recommends for new development on the S7-1200/S7-1500 platform. The cost of option 1 is that the block cannot use symbolic access for its instance data, and a non-optimized block cannot be downloaded into a S7-1200 firmware older than V4.0 - so on a S7-1212C, for example, you may not have a choice on legacy interfaces.
P# Syntax Reference
// Legacy pointer to a data block field
P#DB100.DBX 20.0 BYTE 12 // pointer to 12 bytes starting at DB100.DBX20.0
// Legacy ANY for a full DB area
P#DB 100 BYTE 1000 // ANY pointing at the first 1000 bytes of DB100
// Load pointer into accumulator
LAR1 P##Source
L B [AR1, P#0.0] // indirect load of a byte
Every [AR1, P#x.y] or [MDx] reference is an indirection at runtime. The CPU must load AR1 or MDx, add the offset, and then access. Replace each occurrence with an array element wherever the data layout allows it.
Edge Cases and Caveats
Watchdog and OB1 Overrun
On a S7-1516 (6ES7516-3AN02-0AB0) and below, a tight loop with 200,000 indirect accesses can still overrun OB1. The fix is either to chunk the loop across multiple OB1 cycles, or to lower the number of indirect accesses by uprating the data structure to an array of UDTs. A 1518 has roughly 2x the per-instruction throughput of a 1516, but the same code pattern with the same loop count will still eventually overrun - the math is linear in the loop count.
Multi-Instance Data Blocks and Indirect Access
Multi-instance FBs share a single instance DB. Indirectly indexing into a multi-instance DB requires either a non-optimized block (slower, legacy) or a typed ARRAY of the FB's instance data (the modern approach). The latter requires TIA Portal V14 or later and is supported on S7-1500 firmware V2.0 and above.
Slice Access on S7-1500
The S7-1500 introduced slice access, e.g. %DB1.DBX0.0:3 to read 3 bits starting at DB1.DBX0.0 as a single operand. Slice access is symbolic, fully optimized, and benchmarks faster than any equivalent UW + shift-and-mask pattern. Prefer slices over bit-string operations on the 1518 when the instruction set allows it.
HMI-Triggered Indirect Access
When the index of an array access comes from an HMI tag, you introduce a round trip through the HMI connection. The PLC must read the HMI tag, validate it, and then resolve the array access. If the HMI tag is updated on every screen refresh, you can multiply OB1 load by 2 to 5x during a screen change. The pattern in TIA Portal is to debounce the HMI index input and apply limit checks:
// Defensive HMI index clamp
#idx := LIMIT(1, "HMI".recipeIndex, "dbRecipes".recipe.HIGH);
#temp := "dbRecipes".recipe[#idx].temperature;
Verification and Benchmarking
Before declaring a TIA Portal program "fast enough", run the following checks against the actual hardware on the bench:
- OB1 cycle time: Open the online → diagnostics → cycle time view. Compare the current OB1, OB35 (or whatever cyclic OB is configured), and the longest OB in the project. The S7-1500 exposes the worst-case and last-cycle time on the device web page and in the TIA Portal online view.
-
Indirect access count: Use the SCL compiler listing (Project → Compiler → Generate compiler listing) and grep for
indirectin the assembly output. A well-optimized program should have zero or near-zero indirect accesses in hot paths. - Profiler run: On a 1518 with a valid runtime license, the S7-1500 performance monitor traces per-instruction timing. Use it on a representative scan and confirm the floating-point hot path runs in the expected nanosecond range.
- Watchdog margin: The S7-1500 default OB1 watchdog is 150 ms. The 1518 default is 6000 ms. Confirm the project's worst-case OB1 cycle time is comfortably below the configured watchdog; keep a 50% safety margin for transient events such as Profinet reconfiguration.
Sample Performance Calculation
Suppose a program performs 50,000 indirect word reads and 10,000 indirect real (floating-point) reads per OB1 cycle on a 1518. The estimated contribution to OB1 is:
T_bit = 50,000 * (1 ns + 3 ns indirect) = 50,000 * 4 ns = 0.2 ms
T_real = 10,000 * (6 ns + 12 ns indirect) = 10,000 * 18 ns = 0.18 ms
T_indirect_total ≈ 0.38 ms per OB1 cycle
That is well under the 1 ms OB1 budget for most applications, but on a S7-1510C (6ES7510-1DJ01-0AB0) at roughly half the per-instruction throughput the same code path would be about 0.7 to 0.8 ms, which starts to crowd the budget. This is the practical case for migrating to array/UDT patterns: on the lower-end S7-1500 CPUs the indirect penalty still matters.
Migration Strategy from S7-300/400 to S7-1500
The migration path most teams take is to port the program first, then refactor for optimized blocks. The recommended sequence is:
- Port the program with TIA Portal's migration tool, accepting the default non-optimized access on all migrated blocks. The tool will flag any
P#arithmetic and surface it in the migration report. - Compile and run on the target S7-1500 with the migrated code. Validate functional equivalence against the S7-400 reference.
- For each non-optimized block, switch the access to optimized. Resolve any
POINTER/ANY/P#uses by introducing a UDT and an array parameter. - Re-run the benchmark. Expect a measurable reduction in OB1 time, with the largest gains on floating-point and bit-string hot paths.
- Lock down the project with the Siemens guideline checklist from entry ID 81318674 before commissioning.
Standards and Documentation References
- Siemens "Programming Guideline for S7-1200/S7-1500" - entry ID 81318674
- SIMATIC S7-1500 CPU 1518-4 PN/DP device manual - entry ID 109751826
- SIMATIC S7-1500 S7-1500T Motion Control function manual - entry ID 109749264
- STEP 7 (TIA Portal) SCL programming and operating manual - entry ID 109751784
- SIMATIC S7-1500 communication function manual (Profinet, Modbus, UDP) - entry ID 109744224
FAQ
Does indirect addressing really slow the CPU on a S7-1500?
Yes, but only in absolute terms. The penalty is typically 3 to 12 ns per indirect access on a S7-1518. With optimized blocks and array/UDT access patterns you can reduce indirect accesses to near zero in the hot path, so the practical answer is "indirect addressing slows the CPU by a measurable but very small amount, and you can almost always avoid it on the S7-1200/S7-1500 platform".
How much faster is the S7-1518 than the S7-417 per instruction?
For bit operations the 1518 is roughly 7.5x faster (1 ns vs 7.5 ns). For bit strings and fixed-point math it is about 3.75x faster (2 ns vs 7.5 ns). For floating-point math it is about 2.5x faster (6 ns vs 15 ns). These are typical instruction times at OB1 priority, taken from the Siemens S7-1500 performance data and the programming guideline (entry ID 81318674).
Can I still use POINTER, ANY, and P# on a S7-1500?
Yes, in STL, but only in blocks that are not optimized. If you need symbolic access to instance data and you also need raw pointer arithmetic, you have to keep the block non-optimized, which loses the compiler-driven reordering and symbolic access benefits. The recommended long-term path is to refactor pointer-based code into UDT arrays.
What replaces P# in SCL on the S7-1500?
Use arrays of UDTs for indexed access to structured data, the AT construct for overlaying structured views on byte arrays (for example, parsing a Profinet or Modbus buffer), and VARIANT parameters for type-agnostic block interfaces. All three are documented in the STEP 7 TIA Portal SCL manual and supported on S7-1200 firmware V4.0 and S7-1500 firmware V2.0 and above.
How do I measure the OB1 cycle time impact of indirect access?
Open the project online in TIA Portal, navigate to the device diagnostics, and view the cycle time panel for the active OB. The S7-1500 also exposes worst-case and last-cycle time on the device web page. For per-instruction analysis, use the S7-1500 performance monitor (a licensed add-on) and run a representative scan profile.