S7-1500 CPU Instruction Execution Times: Reference Manual

David Krause14 min read
S7-1200SiemensTechnical Reference
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

S7-1500 CPU Instruction Execution Times: Reference Manual

Execution time on a Siemens S7-1500 CPU is the wall-clock duration a single statement or block of statements consumes inside the OB/FC/FB call stack. Unlike S7-300/400, where every instruction has a published microsecond figure in the Instruction List manual, the S7-1500 family documents timing in categories: bit operations, word operations, fixed-point arithmetic, floating-point arithmetic, and transfer/comparison operations. Per-instruction nanosecond figures are obtained by combining the category baseline with measured RUNTIME samples, or by reading the device-specific data from the Siemens Function Manual.

This reference consolidates the official Siemens cycle and response times manual, the S7-300 Instruction List as a historical baseline, and the RUNTIME instruction to give engineers a complete methodology for estimating and measuring S7-1500 instruction execution time.

S7-1500 instruction timing is CPU-firmware-version dependent. Always re-validate after a firmware upgrade (e.g., V2.5 → V2.9 → V3.0) on the target CPU. Numbers cited from a manual are typical and not worst-case under interrupt load.

1. Reference Documentation

Three primary Siemens documents govern S7-1500 instruction timing. Always cite the exact manual edition used for the calculation.

Document Manual Name Entry ID Contents
S7-1500 / ET 200SP cycle & response times SIMATIC S7-1500, S7-1500R/H, ET 200SP, ET 200pro Cycle and response times 109751809 (support.industry.siemens.com) Table 3-2: operation durations by category; cycle time formulas; response time composition
S7-300 Instruction List S7-300 Instruction List, CPU 31xC, CPU 31x, IM 151-7 CPU, IM 151-8 CPU, IM 154-8 CPU, BM 147-1 CPU, BM 147-2 CPU 13206730 (support.industry.siemens.com) Per-instruction µs figures (e.g., POP 0.03 µs, TAK 0.06 µs on CPU 317); reference for cross-platform timing ratios
S7-1500 Programming Guideline SIMATIC S7-1500 Programming Guideline 81318674 (support.industry.siemens.com) Optimized block design, multi-instance DBs, symbolic vs. absolute access overhead

Open the cycle and response times manual at SIMATIC S7-1500, ET 200SP, ET 200pro Cycle and response times and the S7-300 list at S7-300 Instruction List (A5E00105517-10) for the canonical figures.

2. Why S7-1500 Documents Timing in Categories

The S7-1500 instruction set runs on a different execution engine than the S7-300/400 accumulator-based architecture. Three architectural shifts make per-instruction µs tables impractical:

  1. Optimized compiled code — TIA Portal compiles the entire OB/FC/FB before download. A single STL line may map to multiple micro-operations in the firmware's intermediate language.
  2. Cache effects — the first execution of an instruction may be slower than subsequent executions because the microcode cache is cold.
  3. Data type width is variable — LREAL, LWORD, and LINT operations are typically slower than REAL or DINT, but the ratio is firmware-dependent.

Siemens therefore publishes category execution times in Table 3-2 of the cycle and response times manual. The engineer sums the category figures for the code path to estimate cycle time, then validates with RUNTIME in TIA Portal.

3. Instruction Categories and Default Timing Ranges

The cycle and response times manual groups instructions into the following categories. Values below are typical ranges for a standard CPU 1515-2 PN at firmware V2.9; refer to the manual for the exact CPU in the project.

Category Typical Range (ns) Examples
Bit operations 10 – 60 AND, OR, XOR, assignment, set/reset, edge detection (positive/negative)
Word operations (timers, counters, comparators) 30 – 200 TON, TOF, TP, CTU, CTD, EQ_I, NE_R
Fixed-point arithmetic (16/32-bit) 30 – 100 ADD_I, SUB_D, MUL_I, DIV_D, MOD
Floating-point arithmetic (32/64-bit) 60 – 300 ADD_R, MUL_R, SIN, COS, EXP, LN, SQRT
Transfer / move 20 – 80 MOVE, BLKMOV, FILL, NONDISS, Serialize
Conversion 40 – 200 INT_TO_REAL, REAL_TO_DINT, BYTE_TO_WORD
Program control / block call 200 – 1500 CALL, JMP, JMPN, RET, BE, UC, CC
Indirect addressing 80 – 400 PEEK, POKE, field access with variable index
Floating-point division (DIV_R / DIV_LREAL) is the slowest single arithmetic operation. Replace repeated division with a precomputed reciprocal where the divisor is constant. Use MUL_R/MUL_LREAL whenever the divisor is mathematically convertible.

4. Bit Operation Timing

Bit operations are the fastest class on S7-1500 because they map directly to single-cycle boolean ALU operations. The cycle and response times manual lists the baseline for one logical AND, OR, or assignment at roughly 10 ns on a CPU 1515-2 PN at 100 MHz equivalent throughput.

Edge-detection instructions (FP, FN, R_TRIG, F_TRIG) carry an additional implicit load of the previous bit state. Expect 20–40 ns per edge operation instead of the bare 10 ns baseline. Use R_TRIG/F_TRIG function blocks (single-instance DBs) for clarity; avoid in-line edge logic on hot loops.

Set/reset operations (S, R, RS, SR flip-flops) on individual bit tags cost the same as AND/OR because the underlying microcode is a conditional write. The cost appears in the data access path, not the ALU.

4.1 Bit operation timing example

// Network 1 (STL equivalent of a typical machine interlock)
A   %I0.0         // 10 ns
A   %I0.1         // 10 ns
=   %Q0.0         // 10 ns
A   %I0.2         // 10 ns
FP  %M100.0       // 30 ns (FP loads previous M100.0)
=   %Q0.1         // 10 ns
// Subtotal: ~80 ns excluding I/O image update

5. Word, Timer, and Counter Timing

The S7-1500 timer and counter instructions are IEC 61131-3 system blocks (TON, TOF, TP, TONR, CTU, CTD, CTUD). Internally they maintain an instance DB, so the cost is the function call (~200 ns) plus a few word operations on the time/count tag.

Approximate per-call figures for a CPU 1515-2 PN:

Instruction Typical Cost (ns) Notes
TON / TOF / TP 300 – 600 Includes 1 ms tick, 16/32-bit time base, edge handling
TONR (retentive on-delay) 350 – 700 Adds retentive read/write to instance DB
CTU / CTD 200 – 400 Single counter register, no edge detection overhead
CTUD (up/down) 300 – 500 Two register update paths
EQ_I / NE_I / GT_I / LT_I / GE_I / LE_I 30 – 50 Integer comparator
EQ_R / NE_R / GT_R / LT_R / GE_R / LE_R 50 – 100 Real comparator, NaN handling adds ~20 ns
MOVE (single tag) 20 – 40 Symbolic access slightly slower than absolute
BLKMOV / UMOVE / BMOVE 50 + 5/element Bulk transfer cost scales linearly with element count
FILL_BLK / UFILL_BF 60 + 4/element Fill scales similarly to block move

For TON, TOF, TP, the value range is 0–9999999 ms with a time base of 1 ms by default. This time-base tick is internal to the IEC timer; it does not consume a separate hardware timer slot on the CPU.

6. Fixed-Point and Floating-Point Arithmetic

Fixed-point arithmetic operates on INT, DINT, WORD, DWORD, and BCD. Floating-point operates on REAL and LREAL. Typical CPU 1515-2 PN costs:

Operation 16-bit (ns) 32-bit (ns) 64-bit (ns)
ADD / SUB 30 40 80
MUL 50 60 120
DIV 80 120 240
MOD 100 150 300
SIN / COS / TAN n/a 250 – 500 500 – 900
EXP / LN / SQRT n/a 300 – 700 600 – 1300
ABS / NEG 20 30 50
TRUNC / ROUND / CEIL / FLOOR 30 40 80

Conversion operations have their own cost band:

  • INT_TO_REAL / REAL_TO_DINT: 40–100 ns
  • BYTE_TO_WORD / WORD_TO_DWORD: 30–60 ns
  • BCD_TO_INT / INT_TO_BCD: 60–120 ns
  • CHAR_TO_STRING / STRING_TO_CHAR: 200–500 ns (string header manipulation)

The S7-1500 hardware floating-point unit is IEEE 754 compliant. NaN propagation, denormal handling, and overflow traps follow the IEEE default, so the cost of an exception path can be 2–3× the normal case. Guard division inputs to avoid divide-by-zero NaN results.

7. The RUNTIME Measurement Instruction

The S7-1500 RUNTIME instruction is the canonical way to measure instruction execution time on the target hardware. It captures the CPU's internal high-resolution counter at start and end of a code region, returning elapsed time in nanoseconds.

7.1 Interface

// SCL signature
RUNTIME(MODE := 0, MEM := <instance_DB>);
// MODE = 0: write start timestamp to MEM
// MODE = 1: write end timestamp and compute elapsed time into MEM

7.2 Usage pattern

// FB "PerfProbe" body (SCL)
// At scan start
RUNTIME(MODE := 0, MEM := DB_Perf.runtime);   // 30 ns overhead

// Code to measure
// ...

// At scan end
RUNTIME(MODE := 1, MEM := DB_Perf.runtime);   // 60 ns overhead
// DB_Perf.runtime now holds elapsed nanoseconds as LREAL

The result tag is a LREAL holding the elapsed time. RUNTIME itself consumes 30–90 ns depending on MODE, so subtract the overhead from the measurement when comparing to category baselines.

7.3 Caveats

  • RUNTIME must be called with a valid MEM pointer (instance DB or global DB). Passing a NIL pointer raises OB 121 (programming error).
  • The counter is the CPU's monotonic timestamp counter; it does not roll over within the cycle.
  • RUNTIME is available in all S7-1500 CPUs and the S7-1500R/H redundant CPUs as of firmware V1.8.

8. S7-300 Cross-Reference for Timing Ratios

The S7-300 Instruction List at S7-300 Instruction List (A5E00105517-10) still publishes per-instruction µs figures, e.g. POP 0.03 µs and TAK 0.06 µs on CPU 317. Use it as a ratio baseline when porting an S7-300 program and assessing whether the S7-1500 will hit its cycle budget.

Instruction (CPU 317 baseline) S7-300 Time S7-1500 Time (typical) Speedup
POP (stack pop, no accumulator write) 0.03 µs ~10 ns 3×
TAK (swap top two accumulators) 0.06 µs ~20 ns 3×
+I (16-bit integer add) 0.05 µs ~30 ns 1.7×
+R (32-bit real add) 0.30 µs ~50 ns 6×
/R (32-bit real divide) 0.90 µs ~150 ns 6×

The S7-1500 typically runs 1.5× to 6× faster than the S7-317, with the largest gain on floating-point because the S7-1500 has a dedicated hardware FPU. Integer logic gains are smaller because the S7-317 was already bit-optimized.

Do not assume the S7-1500 always beats the S7-300 by a fixed factor. The 1500's symbolic access path can be slower than the 300's absolute access path for very tight loops. Benchmark with RUNTIME, not by analogy.

9. Cycle Time Calculation

The cycle time of an OB1 is the sum of:

T_cycle = T_image + T_OB1_user + T_comm + T_interrupt + T_diag + T_reserve

Where:

  • T_image: process image update, ~50 µs for 1 KB of I/O at 100 µs PI update rate
  • T_OB1_user: sum of all user-code execution times (the focus of this reference)
  • T_comm: PG/OP/HMI communication, scaled with the configured PG/OP connection count
  • T_interrupt: hardware and time-delay OB execution, weighted by frequency
  • T_diag: system diagnostics and module status reads
  • T_reserve: 5–10% reserve for transient interrupts

The cycle and response times manual provides the exact formulas and coefficient tables per CPU model. Use the manual's cycle time configurator worksheet rather than hand-calculating for production programs.

10. Estimating User-Code Execution Time

Use the following four-step procedure to estimate T_OB1_user before downloading to the CPU:

  1. Inventory the code — count each instruction in the OB1 call tree (OB1 + every FC/FB it invokes).
  2. Classify each instruction — map to a category from Section 3 above.
  3. Sum by category — use the typical range midpoint as the per-instruction figure; multiply by call-site count for reused FBs.
  4. Add block call overhead — every CALL consumes 200–1500 ns for stack frame setup; add 100–300 ns for multi-instance DBs.

10.1 Worked example

Suppose OB1 of a packaging line contains 200 bit ops, 50 word ops, 30 fixed-point math ops, 20 floating-point math ops, and 10 timer calls, plus 5 FC calls.

Class Count Per-Op (ns) Subtotal (ns)
Bit 200 30 6 000
Word (timer/counter) 50 80 4 000
Fixed-point 30 60 1 800
Floating-point 20 150 3 000
Timer (TON) 10 500 5 000
FC call 5 800 4 000
Estimated T_OB1_user 23 800 ns ≈ 24 µs

This is the user-code slice only. Add T_image, T_comm, and the diagnostic reserve per the cycle manual. On a CPU 1515-2 PN the total cycle is typically 1–3 ms for a packaging OB of this size.

11. Measuring Cycle Time Online

Two on-CPU tools measure cycle time directly:

  1. Online → Diagnostic → Cycle time in TIA Portal — shows the last 10 cycle times including min, max, and current.
  2. RD_SYS_T and RD_LOC_T for time-stamping OB1 runs externally; useful for high-precision cycle analysis.

For program-section timing inside the OB, use RUNTIME. For OB-level timing, use the diagnostic buffer or the TIA Portal cycle statistic.

11.1 Verification checklist

  • [ ] Cycle max < 80% of the configured OB1 scan watchdog
  • [ ] RUNTIME peaks documented and within the calculated T_OB1_user ± 15%
  • [ ] Floating-point hot path not in OB35/time-critical OB if cycle > 5 ms
  • [ ] No divide-by-zero possible in floating-point paths (NaN check present)
  • [ ] All TIMERs use the same time base (avoid mixing 1 ms and 10 ms in one loop)

12. Performance Optimization Best Practices

Before optimizing, profile. The S7-1500 spends the majority of its cycle on a small number of hot paths. Use the RUNTIME probe in each suspected hot path before changing code.

  1. Replace DIV with MUL by reciprocal — precompute 1.0/x and multiply. A real divide at 150 ns becomes a real multiply at 60 ns.
  2. Use IEC timer edge-flag alternatives — the S7-1500 system TON is >300 ns; for fast single-shot delays, derive a time flag from a 1 ms cyclic interrupt instead of using TON in the hot path.
  3. Cache real constants — a LREAL constant load is faster than converting from INT at every scan.
  4. Avoid multi-instance DBs for short FBs called millions of times — the DB write-back overhead is real; single-instance DBs in a known hot region can be faster for legacy code.
  5. Use BLKMOV with typed arrays instead of structured field-by-field copy when the layout matches.
  6. Run floating-point math in a slower OB (e.g., OB35 at 100 ms) and use latched results in OB1.
Do not over-engineer. A 2 ms cycle that meets the application is correct engineering. Optimize only when the cycle watchdog fires, a servo axis demands lower jitter, or the OB is migrating to a smaller CPU.

13. Common Errors When Using RUNTIME

Symptom Cause Fix
RUNTIME returns 0 MEM points to an uninitialized instance DB Initialize DB; verify the DB is not optimized (or, if optimized, use a known-offset tag)
OB 121 programming error Invalid MEM pointer Pass a DB number, not a tag address
Reading inconsistent values Two OBs write the same MEM tag Use a per-OB MEM tag
Value rolls over at 32-bit boundary MEM declared as REAL, not LREAL Declare MEM as LREAL
Time appears doubled MODE = 1 called twice in a row Wrap the section with a single MODE = 0 / MODE = 1 pair

14. Field Commissioning Procedure

  1. Pre-load estimate — use Section 10 to compute T_OB1_user.
  2. Download and go online with the project.
  3. Open TIA Portal → Online → Diagnostics → Cycle time and run the machine through a worst-case sequence.
  4. Insert RUNTIME probes in the three suspected hot paths identified by the estimate.
  5. Document cycle min/max/avg in the commissioning report.
  6. Compare with estimate; if delta > 20%, classify the discrepancy (cache effect, interrupt, or estimate error) before re-optimizing.
  7. Lock the firmware version on the CPU (TIA Portal: CPU properties → Protection → disable automatic update) to prevent silent timing changes after commissioning.

15. Frequently Asked Questions

Where do I find the official S7-1500 instruction execution times?

The official figures are in the SIMATIC S7-1500, S7-1500R/H, ET 200SP, ET 200pro Cycle and response times manual, entry ID 109751809, Table 3-2. The manual lists duration by category (bit, word, fixed-point, floating-point), not per individual instruction.

Why does the S7-1500 not publish per-instruction µs like the S7-300?

The S7-1500 instruction set is compiled and cached, so a single STL line can map to multiple micro-operations. Siemens therefore publishes category baselines and recommends the RUNTIME instruction for per-program measurement on the target CPU.

How accurate is the RUNTIME instruction?

RUNTIME uses the CPU's internal timestamp counter, with a resolution of 1 ns on current S7-1500 CPUs. The instruction itself adds 30–90 ns of overhead, which should be subtracted from the measured value when comparing to the category baseline.

Can I use S7-300 instruction list values to estimate S7-1500 times?

Only as a rough ratio. The S7-1500 is typically 3× faster on bit logic and up to 6× faster on floating-point compared to a CPU 317. Validate any critical figure on the target S7-1500 CPU with RUNTIME rather than scaling the S7-300 number.

Does S7-1500 instruction timing change between firmware versions?

Yes. Firmware updates can change the microcode for floating-point and division. Always re-measure with RUNTIME after a firmware upgrade and update the cycle documentation in the project. Lock the firmware version on the CPU to prevent silent changes.

What is the slowest single S7-1500 arithmetic instruction?

64-bit floating-point division and transcendental functions (EXP, LN, SQRT, SIN, COS, TAN) are the slowest, typically 300–1300 ns per call. Replace repeated divides with multiplies by precomputed reciprocals where the divisor is constant.

How do I measure the cycle time of a single FB?

Wrap the FB body with a RUNTIME call pair (MODE 0 at entry, MODE 1 at exit) using a per-FB instance DB to hold the result. Sum the value across scans and divide by call count to get a stable mean.

Back to blog