S7-1500 CPU Instruction Execution Times: Reference Manual
Execution time on a Siemens S7-1500 CPU is the wall-clock duration a single statement or block of statements consumes inside the OB/FC/FB call stack. Unlike S7-300/400, where every instruction has a published microsecond figure in the Instruction List manual, the S7-1500 family documents timing in categories: bit operations, word operations, fixed-point arithmetic, floating-point arithmetic, and transfer/comparison operations. Per-instruction nanosecond figures are obtained by combining the category baseline with measured RUNTIME samples, or by reading the device-specific data from the Siemens Function Manual.
This reference consolidates the official Siemens cycle and response times manual, the S7-300 Instruction List as a historical baseline, and the RUNTIME instruction to give engineers a complete methodology for estimating and measuring S7-1500 instruction execution time.
1. Reference Documentation
Three primary Siemens documents govern S7-1500 instruction timing. Always cite the exact manual edition used for the calculation.
| Document | Manual Name | Entry ID | Contents |
|---|---|---|---|
| S7-1500 / ET 200SP cycle & response times | SIMATIC S7-1500, S7-1500R/H, ET 200SP, ET 200pro Cycle and response times | 109751809 (support.industry.siemens.com) | Table 3-2: operation durations by category; cycle time formulas; response time composition |
| S7-300 Instruction List | S7-300 Instruction List, CPU 31xC, CPU 31x, IM 151-7 CPU, IM 151-8 CPU, IM 154-8 CPU, BM 147-1 CPU, BM 147-2 CPU | 13206730 (support.industry.siemens.com) | Per-instruction µs figures (e.g., POP 0.03 µs, TAK 0.06 µs on CPU 317); reference for cross-platform timing ratios |
| S7-1500 Programming Guideline | SIMATIC S7-1500 Programming Guideline | 81318674 (support.industry.siemens.com) | Optimized block design, multi-instance DBs, symbolic vs. absolute access overhead |
Open the cycle and response times manual at SIMATIC S7-1500, ET 200SP, ET 200pro Cycle and response times and the S7-300 list at S7-300 Instruction List (A5E00105517-10) for the canonical figures.
2. Why S7-1500 Documents Timing in Categories
The S7-1500 instruction set runs on a different execution engine than the S7-300/400 accumulator-based architecture. Three architectural shifts make per-instruction µs tables impractical:
- Optimized compiled code — TIA Portal compiles the entire OB/FC/FB before download. A single STL line may map to multiple micro-operations in the firmware's intermediate language.
- Cache effects — the first execution of an instruction may be slower than subsequent executions because the microcode cache is cold.
- Data type width is variable — LREAL, LWORD, and LINT operations are typically slower than REAL or DINT, but the ratio is firmware-dependent.
Siemens therefore publishes category execution times in Table 3-2 of the cycle and response times manual. The engineer sums the category figures for the code path to estimate cycle time, then validates with RUNTIME in TIA Portal.
3. Instruction Categories and Default Timing Ranges
The cycle and response times manual groups instructions into the following categories. Values below are typical ranges for a standard CPU 1515-2 PN at firmware V2.9; refer to the manual for the exact CPU in the project.
| Category | Typical Range (ns) | Examples |
|---|---|---|
| Bit operations | 10 – 60 | AND, OR, XOR, assignment, set/reset, edge detection (positive/negative) |
| Word operations (timers, counters, comparators) | 30 – 200 | TON, TOF, TP, CTU, CTD, EQ_I, NE_R |
| Fixed-point arithmetic (16/32-bit) | 30 – 100 | ADD_I, SUB_D, MUL_I, DIV_D, MOD |
| Floating-point arithmetic (32/64-bit) | 60 – 300 | ADD_R, MUL_R, SIN, COS, EXP, LN, SQRT |
| Transfer / move | 20 – 80 | MOVE, BLKMOV, FILL, NONDISS, Serialize |
| Conversion | 40 – 200 | INT_TO_REAL, REAL_TO_DINT, BYTE_TO_WORD |
| Program control / block call | 200 – 1500 | CALL, JMP, JMPN, RET, BE, UC, CC |
| Indirect addressing | 80 – 400 | PEEK, POKE, field access with variable index |
4. Bit Operation Timing
Bit operations are the fastest class on S7-1500 because they map directly to single-cycle boolean ALU operations. The cycle and response times manual lists the baseline for one logical AND, OR, or assignment at roughly 10 ns on a CPU 1515-2 PN at 100 MHz equivalent throughput.
Edge-detection instructions (FP, FN, R_TRIG, F_TRIG) carry an additional implicit load of the previous bit state. Expect 20–40 ns per edge operation instead of the bare 10 ns baseline. Use R_TRIG/F_TRIG function blocks (single-instance DBs) for clarity; avoid in-line edge logic on hot loops.
Set/reset operations (S, R, RS, SR flip-flops) on individual bit tags cost the same as AND/OR because the underlying microcode is a conditional write. The cost appears in the data access path, not the ALU.
4.1 Bit operation timing example
// Network 1 (STL equivalent of a typical machine interlock)
A %I0.0 // 10 ns
A %I0.1 // 10 ns
= %Q0.0 // 10 ns
A %I0.2 // 10 ns
FP %M100.0 // 30 ns (FP loads previous M100.0)
= %Q0.1 // 10 ns
// Subtotal: ~80 ns excluding I/O image update
5. Word, Timer, and Counter Timing
The S7-1500 timer and counter instructions are IEC 61131-3 system blocks (TON, TOF, TP, TONR, CTU, CTD, CTUD). Internally they maintain an instance DB, so the cost is the function call (~200 ns) plus a few word operations on the time/count tag.
Approximate per-call figures for a CPU 1515-2 PN:
| Instruction | Typical Cost (ns) | Notes |
|---|---|---|
| TON / TOF / TP | 300 – 600 | Includes 1 ms tick, 16/32-bit time base, edge handling |
| TONR (retentive on-delay) | 350 – 700 | Adds retentive read/write to instance DB |
| CTU / CTD | 200 – 400 | Single counter register, no edge detection overhead |
| CTUD (up/down) | 300 – 500 | Two register update paths |
| EQ_I / NE_I / GT_I / LT_I / GE_I / LE_I | 30 – 50 | Integer comparator |
| EQ_R / NE_R / GT_R / LT_R / GE_R / LE_R | 50 – 100 | Real comparator, NaN handling adds ~20 ns |
| MOVE (single tag) | 20 – 40 | Symbolic access slightly slower than absolute |
| BLKMOV / UMOVE / BMOVE | 50 + 5/element | Bulk transfer cost scales linearly with element count |
| FILL_BLK / UFILL_BF | 60 + 4/element | Fill scales similarly to block move |
For TON, TOF, TP, the value range is 0–9999999 ms with a time base of 1 ms by default. This time-base tick is internal to the IEC timer; it does not consume a separate hardware timer slot on the CPU.
6. Fixed-Point and Floating-Point Arithmetic
Fixed-point arithmetic operates on INT, DINT, WORD, DWORD, and BCD. Floating-point operates on REAL and LREAL. Typical CPU 1515-2 PN costs:
| Operation | 16-bit (ns) | 32-bit (ns) | 64-bit (ns) |
|---|---|---|---|
| ADD / SUB | 30 | 40 | 80 |
| MUL | 50 | 60 | 120 |
| DIV | 80 | 120 | 240 |
| MOD | 100 | 150 | 300 |
| SIN / COS / TAN | n/a | 250 – 500 | 500 – 900 |
| EXP / LN / SQRT | n/a | 300 – 700 | 600 – 1300 |
| ABS / NEG | 20 | 30 | 50 |
| TRUNC / ROUND / CEIL / FLOOR | 30 | 40 | 80 |
Conversion operations have their own cost band:
- INT_TO_REAL / REAL_TO_DINT: 40–100 ns
- BYTE_TO_WORD / WORD_TO_DWORD: 30–60 ns
- BCD_TO_INT / INT_TO_BCD: 60–120 ns
- CHAR_TO_STRING / STRING_TO_CHAR: 200–500 ns (string header manipulation)
The S7-1500 hardware floating-point unit is IEEE 754 compliant. NaN propagation, denormal handling, and overflow traps follow the IEEE default, so the cost of an exception path can be 2–3× the normal case. Guard division inputs to avoid divide-by-zero NaN results.
7. The RUNTIME Measurement Instruction
The S7-1500 RUNTIME instruction is the canonical way to measure instruction execution time on the target hardware. It captures the CPU's internal high-resolution counter at start and end of a code region, returning elapsed time in nanoseconds.
7.1 Interface
// SCL signature
RUNTIME(MODE := 0, MEM := <instance_DB>);
// MODE = 0: write start timestamp to MEM
// MODE = 1: write end timestamp and compute elapsed time into MEM
7.2 Usage pattern
// FB "PerfProbe" body (SCL)
// At scan start
RUNTIME(MODE := 0, MEM := DB_Perf.runtime); // 30 ns overhead
// Code to measure
// ...
// At scan end
RUNTIME(MODE := 1, MEM := DB_Perf.runtime); // 60 ns overhead
// DB_Perf.runtime now holds elapsed nanoseconds as LREAL
The result tag is a LREAL holding the elapsed time. RUNTIME itself consumes 30–90 ns depending on MODE, so subtract the overhead from the measurement when comparing to category baselines.
7.3 Caveats
- RUNTIME must be called with a valid MEM pointer (instance DB or global DB). Passing a NIL pointer raises OB 121 (programming error).
- The counter is the CPU's monotonic timestamp counter; it does not roll over within the cycle.
- RUNTIME is available in all S7-1500 CPUs and the S7-1500R/H redundant CPUs as of firmware V1.8.
8. S7-300 Cross-Reference for Timing Ratios
The S7-300 Instruction List at S7-300 Instruction List (A5E00105517-10) still publishes per-instruction µs figures, e.g. POP 0.03 µs and TAK 0.06 µs on CPU 317. Use it as a ratio baseline when porting an S7-300 program and assessing whether the S7-1500 will hit its cycle budget.
| Instruction (CPU 317 baseline) | S7-300 Time | S7-1500 Time (typical) | Speedup |
|---|---|---|---|
| POP (stack pop, no accumulator write) | 0.03 µs | ~10 ns | 3× |
| TAK (swap top two accumulators) | 0.06 µs | ~20 ns | 3× |
| +I (16-bit integer add) | 0.05 µs | ~30 ns | 1.7× |
| +R (32-bit real add) | 0.30 µs | ~50 ns | 6× |
| /R (32-bit real divide) | 0.90 µs | ~150 ns | 6× |
The S7-1500 typically runs 1.5× to 6× faster than the S7-317, with the largest gain on floating-point because the S7-1500 has a dedicated hardware FPU. Integer logic gains are smaller because the S7-317 was already bit-optimized.
9. Cycle Time Calculation
The cycle time of an OB1 is the sum of:
T_cycle = T_image + T_OB1_user + T_comm + T_interrupt + T_diag + T_reserve
Where:
- T_image: process image update, ~50 µs for 1 KB of I/O at 100 µs PI update rate
- T_OB1_user: sum of all user-code execution times (the focus of this reference)
- T_comm: PG/OP/HMI communication, scaled with the configured PG/OP connection count
- T_interrupt: hardware and time-delay OB execution, weighted by frequency
- T_diag: system diagnostics and module status reads
- T_reserve: 5–10% reserve for transient interrupts
The cycle and response times manual provides the exact formulas and coefficient tables per CPU model. Use the manual's cycle time configurator worksheet rather than hand-calculating for production programs.
10. Estimating User-Code Execution Time
Use the following four-step procedure to estimate T_OB1_user before downloading to the CPU:
- Inventory the code — count each instruction in the OB1 call tree (OB1 + every FC/FB it invokes).
- Classify each instruction — map to a category from Section 3 above.
- Sum by category — use the typical range midpoint as the per-instruction figure; multiply by call-site count for reused FBs.
- Add block call overhead — every CALL consumes 200–1500 ns for stack frame setup; add 100–300 ns for multi-instance DBs.
10.1 Worked example
Suppose OB1 of a packaging line contains 200 bit ops, 50 word ops, 30 fixed-point math ops, 20 floating-point math ops, and 10 timer calls, plus 5 FC calls.
| Class | Count | Per-Op (ns) | Subtotal (ns) |
|---|---|---|---|
| Bit | 200 | 30 | 6 000 |
| Word (timer/counter) | 50 | 80 | 4 000 |
| Fixed-point | 30 | 60 | 1 800 |
| Floating-point | 20 | 150 | 3 000 |
| Timer (TON) | 10 | 500 | 5 000 |
| FC call | 5 | 800 | 4 000 |
| Estimated T_OB1_user | 23 800 ns ≈ 24 µs | ||
This is the user-code slice only. Add T_image, T_comm, and the diagnostic reserve per the cycle manual. On a CPU 1515-2 PN the total cycle is typically 1–3 ms for a packaging OB of this size.
11. Measuring Cycle Time Online
Two on-CPU tools measure cycle time directly:
- Online → Diagnostic → Cycle time in TIA Portal — shows the last 10 cycle times including min, max, and current.
- RD_SYS_T and RD_LOC_T for time-stamping OB1 runs externally; useful for high-precision cycle analysis.
For program-section timing inside the OB, use RUNTIME. For OB-level timing, use the diagnostic buffer or the TIA Portal cycle statistic.
11.1 Verification checklist
- [ ] Cycle max < 80% of the configured OB1 scan watchdog
- [ ] RUNTIME peaks documented and within the calculated T_OB1_user ± 15%
- [ ] Floating-point hot path not in OB35/time-critical OB if cycle > 5 ms
- [ ] No divide-by-zero possible in floating-point paths (NaN check present)
- [ ] All TIMERs use the same time base (avoid mixing 1 ms and 10 ms in one loop)
12. Performance Optimization Best Practices
Before optimizing, profile. The S7-1500 spends the majority of its cycle on a small number of hot paths. Use the RUNTIME probe in each suspected hot path before changing code.
- Replace DIV with MUL by reciprocal — precompute 1.0/x and multiply. A real divide at 150 ns becomes a real multiply at 60 ns.
- Use IEC timer edge-flag alternatives — the S7-1500 system TON is >300 ns; for fast single-shot delays, derive a time flag from a 1 ms cyclic interrupt instead of using TON in the hot path.
- Cache real constants — a LREAL constant load is faster than converting from INT at every scan.
- Avoid multi-instance DBs for short FBs called millions of times — the DB write-back overhead is real; single-instance DBs in a known hot region can be faster for legacy code.
- Use BLKMOV with typed arrays instead of structured field-by-field copy when the layout matches.
- Run floating-point math in a slower OB (e.g., OB35 at 100 ms) and use latched results in OB1.
13. Common Errors When Using RUNTIME
| Symptom | Cause | Fix |
|---|---|---|
| RUNTIME returns 0 | MEM points to an uninitialized instance DB | Initialize DB; verify the DB is not optimized (or, if optimized, use a known-offset tag) |
| OB 121 programming error | Invalid MEM pointer | Pass a DB number, not a tag address |
| Reading inconsistent values | Two OBs write the same MEM tag | Use a per-OB MEM tag |
| Value rolls over at 32-bit boundary | MEM declared as REAL, not LREAL | Declare MEM as LREAL |
| Time appears doubled | MODE = 1 called twice in a row | Wrap the section with a single MODE = 0 / MODE = 1 pair |
14. Field Commissioning Procedure
- Pre-load estimate — use Section 10 to compute T_OB1_user.
- Download and go online with the project.
- Open TIA Portal → Online → Diagnostics → Cycle time and run the machine through a worst-case sequence.
- Insert RUNTIME probes in the three suspected hot paths identified by the estimate.
- Document cycle min/max/avg in the commissioning report.
- Compare with estimate; if delta > 20%, classify the discrepancy (cache effect, interrupt, or estimate error) before re-optimizing.
- Lock the firmware version on the CPU (TIA Portal: CPU properties → Protection → disable automatic update) to prevent silent timing changes after commissioning.
15. Frequently Asked Questions
Where do I find the official S7-1500 instruction execution times?
The official figures are in the SIMATIC S7-1500, S7-1500R/H, ET 200SP, ET 200pro Cycle and response times manual, entry ID 109751809, Table 3-2. The manual lists duration by category (bit, word, fixed-point, floating-point), not per individual instruction.
Why does the S7-1500 not publish per-instruction µs like the S7-300?
The S7-1500 instruction set is compiled and cached, so a single STL line can map to multiple micro-operations. Siemens therefore publishes category baselines and recommends the RUNTIME instruction for per-program measurement on the target CPU.
How accurate is the RUNTIME instruction?
RUNTIME uses the CPU's internal timestamp counter, with a resolution of 1 ns on current S7-1500 CPUs. The instruction itself adds 30–90 ns of overhead, which should be subtracted from the measured value when comparing to the category baseline.
Can I use S7-300 instruction list values to estimate S7-1500 times?
Only as a rough ratio. The S7-1500 is typically 3× faster on bit logic and up to 6× faster on floating-point compared to a CPU 317. Validate any critical figure on the target S7-1500 CPU with RUNTIME rather than scaling the S7-300 number.
Does S7-1500 instruction timing change between firmware versions?
Yes. Firmware updates can change the microcode for floating-point and division. Always re-measure with RUNTIME after a firmware upgrade and update the cycle documentation in the project. Lock the firmware version on the CPU to prevent silent changes.
What is the slowest single S7-1500 arithmetic instruction?
64-bit floating-point division and transcendental functions (EXP, LN, SQRT, SIN, COS, TAN) are the slowest, typically 300–1300 ns per call. Replace repeated divides with multiplies by precomputed reciprocals where the divisor is constant.
How do I measure the cycle time of a single FB?
Wrap the FB body with a RUNTIME call pair (MODE 0 at entry, MODE 1 at exit) using a per-FB instance DB to hold the result. Sum the value across scans and divide by call count to get a stable mean.