1. Problem Statement and Engineering Context
A mid-size process plant is automated with SIMATIC PCS 7 v9 (or compatible) controlling roughly 300 POs (Process Objects). The current runtime uses a single CPU 416-3 with distributed I/O on ET 200M and ET 200S stations over PROFIBUS DP and PROFINET. Because of the cost gap between the existing single-CPU configuration and a full hot-redundant PCS 7 V9 installation (CPU 417-4H with Y-Link pairs and a synchronisation module), the engineering team is evaluating a custom cold-standby architecture using two standard S7-400 CPUs.
The economics are real: a true 417-H redundancy stack requires two synchronisation modules, two fibre-optic synchronisation cables, two CP 443-1 plus matched SCALANCE switches, redundant AS pairs in HW Config, and a PCS 7 license model that flags every PO as redundant. In plants where the process can tolerate a few minutes without closed-loop control, the Y-Link / 417-H budget is not justified.
2. Process Requirements and Tolerable Outage
The constraint list drives the architecture:
- Total PO count: ≈ 300 in the live plant.
- Maximum tolerable control outage: 5 minutes per occurrence.
- Setpoints, tuning constants, operating modes, and recipe data must be present in the standby CPU before a switchover, so that bumpless transfer is achieved as much as the process allows.
- Single PROFIBUS DP and PROFINET address space – the existing field devices (ET 200M, ET 200S, drives, third-party PROFIBUS slaves) must continue to be addressed without re-parameterisation when the standby takes over.
- No Y-Link pairs. No 417-4H. No PCS 7 redundancy license.
A 5-minute outage window is the budget that makes cold standby viable. It is the same envelope Siemens describes for the warm-restart class on S7-400 CPUs in the STEP 7 cold restart functional description, where the CPU is expected to clear the process image, bit memory, timers, counters, and data blocks back to their load-memory start values, then bring the plant back to a defined state.
3. Proposed Cold Standby Architecture
Two mechanically and electrically identical AS racks are built. Each rack carries:
- One PS 405 (or PS 407) power supply, sized for the rack load with ≥ 25 % margin.
- One CPU 416-3 (same MLFB / firmware on both racks to avoid download / signature mismatches).
- One CP 443-1 for the plant / OS Ethernet.
- One DI 32 and one DO 32 in a reserved slot – these are not field I/O, they are the cross-rack life-sign channels.
- The two onboard PROFIBUS DP interfaces of the CPU are configured as follows:
- DP-1: coupled point-to-point between the two CPUs using a PROFIBUS cable with bus terminators powered on both ends. This port carries the data-block synchronisation (master → standby).
-
DP-2: connected to the plant PROFIBUS segment. Both racks carry the same PROFIBUS address on this port (e.g.
2), but only the powered-up SCALANCE and RS-485 repeater attached to the active rack drive the bus. The standby's DP-2 port is held in a passive state (CPU in STOP, or DP master disabled by user program).
The Ethernet ports (CP 443-1 and the PN interface of the CPU) carry the same IP address on both racks. Again, only the rack whose SCALANCE is energised is visible on the OS side.
3.1 Rack-to-rack wiring diagram
The wiring is deliberately small: there is no synchronisation fibre, no Y-Link, no H-CiR. The two racks are electrically independent; the only thing that ties them together is:
- One PROFIBUS segment (DP-1 to DP-1) for DB synchronisation.
- Two wires per rack pair: one DO from rack A to one DI on rack B (and vice versa) for life-sign pulsing.
- One DO from each rack that powers the local SCALANCE + PROFIBUS repeater (24 V routed through a relay or solid-state switch controlled by the DO).
4. Component Selection Reference
| Slot / Function | Module (typical) | Notes |
|---|---|---|
| PS | PS 405 10A (6ES7 405-0KA02-0AA0) or PS 407 10A | Match the second digit of the CPU's supply requirement. |
| CPU | CPU 416-3 (6ES7 416-3XS07-0AB0 or later) | Use the same MLFB and firmware on both racks. CPU 416-3 has two DP interfaces; an additional IF 964-DP can be installed if more DP masters are needed. |
| CP | CP 443-1 (6GK7 443-1EX30-0XE0 or later) | Industrial Ethernet, ISO + TCP, S7 communication for OS and engineering. |
| DI | SM 421 DI 32 × 24 V (6ES7 421-1BL01-0AA0) | Dedicated to life-sign input. |
| DO | SM 422 DO 32 × 24 V / 0.5 A (6ES7 422-1BL00-0AA0) | Dedicated to life-sign output + SCALANCE / repeater power relay. |
| Switch | SCALANCE XC-200 / XR-200 | One per rack, powered through the rack's DO. |
| Repeater | RS 485 Repeater (6ES7 972-0AA01-0XA0) | Active only on the master side. |
5. Master–Standby Synchronisation over PROFIBUS DP
The synchronisation channel is a small DP segment (DP-1) that runs between the two CPUs. Several implementation paths exist; the ones seen in field practice are listed in increasing complexity.
5.1 Option A – S7-300-style software redundancy blocks
Siemens' Software Redundancy package was originally written for S7-300 / S7-400 with a master and a reserve. It uses the SWR_AG / SWR_SEND / SWR_RECV function blocks (FB 1019 family) and supports both MPI and PROFIBUS coupling. On the master side SWR_SEND packages the configured redundancy-relevant data blocks and ships them over the coupling to SWR_RECV in the reserve. SWR_AG drives the life-sign.
This is the only Siemens-blessed mechanism that approaches the architecture in question, and the discussion in the source thread explicitly calls it out: "you should analyse the redundancy through MPI (software redundancy yet implemented) with S7-300". The caveat is that the software redundancy blocks were designed for S7-300 / IM 153-2 stations and S7-400 with a single DP master, and the PCS 7 CFC / SFC libraries do not consume them. You end up writing plain STEP 7 glue around the PCS 7 program.
5.2 Option B – Custom PUT/GET on the DP-1 link
Configure DP-1 on both CPUs as a DP master, and put one of them in master-master cross-traffic mode using DPV1 services. The active CPU runs a cyclic PUT (SFB 15) on the partner's DBs, and the partner runs GET (SFB 14) – or both directions are symmetric, with the standby only writing into a mirror DB. This is robust but it is not redundant traffic; if the link drops, the standby freezes on the last good mirror.
5.3 Option C – BSEND/BRCV with version counters
For 300 POs you are typically looking at 8–25 KB of setpoint, mode, and recipe data. BSEND / BRCV (SFB 12 / SFB 13) on ISO-on-TCP via a small CP 443-1-to-CP 443-1 connection, or raw DP telegrams over DP-1, is the preferred way to ship bulk data. Add a 16-bit monotonic version counter in the payload so the standby can detect a torn packet.
| Parameter | Typical value | Comment |
|---|---|---|
| Sync DB count | 1 (combined) or split per area (recipes, setpoints, modes, alarms) | Keeps the PUT/GET surface small. |
| Sync cycle | 250 ms – 1 s | Must finish well within the 5-min outage budget. |
| Life-sign period | 100 – 500 ms | Drives switchover latency. |
| Watchdog time | 3 × life-sign period | Avoids nuisance switchovers on a single missed pulse. |
| Version counter | UINT, incremented per send | Drops torn or reordered packets. |
6. Life-Sign Monitoring and Switchover Logic
The DO on the master toggles at the configured life-sign period; the partner DI watches for the edge. Symmetrically, the standby pulses its own DO and the master's DI watches that. This gives a 2-out-of-2 (or 2-out-of-3 if you add a third watchdog) health decision.
The switchover logic lives in a dedicated OB 1 cycle – not in OB 35, because the OB 35 time slice is a process-control slice and you do not want a 100 ms scan to be blocked by a PROFIBUS error while the master is healthy. Recommended pseudo-code (STEP 7 STL or SCL):
// Life-sign handling – OB 1
IF DI_LifeA THEN
nMissA := 0;
ELSE
nMissA := nMissA + 1;
END_IF;
IF nMissA > 3 THEN
bPartnerDownA := TRUE; // partner A's DO has gone silent
END_IF;
// Mirror direction
IF bIamMaster AND NOT bPartnerDownA THEN
bIamStandby := FALSE;
bIamActive := TRUE;
else
bIamStandby := TRUE;
bIamActive := FALSE;
end_if;
Priority logic prevents a dual-active condition:
- The rack with the higher rack ID is normally the standby.
- If the master is healthy, the standby holds its DP-2 master function disabled and its SCALANCE powered off.
- If the standby sees the master's life-sign go dark for 3 consecutive periods, it arms the switchover, enables its DP-2 master, energises its SCALANCE, and announces itself on the OS via a redundancy-status tag.
7. Power and Network Switchover
The two SCALANCE switches and the two RS 485 repeaters are powered from 24 V rails that are switched by the master rack's DO. The standby's SCALANCE is unpowered, which means its port LEDs are dark and it is invisible to the OS. When the standby decides to take over, it must:
- Wait one more life-sign period (to be certain).
- Set its own DO high to energise its SCALANCE and repeater.
- Re-initialise its DP-2 master and CP 443-1 connection (CPU 416 DP-2 cold start on the segment).
- Re-establish S7 connections to the OS server(s).
- Announce the switchover via an OS message and a status tag (e.g.
AS_Active = 1on the new active rack,0on the old one).
The 5-minute outage budget covers the cold start of the CPU, the PROFIBUS DP re-parameterisation of the ET 200M / ET 200S stations, and the OS reconnection. ET 200M with IM 153-2 typically needs 3–8 s to come back; ET 200S head stations with 4–8 modules typically need 1–3 s. The OS reconnect with PCS 7 WinCC is generally under 30 s for a single AS pair.
8. Cold Restart vs. Warm Restart on S7-400
The standby CPU on first commissioning runs a cold restart: all data in work memory – process image, bit memory, timers, counters, and the non-retentive data blocks – is reset to the start values stored in load memory. The retentive DBs (those flagged in the DB properties) keep their last values.
Implications for the cold-standby design:
- Any setpoint, recipe, or mode that you want preserved across a switchover must be in a retentive DB, or must be re-sent from the new active CPU after switchover (less attractive, because it pushes the bumpless-transfer goal out of the window).
- The OS-level archive (PCS 7 Process Historian) is the canonical source for long-term values. On switchover, the new active CPU requests the latest archived values and seeds the retentive DBs from there.
- Watch the S7-400
OB 100(warm restart) versusOB 102(cold restart) distinction. The transition from STOP to RUN on the standby after a switchover is a cold restart;OB 102is executed once. Put your mirror-DB validation inOB 102so a torn mirror is detected before the first process OB fires.
For full details, see the STEP 7 functional description of cold restart on S7-400.
9. Risks, Failure Modes, and Unsupported-Scenario Call-Outs
This is the part the field report hammers: a hand-rolled cold standby has more failure modes than a certified 417-H system. Document each one, and test it before the plant is handed over.
| # | Failure mode | Likelihood | Detection / mitigation |
|---|---|---|---|
| 1 | Master CPU goes to STOP (not STOP-7) and the DO driving SCALANCE remains high – dual-active condition. | Low if wired through normally-closed DO contact. | Verify in FAT that DO drops on STOP, RUN-to-STOP transition, and SF / BF. Use a watchdog relay on the DO output. |
| 2 | Life-sign DO contact welds closed (relay fault). Standby never takes over. | Low, increases with vibration and contact age. | Use a solid-state DO rated for ≥ 1 M cycles, or accept a periodic manual test. |
| 3 | DB mirror becomes stale (PUT/GET link drops, no fault raised). Standby switches over with old data. | Medium. | Version counter in payload; standby drops mirror if counter does not advance for > 5 s. |
| 4 | PROFIBUS DP-2 master does not re-parameterise ET 200M after switchover (e.g. wrong GSD revision on the standby). | Medium if racks have diverged. | Pin identical GSD revisions, identical HW Config, identical firmware. Verify in FAT. |
| 5 | OS does not see the new AS (CP 443-1 S7 connection times out). | Medium if WinCC redundancy not configured. | Configure WinCC AS redundancy (logical AS name vs. physical IP). Add connection watchdog. |
| 6 | Both CPUs run their own clock, time stamps diverge, sequence-of-events logs are unusable. | Medium. | Synchronise the standby via NTP or a SIMATIC procedure (time-of-day synchronisation over the CP 443-1). Run time-master only on the active CPU. |
| 7 | Plant-side operator presses Restart on the master during a transient; both racks go to STOP simultaneously. | Low if operator console is locked to one AS. | Disable operator restart on the standby. Route all operator commands to the active AS only. |
| 8 | Y-Link avoidance means there is no PROFIBUS-PA coupling redundancy. Loss of one PA segment is not handled. | Inherent to the design. | Accept the risk or add a separate PA coupler pair (out of scope for the cold-standby budget). |
10. Why PCS 7 Officially Won't Sign This Off
The discussion in the source thread is unambiguous: PCS 7 does not recognise a software-redundancy pair as a redundant AS. PCS 7 AS redundancy is a binary property: the AS is redundant only if the HW Config shows an H-CPU (417-4H or, on the newer platform, 1517H / 1518HF). Anything else is a single AS from the PCS 7 OS perspective, and the OS will not automatically reconnect to a different physical AS on a switchover.
Practical consequences:
- Every PO that was configured as a single PO in PCS 7 will continue to look like a single PO. There is no built-in fail-over for operator faceplates, alarm routes, or SFC steps.
- WinCC must be configured with an AS redundancy project setting pointing to a primary and a standby IP, with the standby's IP identical to the master's at the moment of switchover – this is non-standard and only works because the OS resolves the IP to the active SCALANCE, not to a specific CPU.
- PCS 7 library blocks (CFC, SFC, APL) do not consume life-sign or mirror-status tags. The integrator has to add a status tag and surface it on the OS faceplate as a custom indicator.
11. Modern Alternative: SIMATIC S7-1500 R/H
For new projects where the budget allows it, the SIMATIC S7-1500 R/H CPUs are the supported path to redundancy on the current platform. They provide synchronised CPUs with seamless switchover (R-type: redundancy without hot-standby synchronisation; H-type: full hot-standby with PROFINET ring synchronisation). Compared to the hand-rolled S7-400 cold standby:
| Aspect | S7-400 cold standby (this design) | S7-1500R/H |
|---|---|---|
| Switchover time | 10 s – 5 min (cold start, DP re-init, OS reconnect) | < 50 ms – 300 ms (R), < 50 ms (H) |
| Data synchronisation | Manual PUT/GET or SWR blocks | Built-in, transparent for the user program |
| OS impact | Custom – no PCS 7 redundancy | Native AS redundancy in TIA Portal / WinCC Unified |
| Cost | Low (two standard CPUs, no Y-Link, no redundancy license) | Higher (H-CPU, sync module, ring topology, license) |
| Vendor support | None (non-supported configuration) | Full Siemens warranty and FA |
For a 300-PO plant that is on a tight budget and is already on PCS 7 v9 with a 416-3, the cold-standby architecture is a defensible compromise. For a greenfield plant, the S7-1500R/H is the engineering-rational choice.
12. Commissioning and Verification Checklist
- Hardware identity check. Pull both rack configurations with HW Config and confirm slot-by-slot parity, MLFB, firmware, and GSD revision.
- DI/DO cross-wiring continuity test. Force a DO high on rack A with a hand-held, observe the DI on rack B, and vice versa.
- SCALANCE dropout on master STOP. Stop the master CPU from STEP 7 online. Verify that the SCALANCE loses link within 1 s and that the standby's DI watchdog starts counting.
- Standby cold restart timing. Time the interval from "master STOP" to "standby DP-2 master active and ET 200M in data exchange". Target: < 60 s for the 300-PO load.
- OS reconnection timing. Time the WinCC faceplate update after switchover. Target: < 60 s.
- Mirror freshness. Force a value in the master, read it back from the standby DB, confirm within 2 sync cycles.
- Relay-weld test. Force the master's life-sign DO high for 24 h. Verify the contact resistance is within the data-sheet limit. Repeat on a 6-month maintenance cycle.
- Operator restart isolation. Confirm that an operator "Restart AS" command on the master does not propagate to the standby.
- Clock synchronisation. Confirm that the active CPU is the only time master, and that the OS time stamps are monotonic across the switchover.
- Blackout test. Pull the master power supply. Verify the standby takes over and the plant returns to the last good state within the 5-minute budget.
13. Frequently Asked Questions
Does PCS 7 officially support a software-redundancy cold standby on standard S7-400 CPUs?
No. PCS 7 recognises only the H-system path (CPU 417-4H with synchronisation module, or S7-1500R/H). A two-CPU configuration with shared PROFIBUS / PROFINET addresses is treated as two single AS, not a redundant AS, and the OS will not auto-fail-over.
What is the maximum tolerable switchover time for a 300-PO plant on a CPU 416-3?
For a custom cold standby, plan for 10–60 s on the DP re-initialisation and another 30 s for WinCC reconnection. The 5-minute process budget stated in the source thread is comfortable, but field-tested switchover under load is typically under 90 s.
Can the Software Redundancy (SWR) blocks from S7-300 be reused on a 416-3 PCS 7 project?
The SWR blocks (FB 1019 family) will run on a 416-3, but they are not part of the PCS 7 library set and there is no Siemens guarantee that they will coexist with CFC / SFC generated code. Treat them as a parallel custom code path and validate under load.
How is dual-active (both CPUs driving the bus) prevented?
By routing the SCALANCE and PROFIBUS repeater 24 V supply through a normally-closed contact of the master's life-sign DO. If the master stops, the DO drops, the field network is de-energised, and the standby is the only one able to re-power it.
Why is cold restart behaviour on the standby important for setpoint retention?
A cold restart clears all non-retentive DBs, bit memory, timers, and counters back to the load-memory start values. Any data that must survive a switchover has to live in a retentive DB or be re-seeded from the OS archive after the new active CPU is in RUN – see the STEP 7 description of cold restart on S7-400.
What is the cheapest supported alternative for a 300-PO PCS 7 plant that needs redundancy?
Move to a SIMATIC S7-1500R/H CPU on PROFINET ring topology. It is more expensive than the cold-standby hack but provides certified redundancy with sub-300 ms switchover and full TIA Portal / PCS 7 support.