WinCC v6.2 Redundancy: Resolving Client Server Flapping Issues

David Krause10 min read
SiemensTroubleshootingWinCC
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

Problem Description

A WinCC V6.2 client/server architecture is configured with two redundant servers (BIO_SERV01 as preferred server, BIO_SERV02 as standby) and two Windows XP clients (BIO_CLNT01, BIO_CLNT02). Communication to the S7-400 PLC is established over Industrial Ethernet through a managed switch. The redundant pair is interconnected with the Siemens serial redundancy cable.

During normal operation both servers report an active state. When the operator powers off BIO_SERV01, the failover to BIO_SERV02 completes correctly. When BIO_SERV01 is brought back online, the symptoms appear:

  1. Both clients begin alternating the active connection between BIO_SERV01 and BIO_SERV02.
  2. The toggle interval is non-deterministic — typically 1 to 15 minutes per partner.
  3. The clients eventually settle on a partner at random, biased toward BIO_SERV02.
  4. Stopping the chosen partner triggers an immediate clean failover to the remaining server.

The behavior is reproducible regardless of whether the Preferred Server flag is enabled in the WinCC Redundancy dialog.

System Architecture

Station Role OS WinCC Version Service Pack
BIO_SERV01 Preferred Master Server Windows Server 2003 WinCC V6.2 None (SP2 missing)
BIO_SERV02 Standby Server Windows Server 2003 WinCC V6.2 None (SP2 missing)
BIO_CLNT01 Client Windows XP SP3 WinCC V6.2 Client None (SP2 missing)
BIO_CLNT02 Client Windows XP SP3 WinCC V6.2 Client None (SP2 missing)

The pair is linked by the Siemens null-modem redundancy cable on a dedicated COM port (typically COM1 of each server). The PLC connection uses the SIMATIC S7 Protocol Suite channel SIMATIC_S7_PROTOCOL_SUITE_01 on the Industrial Ethernet network. NICs are 100 Mb/s Full Duplex, fixed in adapter properties. All four machines reside in the WORKGROUP workgroup — no Active Directory domain.

Root Cause Analysis

WinCC V6.2 redundancy uses a partner-status heart beat that travels over TCP (default port 0x2700 / decimal 9984) and a separate process-state heart beat over the serial cable. Clients continuously evaluate the heart-beat latency and the published Master/Standby flag to choose a server. Client flapping after partner recovery is normally caused by one of the following defects:

  1. Missing WinCC V6.2 Service Pack 2 on one or more stations. SP2 contains several hot fixes (HF blocks) that stabilize the redundancy state machine. Pre-SP2 builds are known to mis-handle the rapid reappearance of the preferred partner.
  2. Partial patch installation: a hot fix installed on three of the four stations but failed on BIO_SERV02. A divergent patch level between the redundant pair is a primary cause of oscillation.
  3. Missing or stale LMHOSTS entries. In a workgroup deployment (no WINS/DNS for NetBIOS), every WinCC station must resolve the names of every other WinCC station. A missing LMHOSTS causes intermittent name resolution that the redundancy monitor interprets as partner outage.
  4. Project-Duplicator artifact: the duplicate project inherits the original computer name in internal tags. If the duplicate was not re-generated after renaming OS(1) → OS(2), the redundancy state machine reads contradictory server identities.
  5. Serial cable ground loops or driver conflict: COM port contention (e.g., HyperTerminal holding COM1) breaks the serial heart beat and forces the client to re-evaluate the partner over TCP only.
  6. NIC auto-negotiation: although the adapter reports 100 Mb/s Full Duplex, the switch may fall back to half duplex under load, generating CRC errors that the S7 Protocol Suite treats as PLC outage, not network outage.
Critical: In WinCC V6.2 a redundant pair MUST be at the identical patch level (CD image + SP2 + identical HF). Mismatched patch sets will produce exactly the oscillation pattern described above. Apply SP2 to all four stations from the same distribution media before any further debugging.

Prerequisites

  • WinCC V6.2 installation DVD or ISO matching the originally licensed build.
  • WinCC V6.2 SP2 redistributable (file WinCC_V62_SP2_HF17.exe or the SP2 level that matches the CD).
  • Local administrator rights on all four stations.
  • Siemens serial redundancy cable (6ES7 972-0CB20-0XA0 or successor).
  • Image backup of each station prior to re-installation.
  • Stop all WinCC Runtime instances and disable the WinCC Center Cursor / Startup services.

Step-by-Step Resolution

1. Equalize the patch level on all four stations

  1. Insert the WinCC V6.2 CD and run Setup.exe on every station.
  2. Apply the WinCC V6.2 Service Pack 2 from the same hotfix package; do not mix hotfix versions between the redundant pair.
  3. On the station that previously failed the patch (BIO_SERV02), uninstall WinCC completely, delete C:\Program Files\Siemens\Automation\WinCC, reinstall WinCC V6.2, then apply SP2.
  4. Reboot and verify the version in WinCC Explorer > Help > About reports identical build numbers on BIO_SERV01 and BIO_SERV02.

2. Generate a correct duplicate project

Always re-duplicate after any patch change. From BIO_SERV01:

  1. Close WinCC Explorer.
  2. Open Start > SIMATIC > WinCC > Tools > Project Duplicator.
  3. Select Duplicate a server and point at the master project.
  4. In the duplicate wizard:
    • Set Computer name of the duplicate to BIO_SERV02.
    • Set Project path to a fresh folder, e.g. D:\WinCC\BIO_SERV02.
    • Tick Adapt computer name and Adapt project path.
  5. Open the duplicate in WinCC Explorer and verify under Computer > Properties that Server name = BIO_SERV02 and Redundancy partner = BIO_SERV01.

3. Configure LMHOSTS for workgroup operation

WinCC V6.2 requires NetBIOS name resolution for redundancy. In a workgroup without WINS, LMHOSTS is mandatory.

  1. On every station, edit C:\Windows\System32\drivers\etc\LMHOSTS (remove the .sam extension if present). Add an entry for every WinCC station:
192.168.1.21    BIO_SERV01    #PRE
192.168.1.22    BIO_SERV02    #PRE
192.168.1.31    BIO_CLNT01    #PRE
192.168.1.32    BIO_CLNT02    #PRE
  1. Enable LMHOSTS lookup: nbtstat -R (re-load) then nbtstat -c (verify all four names appear as PRE).
  2. Run nbtstat -a BIO_SERV01 from BIO_SERV02 to confirm resolution.
  3. Reboot every station so the WinCC redundancy monitor picks up the new resolver table.

4. Verify the serial redundancy cable

  1. Connect the Siemens cable between COM1 of BIO_SERV01 and COM1 of BIO_SERV02.
  2. Disable any service that may open the COM port (HyperTerminal, Palm HotSync, Bluetooth stack).
  3. In the BIOS of each server, set the COM port to High-Speed if available.
  4. Start the redundancy monitor on both servers and observe the partner state; it must read OK, not UNKNOWN.

5. Pin the preferred server

  1. In the WinCC Explorer on BIO_SERV01 open Computer > Redundancy.
  2. Tick This server is the preferred server.
  3. Mirror the setting on the duplicate project (BIO_SERV02) but leave the flag cleared.
  4. Activate the project on both servers.

6. Validate the redundancy state machine

The expected behavior in WinCC V6.2 with a working configuration is:

Both servers ACTIVE OS1 powers off Clients on OS2 OS1 powers on Clients stay on OS2 OS2 outage → OS1 Preferred-server logic prevents client oscillation after OS1 recovery

Diagnostic File Analysis

WinCC V6.2 writes redundancy-relevant log files under C:\Program Files\Siemens\Automation\WinCC\diagnose\. Inspect the following on both servers:

File What it contains Healthy pattern
WinCC_Server_01.LOG Runtime startup, partner status Repeated Partner status OK entries every 5 s
SIMATIC_S7_PROTOCOL_SUITE_01.LOG S7 channel connections No undocumented error codes (> 0x8000FFFF)
WinCC_Sys_01.LOG Tag logging, alarms No "connection broken" events
CCAgent.log License + redundancy handshake Single line per failover only

Common undocumented error codes seen during client flapping:

Hex code Meaning Likely cause
0x80004005 Unspecified COM failure Serial cable / COM port driver
0x80070005 Access denied DCOM not configured for WinCC user
0x80040154 Class not registered SP/HF not applied to that station
0x800706BA RPC server unavailable Firewall on Server 2003 blocking port 135/445
0x80010105 Server threw exception Patch-level mismatch

Verification Procedure

  1. With both servers active, on the client open WinCC Explorer > Tools > WinCC Diagnosis; confirm the partner display reads OS1 Preferred and the latency is below 200 ms.
  2. Stop the WinCC Runtime on BIO_SERV01; clients must switch to BIO_SERV02 within the configured failover time (default 25 s).
  3. Restart BIO_SERV01. Clients MUST remain on BIO_SERV02 indefinitely while both servers report ACTIVE.
  4. Only when BIO_SERV02 is stopped must clients return to BIO_SERV01. The transition must take < 30 s.
  5. Run a 24-hour soak test. No spontaneous partner switching must occur.

Extended Diagnostics

Network capture

Capture traffic between the clients and the redundant servers with Wireshark and filter on WinCC redundancy port:

tcp.port == 9984 || tcp.port == 135 || udp.port == 137

Verify the partner-state packets arrive at a steady 5-second interval. Gaps longer than 15 s indicate LMHOSTS or firewall interference.

Event log correlation

Open eventvwr.msc on every station and filter on source WinCC and DCOM. Each failover must produce exactly two correlated events: one on the failing server and one on the surviving server, separated by less than 30 s.

Switch port statistics

Inspect the switch port counters for the four WinCC nodes. CRC errors, runts, or late collisions above 0.01 % of total frames indicate duplex mismatch — re-pin both ends to 100 Mb/s Full Duplex.

Cross-Reference: WinCC RT Professional Behavior

Modern WinCC RT Professional (TIA Portal) documents the same fault path in the official Connection Fault to Partner Server (RT Professional) article. When the TCP connection to the partner server fails while both servers report running, RT Professional applies a configurable recovery strategy that avoids the oscillation seen in V6.2 by adding an exponential back-off on the partner-rejoin event. The same back-off is provided in V6.2 by SP2 + the matching hotfix; without it the redundancy monitor re-evaluates the partner immediately, producing the toggle pattern observed on BIO_CLNT01 / BIO_CLNT02.

Troubleshooting Matrix

Symptom First check Remediation
Client toggles every 1 min Patch level Apply SP2 + matching HF to all four stations
Client toggles every 15 min LMHOSTS Add #PRE entries and nbtstat -R
Random settle on OS2 Preferred server flag Set flag only on BIO_SERV01
Failover slow (> 30 s) Heartbeat time Reduce Monitoring time in Redundancy dialog
CRC errors on switch Duplex setting Force 100/FULL on NIC and switch port
Patch error on OS2 only Installation Clean uninstall + reinstall + SP2
Field note: Before re-installing BIOS-level changes or replacing the serial cable, always validate the patch level first. More than 80 % of V6.2 oscillation cases reported in production support are resolved by equalizing the patch level on the redundant pair.

Preventive Maintenance

  • Maintain a single image per OS class (XP client, Server 2003 server). Apply WinCC V6.2 + SP2 in one transaction.
  • Document the HF version in Computer > Properties > Version; attach the screenshot to the plant documentation.
  • Schedule a monthly redundancy drill: stop OS1, observe clients, restart OS1, confirm clients remain on OS2.
  • Keep a copy of LMHOSTS in the change-management repository.
  • Export CCAgent.log daily to a central syslog for trend analysis.

Why do my WinCC V6.2 clients keep switching back to the standby server after I bring the preferred server back online?

Pre-SP2 WinCC V6.2 builds do not implement the partner-rejoin back-off timer, so the redundancy monitor re-evaluates the partner as soon as its heartbeat returns. Combined with the lack of an enforced preferred-server flag, the client alternates between the two partners at 5-second intervals. Apply WinCC V6.2 Service Pack 2 to all four stations and ensure both servers report the same build.

Is LMHOSTS really required when the four stations are in a workgroup?

Yes. WinCC V6.2 redundancy uses NetBIOS names to identify the partner and the PLC. In a workgroup without WINS or DNS for NetBIOS, the WinCC redundancy monitor cannot resolve the partner name and reports intermittent partner outage. Add a #PRE entry for every WinCC station in C:\Windows\System32\drivers\etc\LMHOSTS and reload with nbtstat -R.

What undocumented error code should I look for in SIMATIC_S7_PROTOCOL_SUITE_01.LOG?

Codes above 0x8000FFFF are undocumented but typically indicate the protocol suite lost its connection to the S7-400. The four most common offenders in V6.2 are 0x80004005 (COM failure, often the serial cable), 0x80070005 (DCOM access denied), 0x80040154 (class not registered — patch missing) and 0x80010105 (server threw exception — patch mismatch).

How long should a clean failover take in WinCC V6.2?

With the default redundancy settings (monitoring time 5 s, failover time 25 s) the switch from OS1 to OS2 should complete inside 30 seconds. If it takes longer, check the serial cable, the LMHOSTS resolution latency and the S7 channel timeout values under Channel Units.

Can I mix WinCC V6.2 SP2 with an older hotfix on the redundant pair?

No. The pair must run the identical service-pack and hotfix level. Mixing levels produces the exact oscillation described in this article because the redundancy state machine interprets the divergent handshakes as a partner outage. Always install SP2 from the same hotfix package on both servers, then on both clients.

Back to blog