Troubleshooting Ignition 7.9 Gateway Memory Growth and GC Stalls

Karen Mitchell9 min read
HMI / SCADAOther ManufacturerTroubleshooting
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

An Ignition 7.9 gateway that runs out of heap, thrashes the disk, and knocks other services off the same server is not always leaking. In this installation the cause was a rarely triggered full garbage collection on a heap whose old-generation data had been paged out to the HDD, made worse by a table-update IO bug fixed in 7.9.1. Moving from the CMS collector to G1 flattened the memory trend. Work through the steps below in order; each ends with a check.

What does the operator see when the gateway heap runs out?

The operator sees a UI that stops responding, then repeats the same command several times because nothing acknowledges it. On the server, disk activity pegs and unrelated services drop out. Those are three separate layers of one event:

Observed at Symptom Underlying cause
Client / operator Frozen screen, repeated commands Gateway threads stalled during a long collection or IO wait; repeated commands add load
Server OS Heavy HDD swapping, other services dropping Pages holding old JVM heap data were swapped out and must be read back
Gateway memory graph Sawtooth plus occasional large drop Small spikes are young-generation collections; the large drop is a full collection of tenured data
Gateway status Pause and clock drift warnings The ClockDriftDetector warning and the 7.9 status overview pause list report pauses from any cause, not only GC

Check: open the gateway memory graph and the status overview pause list. Note the time of the last stall and match it to a large drop in the graph and to any pause entry. If they line up, continue; if the stall has no matching pause, look at the OS-level cause (disk, other processes) first.

Is it a leak or a delayed full collection?

Read the low-water mark, not the peaks. A leak shows as a low-water mark (the value just after a big collection) that rises from one big collection to the next. A delayed full collection shows as tenured data that grows for a long time, then drops back to roughly the same baseline. A healthy application should bottom out at roughly 10-20% of its maximum allowed heap after a big cycle, which leaves the collector room to choose when to compact.

Two points from this case are worth applying to any similar graph:

  • The regular small spikes are normal young-generation collections. They say nothing about a leak.
  • A leak in the custom web page served through the web-browser module would normally show up as client memory growth, not gateway heap growth. Suspecting the module is reasonable, but the graph decides it: if the baseline after big drops returns to the same level, the module is not leaking on the gateway.

In this installation the big collection turned out to happen only seldom, and the graph after it showed a return to baseline. That ruled out a leak as the primary cause.

Check: record the low-water mark after at least two large drops. Equal values mean tenured growth, not a leak. Rising values send you to the histogram procedure in the section on a rising low-water mark below.

Is the gateway actually running with the heap you configured?

A graph that shows out-of-memory with less than 1 GB occupied while the config says 2 GB means the running JVM heap does not match the file. Here the explanation was timing: the heap was 1 GB when the failure happened and was raised to 2 GB afterward. The configuration is in the gateway's wrapper configuration file (typically ignition.conf in the data folder of a standard install):

wrapper.java.initmemory=2048   # was 1024 MB
wrapper.java.maxmemory=2048    # was 1024 MB

Changes to these lines take effect only after the gateway service restarts. Setting initial equal to maximum removes heap resizing as a variable. If the graph still tops out below the configured maximum on a machine you expect to support it, check whether the JVM is 32-bit; a 32-bit JVM cannot address large heaps regardless of the file value.

Raising the heap buys time but does not remove the failure. A larger heap lets tenured data grow larger before the full collection, which can make the eventual pause longer and the paging worse if the server is short on physical RAM.

Check: after a restart, confirm the maximum heap shown in the gateway memory graph matches wrapper.java.maxmemory, and confirm physical RAM on the server exceeds the heap plus the other applications resident on it.

Why does a full collection on a paged-out heap freeze other services?

A full collection has to touch every tenured page to find unreferenced objects. If the gateway sat idle in the old generation for hours while other applications opened and closed, the OS paged out much of that unreferenced data. The collector then reads it back from disk, generating a burst of IO that starves other threads and applications. The result is timeouts, an unresponsive UI, and operators re-issuing commands, which adds more load.

A second contributor applied to this system: a bug in table updates caused excessive IO and was fixed in 7.9.1. The full-collection paging and the table-update IO combined to produce the outage. Neither alone was described as fatal.

The fix strategy follows from the mechanism: either touch and clean tenured data more often so it stays resident, or collect it in chunks so no single operation demands the whole paged-out heap at once.

Setting / item Location Effect
Gateway version 7.9.1 or later Gateway install / upgrade Removes the table-update IO bug that amplified the disk load
Physical RAM vs. heap Server hardware Prevents the OS from paging out JVM pages in the first place
GC algorithm Wrapper config Determines whether tenured data is scanned continuously or in one large full collection

Check: read the gateway version in the status page. If it is below 7.9.1, plan the upgrade alongside the GC change. During the next stall window, watch the server's disk queue and page-file usage; a spike that coincides with the large heap drop confirms this mechanism.

Why was CMS the collector, and what does it cost?

The gateway in this case ran the concurrent mark-sweep collector, set by -XX:+UseConcMarkSweepGC with -XX:+CMSClassUnloadingEnabled. It was chosen over the parallel collector to avoid frequent pauses of 500 ms or more.

CMS does not compact. In practice that does not mean the heap becomes permanently fragmented and unusable. It means the tenured area grows until a stop-the-world collection is forced. That growth-then-forced-pause pattern is exactly what the graph showed, and it becomes worse when tenured pages are on disk.

Changing the collector does not repair a real leak. It only changes how and when the memory is reclaimed, so decide leak versus tenured growth first (see the section on leak versus delayed collection).

Check: open the wrapper config and confirm which collector flags are present before changing anything. Copy the file as a backup.

How do I move the gateway from CMS to G1?

Replace the CMS flags with the G1 flag and keep the rest. The starting configuration was:

wrapper.java.additional.1=-XX:+UseConcMarkSweepGC
wrapper.java.additional.2=-XX:+CMSClassUnloadingEnabled
wrapper.java.additional.3=-Ddata.dir=data
wrapper.java.additional.4=-Dorg.apache.catalina.loader.WebappClassLoader.ENABLE_CLEAR_REFERENCES=false
#wrapper.java.additional.5=-Xdebug
#wrapper.java.additional.6=-Xrunjdwp:transport=dt_socket,server=y,suspend=n,address=8000
  1. Stop nothing yet; copy the wrapper config file to a backup.
  2. Replace the two CMS lines with one G1 line and renumber so the wrapper.java.additional.N sequence has no gaps (the Java service wrapper stops reading at a missing number). Remove -XX:+CMSClassUnloadingEnabled, since it is a CMS option.
  3. Keep -Ddata.dir=data and the ENABLE_CLEAR_REFERENCES line unchanged.
  4. Restart the gateway service.
wrapper.java.additional.1=-XX:+UseG1GC
wrapper.java.additional.2=-Ddata.dir=data
wrapper.java.additional.3=-Dorg.apache.catalina.loader.WebappClassLoader.ENABLE_CLEAR_REFERENCES=false
#wrapper.java.additional.4=-Xdebug
#wrapper.java.additional.5=-Xrunjdwp:transport=dt_socket,server=y,suspend=n,address=8000

G1 works here because it runs a multi-phase concurrent marking cycle that assesses the liveness of tenured data during normal operation, so tenured pages are read regularly and are less likely to sit in the page file. Its mixed collections then reclaim tenured regions in chunks rather than in one pass, which leaves IO time for other threads even if some data is paged out.

Weigh two caveats before committing:

Consideration Detail
Heap size G1 is not recommended for heaps smaller than 4-6 GB. This gateway ran at 2 GB and behaved well, but treat it as outside the recommended range and watch it.
CPU G1 raised baseline gateway CPU in testing. Record CPU before the change so you can compare.
Pause behavior G1 has been reported to hold latencies under 100 ms on high-packet-rate collection; confirm against the pause list on your own system.

G1 becomes the default collector from Java 9 onward, but this gateway's JVM is whichever runtime the 7.9 install ships, so set the flag explicitly.

Check: after restart, confirm the gateway starts cleanly (a bad flag or numbering gap stops the JVM from launching), and that the memory graph begins a new trace.

What if the low-water mark still climbs after the change?

A baseline that still rises after switching to G1 points to a real leak, and the JVM tools identify the owning class. Run a class histogram against the gateway JVM and compare two snapshots taken hours apart:

jmap -histo:live <gateway-jvm-pid> > histo_1.txt
(wait several hours of normal operation)
jmap -histo:live <gateway-jvm-pid> > histo_2.txt

Find the gateway JVM process ID with jps or Task Manager. The :live option forces a full collection before counting, so it causes a pause; run it in a low-impact window, not during production peaks. Classes whose instance counts and bytes grow steadily between snapshots are the leak candidates. Map the owning class to a module or script. Since this was the first use of the web-browser module with a custom page, disabling or isolating that page and re-comparing the histograms separates a module problem from a project-script problem.

Check: the leak is confirmed only if the same classes grow between two :live histograms; a one-off increase that disappears after a full collection is tenured growth, not a leak.

How do I prove the fix end to end?

After the G1 change on this gateway, the memory footprint held between 250 and 1400 MB with no rising trend and no large collections that could thrash the disk. Reproduce that result on your own system with this sequence:

  1. Confirm the gateway version is 7.9.1 or later so the table-update IO bug is not masking the result.
  2. Confirm wrapper.java.maxmemory matches the heap shown in the memory graph.
  3. Confirm the G1 flag is present and the additional-parameter numbering has no gaps.
  4. Watch the memory graph for a period longer than the previous interval between big collections. The low-water mark should stay flat and there should be no rising trend.
  5. Watch server page-file usage and disk queue. They should stay flat while the gateway runs.
  6. Check the status overview pause list and the ClockDriftDetector warnings for the same window. Pauses should be short and infrequent.
  7. Compare gateway CPU with the pre-change baseline to quantify the G1 overhead.

FAQ

How do I tell a memory leak from normal garbage collection in the Ignition gateway?

Compare the memory low-water mark after two or more large drops. A rising low-water mark indicates a leak; a return to the same baseline indicates tenured growth that a full collection reclaims. Confirm with two jmap -histo:live snapshots taken hours apart.

How do I switch the Ignition 7.9 gateway from CMS to the G1 garbage collector?

Edit the wrapper config: replace -XX:+UseConcMarkSweepGC and -XX:+CMSClassUnloadingEnabled with -XX:+UseG1GC, renumber the wrapper.java.additional.N lines with no gaps, and restart the gateway service. Expect somewhat higher baseline CPU, and note that G1 is not recommended for heaps below 4-6 GB.

How do I confirm the G1 change fixed the gateway memory problem?

Watch the gateway memory graph for longer than the old interval between big collections and confirm the footprint stays in a stable band (here 250-1400 MB) with no rising trend, while server page-file usage, the status overview pause list, and ClockDriftDetector warnings stay quiet.

Back to blog