The Ignition timer script’s synchronous system.net.httpClient().get(...) call was still waiting in Java’s HTTP client when the site internet dropped; the thread dump does not by itself show whether the wait was caused by a dead network path, a slow response, or a timeout that was not applied to that request.
Which hop stops carrying the HTTP request?
Trace the request from the timer script outward: the script calls the Ignition HTTP wrapper, the wrapper calls Java’s HTTP client, and the client sends the request across the host’s network path toward the URL. Replies must traverse the path back to the host before the synchronous call can return. A brief internet outage can interrupt any hop after the request leaves the host.
The stack shows execution inside HttpClientImpl.send, beneath Ignition’s JythonHttpClient.get. A synchronous send waits for a response. If a network device or remote endpoint leaves the TCP connection looking open while no response arrives, the calling thread can remain parked until the client, operating system, or network path reports a failure. A device doing deep packet inspection or other connection handling is one possible path issue, but the stack trace does not identify such a device as the cause.
Before changing the script, correlate a stuck invocation with gateway-host network evidence: check whether the host can resolve the URL’s name, establish a connection to the destination, and receive a response; inspect firewall, proxy, and packet-capture evidence along the route if those facilities are available. Record the destination and time of the failed request so the network team can inspect the same interval. Check: determine whether the host received an HTTP response, a connection error, or no transport-level completion before moving on.
What does the parked stack frame prove?
The important frames are CompletableFuture.get, HttpClientImpl.send, and Ignition’s JythonHttpClient.get. Together they indicate that the script thread is blocked waiting for the synchronous HTTP operation to complete. They do not prove that Python is executing stale code, nor do they reveal the remote server’s state or a particular network appliance’s behavior.
The stack’s generated Python frame names the script library function and reports a line number. Generated runtime frames can make source-to-line mapping less intuitive, especially when the active project source has changed since the call began. A reported line within getForecastData can represent the call path into a helper such as extractMultiTableCSV; the HTTP call may be elsewhere in the same function or below it. It is not evidence by itself that the gateway loaded a different compiled file.
| Observation | What it indicates | Next check |
|---|---|---|
Thread parked under CompletableFuture.get and HttpClientImpl.send
|
The synchronous HTTP send has not returned. | Inspect timeout configuration for the actual request. |
| Thread dump points into a function but source line appears unrelated | The line mapping or displayed source may not identify the immediate HTTP statement. | Compare the deployed project source and call path for that function. |
| Slow test endpoint returns a Java timeout exception | That test exercised delayed response handling, not necessarily an interrupted or blackholed network path. | Test the request deadline and diagnose the real network path separately. |
Check: capture a fresh thread dump while the timer is stuck and confirm the same HTTP wait frames before treating it as a Python logic or compilation problem.
Which timeout applies to connection setup and response wait?
Separate the timeout used when constructing the HTTP client from the timeout attached to an individual request. A client connect timeout bounds connection establishment; it does not necessarily bound the full time spent waiting for the response after a connection is established. A request-level timeout is the control to inspect for the complete operation. Do not rely on a remembered default: the discussion around this installation mentioned a presumed 60-second default, but the observed multi-hour wait makes the actual behavior in the deployed API and code the deciding evidence.
| Setting or operation | What it bounds | What to verify |
|---|---|---|
| HTTP client connect timeout | Attempt to establish a connection | Check the value used when the client is created; do not treat it as the response deadline. |
| Timeout supplied to the GET request | The request’s wait for completion, according to the API’s timeout semantics | Confirm the timeout is passed on the exact request that can hang and that the installed Ignition method supports that argument. |
| Async operation with a bounded wait | How long the script waits for the asynchronous result | Check the Promise/future wait limit and define how timeout or failure is handled. |
Adding a timeout to client creation alone does not prove that the call to CLIENT.get(reportUrl) is bounded. Likewise, adding a timeout parameter in source only helps if the running request uses it and the installed Ignition wrapper applies it as expected. Inspect the version-specific method signature or API documentation available with the installation rather than guessing the argument name, units, or default.
Check: confirm the active GET call contains the request-level timeout and identify its configured value and units from the deployed code and API definition.
How should the interruption test differ from a slow-server test?
A delayed endpoint that eventually returns tests a server that accepts a connection and responds late. Stopping the Ignition service during that delay tests service shutdown behavior as well. Neither necessarily reproduces a live client whose network path silently stops forwarding packets while the TCP session appears open. That distinction explains why the development test can consistently end in an IOException while the production timer remains blocked.
- Run a controlled request against the delayed endpoint with a request-level timeout explicitly set.
- Measure elapsed time and record whether the request returns normally, raises a timeout, or reports another I/O failure.
- In a separate controlled test, interrupt the relevant network path if the environment permits, without changing the request code.
- Compare the client result and thread state across both tests; do not use a service stop as a substitute for a network interruption.
Use packet capture or firewall/proxy logs when available to distinguish a request that never connects, a connection reset, and a connection that remains open with no response. If an intermediary preserves apparent TCP state through an upstream outage, the HTTP layer may not receive a prompt failure from that hop; an explicit request deadline limits how long the script waits regardless.
Check: the delayed-response test should terminate within the configured request deadline, and the interruption test should produce a bounded outcome rather than leave the timer call waiting indefinitely.
How can the timer call be bounded and handled?
Apply a finite request-level timeout to every potentially blocking GET, including the call to CLIENT.get(reportUrl). Choose a deadline appropriate to the operation’s allowed response time, then handle timeout and I/O exceptions at the script boundary. Log the destination, elapsed time, and exception category; avoid logging sensitive request data. Decide whether a failed fetch should skip the current timer cycle, retry under a bounded policy, or mark the data unavailable. Do not add unbounded retries, because they can keep a timer busy after connectivity fails.
If the synchronous method does not provide the required bound in the installed environment, consider the asynchronous HTTP path and give the Promise/future wait its own finite timeout, as suggested for this case. A timeout on the caller’s wait is useful only if the script can stop waiting and handle the resulting condition; verify what happens to the underlying request after that wait expires. Do not assume cancellation occurred unless the API documents it and the test confirms it.
Use a small test first: invoke the endpoint under normal conditions, then under a controlled delay exceeding the selected deadline. Verify that the timer logs and exits its request-handling path on timeout rather than holding a script thread. Check: confirm the measured wait is bounded by the configured policy and subsequent timer executions continue.
How can the active script source be matched to the stack?
Inspect the project source actually deployed to the gateway that produced the dump, not only an editor copy or a development project. Follow the call path from getForecastData through its helper functions and identify every HTTP request. Compare that source with the code where CLIENT is created and where the timeout argument is supplied. If a timer invocation was already running during a project change, use a new invocation after confirming the deployment so the next stack reflects the intended code.
Check: start a new timer cycle after deployment and confirm its observed call path includes the request-level timeout in the code currently running on the gateway.
What final end-to-end check proves the fix?
Run the deployed timer through a normal request and a controlled failure condition while observing its logs and a thread dump if it stalls. Confirm normal responses still reach the parsing logic, timeout or network failure follows the chosen exception path, and no timer thread remains parked in the HTTP send beyond the configured deadline. Then verify a later timer cycle can run after connectivity returns.
Final check: reproduce the interruption, measure the request duration, and verify that the HTTP call exits within the configured request timeout and the next scheduled invocation completes.
FAQ
What happens if the HTTP client has only a connect timeout?
It limits connection establishment, not necessarily the wait for a response after a connection is made. Add and verify a finite request-level timeout on the GET call.
What happens if an internet outage leaves the TCP connection open?
The synchronous send can continue waiting because it has not received a response or transport failure. A request deadline bounds the script’s wait; use network-path diagnostics to find where packets stop.
What happens if the stack trace line does not show the HTTP call?
The generated Python frame identifies the function and a mapped line, but it may not name the immediate HTTP statement. Check the deployed source and follow calls from getForecastData into its helpers.
What happens if a delayed endpoint test always raises an IOException?
That result shows the tested delay and shutdown path reports an I/O failure; it does not prove the test reproduced a silent network interruption. Repeat with a controlled path interruption and verify the request exits within its configured timeout.