Log signature of an exhausted scheduler pool
A Mango free-m2m2-core-3.0.1 instance publishing about 20 points over BACnet exposes only 8 of them to external BACnet software, and its log shows the scheduler pool pinned at its ceiling. Treat the missing points as a symptom of that saturation until the log proves otherwise.
| Log entry | What it means | Action |
|---|---|---|
FATAL ... RejectedRunnableEventGenerator.rejectedExecution with RejectedExecutionException, pool size = 100, active threads = 100, queued tasks = 0
|
All 100 workers of the OrderedThreadPoolExecutor are busy and there is no queue depth to absorb new work, so every new TimeoutTask is rejected. completed tasks = 11022 shows the pool ran normally before it saturated. |
Raise the High Priority Thread Pool maximum, then find what holds the threads. |
| Only 8 of about 20 points visible in the external client | Either the publisher never registered all objects, or the client's read of the object list was incomplete. | Compare the publisher point list against the client's object list. |
Save the log excerpt with the first FATAL timestamp. Every later step is judged against it.
Baseline of publisher points against the client object list
Record which points are missing before changing any setting, so the final check has something to compare with. The prerequisite is a BACnet client that can discover the Mango device and read its object list.
- Open the BACnet publisher in Mango and list every published point with its object type and instance number. Confirm the count matches the roughly 20 points under test.
- Confirm each source data point behind a published point is enabled and holds a current value. A publisher can only expose what its source points supply.
- Run device discovery from the external client and read the Mango device's object list. Record the 8 objects that appear.
- Compare the two lists. Look for a shared trait among the 12 missing points: object type, instance range, or source data source.
Gate: you hold a table of published points versus visible points. Duplicate object instance numbers within one object type collapse into fewer visible objects, so check for them now.
Raising the High Priority Thread Pool maximum
The fatal entries show a 100-thread pool fully active with an empty queue, so more headroom is the first change. Change the maximum only; leave everything else as found.
- Open the system settings page in Mango and locate the High Priority Thread Pool maximum pool size. Note the current value; the FATAL lines show 100 at the time of the fault.
- Raise the maximum in a moderate step, not to an arbitrary large figure. Each thread is an OS thread with its own stack.
- Save the setting. If the page indicates a restart is required, restart Mango.
- Read the next executor line in the log. The
pool sizeceiling must reflect the new maximum.
Gate: the new ceiling appears in the log line. If it still shows 100, the setting did not apply, and nothing downstream is worth diagnosing.
Confirming the pool is no longer saturated
A larger ceiling helps only if the active count stays below it. Watch the pool for several polling periods of the busiest data source.
- Tail the Mango log and filter for
RejectedExecutionException. No newFATALentries may appear. - If a rejection line appears, read
active threadsagainst the newpool size. If active equals the new maximum again, the threads are blocked rather than under-provisioned, and raising the maximum again only delays the failure. - For blocked threads, take a thread dump of the Mango JVM with the JDK's
jstack <pid>(the log shows a Java 8 runtime). Look for many threads waiting in the same BACnet request or I/O call.
Gate: no FATAL rejections across several poll periods, with active threads below the ceiling. A stack full of threads waiting on BACnet responses points to the transport and remote-device problems in the next two sections.
Clearing the bacnet4j expire NullPointerException
After the pool change, the log shows Error during expire messages from DefaultTransport.run:462 instead of the FATAL entries. The transport keeps a list of outstanding confirmed requests. On timeout, expire calls sendForResponse to retry, which calls Network.sendAPDU, and line 103 dereferences a null. Timeouts pile up when remote devices do not answer, and the loop then repeats the error at a high rate.
- Note which BACnet component owns the transport: the BACnet publisher's local device, or a BACnet IP data source polling remote devices. Check the log for the data source or publisher name near the first NPE.
- Disable, then re-enable that publisher or data source to rebuild its transport and network objects. Watch whether the NPE recurs.
- If it recurs, check whether the requests target a remote device that is offline or unreachable from the Mango host. Confirm reachability at the IP level first, then at BACnet level with an independent client.
- If the NPE persists with the pool healthy and every target reachable, treat it as a defect in the bundled
bacnet4j-4.0.1-SNAPSHOTlibrary. Snapshot builds carry no release guarantee. Read the release notes of a newer Mango core build for a bacnet4j fix before upgrading.
Gate: the ERROR line is absent for a full observation window after the restart of the owning component.
Clearing Task Queue Full on the polling data source
TaskRejectionHandler warnings mean a scheduled poll was dropped because no worker was free at fire time. The identifier after DS_ is the data source XID, so it maps to one specific data source.
- Find the data source whose XID is
DS_44f5a17c-f4a1-40e0-a221-390bb181bee4on the Data Sources page and record its type and polling period. - Read the settings page for every pool it lists. The polling job may run in a pool other than the one you raised.
- If the data source polls remote BACnet devices, check whether a slow or offline device holds workers until the request times out. Lengthen the poll period or reduce points per device until one poll completes well inside its period.
Gate: no Task Queue Full warning for that XID across at least several of its polling periods.
Re-registering the publisher points and re-reading the object list
Once the pool and the transport are quiet, force the publisher to rebuild its state and re-read from the client. Do this only after the previous gates pass, because a publisher restarted on a saturated pool fails the same way.
- Disable, then enable the BACnet publisher.
- Re-run discovery in the external client and read the object list again. Clients cache object lists, so clear the client's cached device entry or rescan.
- Compare against the baseline table. Every published object must appear with the recorded type and instance.
If the count is still short with a clean log, check three items: duplicate instance numbers within an object type, source data points that are disabled, and the client's handling of the object-list read. A client that cannot handle a segmented response may report a partial list; read individual array indexes of the object list to separate a client limit from a publisher gap. With only about 20 objects, expect saturation, not segmentation, to have caused the original shortfall.
End-to-end verification of all published points
- Confirm the log carries no
FATALRejectedExecutionException, noError during expire messages, and noTask Queue Fullwarning forDS_44f5a17c-f4a1-40e0-a221-390bb181bee4across a period covering several polls. - Confirm the executor line shows active threads below the raised pool ceiling.
- Read the Mango device's object list from the external BACnet client and count all published objects. The count must match the publisher's point list, not the earlier 8.
- Read the present value of a changing point twice, one poll period apart, from the external client. The value must update between reads and match the value shown in Mango.
FAQ
What happens if I raise the High Priority Thread Pool maximum and active threads still reach the new ceiling?
The threads are blocked, not under-provisioned, so a larger pool only delays the next rejection. Take a jstack thread dump and look for many threads waiting on the same BACnet request or I/O call, then fix the unreachable device or slow poll holding them.
What happens if the NullPointerException at Network.sendAPDU keeps appearing after the pool is healthy?
Restart the owning publisher or data source, then check that every remote target is reachable. If the error persists at Network.java:103 under those conditions, it is a defect in the bundled bacnet4j-4.0.1-SNAPSHOT library and needs a newer Mango core build.
What happens if Task Queue Full warnings continue for only one polling data source?
That data source's polls are still being dropped at fire time. Match its XID (DS_44f5a17c-f4a1-40e0-a221-390bb181bee4 in the log) to the Data Sources page, compare its poll period to the warning spacing, and lengthen the period or cut points per device until a poll completes inside its period.