The Dabao SDK millis() assembles a 64-bit timestamp from two separate 32-bit register reads, TICKTIMER_TIME0 (low) then TICKTIMER_TIME1 (high). No single access returns a coherent 64-bit value. If the low word rolls from to between the two reads, the function returns the old low word paired with the new high word, a value 232 counts ahead of true time. The same SDK also exposes UART transfers only through a polling pair, uart_write_async() and uart_write_done(). Both issues are fixable in software without touching the hardware.
Torn 64-bit reads on a split TIME0/TIME1 counter
The failure window is a few bus cycles wide and opens once per low-word rollover. If millis()ticktimer_init() configures), 232 ms is about 49.7 days. A bench session never spans a rollover. A fleet of devices in the field does, and each affected unit shows one inexplicable timing event.
With the read order low-then-high, the torn result is always ahead of true time by 232 counts. What that does to delay_ms() depends on how the interval is compared:
| Observed behavior | Mechanism |
|---|---|
millis() returns one value ~232 counts ahead, then correct values again |
Low word read before rollover, high word read after |
delay_ms() returns early |
A read taken with the +232 error inside the loop, or a torn start captured too high: the uint64_t subtraction millis() - start wraps to a huge value and fails the < ms test |
Any code that stores a millis() value as a deadline or timestamp inherits the same error. Fix millis() and every consumer is fixed.
Read methods compared
Four approaches exist for reading a 64-bit counter split across two 32-bit registers. Compare them on register reads per call, the condition under which each is correct, and how each fails.
| Method | Reads | Correct when | Failure mode |
|---|---|---|---|
| Low, then high (current) | 2 | Never guaranteed | +232 error at each low-word rollover |
| Rely on hardware latching the high word on a low-word read | 2 | Only if the RTL proves the latch exists and applies to TIME1
|
Silent error if the assumption is wrong; ties software correctness to an unverified hardware detail |
| High, low, high; retry if high changed | 3 | Counter is monotonic, at any tick rate | At most one retry per rollover |
| Low, high, low, high; retry unless both pairs match | 4 | The counter does not tick between the two pairs | Retries whenever a tick lands between the pairs; can spin repeatedly if the counter advances faster than the four reads complete |
Recommendation: high, low, high. Its exit condition is "the high word did not change", which is independent of tick rate. The four-read compare-both-pairs variant exits only when the counter did not advance at all across its reads. That is harmless at a slow tick and turns into a livelock risk if the counter clocks faster than the read sequence. Skip the latch shortcut unless the RTL proves it, and even then the three-read loop costs one extra bus read and stays correct either way.
The high-low-high loop covers both rollover placements. Rollover between hi_a and the low read yields hi_b = hi_a + 1. Rollover between the low read and hi_b also yields hi_b = hi_a + 1. Rollover before hi_a leaves all three reads consistent.
Register macro prerequisites before changing the loop
The retry loop is only as good as the compiler's respect for read order and count. Confirm each item before editing millis():
- Open the register definition header that defines
TICKTIMER_TIME0andTICKTIMER_TIME1. Confirm each expands to avolatile-qualified 32-bit dereference. Gate: if either is a plain load, the compiler can merge or reorder the reads and the loop degenerates; fix the macro first. - Confirm
ticktimer_init()runs once before the firstmillis()call. The lazys_ticktimer_initializedcheck insidemillis()adds a branch and a first-call race ifmillis()can be called from both main context and an interrupt; callticktimer_init()during startup. - Check the RTL for the timer block. Record whether reading
TIME0latchesTIME1, and record the tick rate. The tick rate sets the real rollover period, and the latch answer tells you whether the hardware already protects a low-then-high read. Do not remove the retry loop on that basis.
If the macros are volatile, accesses to them keep program order and no barrier is needed. If you keep a compiler barrier as belt-and-braces, write it as a complete statement: __asm__ volatile ("" ::: "memory");.
Replacement millis() and delay_ms()
Drop this in for the existing function. It reads high, low, high and retries until the two high words match:
uint64_t millis(void)
{
if (!s_ticktimer_initialized) {
ticktimer_init();
}
uint32_t hi_a;
uint32_t lo;
uint32_t hi_b;
do {
hi_a = TICKTIMER_TIME1;
lo = TICKTIMER_TIME0;
hi_b = TICKTIMER_TIME1;
} while (hi_a != hi_b);
return ((uint64_t)hi_b << 32) | lo;
}
The delay_ms() body needs no structural change once millis() is coherent:
void delay_ms(uint32_t ms)
{
uint64_t start = millis();
while ((millis() - start) < ms) { }
}
Keep the subtraction in unsigned 64-bit. With a coherent, monotonic millis(), millis() - start is the true elapsed time, and unsigned wraparound only matters after the 64-bit counter itself wraps. Do not cast the difference to a signed type.
Host-side rollover test
A bench run cannot reach the rollover, so test the algorithm on a PC with simulated registers. The simulated counter advances on every register read, which is the worst case for a torn read. Sweep start values across the 32-bit boundary and require every result to lie between the counter value before the call and after it:
#include <stdint.h>
#include <stdio.h>
static uint64_t sim;
static uint32_t T0(void){ uint32_t v = (uint32_t)sim; sim++; return v; }
static uint32_t T1(void){ uint32_t v = (uint32_t)(sim >> 32); sim++; return v; }
static uint64_t millis_naive(void){
uint32_t lo = T0(); uint32_t hi = T1();
return ((uint64_t)hi << 32) | lo;
}
static uint64_t millis_hlh(void){
uint32_t a, lo, b;
do { a = T1(); lo = T0(); b = T1(); } while (a != b);
return ((uint64_t)b << 32) | lo;
}
int main(void){
int bad_n = 0, bad_h = 0;
for (uint64_t s = 0xFFFFFFF0ull; s < 0x100000010ull; s++){
uint64_t r;
sim = s; r = millis_naive(); if (r < s || r > sim) bad_n++;
sim = s; r = millis_hlh(); if (r < s || r > sim) bad_h++;
}
printf("naive bad=%d hlh bad=%d\n", bad_n, bad_h);
return bad_h != 0;
}
Gate: bad_n is non-zero (the test reproduces the defect) and bad_h is zero. A test that cannot fail the original code proves nothing about the fix. This validates the algorithm only; it says nothing about the macros' volatility or the real bus timing.
UART completion callback in place of polling
The UDMA-driven peripheral model has no simple UART data register flow: the CPU hands a buffer to the DMA engine and the transfer completes on its own. Polling uart_write_done() exposes that mechanism clearly for bring-up, but production firmware should not spin waiting for it. Add a callback layer that leaves the low-level UDMA mechanism visible and lets higher-level code avoid polling. An illustrative shape (not the SDK's current signature):
typedef void (*uart_done_cb_t)(void *ctx);
int uart_write_async(/* port */, const uint8_t *buf, size_t len,
uart_done_cb_t cb, void *ctx);
Apply these constraints when implementing it:
- Store
cbandctxper port before starting the DMA transfer, so a fast completion cannot fire against an unset callback. Gate: callback pointer is written before the start-transfer register access in the code path. - Invoke the callback from the DMA/interrupt completion handler after clearing the port's busy state, so the callback can queue the next transfer directly.
- Return a busy error if a transfer is already outstanding on that port, unless you add a queue. One outstanding transfer per channel is the simplest correct contract.
- Document that
bufmust stay valid and unmodified until the callback runs; the DMA engine reads it after the call returns. Confirm from the memory map in the RTL or datasheet that the buffer lives in memory the UDMA can address. - Document that the callback runs in interrupt context: no
delay_ms(), no blocking UART call, no long loops. - Keep
uart_write_done()as the polling path for simple examples and bring-up code; it becomes a thin wrapper over the same busy flag.
Verification on the target
- Run the host test from the previous section and confirm
bad_h == 0before flashing. - On hardware, call
millis()in a tight loop and count any result lower than the previous one. Gate: zero violations over a run long enough to cover many ticks. If the RTL exposes a way to preload the counter, set it just below a 32-bit boundary (TIME0near ) and repeat the loop across the rollover; otherwise the host test carries the rollover coverage. - Send a known payload with the callback-based UART write and capture it on a receiver. Gate: the receiver sees the payload intact and the completion callback fires exactly once per transfer, with the count matching the number of writes issued.
FAQ
Can I read TICKTIMER_TIME0 and TICKTIMER_TIME1 without a retry loop?
Only if the RTL proves that reading the low word latches the high word. Without that proof, use the high-low-high loop; it costs one extra register read and is correct with or without a hardware latch.
Does the high-low-high loop ever spin indefinitely?
No. It retries only when the high word changes between hi_a and hi_b, which happens at most once per 232 counts, so the second pass always succeeds. The compare-both-pairs variant is different: it retries whenever the counter ticks between its reads.
Does a torn millis() read always make delay_ms() hang?
No. With the unsigned subtraction shown, a value that is 232counts ahead usually ends the delay early.
Can I call millis() from an interrupt handler?
Yes, the loop keeps no shared state after initialization, so it is reentrant. Call ticktimer_init() at startup rather than relying on the lazy s_ticktimer_initialized check from interrupt context.
Can I keep uart_write_done() after adding a completion callback?
Yes. Keep it as the polling path for bring-up and simple examples, backed by the same per-port busy flag the DMA completion handler clears before invoking the callback.