TL;DR — Biometric terminals at a canteen push each scan over plain HTTP to a small agent on the same network. The agent writes the scan to disk, acknowledges it, and forwards it to the backend when it can. Before forwarding, it applies four time-based rules: refuse scans older than 3 minutes, refuse scans stamped more than 5 minutes in the future, drop the same person on the same terminal within 5 seconds, and cap each terminal at 120 scans a minute. Every one of those rules has to measure time from when the scan was captured and queued, not from when it is delivered. Measured at delivery, a 10-minute internet outage drains as a burst of scans that are all “stale”, arrive within the same second, and look like one terminal firing hundreds of duplicates — so a network blip would refuse every queued meal. An audit caught that before the first site went live.

Key takeaways

  • An offline queue has three clocks: the device’s clock at capture, the agent’s clock at queue write, and the clock at delivery. Delivery time is the only one that is wrong for every business rule.
  • Acknowledge only after the write is durable. A device told “got it” deletes its copy.
  • Staleness exists for one case: a device replaying history. Judge it against the queue write, and a real outage stops looking like a replay.
  • Run the stale check first and return early, so a replayed log does not spend the rate budget or poison the duplicate window.
  • A strictly ordered queue turns one permanently failing item into an outage. Cap item size and never retry what cannot succeed.

Why there is an agent at all

The terminals — face and fingerprint readers bolted next to the serving counter — are HTTP push clients with firm opinions. You configure a bare IP address and a port. They send plaintext HTTP/1.0, with no TLS and no authentication, using one of two protocol families: the iclock/ADMS push protocol used by ZKTeco-family devices, and a length-prefixed JSON protocol used by another vendor’s line. They cannot call an HTTPS webhook, and a cloud host that only accepts HTTPS cannot accept them.

So a small agent runs on a machine at the site — a Windows PC or an Android POS tablet — listening on the local network. It does three jobs: speak each terminal’s protocol correctly, keep every scan safe through an internet outage, and forward scans one at a time as signed HTTPS requests to the backend endpoint described in the serving webhook post.

  terminal ──plain HTTP──▶ agent (LAN) ──▶ disk queue ──HTTPS + HMAC──▶ backend
     ▲                        │                                   │
     └──── ack after fsync ◀──┘                          200 {served | rejected}

  internet down?  terminal notices nothing; the queue grows; drain on recovery

Honest framing: this is a design and build write-up. The agent has passed its unit and end-to-end suites and been run against real terminals on a test bench; the first site is scheduled to go live in October. There are no production numbers here yet.

Acknowledge only after the write is durable

The first rule is about the device, not the backend. When a terminal gets an acknowledgement, it considers the scan delivered and stops holding it. If the agent acks first and writes second, a power cut between the two loses a real person’s meal with no record anywhere.

So the write is synchronous and happens before the reply:

// Illustrative. The listener's hot path.
function onScan(req, res, scan) {
  try {
    queue.insert({ ...scan, queuedAt: Date.now() });  // synchronous, fsync'd
  } catch (err) {
    // Disk full, permissions, corrupted file. Do NOT ack:
    // the device keeps the record and resends it.
    res.writeHead(503).end();
    return;
  }
  res.writeHead(200).end(protocol.ackFor(scan));
  void drain();                                        // nudge delivery, don't wait for it
}

The queue is SQLite in WAL mode with synchronous = FULL. The distinction matters: the SQLite documentation is explicit that WAL with synchronous = NORMAL is consistent but “does lose durability”, and “a transaction committed in WAL mode with synchronous=NORMAL might roll back following a power loss”. A PC under a canteen counter loses power more often than a server in a rack does. FULL costs an fsync per scan, and at one scan every few seconds that is nothing.

The listener never waits on the internet. That is the property that makes an outage invisible to the hardware: the terminal’s request completes in milliseconds whether the backend is reachable or not.

Three clocks, and which one each rule uses

Every scan carries three timestamps by the time it is judged:

  capture        the device's clock when the finger touched the reader   (scan instant)
  queue write    the agent's clock when the scan hit disk                 (queuedAt)
  delivery       the agent's clock when the rules run before forwarding   (now)

  normally:      capture ≈ queue write ≈ delivery      (milliseconds apart)
  after outage:  capture ≈ queue write  ≪  delivery    (minutes apart)

This is the same split Tyler Akidau draws between event time and processing time in Streaming 101, and it has the same moral: rules about what happened must be evaluated in event time. Delivery time only tells you when the network came back.

The agent’s four rules, and the clock each uses:

RuleThresholdMeasured asWhy this clock
Stale≥ 3 minqueue write − scan instantCatches replayed history; an outage delays delivery, not the queue write
Clock skew> 5 min in the futurequeue write − scan instantA scan stamped after it was received means a wrong device clock
Duplicatesame device + person within 5 sscan instant vs previous scan instantTwo touches seconds apart are one person; two deliveries seconds apart are not
Rate limit120 per device per minutesliding window over scan instantsA terminal cannot physically scan faster; a drain can deliver much faster

The reference point for age is the queue write, clamped so it can never be later than now:

// Illustrative.
const instant    = toInstant(scan.scannedAtLocal, site.timezone);   // device clock
const observedAt = Math.min(scan.queuedAt ?? Date.now(), Date.now()); // never trusted into the future
const age        = observedAt - instant;

if (age >= 3 * 60_000)  return reject(scan, 'stale');
if (age < -5 * 60_000)  return reject(scan, 'clock_skew');
if (rateLimiter.exceeded(scan.device, instant))  return reject(scan, 'rate_limited');
if (dedup.seenWithin(scan.device, scan.person, instant, 5_000)) return reject(scan, 'duplicate');

Nothing in that block reads the delivery time. That is the point of it.

What an outage drain looks like on the wrong clock

Take a 10-minute internet outage during lunch, one terminal, 200 scans queued. The network comes back; the agent drains the backlog in a couple of seconds. Now run the rules against delivery time:

  scans captured:   12:30:00 ──────────────────────────── 12:40:00   (200 people)
  delivered:                                             12:40:02 ── 12:40:04

  stale?        age = delivery − capture = 0 to 10 minutes  → ~140 of 200 "stale"
  rate limit?   200 scans in 2 seconds                        → "rate_limited"
  duplicate?    windows keyed on delivery time collapse        → "duplicate"

  ▼  every queued meal refused — by the rules meant to stop a replay

The three rules fail together, and each failure looks plausible on its own. Stale scans are exactly what a replay produces. A burst from one terminal is exactly what a stuck device produces. That is why the bug is easy to write and hard to see in a log: every individual rejection reason is a reason the rule was designed to give.

Measured on capture and queue-write time instead, the same drain passes cleanly. Every scan was queued within milliseconds of its capture, so every age is near zero. Scan instants are spread across ten minutes, so the rate limiter sees twenty a minute. Duplicate windows are compared between capture instants, so two different people are two different people.

The case staleness actually exists for

The stale rule is not paranoia. Terminals replay their history, and a replayed punch from last Tuesday must not open today’s lunch.

The two protocol families do it differently. One vendor’s terminals, on first connection, replay their entire stored log unless every reply carries exactly the acknowledgement they expect — anything else and they “re-send the same record forever, or stop talking entirely”. The iclock family replays the whole attendance log on every reconnect unless the handshake tells it ATTLOGStamp=None, and batches punches unless told Realtime=1. Getting those replies byte-exact is most of the protocol work.

When a replay does happen, the scans arrive now, were queued now, and were captured days ago. Queue write minus capture is large, so the stale rule fires. That is the one signal that separates a replay from an outage drain: in a drain, capture and queue write are close together and delivery is late; in a replay, capture is old and queue write is current.

Two details make it work. The stale check runs first and returns early, so a replayed log never reaches the rate limiter or the duplicate window — otherwise a thousand old punches would spend the terminal’s rate budget and plant dedup entries that block real scans. And stale scans are still acknowledged to the device. An ack is not agreement; it tells the terminal to stop resending. Refusing to ack a stale record would make the device replay it forever.

Two layers of duplicate detection

The five-second window lives in memory, in a map keyed by device and person, holding up to 50,000 entries. Each entry stores the event id as well as the instant, so a scan being retried after a failed delivery is never mistaken for its own duplicate.

Memory does not survive a restart, and a restart in the middle of lunch is exactly when a double-serve would happen. So there is a second, durable layer: a served record written to a separate SQLite file, also synchronous = FULL, after the backend returns 2xx and before the scan is removed from the queue. The comment on that table in the code says what it is for: it “must survive a power cut”.

The backend has its own idempotency claim and its own unique constraint on the served meal, covered in the webhook post. The agent’s layers exist so that the common duplicates never leave the building, and the backend’s layers exist so that the rare ones cannot do damage when they do.

A strictly ordered queue, and what that costs

Scans are delivered first-in, first-out, in batches of 50, so the backend sees them in the order people stood at the counter. The price of strict ordering is head-of-line blocking: one scan that keeps failing holds up everything behind it. Three decisions keep that price bounded.

Retry only what can succeed. A timeout (10 seconds), a network error, 408, 429 or any 5xx is retried with exponential backoff: 2 seconds after the first failure, doubling, capped at 60 seconds. A new scan arriving skips the wait and triggers an immediate drain, so recovery is noticed within one scan. Any other 4xx is dead-lettered with the response body, because resending the same bytes will get the same answer. Redirects are not followed.

Business refusals are not failures. “No meal plan” is a 200 with a reason, for the reason the webhook post spends a section on. From the agent’s side the argument is sharper: if a business refusal were a 5xx, one person without a plan would block every person behind them until the backoff cap, and then again.

Cap the size of anything that enters the queue. Frames over 256 KB are refused before they enter the queue, well under the backend’s 1 MB body limit. More on why below.

What broke — in an audit, not in production

On August 14, before any site was live, a full review of the agent and its cloud predecessor turned up ten bugs. They were fixed in one commit, verified by 62 unit tests, 29 end-to-end tests and a 29-point live smoke run. Three of them belong in this post.

The outage drain. The first version judged staleness as now − capture, fed the rate limiter now, and ran the duplicate window on now. The database-side duplicate check read the row’s received-at time. The review’s severity table put it plainly: “any internet outage longer than 3 minutes meant every queued scan was rejected as stale on drain”. Nobody had hit it, because nobody had pulled the network cable during a test. The fix added queuedAt to every queued scan and moved every rule onto capture and queue-write time. The database check moved to the scan’s own timestamp, with a comment that “received_at would reject all but the first”.

I want to be precise about one thing: the regression test for this arrived five days later, with the rewrite of the agent, not with the fix. It queues a scan 600 seconds old with a queue write 570 seconds ago and asserts it is still fresh. A fix without that test was a fix I was trusting from reading the diff.

One oversized frame wedged the queue forever. It was rated the most severe finding. A single malformed or oversized item failed delivery on every attempt, and the queue could not move past it. Every later scan piled up behind it “and silently aged out” — which, with the stale rule, means refused. The fix was the 256 KB cap at the listener and dead-lettering permanent failures one item at a time, so a bad item costs one scan instead of the queue.

A config flag that could not be turned off. The configuration schema coerced environment strings to booleans, and in that coercion the string "false" is truthy. A flag set to false in the deployment file, to keep a safety check on, was read as true and turned the check off. It had nothing to do with clocks, and I include it because it is the same class of bug: a value that reads correctly and means the opposite. Booleans from the environment are now parsed explicitly, with a test for exactly this string.

FAQ

Should an offline queue use event time or processing time? Event time — the capture timestamp, or the moment the event was durably queued — for every rule about what happened: staleness, duplicate windows, rate limits. Processing time only tells you when delivery ran, which after an outage is minutes late for every event. Rules evaluated on processing time will reject a legitimate backlog as stale and duplicated.

How do you tell a replayed event from a delayed one? Compare the capture time with the time the event was first queued. A delayed event was queued moments after capture and delivered late. A replayed event was captured long ago and queued just now. Staleness measured as queue write minus capture flags replays and ignores outages.

When should an edge device acknowledge a message? Only after the message is durably written. The sender deletes or stops retrying once acknowledged, so an ack before the write risks losing the message on a crash. If the write fails, do not ack; return an error so the sender keeps its copy and resends.

Is SQLite synchronous=NORMAL safe for a durable queue? Not if you acknowledge on commit. In WAL mode, synchronous = NORMAL keeps the database consistent but a committed transaction can roll back after a power loss. For a queue that acknowledges a sender on commit, use synchronous = FULL; at a few writes per second the cost is negligible.

How do you rate-limit events that arrive in a burst after an outage? Rate-limit on the events’ capture timestamps, not their arrival. A device that produced twenty events a minute for ten minutes is within its limit even if all two hundred arrive in two seconds. A limit on arrival time punishes the recovery.

Why should stale events still be acknowledged? Because an acknowledgement tells the sender to stop sending, not that you accepted the event. Refusing to acknowledge a stale event makes most devices resend it indefinitely. Acknowledge, record the rejection, and move on.

How do you prevent one bad message from blocking an ordered queue? Separate retryable failures (timeouts, 429, 5xx) from permanent ones (other 4xx), and dead-letter the permanent ones immediately instead of retrying. Cap message size at the point of entry. Return business refusals as successes so they never enter the retry path.

What I’d still improve

  • The 503-means-resend behaviour is asserted, not tested on hardware. The protocols say a terminal that does not get an ack keeps the record. I have not pulled a disk out from under the agent on a real terminal to watch it happen, and I would like to before the first site depends on it.
  • The queue drops its oldest scans past 100,000. A week-old scan is useless for serving lunch, so this is the right policy, and it is counted. It also contradicts a line in the agent’s own docs that says nothing is lost. The docs are wrong and need fixing.
  • A comment that describes code that does not exist. The durable-dedup comment says a crash between recording “served” and removing the scan from the queue re-attempts an idempotent delivery. Reading the code, after a restart the scan would match its own served record and be dropped as a duplicate. The outcome is harmless — the person was served — and the comment is still wrong about how.
  • The in-memory window and rate limiter reset on restart. The durable layer covers the case that matters, a double serve. A restart can still let one burst through the rate limiter.
  • The device’s clock is trusted for the capture instant, checked only for gross skew against the agent. A device clock that is wrong by a consistent amount under five minutes would pass. That is the next thing I would monitor: the distribution of queue write minus capture, per terminal.

The one idea to take away

If a rule is about when something happened, never measure it with the clock of when you found out.

An outage changes only one thing about a queued event: when it is delivered. Every rule that reads delivery time inherits the outage as if it were a property of the event, and turns a network problem into a business decision. Stamp the event when you first hold it, judge it by that stamp, and the backlog after an outage is just a backlog.


I write about backend reliability, webhook and payment correctness, and Postgres data modelling from work on food-tech and healthcare platforms. More on what I build and how I work, the webhook where every business outcome is a 200, and a durable queue in Postgres instead of Kafka. If your offline queue has a staleness rule, check which clock it reads.