TL;DR — A person puts a face or a finger on a terminal at a canteen; the vendor records the scan and POSTs it to the platform; within the time it takes to pick up a plate, the response decides whether that person gets this meal at this outlet right now. The endpoint is a webhook, and webhooks come with a rule the caller enforces: non-2xx means retry. That rule shapes everything. A business “no” — no plan, wrong window, already ate — is a 200 with a reason, never a 500, because a retry cannot change a business fact. The rest of the design follows from taking that seriously: verify the signature over the raw bytes before parsing, claim the idempotency key before validating anything so a bad delivery is searchable, anchor the meal to when the scan happened rather than when it arrived, and answer from one gather-only database call that makes no decisions.
Key takeaways
- Separate the two questions a webhook response answers: did you receive this (HTTP status) and what did you decide (body). Conflating them turns every business rejection into a retry storm.
- Verify HMAC over the bytes you were sent. Parsing and re-serialising produces different bytes, and every check fails.
- Claim the idempotency key first, validate second. The failed, malformed and rejected deliveries are the ones support needs to find.
- Two idempotency defences are better than one when they fail differently: a unique key on the delivery, and a unique constraint on the business fact.
- A gather-only RPC collapses N round trips into one without moving decisions into the database, so the serve logic keeps exactly one tested implementation.
The problem
person ──face/finger──▶ terminal ──▶ vendor cloud ──POST──▶ platform
│
decide: serve?
│
person ◀──── green light / "no plan" ◀──── terminal ◀── 200 {status, reason}
budget: the time it takes to pick up a plate
The vendor knows who scanned, where, and when. Everything after that is the platform’s: meal windows, subscription plans, penalties for unclaimed bookings, wallet balance, whether this person already ate this meal today. The request that arrives looks roughly like this:
POST /serve
X-Terminal-Idempotency-Key: 9f2c1b04-…:4471
X-Terminal-Signature: t=1785995030,v1=1560a263…
{ "userId": "10234", "outletId": "…", "deviceSn": "…",
"scannedAt": "2026-08-06T05:43:50.000Z", "method": "face", "scanId": "4471" }
Honest framing before the design: this was built for a large client site, integrated with their counter frontend and the terminal vendor, and tested against 62 endpoint cases plus 24 more on the decision and gather layers. It is scheduled to go live in October and has not served a live lunch at the time of writing. This is a design post; there are no production numbers in it, and there will not be until there are.
Every business outcome is a 200
The vendor’s delivery contract is the standard one: 2xx means delivered, anything else is retried on a backoff schedule. That is the right contract for a webhook, and it means the HTTP status can only answer one question — did the platform receive and record this scan?
The business question — does this person get lunch? — has to live in the body:
{ "status": "rejected", "reason": "no_active_plan_at_outlet", "scanId": "4471" }
status is one of served, already_served, rejected. reason is a stable string from
a fixed list of about a dozen: unknown_user, outside_meal_window,
no_active_plan_at_outlet, plan_expired, meal_not_in_plan, blocked_by_penalty,
insufficient_wallet_balance, scan_too_old, and so on. Those strings land in the vendor’s
dashboard next to the scan and are what a canteen supervisor searches on, so they are part
of the contract: never renamed once live.
The rule I would put at the top of any webhook spec: a person with no plan is never a 500. If they were, the vendor would retry — six more times over the next hour — asking the same question about the same person, and the answer would be the same every time. A retry is a request to try again in case something transient got in the way. Nothing transient is in the way of “you don’t have a plan here”.
Non-2xx is reserved for the platform’s own trouble, and each code says what the vendor should do:
| Status | Meaning | Vendor’s action |
|---|---|---|
200 | Received and decided (served / already_served / rejected + reason) | Show the result. Done. |
400 | Malformed — missing fields, bad timestamp | Dead-letter immediately. A retry sends the same bytes. |
401 | Signature invalid | Dead-letter and alarm on the vendor side — a key rotation went wrong. |
429 | Twin delivery still in flight | Retry — the first copy will have finished. |
500 / 503 | Database unreachable, or the shared secret is not configured on our side | Retry. This is the case retries are for. |
Each reason also maps, in the integration document, to who fixes it: the outlet
(window misconfigured), the customer (plan expired, balance low), nobody (already served
today), or the platform (unknown user — an enrolment is missing). That column is what turns a
rejection from a complaint into a ticket.
HMAC over the raw bytes, before the JSON parser
The signature is t=<unix seconds>,v1=<hex>, an HMAC-SHA256 over <t>.<raw body>. The
detail that matters: it has to be computed over the bytes the vendor sent, not over
JSON.stringify(req.body). Re-serialising a parsed object reorders nothing in theory and
changes whitespace, number formatting and key order in practice, and every check fails.
// Illustrative. Mount the raw parser on this route BEFORE the app-wide JSON parser.
app.post('/serve', express.raw({ type: 'application/json' }), verifySignature, handleServe);
app.use(express.json()); // everything else
function verifySignature(req, res, next) {
const { t, v1 } = parseSignatureHeader(req.get('X-Terminal-Signature'));
if (Math.abs(Date.now() / 1000 - Number(t)) > 300) return res.status(401).end();
const expected = crypto.createHmac('sha256', SECRET).update(`${t}.`).update(req.body).digest();
const given = Buffer.from(v1, 'hex');
if (given.length !== expected.length || !crypto.timingSafeEqual(given, expected)) {
return res.status(401).end();
}
req.scan = JSON.parse(req.body); // parse only after the bytes are trusted
next();
}
timingSafeEqual because a byte-by-byte string compare leaks how many leading bytes
matched. The 300-second tolerance on t is the replay defence: a captured request replayed
at dinner is what it rejects. And the parse happens after verification, so a malformed body
from an unauthenticated sender never reaches the JSON parser at all.
Claim before you validate
Idempotency for a webhook usually means: look up the key, replay the stored response if you have one. The payment webhook post covers the layers of that for money. What this endpoint adds is ordering: the key is claimed — a row inserted into the delivery log — before the body is validated and before the meal decision runs.
-- Illustrative. The claim. UNIQUE on idempotency_key does the work.
INSERT INTO serve_log (idempotency_key, scan_id, received_at, status)
VALUES ($1, $2, now(), 'processing')
ON CONFLICT (idempotency_key) DO NOTHING
RETURNING id;
Why claim first? Because the deliveries that matter for support are the bad ones. A scan from a terminal with the wrong outlet id, a delivery with a garbage timestamp, a person the terminal knows and the platform does not — if those are rejected before anything is written, they exist only in a log file. Claimed first, they are rows in a table that a supervisor can search by scan id, with the reason next to them.
The claim also resolves the concurrent-twin case without a lock. If the insert returns no row, a delivery with this key already exists, and there are two possibilities:
- It has a stored response → replay it, with
X-Idempotent-Replay: trueso the vendor’s dashboard can tell a replay from a first answer. - It is still
processing→429, and the vendor retries after its backoff. By then the twin has finished and the retry gets the replay.
And one more: a claim that has been processing for more than 90 seconds belongs to a
request that crashed between claiming and completing. The next delivery with that key takes
the claim over rather than waiting forever behind a ghost. Ninety seconds is far longer than
any real request and shorter than a typical retry backoff, which is the only property the
number needs.
Two defences that fail differently
The delivery-key uniqueness above catches the same delivery twice. It does not catch the same person at two terminals in the same outlet, seconds apart, which produces two deliveries with two different keys and, on a busy counter, is not hypothetical.
That is a different fact, so it gets a different constraint:
-- Illustrative. One person, one meal, one day, one outlet — enforced by the index.
CREATE UNIQUE INDEX served_meals_once
ON served_meals (served_date, menu_id, customer_id);
menu_id encodes outlet, meal and day, so the index says exactly what the business rule
says. The second terminal’s insert loses the race and the response is already_served —
a 200, with a reason — not a constraint error surfacing as a 500. Uniqueness is enforced by
the index, not by a read-then-write in application code, for the
same reason it is everywhere else on this platform:
two requests can both read “not served yet” and both proceed; they cannot both insert.
What happens when a table skips that constraint
is 2,008 duplicate bookings.
The two defences fail in different ways, which is why both exist. The delivery log can be taken over after 90 seconds; the meals-served index cannot be. The delivery log is per delivery; the index is per fact. A bug that broke one would not break the other.
The claim-then-complete shape turned out to be generic enough that it was later lifted into
a webhook_deliveries table shared by other inbound integrations. The property it buys is
worth stating plainly: a delivery that arrives during an outage is recorded even if it
cannot be applied, so the worst case after an incident is “replay these” rather than
“those meals are gone”.
Anchor the meal to scan time, not arrival time
A delivery that arrives after a retry backoff — or is replayed from the vendor’s dashboard —
can land minutes after the scan. If the meal window is resolved against now(), a lunch
scan that arrives at 14:31 resolves to no window, or to the wrong one. So the window is
resolved against scannedAt from the body: the meal the person actually stood in front of
the terminal for.
Windows are compared as local wall-clock times, using the same comparison the existing RFID
path uses, so a face terminal and a card reader at the same counter agree about what “lunch”
means. And a maximum scan age — 24 hours by default — refuses a dead delivery hand-replayed
days later, with scan_too_old as the reason. That is the one place where “the scan is
real but we will not act on it” is the right answer.
The same rule applies one hop earlier. The agent at the site that buffers scans through an internet outage judges staleness, duplicates and rate limits by capture time too, for the reasons in the offline-queue post.
Identity is scoped to the outlet
The terminals enrol people under their own numeric ids, and those ids are namespaced per
outlet. User 10234 at one canteen is a different human from 10234 at the next. So the
lookup order is: resolve the outlet from the request, then resolve the person from
(outlet_id, terminal_user_id) on the row that whitelists that customer at that outlet.
-- Illustrative. Two people at one outlet can never share a terminal id.
CREATE UNIQUE INDEX whitelist_terminal_id_once
ON outlet_whitelist (outlet_id, terminal_user_id)
WHERE terminal_user_id IS NOT NULL;
The partial unique index is the guard against the failure that would actually hurt: an enrolment typo giving two people the same id at one outlet, which would hand one of them the other’s meal. The customer record is then read by its own primary key, never by the terminal id.
Terminal devices get the opposite scope. A device key is unique globally, because a terminal is bolted to exactly one wall, and the same key arriving for two outlets means somebody cloned a device record rather than two sites legitimately sharing one.
Six round trips to one — without moving the decision
The first working version of the handler made six sequential calls through the API gateway to Postgres before it could answer: the outlet, the enrolment, the day’s windows and menus, the already-served check, and a booking lookup. Each costs roughly 200 ms through the proxy regardless of how trivial the query is, so a person at the counter waited about 1.4 seconds for work that takes Postgres single-digit milliseconds.
before: outlet ──▶ enrolment ──▶ windows ──▶ menus ──▶ served? ──▶ booking?
200ms 200ms 200ms 200ms 200ms 200ms ≈ 1.4 s
after: gather_serve_context(outlet, terminal_user, date, week, dow) ──▶ one JSONB
≈ 200ms + a few ms of SQL
The replacement is one function that returns everything as a single JSONB document. The design constraint on that function is the part I would defend hardest: it makes no decisions. No window matching, no entitlement evaluation, no already-served ruling. It gathers. Every decision stays in the service layer, where the 62 tests are, so the serve logic has exactly one implementation and moving the data fetch did not create a second one.
That is a deliberate tension with the advice to move business logic into RPC functions, which I stand by for aggregations and filters — work the database is better at. A decision with a dozen reason codes and a test suite is not that. The database is better at fetching five things in one round trip; the service is better at deciding what they mean.
Two smaller choices in the same spirit. Week number and day-of-week are parameters to the function, computed once in the service against the local timezone, because duplicating that arithmetic in SQL would create a second source of truth for which menu is today. And the response shape to the counter frontend could not change, because it was already integrated — so the tests lock the outgoing shape, and the refactor was measured by the tests not moving. The overload postmortem is what happens when per-call round-trip cost is ignored at scale; this is the same arithmetic applied before the scale arrives.
Record the fact the tables cannot reconstruct
The client’s rule: nobody is turned away for not booking, but the kitchen needs to know afterwards who ate without one. That sounds derivable — join meals served to bookings — and it is not. The served row says a meal went out; the booking row says a booking exists; neither says whether the booking existed at the moment of serving, and bookings can be created or edited after the fact.
So the served row records it directly, as a three-state column:
true— a booking existed when the meal was servedfalse— it did not; this is the number the kitchen wantsNULL— booking did not apply to this meal, or the row predates the column
Nullable with no default makes it a catalogue-only change on a table that will carry on the order of 100,000 rows a day at that site: no rewrite, no lock, no backfill of a value that cannot be reconstructed anyway. The general rule is the one from the audit-trail post: a fact about the state of the world at the time of an event is an observation, not a derivation, and if you do not write it down when it happens, it is gone.
What actually broke — in the build, not in production
There is no live incident to report, so the honest version of this section is what broke on the way.
The six round trips broke the latency budget before anything else did. The first version was correct and took about 1.4 seconds, which at a counter serving a person every few seconds is a visible pause and at a rush is a queue. That measurement is what produced the gather function, and it came after the endpoint was already integrated with the counter frontend — which is why the response shape was locked by tests before the refactor rather than designed for it.
The booking fact was the second thing. The first version recorded that a meal was served and relied on the bookings table to say whether it had been booked. The client’s question — who ate without booking? — turned out to be unanswerable from those two tables, for the reason above, and the column went in a day after the gather function. A fact you will be asked about later is a fact to write down now.
Go-live checklist
Written for the first site, kept for the next:
- Apply the migration; confirm the function and the unique indexes exist.
- Backfill terminal ids per outlet from the vendor’s enrolment export — a SQL job, since no API exists for it. Deleting a whitelist row deletes the enrolment with it, on purpose.
- Set the shared secret; confirm a request with a bad signature gets a 401 and a request with no secret configured gets a 503, not a 200.
- Send the vendor the outlet id for each terminal serial.
- Fire a test scan from the vendor’s dashboard at staging; confirm the replay header on the second delivery.
- Watch the first live lunch with one query:
SELECT status, reason, count(*)
FROM serve_log
WHERE received_at > current_date
GROUP BY 1, 2
ORDER BY 3 DESC;
A table full of served with a short tail of outside_meal_window is a good first lunch. A
table full of unknown_user is step 2 done wrong.
FAQ
Should a webhook return an error status for a business rejection?
No. A webhook’s HTTP status tells the sender whether the delivery was received; the sender
retries on non-2xx. A business rejection — no plan, outside the window, already served — is
a fact that a retry cannot change, so return 200 with a structured status and reason in
the body. Reserve 4xx for deliveries that can never succeed and 5xx for your own transient
failures.
Why must webhook signatures be verified over the raw request body? Because the HMAC was computed over the exact bytes the sender transmitted. Parsing the JSON and re-serialising it changes whitespace, key order or number formatting, so the digest no longer matches. Mount a raw-body parser on the webhook route ahead of the JSON parser and verify before parsing.
Why claim the idempotency key before validating the request? So that malformed, misconfigured and rejected deliveries are recorded as rows, searchable by scan id with a reason, rather than existing only in a log file. Claiming first also resolves concurrent duplicates: a second delivery with the same key finds the claim, gets a 429 while the first is in flight, and a replay of the stored response afterwards.
How do you handle a webhook delivery whose processing crashed midway?
Give the claim a stale timeout. A delivery row left in processing for longer than any
real request could take (90 seconds here) is treated as abandoned, and the next delivery with
the same key takes it over. Choose the timeout to be far shorter than the sender’s retry
interval so the retry, not a human, does the recovery.
Should business logic live in the database function or the application? Fetching should move to the database when it saves round trips; deciding should stay where it is tested. A gather-only function that returns everything the decision needs in one call gets the latency win without creating a second implementation of the rules.
How do you prevent the same person being served twice at two terminals?
With a unique index on the business fact — (date, menu, customer) — not with a
read-then-write check in code. Two requests can both read “not yet served”; only one can
insert. The loser gets already_served as a normal 200 response.
Why resolve the meal window against the scan timestamp rather than the current time?
Because a delivery can arrive minutes late after a retry or a replay, and the person was
standing at the terminal at scan time, not at arrival time. Resolve against scannedAt, and
enforce a maximum age so a very old delivery is refused rather than served against a
yesterday’s window.
What I’d still improve
- Overlapping meal windows. Where two windows overlap, the resolver takes the first candidate, which is arbitrary at the boundary. The launch site’s windows do not overlap, so this is latent rather than live; the fix is a tie-break rule (closest window start, or the one the person has a plan for) before the second site.
- A supervisor view over the serve log, so “why was I rejected?” is answered at the counter from the reason column rather than by a support ticket.
- Load test the gather function at the site’s peak — one scan every few seconds across several counters — before the first live lunch, not after.
- Alert on
unknown_userrate. A spike means an enrolment sync is broken, and it will present as a queue of angry people before it presents as a log line.
The one idea to take away
A retry cannot change a business fact, so do not ask for one.
A webhook’s status code is a delivery receipt, and the sender will act on it mechanically. The moment a business “no” leaks into that channel, you have told a machine to keep asking a question whose answer is fixed. Put the decision in the body with a stable reason, keep the status code for your own failures, and everything downstream — the dashboards, the retries, the support search — starts meaning what it says.
I write about backend reliability, webhook and payment correctness, and Postgres data modelling from work on food-tech and healthcare platforms. More on what I build and how I work, the webhook that decides an order is paid, and why round trips matter before you have scale. If your webhook returns a 4xx for a business rejection, check how many times the sender is asking.