TL;DR — A person puts a face or a finger on a terminal at a canteen; the vendor records the scan and POSTs it to the platform; within the time it takes to pick up a plate, the response decides whether that person gets this meal at this outlet right now. The endpoint is a webhook, and webhooks come with a rule the caller enforces: non-2xx means retry. That rule shapes everything. A business “no” — no plan, wrong window, already ate — is a 200 with a reason, never a 500, because a retry cannot change a business fact. The rest of the design follows from taking that seriously: verify the signature over the raw bytes before parsing, claim the idempotency key before validating anything so a bad delivery is searchable, anchor the meal to when the scan happened rather than when it arrived, and answer from one gather-only database call that makes no decisions.

Key takeaways

  • Separate the two questions a webhook response answers: did you receive this (HTTP status) and what did you decide (body). Conflating them turns every business rejection into a retry storm.
  • Verify HMAC over the bytes you were sent. Parsing and re-serialising produces different bytes, and every check fails.
  • Claim the idempotency key first, validate second. The failed, malformed and rejected deliveries are the ones support needs to find.
  • Two idempotency defences are better than one when they fail differently: a unique key on the delivery, and a unique constraint on the business fact.
  • A gather-only RPC collapses N round trips into one without moving decisions into the database, so the serve logic keeps exactly one tested implementation.

The problem

  person ──face/finger──▶ terminal ──▶ vendor cloud ──POST──▶ platform
                                                                 │
                                                          decide: serve?
                                                                 │
  person ◀──── green light / "no plan" ◀──── terminal ◀── 200 {status, reason}

  budget: the time it takes to pick up a plate

The vendor knows who scanned, where, and when. Everything after that is the platform’s: meal windows, subscription plans, penalties for unclaimed bookings, wallet balance, whether this person already ate this meal today. The request that arrives looks roughly like this:

POST /serve
X-Terminal-Idempotency-Key: 9f2c1b04-…:4471
X-Terminal-Signature: t=1785995030,v1=1560a263…

{ "userId": "10234", "outletId": "…", "deviceSn": "…",
  "scannedAt": "2026-08-06T05:43:50.000Z", "method": "face", "scanId": "4471" }

Honest framing before the design: this was built for a large client site, integrated with their counter frontend and the terminal vendor, and tested against 62 endpoint cases plus 24 more on the decision and gather layers. It is scheduled to go live in October and has not served a live lunch at the time of writing. This is a design post; there are no production numbers in it, and there will not be until there are.

Every business outcome is a 200

The vendor’s delivery contract is the standard one: 2xx means delivered, anything else is retried on a backoff schedule. That is the right contract for a webhook, and it means the HTTP status can only answer one question — did the platform receive and record this scan?

The business question — does this person get lunch? — has to live in the body:

{ "status": "rejected", "reason": "no_active_plan_at_outlet", "scanId": "4471" }

status is one of served, already_served, rejected. reason is a stable string from a fixed list of about a dozen: unknown_user, outside_meal_window, no_active_plan_at_outlet, plan_expired, meal_not_in_plan, blocked_by_penalty, insufficient_wallet_balance, scan_too_old, and so on. Those strings land in the vendor’s dashboard next to the scan and are what a canteen supervisor searches on, so they are part of the contract: never renamed once live.

The rule I would put at the top of any webhook spec: a person with no plan is never a 500. If they were, the vendor would retry — six more times over the next hour — asking the same question about the same person, and the answer would be the same every time. A retry is a request to try again in case something transient got in the way. Nothing transient is in the way of “you don’t have a plan here”.

Non-2xx is reserved for the platform’s own trouble, and each code says what the vendor should do:

StatusMeaningVendor’s action
200Received and decided (served / already_served / rejected + reason)Show the result. Done.
400Malformed — missing fields, bad timestampDead-letter immediately. A retry sends the same bytes.
401Signature invalidDead-letter and alarm on the vendor side — a key rotation went wrong.
429Twin delivery still in flightRetry — the first copy will have finished.
500 / 503Database unreachable, or the shared secret is not configured on our sideRetry. This is the case retries are for.

Each reason also maps, in the integration document, to who fixes it: the outlet (window misconfigured), the customer (plan expired, balance low), nobody (already served today), or the platform (unknown user — an enrolment is missing). That column is what turns a rejection from a complaint into a ticket.

HMAC over the raw bytes, before the JSON parser

The signature is t=<unix seconds>,v1=<hex>, an HMAC-SHA256 over <t>.<raw body>. The detail that matters: it has to be computed over the bytes the vendor sent, not over JSON.stringify(req.body). Re-serialising a parsed object reorders nothing in theory and changes whitespace, number formatting and key order in practice, and every check fails.

// Illustrative. Mount the raw parser on this route BEFORE the app-wide JSON parser.
app.post('/serve', express.raw({ type: 'application/json' }), verifySignature, handleServe);
app.use(express.json()); // everything else

function verifySignature(req, res, next) {
  const { t, v1 } = parseSignatureHeader(req.get('X-Terminal-Signature'));
  if (Math.abs(Date.now() / 1000 - Number(t)) > 300) return res.status(401).end();

  const expected = crypto.createHmac('sha256', SECRET).update(`${t}.`).update(req.body).digest();
  const given = Buffer.from(v1, 'hex');
  if (given.length !== expected.length || !crypto.timingSafeEqual(given, expected)) {
    return res.status(401).end();
  }
  req.scan = JSON.parse(req.body); // parse only after the bytes are trusted
  next();
}

timingSafeEqual because a byte-by-byte string compare leaks how many leading bytes matched. The 300-second tolerance on t is the replay defence: a captured request replayed at dinner is what it rejects. And the parse happens after verification, so a malformed body from an unauthenticated sender never reaches the JSON parser at all.

Claim before you validate

Idempotency for a webhook usually means: look up the key, replay the stored response if you have one. The payment webhook post covers the layers of that for money. What this endpoint adds is ordering: the key is claimed — a row inserted into the delivery log — before the body is validated and before the meal decision runs.

-- Illustrative. The claim. UNIQUE on idempotency_key does the work.
INSERT INTO serve_log (idempotency_key, scan_id, received_at, status)
VALUES ($1, $2, now(), 'processing')
ON CONFLICT (idempotency_key) DO NOTHING
RETURNING id;

Why claim first? Because the deliveries that matter for support are the bad ones. A scan from a terminal with the wrong outlet id, a delivery with a garbage timestamp, a person the terminal knows and the platform does not — if those are rejected before anything is written, they exist only in a log file. Claimed first, they are rows in a table that a supervisor can search by scan id, with the reason next to them.

The claim also resolves the concurrent-twin case without a lock. If the insert returns no row, a delivery with this key already exists, and there are two possibilities:

  • It has a stored response → replay it, with X-Idempotent-Replay: true so the vendor’s dashboard can tell a replay from a first answer.
  • It is still processing → 429, and the vendor retries after its backoff. By then the twin has finished and the retry gets the replay.

And one more: a claim that has been processing for more than 90 seconds belongs to a request that crashed between claiming and completing. The next delivery with that key takes the claim over rather than waiting forever behind a ghost. Ninety seconds is far longer than any real request and shorter than a typical retry backoff, which is the only property the number needs.

Two defences that fail differently

The delivery-key uniqueness above catches the same delivery twice. It does not catch the same person at two terminals in the same outlet, seconds apart, which produces two deliveries with two different keys and, on a busy counter, is not hypothetical.

That is a different fact, so it gets a different constraint:

-- Illustrative. One person, one meal, one day, one outlet — enforced by the index.
CREATE UNIQUE INDEX served_meals_once
  ON served_meals (served_date, menu_id, customer_id);

menu_id encodes outlet, meal and day, so the index says exactly what the business rule says. The second terminal’s insert loses the race and the response is already_served — a 200, with a reason — not a constraint error surfacing as a 500. Uniqueness is enforced by the index, not by a read-then-write in application code, for the same reason it is everywhere else on this platform: two requests can both read “not served yet” and both proceed; they cannot both insert. What happens when a table skips that constraint is 2,008 duplicate bookings.

The two defences fail in different ways, which is why both exist. The delivery log can be taken over after 90 seconds; the meals-served index cannot be. The delivery log is per delivery; the index is per fact. A bug that broke one would not break the other.

The claim-then-complete shape turned out to be generic enough that it was later lifted into a webhook_deliveries table shared by other inbound integrations. The property it buys is worth stating plainly: a delivery that arrives during an outage is recorded even if it cannot be applied, so the worst case after an incident is “replay these” rather than “those meals are gone”.

Anchor the meal to scan time, not arrival time

A delivery that arrives after a retry backoff — or is replayed from the vendor’s dashboard — can land minutes after the scan. If the meal window is resolved against now(), a lunch scan that arrives at 14:31 resolves to no window, or to the wrong one. So the window is resolved against scannedAt from the body: the meal the person actually stood in front of the terminal for.

Windows are compared as local wall-clock times, using the same comparison the existing RFID path uses, so a face terminal and a card reader at the same counter agree about what “lunch” means. And a maximum scan age — 24 hours by default — refuses a dead delivery hand-replayed days later, with scan_too_old as the reason. That is the one place where “the scan is real but we will not act on it” is the right answer.

The same rule applies one hop earlier. The agent at the site that buffers scans through an internet outage judges staleness, duplicates and rate limits by capture time too, for the reasons in the offline-queue post.

Identity is scoped to the outlet

The terminals enrol people under their own numeric ids, and those ids are namespaced per outlet. User 10234 at one canteen is a different human from 10234 at the next. So the lookup order is: resolve the outlet from the request, then resolve the person from (outlet_id, terminal_user_id) on the row that whitelists that customer at that outlet.

-- Illustrative. Two people at one outlet can never share a terminal id.
CREATE UNIQUE INDEX whitelist_terminal_id_once
  ON outlet_whitelist (outlet_id, terminal_user_id)
  WHERE terminal_user_id IS NOT NULL;

The partial unique index is the guard against the failure that would actually hurt: an enrolment typo giving two people the same id at one outlet, which would hand one of them the other’s meal. The customer record is then read by its own primary key, never by the terminal id.

Terminal devices get the opposite scope. A device key is unique globally, because a terminal is bolted to exactly one wall, and the same key arriving for two outlets means somebody cloned a device record rather than two sites legitimately sharing one.

Six round trips to one — without moving the decision

The first working version of the handler made six sequential calls through the API gateway to Postgres before it could answer: the outlet, the enrolment, the day’s windows and menus, the already-served check, and a booking lookup. Each costs roughly 200 ms through the proxy regardless of how trivial the query is, so a person at the counter waited about 1.4 seconds for work that takes Postgres single-digit milliseconds.

  before:  outlet ──▶ enrolment ──▶ windows ──▶ menus ──▶ served? ──▶ booking?
           200ms      200ms         200ms       200ms     200ms       200ms   ≈ 1.4 s

  after:   gather_serve_context(outlet, terminal_user, date, week, dow) ──▶ one JSONB
           ≈ 200ms + a few ms of SQL

The replacement is one function that returns everything as a single JSONB document. The design constraint on that function is the part I would defend hardest: it makes no decisions. No window matching, no entitlement evaluation, no already-served ruling. It gathers. Every decision stays in the service layer, where the 62 tests are, so the serve logic has exactly one implementation and moving the data fetch did not create a second one.

That is a deliberate tension with the advice to move business logic into RPC functions, which I stand by for aggregations and filters — work the database is better at. A decision with a dozen reason codes and a test suite is not that. The database is better at fetching five things in one round trip; the service is better at deciding what they mean.

Two smaller choices in the same spirit. Week number and day-of-week are parameters to the function, computed once in the service against the local timezone, because duplicating that arithmetic in SQL would create a second source of truth for which menu is today. And the response shape to the counter frontend could not change, because it was already integrated — so the tests lock the outgoing shape, and the refactor was measured by the tests not moving. The overload postmortem is what happens when per-call round-trip cost is ignored at scale; this is the same arithmetic applied before the scale arrives.

Record the fact the tables cannot reconstruct

The client’s rule: nobody is turned away for not booking, but the kitchen needs to know afterwards who ate without one. That sounds derivable — join meals served to bookings — and it is not. The served row says a meal went out; the booking row says a booking exists; neither says whether the booking existed at the moment of serving, and bookings can be created or edited after the fact.

So the served row records it directly, as a three-state column:

  • true — a booking existed when the meal was served
  • false — it did not; this is the number the kitchen wants
  • NULL — booking did not apply to this meal, or the row predates the column

Nullable with no default makes it a catalogue-only change on a table that will carry on the order of 100,000 rows a day at that site: no rewrite, no lock, no backfill of a value that cannot be reconstructed anyway. The general rule is the one from the audit-trail post: a fact about the state of the world at the time of an event is an observation, not a derivation, and if you do not write it down when it happens, it is gone.

What actually broke — in the build, not in production

There is no live incident to report, so the honest version of this section is what broke on the way.

The six round trips broke the latency budget before anything else did. The first version was correct and took about 1.4 seconds, which at a counter serving a person every few seconds is a visible pause and at a rush is a queue. That measurement is what produced the gather function, and it came after the endpoint was already integrated with the counter frontend — which is why the response shape was locked by tests before the refactor rather than designed for it.

The booking fact was the second thing. The first version recorded that a meal was served and relied on the bookings table to say whether it had been booked. The client’s question — who ate without booking? — turned out to be unanswerable from those two tables, for the reason above, and the column went in a day after the gather function. A fact you will be asked about later is a fact to write down now.

Go-live checklist

Written for the first site, kept for the next:

  1. Apply the migration; confirm the function and the unique indexes exist.
  2. Backfill terminal ids per outlet from the vendor’s enrolment export — a SQL job, since no API exists for it. Deleting a whitelist row deletes the enrolment with it, on purpose.
  3. Set the shared secret; confirm a request with a bad signature gets a 401 and a request with no secret configured gets a 503, not a 200.
  4. Send the vendor the outlet id for each terminal serial.
  5. Fire a test scan from the vendor’s dashboard at staging; confirm the replay header on the second delivery.
  6. Watch the first live lunch with one query:
SELECT status, reason, count(*)
  FROM serve_log
 WHERE received_at > current_date
 GROUP BY 1, 2
 ORDER BY 3 DESC;

A table full of served with a short tail of outside_meal_window is a good first lunch. A table full of unknown_user is step 2 done wrong.

FAQ

Should a webhook return an error status for a business rejection? No. A webhook’s HTTP status tells the sender whether the delivery was received; the sender retries on non-2xx. A business rejection — no plan, outside the window, already served — is a fact that a retry cannot change, so return 200 with a structured status and reason in the body. Reserve 4xx for deliveries that can never succeed and 5xx for your own transient failures.

Why must webhook signatures be verified over the raw request body? Because the HMAC was computed over the exact bytes the sender transmitted. Parsing the JSON and re-serialising it changes whitespace, key order or number formatting, so the digest no longer matches. Mount a raw-body parser on the webhook route ahead of the JSON parser and verify before parsing.

Why claim the idempotency key before validating the request? So that malformed, misconfigured and rejected deliveries are recorded as rows, searchable by scan id with a reason, rather than existing only in a log file. Claiming first also resolves concurrent duplicates: a second delivery with the same key finds the claim, gets a 429 while the first is in flight, and a replay of the stored response afterwards.

How do you handle a webhook delivery whose processing crashed midway? Give the claim a stale timeout. A delivery row left in processing for longer than any real request could take (90 seconds here) is treated as abandoned, and the next delivery with the same key takes it over. Choose the timeout to be far shorter than the sender’s retry interval so the retry, not a human, does the recovery.

Should business logic live in the database function or the application? Fetching should move to the database when it saves round trips; deciding should stay where it is tested. A gather-only function that returns everything the decision needs in one call gets the latency win without creating a second implementation of the rules.

How do you prevent the same person being served twice at two terminals? With a unique index on the business fact — (date, menu, customer) — not with a read-then-write check in code. Two requests can both read “not yet served”; only one can insert. The loser gets already_served as a normal 200 response.

Why resolve the meal window against the scan timestamp rather than the current time? Because a delivery can arrive minutes late after a retry or a replay, and the person was standing at the terminal at scan time, not at arrival time. Resolve against scannedAt, and enforce a maximum age so a very old delivery is refused rather than served against a yesterday’s window.

What I’d still improve

  • Overlapping meal windows. Where two windows overlap, the resolver takes the first candidate, which is arbitrary at the boundary. The launch site’s windows do not overlap, so this is latent rather than live; the fix is a tie-break rule (closest window start, or the one the person has a plan for) before the second site.
  • A supervisor view over the serve log, so “why was I rejected?” is answered at the counter from the reason column rather than by a support ticket.
  • Load test the gather function at the site’s peak — one scan every few seconds across several counters — before the first live lunch, not after.
  • Alert on unknown_user rate. A spike means an enrolment sync is broken, and it will present as a queue of angry people before it presents as a log line.

The one idea to take away

A retry cannot change a business fact, so do not ask for one.

A webhook’s status code is a delivery receipt, and the sender will act on it mechanically. The moment a business “no” leaks into that channel, you have told a machine to keep asking a question whose answer is fixed. Put the decision in the body with a stable reason, keep the status code for your own failures, and everything downstream — the dashboards, the retries, the support search — starts meaning what it says.


I write about backend reliability, webhook and payment correctness, and Postgres data modelling from work on food-tech and healthcare platforms. More on what I build and how I work, the webhook that decides an order is paid, and why round trips matter before you have scale. If your webhook returns a 4xx for a business rejection, check how many times the sender is asking.