TL;DR — Splitting a customer’s payment between platform and vendor works until a subsidy makes the customer’s payment tiny: on a 99%-subsidised ₹100 order the customer pays ₹1, the vendor’s share is 81 paise, and the gateway refuses to transfer less than ₹1 — so the order could not be created at all. The fix is to stop splitting the payment and instead transfer the vendor its full net (customer share plus employer share) from the platform’s own balance, which is safe precisely because the employer’s money was collected up front. The harder half is paying that amount exactly once when both a payment webhook and a reconciliation job are allowed to pay: claim the row with a compare-and-swap, cap the attempts, put a recovery key in the transfer’s notes, and ask the provider what it already did before re-firing a claim a dead worker left behind.
Key takeaways
- A payment-linked split is capped at the payment and subject to the gateway’s minimum. Both limits are invisible until a subsidy shrinks the customer’s share below them.
- Once a third party has pre-funded most of a meal, the vendor payout is not a fraction of the customer’s payment any more. It is a transfer of money you already hold.
- Fire a payout only after capture is confirmed. A transfer for a payment that never completes is real money out the door.
- When two workers may pay the same row, “pay once” has to be a property of the row — a status transition that only one of them can win — not a property of the code path.
- Before retrying a payout that a crashed worker left half-done, ask the provider whether it happened. The recovery key is whatever you put in the transfer that lets you find it again.
The order nobody could pay
The platform runs corporate meal programs: an employer pre-funds a wallet, and each employee’s order is mostly paid from that wallet with the employee paying the remainder online. The subsidy engine decides the split; the wallet top-up is where the employer’s money comes in. This post is where money goes out, to the outlet that cooked the meal.
The original payout model was the one every marketplace starts with: the customer pays, and the gateway routes a slice of that payment to the vendor’s linked account. Razorpay calls this Route, and it has two constraints that are entirely reasonable and entirely fatal here. A transfer linked to a payment cannot exceed that payment’s amount, and no transfer can be smaller than ₹1.
₹100 order, 99% employer subsidy, 19% platform commission on the vendor's take
customer pays online ₹ 1.00
employer wallet covers ₹ 99.00
───────────────────────────────────
vendor's net for the meal ₹ 81.00 (₹100 − 19% commission)
Route split of the customer's payment:
vendor's share of ₹1.00 = ₹0.81 ← below the ₹1 minimum → transfer rejected
→ order creation fails
→ customer cannot pay ₹1 for a ₹100 meal
The people with the most generous subsidy were the people who could not check out. The failure was in order creation, before any money moved, so nothing was lost — but nothing could be bought either.
Why the first failure logged nothing
This deserves its own paragraph because it is why the first incident report was useless.
The gateway SDK’s errors, in the version we ran, do not carry a .message. The shape is:
// Illustrative. What the SDK actually rejects with.
{
statusCode: 400,
error: {
code: 'BAD_REQUEST_ERROR',
description: 'The amount must be at least INR 1.00',
field: 'transfers[0].amount',
},
}
The handler did logger.error('order create failed', err.message), which logged an empty
string. The first incident report was “orders fail for some users, no error”. If you log a
provider error, log the whole object — or at least err.error?.description ?? err.message.
An empty log line is worse than no log line, because it looks like coverage.
Stop splitting the payment; pay from the balance
The insight that unlocked the redesign is that the split model was answering the wrong question. Route asks: how should this ₹1 be divided? The right question is: what does the platform owe the vendor for this order? — and the answer is ₹81, regardless of how the ₹100 was funded.
Razorpay’s Direct Transfer moves money from the platform’s own account balance to a linked account, untied to any payment. So instead of sending the vendor 81 paise from the customer’s ₹1, the platform sends the vendor ₹81 from its balance. That amount is always comfortably above ₹1 for any real order, and the payment-cap constraint no longer exists because the transfer is not linked to a payment.
before: customer ──₹1──▶ gateway ──split──▶ ₹0.81 to vendor ✗ (min ₹1)
₹0.19 to platform
after: customer ──₹1──▶ platform balance
employer ──₹99─▶ platform balance (already there, from the wallet top-up)
platform balance ──transfer──▶ ₹81 to vendor ✓
Why it is safe to pay ₹81 when only ₹1 arrived with the order: the other ₹99 is already in the platform’s hands. The employer topped the wallet up before the employee could order against it, and that top-up was captured and reconciled at the time. The vendor is being paid from money that genuinely exists, not from money the platform expects to receive.
This is the distinction that decides which orders the direct transfer applies to:
| Condition | Direct transfer? | Why |
|---|---|---|
| Subsidy > 0 and vendor has a linked account | Yes | The employer’s share is pre-funded; the vendor’s net is real money the platform holds |
| Customer paid online or from wallet | Yes | Capture confirms the customer’s share is in hand too |
| Customer paid cash at the counter | No | The vendor physically holds the cash. What the platform owes is the net of that — often small, sometimes negative (the vendor owes commission). Stays on the offline settlement ledger |
| No subsidy | No | Nothing pre-funded; the ordinary payment split is fine and the minimum is never hit |
The cash row is the one that needed the most thought. A cash order with a subsidy means the employer’s share still has to reach the vendor, but the vendor already has the customer’s share in the till. The platform’s obligation is subsidy minus commission, which on a small order can be a few rupees or below zero. Trying to express that as a transfer means sometimes transferring nothing and sometimes needing money back; the offline ledger already handles both directions and was left to do so.
Fire only after capture
A direct transfer is not reversible by the customer’s bank. If it fires for an order whose payment is later declined, the platform has paid ₹81 for a meal nobody bought. So the transfer is fired from exactly two places: the payment webhook, after the gateway confirms capture, or immediately after a wallet debit succeeds — never from the checkout request itself.
That is the same rule as treating the webhook as the source of truth for the order’s paid status, extended one step: the event that marks the order paid is the event that makes the vendor payable. The order row is updated first, in the same transaction that records the capture, and the transfer is attempted after that commit. If the transfer attempt fails, the order is still paid and the row still says the vendor is owed.
Pay once, across two workers
Here is the problem the redesign created. Two things are allowed to attempt the transfer: the webhook handler, inline, and a reconciliation job that sweeps for rows still owed. Both are necessary — the inline attempt is fast, the sweep catches everything the inline attempt misses — and both must never pay the same order.
The property “this order’s vendor has been paid at most once” cannot live in the code, because there are two code paths and they run in different processes. It has to live in the row:
-- Illustrative. Claim the payout. Exactly one caller sees a row come back.
UPDATE orders
SET payout_status = 'processing',
payout_claimed_at = now(),
payout_attempts = payout_attempts + 1
WHERE id = $1
AND payout_status IN ('pending', 'failed')
AND payout_attempts < 5
RETURNING id, vendor_account_id, vendor_net_amount;
This is a compare-and-swap on payout_status. If the webhook handler and the sweep both run
it for the same order, one of them updates the row and gets it back; the other matches zero
rows and does nothing. There is no read-then-write window, no lock to hold across the
network call, and no coordination between the two workers beyond the database’s own row
versioning. The outbox post uses FOR UPDATE SKIP LOCKED for the same guarantee inside a queue; a CAS is the right shape when the two
contenders are not both queue consumers.
The attempt cap matters as much as the claim. A transfer can fail for reasons that will not
improve on retry — a vendor’s linked account suspended, a currency mismatch, a malformed
account id — and a sweep that retries every five minutes forever produces a very expensive
log. After five attempts the row is marked parked and goes to a human. Parked rows are the
number I would put on a dashboard first.
Ask the provider before re-firing a stale claim
The claim has one hole. A worker claims the row, calls the gateway, and dies — process
killed, network partition, deploy in the wrong second — before writing the result back. The
row says processing forever. Was the transfer made? The database cannot tell you; only the
gateway knows.
The naive fix is a stale-claim rule: a processing row older than ten minutes belongs to a
dead worker, so reset it and retry. That is also the double-payout bug. If the dead worker’s
transfer did go through, the retry pays the vendor twice, and the platform’s balance is the
thing that shrinks.
So the rule is: before retrying a stale claim, ask the provider what it already did.
// Illustrative. Recover a stale claim without paying twice.
async function recoverStaleClaim(order) {
// The transfer was created with notes: { transaction_id } — that is the
// recovery key. Never drop it; it is the only thing that lets us find the
// transfer again without our own record of it.
const window = { from: order.payout_claimed_at - 5 * MIN, to: order.payout_claimed_at + 15 * MIN };
const existing = await gateway.transfers.list(window);
const match = existing.find(t => t.notes?.transaction_id === order.transaction_id);
if (match) {
// The dead worker's transfer happened. Record it; do not pay again.
return markPaid(order.id, { transfer_id: match.id, recovered: true });
}
// Nothing on the provider's side. The claim can be released and retried.
return releaseClaim(order.id);
}
The transfer is created with the platform’s own transaction id in the provider’s free-form
notes field. That is the recovery key, and it is the one line in the create call that must
never be refactored away, because it is the only link between a transfer the provider knows
about and an order the platform is unsure about. Listing transfers by time window and
matching on the note is slower than a lookup by id — but there is no id to look up, because
the worker died before it could save one.
This is the ask half of reconciliation rather than sync: two systems that should agree, and a question asked of the authoritative one before acting on a guess. The difference from a scheduled sweep is that here the question is asked at the moment of retry, per row, because the cost of guessing wrong is a duplicate payment rather than a stale status.
The sweep that catches what the inline call misses
The inline attempt from the webhook can miss for reasons that have nothing to do with the
order: the payout feature was flagged off during a rollout; the gateway was down; the
platform’s balance was short at that moment; the process died mid-call. Each of those leaves
a paid order with payout_status = 'pending' or 'failed', and each is invisible unless
something looks.
The sweep runs every five minutes, looks back 72 hours, takes 100 rows at a time, and runs the same claim-then-transfer path as the webhook. Because the claim is a CAS, running the sweep on several instances at once cannot double-pay; the worst case is that instances contend for the same rows and most of them match zero. The 72-hour lookback is deliberately longer than any gap I expect an outage or a flag rollout to leave, and short enough that the query stays on the recent slice of the index.
Two guards in the sweep are worth stating because they are easy to forget:
- A row is only settled if its order exists. A captured payment whose order failed to create is refunded to the customer, not paid out to the vendor. The sweep joins to the order and skips anything orphaned; refunds are a different job’s problem.
- The offline ledger stays populated. Every order still writes its offline settlement amount, and the ledger view excludes rows that direct transfer has settled or is settling. If direct transfer had to be switched off tomorrow, the offline process would pick up where it left off with no backfill.
One setting per employer-vendor pair, not one per platform
The last piece was not a technical requirement but a commercial one. A large client wanted to settle its vendors directly — they had their own agreements and their own finance team, and did not want the platform in the middle of that money. So the payout mode is a property of the (employer, vendor) pair, with a check constraint and a default:
-- Illustrative. Three modes; the default is the automated one.
ALTER TABLE employer_vendor
ADD COLUMN settlement_mode text NOT NULL DEFAULT 'platform_transfer'
CHECK (settlement_mode IN ('platform_transfer', 'corporate_direct', 'manual_offline'));
platform_transfer is everything above. corporate_direct means the employer pays the
vendor itself and the platform records the obligation but never moves money.
manual_offline is the original ledger process, for vendors without a linked account. The
mode is read at the moment the transfer would fire, so switching a pair from one mode to
another takes effect on the next order and never re-opens settled ones.
What actually broke in production
The 81-paise rejection was found in production, by the people it affected, not by a test. Nothing in the suite exercised a customer share small enough to hit the gateway’s minimum, because the boundary that mattered was the gateway’s, not the platform’s, and no test knew the gateway had one.
The empty error message meant the first report said “orders fail, no error” and the diagnosis started from nothing. The fix to the logging went in the same change as the fix to the money.
What I cannot tell you is how often the pay-once guard has actually been exercised. Both contenders are live — the inline path on the webhook and the five-minute sweep — and no vendor has been reported paid twice, but the zero-row branch of the claim was never instrumented, so I have no count of how often the two have reached for the same row. “No complaints” is weaker evidence than a counter, and I know it.
FAQ
What is the minimum transfer amount for Razorpay Route? A transfer cannot be smaller than ₹1, and a transfer linked to a payment cannot exceed that payment’s amount. Both limits are irrelevant for ordinary marketplace orders and become fatal when a subsidy or discount reduces the customer’s payment to a small amount whose vendor share falls below ₹1.
What is the difference between a Razorpay Route split and a direct transfer? A Route split divides a specific customer payment between linked accounts and is capped at that payment. A direct transfer moves money from the platform’s own account balance to a linked account, untied to any payment. Direct transfers are the right tool when the platform already holds the money it owes the vendor — for example, when an employer pre-funded the subsidised part of an order.
How do you prevent paying a vendor twice when a webhook and a cron can both pay?
Make the payout a status transition on the order row and claim it with a compare-and-swap:
UPDATE … SET status = 'processing' WHERE status IN ('pending', 'failed') RETURNING ….
Exactly one caller gets the row back. Cap the attempts so a permanently failing row is
parked for a human instead of retried forever.
How do you recover a payout when the worker crashed mid-transfer?
Put your own transaction id in the transfer’s notes when you create it. When a claim goes
stale, list the provider’s transfers in a window around the claim time and match on that
note. If a match exists, record it and do not pay again; only if nothing is found is it safe
to release the claim and retry.
When should a vendor payout be triggered? Only after the customer’s payment is confirmed captured — from the payment webhook or after a successful wallet debit — never from the checkout request itself. A transfer for a payment that later fails is money the platform cannot get back from the customer.
Why exclude cash orders from automated vendor transfers? Because the vendor already holds the customer’s cash. What the platform owes is the subsidy minus its commission, which can be tiny or negative. Expressing that as a one-way transfer means sometimes sending nothing and sometimes needing money back; an offline settlement ledger that nets both directions handles it better.
Why do Razorpay SDK errors log as empty strings?
In the SDK version we used, rejected promises carry { statusCode, error: { code, description, field } } and no top-level message. Logging err.message prints nothing.
Log the whole error object, or err.error?.description, so the first failure is diagnosable.
What I’d still improve
- Instrument the lost-CAS branch. A counter for “claim matched zero rows” is the evidence that the pay-once guarantee is doing work, and it costs one line.
- Alert on parked rows. Five failed attempts is a vendor not getting paid; it should page, not wait for the weekly finance review.
- Add a boundary fixture — a 99% subsidy on a small order — to the payment test suite, so the gateway’s minimum is a test the platform owns rather than a fact it learned from a customer.
- Use an idempotency key on transfer creation if the provider supports one for that endpoint. The recovery-by-notes path works, but a key would make the crash case a no-op instead of a lookup.
The one idea to take away
Once someone else has pre-funded the order, the vendor payout is not a slice of the customer’s payment. It is money you already hold, and you should move it as such.
That reframing dissolves the minimum-amount bug, and it changes what “pay once” has to mean. A split of a payment is idempotent by construction — the gateway ties it to one payment. A transfer from your own balance is idempotent only if you make it so: a claim only one worker can win, a cap on attempts, a recovery key in the transfer itself, and the discipline to ask the provider what it did before you do it again.
I write about backend reliability, payments correctness and Postgres data modelling from work on food-tech and healthcare platforms. More on what I build and how I work, the money-in half of this system, and why the webhook, not the redirect, decides that an order is paid. If your payout path can be reached from two places, go and check which one wins.