Handling payment webhooks in production SaaS
How we kept Shoprocket payment records correct across Stripe, PayPal, PayU and crypto when webhooks arrived late, duplicated or out of order.
Based on production work on Shoprocket . This is engineering perspective, not product documentation.
A customer can finish checkout, close the tab, and still leave your system unsure whether money actually moved. That gap between UX and settlement is where payment webhooks do their work, and where things get messy if you treat a redirect as proof of payment.
On Shoprocket we routed money through Stripe, Stripe Connect, PayPal, PayU and a few crypto flows. Wiring each provider was the easy part. Keeping local payment records honest when callbacks showed up late, duplicated, or in an order that did not match checkout took most of the effort.
Keep payments separate from orders
Orders and payments move together in the business, but they should not share one lifecycle in code. I have seen teams collapse them into a single status field and regret it the first time a capture succeeded while fulfillment was blocked, or a partial refund needed to leave the rest of the order untouched.
A provider-agnostic payment record, linked to the order, gives you room to reconcile, retry, and explain state to merchants without rewriting checkout every time a gateway behaves differently.
Verify signatures on the raw body
Signature verification has to run against the raw request body, before you parse JSON into arrays and lose the bytes the provider signed. Under pressure it is tempting to trust the Stripe dashboard because it shows paid while your database still says pending. The webhook is the system of record, not the admin UI.
Our default path was simple: verify, persist the payload, then hand work to a queue. That persistence saved us more than once when support needed to replay what PayPal or Stripe actually sent during a dispute.
Assume every webhook will arrive twice
Providers retry. Queues retry. Recovery jobs retry. Without idempotency, a duplicate
payment_intent.succeeded is not a log line, it is a finance problem.
What worked for us was recording a stable external event ID before changing local state, then exiting early if that ID was already processed. Outbound idempotency keys matter too, but inbound deduplication is what stops a retry storm from double-charging or double-fulfilling.
Translate provider events on purpose
Stripe, PayPal and regional gateways use different names for similar ideas. Internal states like authorized, captured, failed, refunded, disputed and cancelled need an explicit mapping table, not tribal knowledge in one senior engineer's head.
Assumptions bite quickly. Some providers emit success before funds settle. Others emit three events for what feels like one refund to the merchant. The mapping work is dull. So are the weeks of support it prevents.
Schedule reconciliation, not prayer
Webhooks get dropped. Customers abandon tabs. Gateways have outages. We ran reconciliation jobs that compared provider state to local records and flagged mismatches before a merchant noticed. Perfect sync forever is unrealistic. Catching drift in hours instead of weeks is not.
Return 200 quickly, do the work in workers
Slow webhook handlers invite timeouts and more retries, which makes idempotency even more important. Acknowledge fast, process in a worker where logging, retries and alerts are easier to control.