webhook-gateway — durable webhook ingestion and delivery

A gateway that verifies, deduplicates and stores inbound webhooks, then delivers them with retries, circuit breaking and a dead-letter queue — with PostgreSQL as the queue.

Outcome

154 tests, every one against a real PostgreSQL — and a benchmark whose least flattering row is the one that gets explained.

The problem

A third-party webhook arrives once. If it is lost between the socket and the database, nobody can ask Stripe or GitHub to send it again, and the failure is silent — the producer got its 200. Meanwhile the subscriber on the other side is a machine you do not control: it times out, it returns a 500, it goes down for an hour, it returns 400 because the contract changed. Each of those needs a different answer, and "retry everything with exponential backoff" is the wrong answer to most of them.

The guarantee is stated honestly, because the honest version is the useful one: exactly once at the persistence boundary, at least once to the subscriber. The first half is enforced by a unique constraint and proved by a concurrency test. The second half cannot be enforced by anybody — the last step is an HTTP request to someone else's machine, and a worker killed between their 200 and our commit must retry rather than guess. So every outbound request carries a stable Idempotency-Key and an attempt number, which is exactly what a subscriber needs to close the gap on their side.

PostgreSQL is the queue, so the queue and the data it refers to are in one transaction. Fan-out cannot commit an event whose work item was lost, a claim cannot survive a rolled-back attempt, and there is no reconciliation job between two stores that disagree. A broker would buy cross-language fan-out and cost exactly the property this service exists to have. FOR UPDATE SKIP LOCKED lets N workers claim disjoint batches with no coordinator, and a lease means a worker killed mid-flight has its rows swept back rather than stranded.

Every failure mode has a considered answer rather than a generic retry. A duplicate is 200 with duplicate: true, not a 409 that makes the producer retry harder and page someone. A 400 from a subscriber is dead-lettered immediately, because eight identical rejections teach nobody anything. A 429 with Retry-After is honoured, because a subscriber saying when to come back knows better than our curve. Backoff is full jitter rather than "exponential plus noise", which re-synchronises the whole herd onto one instant and re-kills the endpoint that just recovered. A circuit opens after N failures and reschedules without spending an attempt, so a 30-second outage does not exhaust an eight-attempt budget on requests that were never sent — and exactly one probe is admitted on half-open, or the backlog stampedes the endpoint the moment it comes back.

Outbound delivery is where a gateway is a confused deputy by construction, so every resolved address is validated and then the connection is pinned to the address that passed, with Host and TLS SNI preserved: validating a DNS answer and then letting the client resolve again is a TOCTOU window, and 169.254.169.254 hands out cloud credentials.

The centrepiece test fires 24 genuinely concurrent identical webhooks — separate sessions, separate connections, separate backends racing on one constraint — and asserts one event, one winner and one delivery. Every test runs against a real PostgreSQL, because the guarantees are PostgreSQL's and a suite passing against SQLite would be testing a fiction. The load harness found two real bugs, including a fail-open path that was paying a Redis connect timeout on every request; the benchmark table reports the machine and the unflattering rows along with the rest.