webhook-gateway — durable webhook ingestion and delivery
A gateway that verifies, deduplicates and stores inbound webhooks, then delivers them with retries, circuit breaking and a dead-letter queue — with PostgreSQL as the queue.
154 tests, every one against a real PostgreSQL — and a benchmark whose least flattering row is the one that gets explained.
The problem
A third-party webhook arrives once. If it is lost between the socket and the
database, nobody can ask Stripe or GitHub to send it again, and the failure is
silent — the producer got its 200. Meanwhile the subscriber on the other side
is a machine you do not control: it times out, it returns a 500, it goes down
for an hour, it returns 400 because the contract changed. Each of those needs
a different answer, and "retry everything with exponential backoff" is the wrong
answer to most of them.
The guarantee is stated honestly, because the honest version is the useful one:
exactly once at the persistence boundary, at least once to the subscriber.
The first half is enforced by a unique constraint and proved by a concurrency
test. The second half cannot be enforced by anybody — the last step is an HTTP
request to someone else's machine, and a worker killed between their 200 and
our commit must retry rather than guess. So every outbound request carries a
stable Idempotency-Key and an attempt number, which is exactly what a
subscriber needs to close the gap on their side.
PostgreSQL is the queue, so the queue and the data it refers to are in one
transaction. Fan-out cannot commit an event whose work item was lost, a claim
cannot survive a rolled-back attempt, and there is no reconciliation job between
two stores that disagree. A broker would buy cross-language fan-out and cost
exactly the property this service exists to have. FOR UPDATE SKIP LOCKED lets
N workers claim disjoint batches with no coordinator, and a lease means a worker
killed mid-flight has its rows swept back rather than stranded.
Every failure mode has a considered answer rather than a generic retry. A
duplicate is 200 with duplicate: true, not a 409 that makes the producer
retry harder and page someone. A 400 from a subscriber is dead-lettered
immediately, because eight identical rejections teach nobody anything. A 429
with Retry-After is honoured, because a subscriber saying when to come back
knows better than our curve. Backoff is full jitter rather than "exponential
plus noise", which re-synchronises the whole herd onto one instant and re-kills
the endpoint that just recovered. A circuit opens after N failures and
reschedules without spending an attempt, so a 30-second outage does not
exhaust an eight-attempt budget on requests that were never sent — and exactly
one probe is admitted on half-open, or the backlog stampedes the endpoint the
moment it comes back.
Outbound delivery is where a gateway is a confused deputy by construction, so
every resolved address is validated and then the connection is pinned to the
address that passed, with Host and TLS SNI preserved: validating a DNS
answer and then letting the client resolve again is a TOCTOU window, and
169.254.169.254 hands out cloud credentials.
The centrepiece test fires 24 genuinely concurrent identical webhooks — separate sessions, separate connections, separate backends racing on one constraint — and asserts one event, one winner and one delivery. Every test runs against a real PostgreSQL, because the guarantees are PostgreSQL's and a suite passing against SQLite would be testing a fiction. The load harness found two real bugs, including a fail-open path that was paying a Redis connect timeout on every request; the benchmark table reports the machine and the unflattering rows along with the rest.