← blog · September 14, 2026

Webhook Design: Retries, Idempotency, and Signature Verification

Webhooks look simple until production traffic exposes three hard questions: delivery guarantees, duplicate events, and sender authenticity. This post covers retry strategy, idempotency keys, and signature verification with practical examples.

How do you tell another system that something happened? There are two basic approaches: the other side asks you periodically (polling), or you tell it the moment something happens (a webhook). Polling is simple but trades latency against load: poll often and you burn resources, poll rarely and events are noticed late. Webhooks flip that trade: notification is instant, but now you own an inbound HTTP endpoint, you have to validate what arrives at it, and you have to handle delivery failures.

A webhook sounds like a one-line idea: "POST to this URL when something happens." Three questions break that simplicity in production: what delivery guarantee do you offer, what happens if the same event arrives twice, and how does the receiver know the request actually came from you. An integration built without answering these three questions will either lose events or process the same event multiple times the first time traffic spikes.

Delivery guarantee: at least once or at most once

There are two basic models. At-most-once delivery means the sender tries once and gives up on failure; it is simple, but a brief network blip or a five-second maintenance window on the receiver's side means the event is gone for good. At-least-once delivery means the sender retries until it gets a successful 2xx response, or until it hits a retry limit; events are not lost, but receiving the same event more than once becomes a normal, expected occurrence rather than an edge case.

In production, at-least-once is almost always the right default. A lost payment notification is far more expensive than a duplicate order update. But that choice is not free: it makes idempotency on the receiving side mandatory, not optional.

Retrying at a fixed interval, say every 30 seconds, creates two problems. If the receiver is already overloaded, a steady drumbeat of retries keeps it from recovering, and if many senders' retries line up on the same schedule you get a thundering herd. The standard fix is exponential backoff with jitter: wait 1 second after the first failure, then 2, then 4, doubling each time up to a cap (an hour, say), and add 10-20% random jitter to each wait so clients do not all retry in lockstep. You also need a ceiling on total retry duration. Rather than retrying forever, giving up after 24-72 hours and moving the event to a dead letter queue avoids wasting resources and makes the failure visible instead of silent.

Idempotency: what happens when the same event lands twice

At-least-once delivery necessarily produces duplicate requests. The usual fix on the receiving end has two parts: every event needs a unique identifier, and the receiver needs to be able to check whether it has already processed that identifier.

The simplest implementation keeps a table of processed event IDs: when a request comes in, check whether its ID is already in the table. If it is, return 200 and do nothing. If it is not, do the work and write the ID to the table inside the same database transaction. That last detail matters: if you run the business logic and record the ID as two separate steps, a crash between them means the next retry processes the event again. Retention matters too. There is no need to keep event IDs forever; keeping them for at least as long as the sender's longest retry window (72 hours in the example above) is enough.

Some systems ask for a client-generated "idempotency key" instead of, or alongside, an event ID; that protects against the client itself retrying the same logical operation, say, a payment request. Do not conflate the two: an event ID is generated and guaranteed unique by the sender, while an idempotency key is supplied by the client so the server can recognize a repeated request. In a webhook system, an event ID is usually enough, because you are the one generating the events.

Signature verification: how the receiver knows it is really you

A webhook endpoint is a public HTTP address. Without authentication, anyone who learns that address can send a forged "payment completed" or "user deleted" event. The fix is to attach a signature computed by the sender to every request, and have the receiver recompute it and compare.

The common approach is HMAC-SHA256: sender and receiver agree on a shared secret in advance, the sender computes the HMAC of the request body (usually along with a timestamp) and sends it in a header such as X-Signature, and the receiver runs the same computation and compares the two values. In Node.js, that check looks like this:

const crypto = require('crypto');

function verifySignature(payload, signatureHeader, secret) {
  const expected = crypto
    .createHmac('sha256', secret)
    .update(payload)
    .digest('hex');
  return crypto.timingSafeEqual(
    Buffer.from(expected),
    Buffer.from(signatureHeader)
  );
}

Two details matter here. First, compare with timingSafeEqual, not ===. A normal string comparison short-circuits character by character, so how long it takes leaks how many leading characters of the correct signature you already have, which opens the door to a timing attack. Second, sign the raw body, not a re-serialized version of it. Parsing the JSON and comparing against a re-encoded copy is a mistake, because key ordering or whitespace differences produce a completely different signature, and that usually shows up as a confusing "works for some requests, fails for others" bug.

Including a timestamp in the signed payload solves a separate problem: replay attacks. Even a correctly signed request is dangerous if an attacker can capture it once and resend it later, and the receiver has no way to tell it apart from a fresh one. Adding a timestamp to what gets signed, and rejecting requests where that timestamp is older than a few minutes, closes that gap.

Common pitfalls

Retrying forever on 4xx errors (400, 401, 422) is a common mistake. These usually indicate a problem with the request itself, a malformed payload, an invalid signature, and retrying will not fix that; it just burns resources. Retry logic should be reserved for 5xx errors and timeouts; 4xx responses should send the event straight to a dead letter queue instead.

If processing takes a while on the receiving side, a heavy computation or a third-party call, handling the webhook synchronously and returning 200 only once that work is done is a risky design. Senders typically enforce a timeout of just a few seconds; if you exceed it, the request is marked failed and resent, and you end up processing the same event from both the original and the retry. The correct pattern is to accept the request, push the event onto a queue, and return 200 immediately, doing the actual work asynchronously.

Secret rotation is another thing that gets overlooked. When you rotate the shared secret, you need a plan for requests still in flight that were signed with the old one. Usually that means accepting both the old and new secret during a transition window; otherwise a batch of events gets permanently rejected at the moment of rotation.

Finally, do not assume ordering. Network delays and retries mean events can arrive out of the order they were produced in. If order matters, say a "created" event needs to be processed before an "updated" one for the same record, attach a sequence number or timestamp to each event, and let the receiver either discard or requeue anything that arrives out of turn.

When not to reach for a webhook

If the receiving system cannot reliably run an accessible HTTP endpoint, a desktop application, a mobile client, or something sitting behind a closed corporate network, a webhook is a poor fit; polling, or a model where the client connects out and pulls from a queue, works better there. Likewise, if many consumers need to process the same event stream at different speeds, replay past events, or retain a long history, an event log like Kafka is the more appropriate tool; a webhook is inherently closer to fire-and-forget and does not give you a durable history. Webhooks earn their keep when there is a single consumer that needs near real time notification and that consumer is normally reachable. Outside those conditions, another pattern will save a lot of pain.