Direct Answer

A reliable webhook idempotency design accepts repeated deliveries without applying the same business action more than once. This matters because webhook providers commonly describe delivery as “at least once,” not “exactly once”: a timeout can occur after your server commits a change but before it returns a successful HTTP response, causing the sender to retry. A retry is therefore not evidence that the first attempt failed; it may simply be an indistinguishable duplicate. The correct unit of protection is usually the provider’s event identifier, such as Stripe’s event.id, combined with protection against accidental reuse of that identifier for different payloads.

Also worth reading: How Should B2B Teams Design an Issue Operations Workflow in 2026? · How Do You Choose B2B Case Management Software for Complex Support, Compliance, and Public-Affairs Operations? · How Should Teams Manage Case Access Governance Without Slowing Down Case Operations?

For a B2B issue-operations platform, idempotency must cover more than inserting one database row. Creating a case, assigning it to a queue, sending an email, updating a customer record, starting a compliance review, and recording an audit entry can all have externally visible effects. A duplicate webhook can produce duplicate cases, duplicate notifications, conflicting assignments, or audit records that make automation appear unsafe. Design the endpoint so that each logical event has a durable processing record, the intended state change occurs within a controlled transaction, and every side effect has a stable deduplication key.

How Idempotency Works

The receiver extracts a provider event ID and the relevant account or tenant ID before doing business work. It stores those values with a processing status, payload hash, event type, and receipt time. If the same ID is received again while the original is queued, running, complete, or temporarily failed, the endpoint returns the recorded result rather than executing the workflow again. A unique database constraint on the provider ID—or on a compound key such as provider plus account plus event ID—closes race conditions that an application-level “check, then insert” cannot reliably prevent by itself.

The system should distinguish several states: newly accepted, successfully processed, safely ignorable, retryable failure, and rejected after investigation. “Processed” should mean that the authoritative state change has committed, not merely that a worker placed the event on a queue. If work is queued asynchronously, the intake transaction must durably record both the event and the queue job before returning success. This is commonly described as the transactional outbox pattern, and it prevents the opposite failure: accepting a webhook successfully but losing the job before a worker can perform it.

A useful implementation sequence takes approximately 10–20 milliseconds for local database deduplication, although real-world latency depends on network distance, database load, and verification requirements. The receiver checks the provider signature, inserts the idempotency record with a unique constraint, validates tenant context, persists the domain change and outbox job in one transaction, and returns a 2xx response. Retries then receive the stored response even if they arrive seconds or days later. Keep the receipt record longer than the provider’s maximum retry period; indefinite retention is safest for audit-heavy systems, while teams operating under storage constraints can often retain event IDs for at least 90 days and preserve a separate audit record longer.

Delivery, Queues, and Exactly-Once Claims

Idempotency does not make the network exactly once. The sender and receiver remain separate systems, and an HTTP response can be lost after processing. What idempotency provides is exactly-once business effect under at-least-once delivery, provided every consumer and external destination is also designed to tolerate retries. If the workflow sends an email through a third party, record an operation key such as the event ID plus recipient and notification type before sending. If the provider supports its own idempotency key, pass the same key to that provider. If it does not, accept that an ambiguous timeout may require reconciliation rather than pretending duplicate suppression is guaranteed.

Workers need leases and bounded retries. A visibility timeout of roughly 60 seconds is common, but it should be shorter than the worker’s expected processing time or be accompanied by heartbeat-based lease renewal. A job can be retried after exponential delays such as 1 second, 5 seconds, 30 seconds, 2 minutes, and 10 minutes, while permanent errors such as malformed JSON or an invalid account mapping should go to a review state. After 5–10 failed attempts, alert an operator and retain the payload for manual inspection rather than retrying forever. The acceptance threshold should not simply be “we returned 200”: for regulated workflows, it should be “the durable domain change and all required follow-up jobs are committed.”

Ordering is a separate concern. An older event may arrive after a newer one, so dedupe does not establish sequence. Compare provider creation timestamps or monotonic provider sequence numbers within the same object and reject or quarantine events that are stale. Do not compare timestamps across unrelated objects, and do not assume timestamps alone prove ordering because clocks and serialization can introduce ambiguity. For case automation, the authoritative state machine should decide whether an older transition is still valid; for example, a “case closed” event received after “case reopened” cannot simply overwrite the current case.

Practical Database Design

A minimal idempotency table normally contains an event ID, integration ID, tenant ID, event type, payload hash, status, attempt count, first and last received timestamps, and the eventual result. Add a unique constraint such as (integration_id, external_event_id). If providers can reuse IDs across sandbox and production environments, include the environment or account scope in that key. Storing a SHA-256 payload hash helps detect the unusual but serious case where the same event ID appears with a changed body; it should trigger investigation instead of silently treating the second request as identical.

The domain transaction should atomically update the case and write an outbox entry. For a new issue received from a partner, that transaction might insert the case, initialize its status, create the first audit event, and insert an email job with a unique operation key. Once committed, a separate worker sends the email and marks the job complete. This arrangement avoids the dual-write problem between your database and a queue broker. It also makes the webhook response fast: p95 intake latency below 500 ms is a reasonable target for many B2B APIs, while expensive validation, enrichment, or third-party calls should occur asynchronously.

Do not hold a database transaction open while calling Stripe, an identity provider, a messaging service, or another external API. A 30-second network timeout can consume a connection, exhaust a pool, and cause more failures than it prevents. Commit durable local state first, then perform external work through retryable operations with their own keys. For operations that cannot be rolled back, model “intent to perform,” “started,” and “confirmed” states so an operator can reconcile an ambiguous result. In compliance workflows, the audit trail should retain the original event ID, payload digest, actor, processing result, and reason for any suppression.

Comparison of Implementation Approaches

The right approach depends on whether the endpoint performs the work synchronously, whether the provider supports idempotency keys, and how much audit evidence the business needs. A database receipt table is the baseline for nearly all production webhooks; an outbox improves reliability but introduces worker and operational responsibilities.

FeatureDatabase receipt tableDatabase receipt plus transactional outbox
Duplicate suppressionStrong for the same provider event IDStrong for intake and every generated job
External side effectsRequires separate side-effect keysJobs use stable operation keys and retry independently
Response timeFast when work is localFast because queued work follows acceptance
Failure recoveryManual for failed downstream actionsAutomatic through durable worker retries
ComplexityLow to moderateModerate, requiring outbox cleanup and monitoring
Best fitSimple status updates or bounded local writesCases, billing, notifications, and audit-sensitive workflows
A queue-only design is tempting because it makes the endpoint appear thin: verify, publish, return 200. It is still safe only when deduplication and durable recording happen before acknowledgment, and when publication cannot silently fail. Publishing to a queue without a database receipt can duplicate work on every redelivery; publishing directly to an external email or ticketing API creates the same risk plus a database-to-API consistency gap. A transactional outbox addresses this by recording the job atomically with the local state, but it is not a substitute for idempotent workers.

For very low-risk internal events, an in-memory cache may be adequate for a single process, but it fails across instances, deployments, restarts, and cache eviction. Redis can provide fast atomic deduplication, often with sub-millisecond operations under ordinary load, but a volatile cache should not be the only record if replay prevention and audit evidence matter. Use Redis with persistence and replication when it is the coordination layer, or combine it with a durable database key. Managed infrastructure can simplify operations, but it does not eliminate the need to define ownership, retention, and reconciliation.

Security and Multi-Tenant Boundaries

Idempotency and authentication must be designed together. Verify the webhook signature before parsing or trusting business fields, and use the provider’s raw request bytes when signature verification depends on exact serialization. For example, Stripe signs the raw payload and provides signing and timestamp tolerance guidance; tolerance is measured in seconds and should be enforced by the SDK rather than recreated informally. Reject unsigned or incorrectly signed traffic before it can reserve an event ID, because otherwise an attacker could poison the record and block a genuine delivery.

In a multi-tenant case system, derive tenant context from verified provider metadata or an authenticated integration mapping, not from an untrusted tenant_id in the payload. A global event-ID key may be safe for a provider that issues globally unique IDs, but an account-scoped key is necessary when uniqueness is only guaranteed within an account. When one integration can send events for hundreds of customer organizations, store the external integration account ID alongside the internal organization ID and enforce that relationship. A hash mismatch, signature failure, or tenant mismatch should result in rejection or quarantine with a reason code, not a false 200 that makes the sender stop retrying.

Replay protection must account for legitimate provider retries. A timestamp-only rule that rejects every request outside a five-minute window can reject valid retries after an outage, while accepting every request indefinitely makes captured payloads useful for replay attacks. Verify timestamps, retain deduplication keys for the provider’s retry horizon, and require valid signatures. Stripe’s standard guidance commonly uses a five-minute tolerance, while systems with long provider retry schedules may retain receipt keys for 30–90 days or longer. These are design thresholds, not universal guarantees; provider documentation and your own incident history should determine the final retention period.

Common Mistakes and Operational Thresholds

The most common mistake is checking whether an event exists and then inserting it without a database uniqueness constraint. Two concurrent deliveries can both pass the check. The second common mistake is returning success before durable processing has occurred, which loses work; conversely, returning a 500 after the domain transaction commits creates duplicates on retry. Record the result and return it consistently, including for accepted duplicates, while using 4xx only for permanent request errors and 5xx only when a retry could succeed.

Another mistake is assuming HTTP 200 means the entire workflow succeeded. Monitor separate metrics for received, verified, accepted, duplicates, queued, processed, retried, permanently failed, and quarantined events. A duplicate rate is not inherently bad: after an outage, 15–30% duplicate deliveries can be reasonable if the original successes were committed. A sudden 100% duplicate rate, an event count that rises without corresponding cases, or a p95 intake latency above 1 second should prompt investigation. Set alerts around error budgets rather than noisy universal rules, such as more than 1% permanently failed events over 15 minutes or a backlog older than 5 minutes for a fast workflow.

Do not make business actions depend on a display name or free-text subject. Use immutable provider IDs and explicit operation keys. Do not retry permanent 400-class validation failures indefinitely, and do not expose provider payloads, personal data, or signatures in general-purpose logs. Redact sensitive fields and use correlation IDs that operators can follow across intake, outbox, worker, and downstream provider logs. Finally, test duplicates, reordered deliveries, worker crashes after commit, lost responses, expired signatures, cross-tenant collisions, and payload-hash mismatches before launch.

When to Act, and What It May Cost

Act before enabling a webhook that creates cases, sends external messages, changes compliance state, or triggers billing-related operations. For a low-risk status notification, a narrow idempotency record may take one afternoon; a multi-tenant case workflow with outbox workers, dashboards, replay tooling, and retention policies may take several engineering days to two weeks. The cost is mainly engineering and observability rather than a separate license. A managed database can keep the incremental spend low, while high-volume ingestion may require partitioning, read replicas, or dedicated queue workers.

Estimate capacity from peak events rather than average events. If a partner sends 1,000 events per minute and each receipt row is roughly 1–3 KB including indexes and metadata, raw storage can reach about 1.4–4.3 GB per day before replication and retention overhead. A 30-day retention policy therefore needs tens of gigabytes to more than 100 GB depending on row size, compression, and database overhead. Partitioning by month and archiving payloads while retaining compact event IDs can control growth. Outbox rows should also have a defined deletion policy after confirmed completion, subject to audit requirements.

The best rollout is staged. First, log verified events and establish baseline volume and failure rates. Next, add a durable unique key, then move work behind an outbox, and finally enable automatic downstream retries with alerts. Maintain a replay command that requires authorization, displays the affected tenant and event scope, and preserves the original receipt. Idempotency is not a feature you “turn on” once; it is a contract among the sender, receiver, database, workers, and external providers. In B2B issue operations, that contract is what lets a case house automate high-volume intake without making duplicate incidents, contradictory compliance history, or repeated customer notifications ordinary operating risks.