Direct answer: treat every webhook as an at-least-once command
A reliable webhook idempotency design assumes that a provider may send the same event more than once, deliver events out of order, retry after a transient failure, and occasionally use slightly inconsistent timestamps or identifiers. Your system should therefore make repeated processing safe rather than trying to prevent every duplicate at the network edge. The standard approach is to assign each logical event a stable provider event identifier, store that identifier in a database table with a unique constraint, and commit the event record together with the business-state change in one transaction. If the identifier already exists, acknowledge the duplicate without repeating the side effect. This model is particularly important for issue-ops platforms because one payment, identity-check, case, or compliance event can trigger several operational actions, including case updates, customer notices, audit records, and downstream accounting entries. The exact language differs across providers, but the reliability principle is stable: transport delivery is probabilistic, while business processing must be deterministic. For an issue house managing support, compliance, and public-affairs work, the durable event record should also be retained beyond the duplicate-detection window so teams can reconstruct why a case changed.
Also worth reading: What Enterprise Webhook Idempotency Patterns Actually Prevent Duplicate Case Processing in 2026? · How Should Modern B2B Teams Architect Case Access Control Design for Secure Operations? · How Do You Build Reliable Webhook Replay Protection Without Breaking Retries?
The transaction boundary that prevents duplicate work
The safest design places the idempotency key, the processed-event metadata, and the local state transition inside the same database transaction. Suppose a payment provider sends evt_123; the receiving service inserts that key into an webhook_events table and marks the relevant invoice as paid before committing. If payment succeeds but the response to the provider is lost, the provider retries evt_123; the next attempt violates the unique key, so the service recognizes the event and returns a successful acknowledgment. This prevents the invoice from being marked paid twice, a downstream account from being credited twice, or two customer communications from being created. The key should identify the logical event, not merely the request payload, because a provider may resend identical content with a new HTTP delivery identifier. A strong implementation also records the provider, event type, first-seen time, last-seen time, attempt count, payload hash, processing status, and application version. Atomicity matters more than an application-level “check, then insert” performed without a database constraint, because two workers can pass the check simultaneously. A unique index or equivalent database mechanism turns the race into a controlled conflict instead of duplicate business processing.
Choosing keys, fingerprints, and useful retention periods
Use the sender's immutable event ID whenever it is documented and globally unique within that provider. Prefix stored keys with the provider or source account, such as stripe:evt_123, so two providers cannot collide. Do not use a timestamp, arrival time, or current timestamp as the primary key because retries can contain altered timestamp fields and concurrent requests can share the same second. A payload hash can help detect conflicting reuse of a key, but it should not replace the event ID: retries can differ in transport metadata, and semantic equivalents may be serialized differently. If the provider offers no stable ID, generate one only after identifying a stable business tuple, and test that design against real retry behavior rather than assuming the tuple is unique forever. A practical starting retention period is at least the provider's maximum retry window plus a safety margin, often 72 hours to 30 days; high-value case or compliance workflows may warrant 1-5 years for audit evidence, with the operational duplicate ledger separated from the longer-lived audit archive. Retention is not free, so teams should measure event volume and choose a period based on retry obligations, contractual rules, and investigation needs.
| Design choice | Simple key lookup | Database-backed idempotency | Distributed processing platform |
|---|---|---|---|
| Duplicate control | In-memory cache | Unique event table and transaction | Broker offsets plus durable store |
| Survives restart | Usually no | Yes | Yes, if configured persistently |
| Handles provider retries | Partly | Yes | Yes |
| Suitable for billing or case state | Rarely | Recommended | Possible for high-volume pipelines |
| Main weakness | Process-local state | Requires schema and transaction discipline | More operational complexity |
| Typical cost | Low engineering effort | Database storage and modest compute | Broker, workers, monitoring, and on-call cost |
Acknowledgment, retries, ordering, and poison events
The endpoint should validate the request quickly, enqueue or durably record it, and return a success status only when the system has accepted responsibility for processing. Returning 2xx after validation but before durable persistence creates a misleading success signal: a crash can lose the event while the sender believes it was handled. If the event is written to a durable queue within the request, a 2xx response is often reasonable; otherwise, return success only after the database transaction commits. Use 4xx responses for malformed or permanently invalid payloads, 5xx responses for temporary database or dependency failures, and configure exponential backoff with jitter according to the provider's rules. A retry schedule such as 1 second, 5 seconds, 30 seconds, 2 minutes, 10 minutes, and 1 hour is illustrative, not universal; the provider's contract is authoritative. Do not assume that processing order is guaranteed even when delivery appears ordered, because separate event types can move through queues and workers differently. A refund can arrive before the original charge is applied, so the consumer should fetch current state, tolerate version conflicts, and reconcile late events rather than simply discarding them.
A poison event should enter a dead-letter or quarantine path after a bounded number of attempts, not loop forever. Teams need an operational view showing the event ID, source, event type, failure category, first and last attempts, and the case or account affected. Sensitive payloads such as identity documents, payment data, or public-affairs correspondence should be redacted or tokenized in logs according to privacy obligations. A 2026 architecture should also account for changing event schemas: validate required fields, preserve the original payload in encrypted storage when policy permits, and version parsers so a deployment does not make previously accepted events unreadable. Monitoring should distinguish “provider retried” from “our worker failed,” because the two require different incident responses. The provider's retry count does not tell you whether an event was successfully committed, while your internal attempt count does not tell you whether the sender's deadline is about to expire.
Where B2B issue operations adds special requirements
For ordinary consumer notifications, duplicate emails are inconvenient; for support, compliance, and public-affairs operations, duplicates can corrupt evidence, create inconsistent case timelines, or trigger repeated escalation. A single identity-verification event might update a case status, notify a reviewer, start a screening task, and record an audit entry. The idempotent operation must cover all of those effects, either through one transaction or through a durable workflow with explicit step-level keys. A case-management system should also distinguish “the event was received” from “the business action was completed.” If the event record is committed before all downstream work succeeds, a later reconciliation job can continue safely; if all effects must be atomic, the domain boundary may need to be redesigned. This is why an issue house should model idempotency around the business aggregate, such as a case, application, invoice, or account, rather than around a generic HTTP endpoint.
A useful threshold is to alert when the same source event has more than one delivery within 24 hours, when the duplicate rate rises above a negotiated baseline, or when any duplicate appears to accompany a financial, identity, or regulatory state change. Baselines vary too much to prescribe one universal percentage; many mature systems operate below 1% unexpected duplicates, while a newly integrated source can temporarily exceed 5% during retries or schema problems. Teams should measure their own normal rate over at least 14 days and alert on a sustained deviation, such as three times the rolling 30-day median. Every case mutation should record the event ID that caused it, allowing an investigator to answer whether a second update was a true upstream correction or an accidental replay. That audit link is more valuable than a generic “processed successfully” log because it connects external evidence to the internal case history.
Practical implementation sequence
Begin by writing down the provider's delivery guarantees, retry window, signature format, event ID rules, and schema version policy. Add a receiver that verifies the signature against the raw request body before parsing, because parsing and then verifying a reconstructed payload can introduce inconsistencies. Persist the verified event in a transaction that includes the unique key and any local state update, then return the appropriate HTTP response. Build a second path for events that are valid but not yet processable because a related invoice, case, or account is temporarily unavailable; this path should preserve the event and its original business key instead of discarding it. Add background reconciliation for late events and a controlled replay tool that requires an event ID, expected current version, and operator reason. Replay should not bypass the same idempotency checks; a manual replay is another command that can fail between its external call and its local acknowledgment.
Test concurrency, not just sequential retries. Send the same event simultaneously from 2, 10, and 100 workers, then verify that exactly one business transition occurs and every request receives a safe response. Test failure injection after the first external call, before the database commit, after the commit but before the HTTP response, and during dead-letter movement. Include cases where the provider sends an old version after a newer one, sends the same event with different transport headers, or sends a newly added field. Track metrics for receipt latency, processing latency, retry volume, duplicate suppression, dead letters, reconciliation backlog, and signature failures. A practical operational target is to process 95% of valid events within 60 seconds and 99% within 5 minutes for a typical transactional workflow, but regulated or high-risk processes may require stronger service objectives. The target should be measured from durable acceptance, not from the sender's timestamp, which may be delayed or clock-skewed.
Common mistakes and costly alternatives
The most common mistake is treating a successful HTTP response as proof that every downstream action happened. The second is using a cache with a short TTL and assuming that it can coordinate multiple application instances. The third is checking for an existing record without a unique constraint, which leaves a race between concurrent requests. Another mistake is making the event ID the only record, because a duplicate key cannot explain what was processed or whether an audit entry was lost. Teams also err by acknowledging invalid signatures, logging entire identity or payment payloads, retrying permanent validation failures indefinitely, and exposing a replay button without authorization. A high-volume queue can improve intake throughput, but it can make ordering and debugging harder if events are partitioned incorrectly. A fully synchronous design is easier to understand for low volume, yet it can hold open connections and cause sender timeouts during database degradation.
Idempotency is not a substitute for authentication, authorization, validation, or reconciliation. A replayed event can be authentic yet stale, so confirm that its business version and source account match the current record. A request can be valid but unauthorized for a particular tenant, so verify tenant boundaries before storing or processing it. Rate limits should be applied by source and tenant without rejecting legitimate retries; a duplicate should normally be acknowledged quickly after validation rather than challenged with a CAPTCHA or an interactive approval. If a provider offers its own event replay API, use it for recovery, but still keep local duplicate protection because the provider's behavior and your internal retries can overlap. The cost should be modeled as storage, database indexes, worker capacity, monitoring, engineering time, and incident recovery rather than as a single line item. At low volume, a managed relational database may cost tens of dollars monthly; at high volume, the same design can require dedicated partitions, a queue, a data warehouse, and on-call staffing, but no exact price is defensible without provider, volume, retention, and compliance assumptions.
When to act and how far to go
Act before production launch if webhooks can modify invoices, identity status, case priority, compliance deadlines, or externally visible customer communications. For an internal prototype that only displays a non-destructive status badge, a lightweight database record may be enough, provided the limitation is explicit and the prototype cannot become a system of record accidentally. Act immediately when an integration has already produced duplicate notifications, inconsistent case histories, double credits, unexplained record changes, or a support ticket that depends on knowing which event caused an update. During vendor selection, ask whether event IDs are stable, what the maximum retry period is, whether replay is supported, how schema changes are announced, and whether delivery can be delayed. For case operations, also ask whether the integration can preserve a tenant-specific event identifier and whether the vendor will sign payloads in a way compatible with your rotation schedule.
There is no need to build a globally distributed exactly-once system for every integration. “Effectively once” behavior, achieved through durable keys and transactional business updates, is usually more practical and easier to operate. Expand the design as volume, risk, or tenant count increases: first add a database table and unique constraint, then add a queue when intake latency becomes a problem, then add partitioning or a dedicated event platform when measured throughput requires it. Keep a kill switch for a new event type, a feature flag for a changed consumer, and a documented rollback that does not erase event history. Revisit the retention and access policy at least twice a year because privacy obligations, contractual requirements, and provider retry behavior can change. As of 30 September 2026, teams should treat webhook idempotency as an operational control with an owner, measurable service levels, and an audit trail—not as a one-time coding trick.