The Direct Answer: Use Durable, Scheduled Retries With Idempotent Consumers

A reliable webhook retry architecture should not simply resend a failed HTTP request until it succeeds. It should create a durable event record, distinguish temporary failures from permanent ones, schedule retries according to an explicit backoff policy, preserve delivery history, and require the receiving system to process each event idempotently. A practical baseline is 8 attempts over approximately 24 hours, with delays near 1 minute, 5 minutes, 30 minutes, 2 hours, 6 hours, 12 hours, and 24 hours. The exact schedule depends on the business process, but a bounded policy is easier to operate than unlimited retries.

Also worth reading: How to configure webhook retry policies in issues.house for reliable event delivery across support, compliance, and public‑affairs workflows? · How Should Enterprise Case Management Platforms Design a Secure Architecture in 2026? · How Should Enterprises Optimize Compliance Workflow Architecture Without Losing Control?

The sender should return a normal 2xx response only after it has durably accepted the event, not necessarily after downstream processing is complete. Timeouts, connection failures, HTTP 408, 429, and most 5xx responses normally justify another attempt. Authentication failures, malformed payloads, and unsupported operations usually require correction rather than blind repetition. For issue operations, the same architecture can connect case updates, compliance events, and public-affairs notifications without making those teams operate several unrelated delivery mechanisms.

As of 2 October 2026, webhook delivery remains a simple integration contract, but reliability requires more than an HTTPS endpoint. AWS recommends treating event notifications as asynchronous messages and designing consumers for retries and duplicate delivery. The central design principle is that at-least-once delivery creates duplicates; therefore, exactly-once business effects must come from idempotency, transactional state changes, and deduplication rather than an unrealistic assumption that the network never repeats a message.

Core Components and How Data Moves Through the System

A production architecture has six functional layers: event capture, durable storage, delivery workers, a scheduling mechanism, receiver controls, and operational observability. When an application changes a case or sends a notification, it first writes an event with a globally unique event ID. That event should contain a stable object version, event type, creation time, destination ID, payload, and the number of the current delivery attempt. For example, a case might move from “Awaiting evidence” to “Ready for review,” while the payload records the case ID, prior state, new state, actor, and correlation ID.

The event can initially be placed in a transactional outbox if the business record and webhook must be committed consistently. Otherwise, a process crash between “case updated” and “webhook queued” can silently lose the notification. An outbox or equivalent durable event table closes that gap, while a background dispatcher claims records using leases or database row locking. It then makes the HTTP request with a strict connection and total-response timeout. A 10-second timeout is a common starting point, although customer systems may require higher limits; exceeding 30 seconds without a deadline can exhaust workers and worsen congestion.

Each attempt should be logged separately with its HTTP status, latency, response excerpt, error class, and next scheduled time. The canonical event should remain immutable so support engineers can reconstruct what was sent. Aggregate dashboards can instead track delivery rate, unique-event success rate, duplicate suppression, p50, p95, and p99 latency, plus the oldest pending event. These distinctions matter because 99% of HTTP calls returning 2xx does not prove that 99% of business notifications were processed exactly once.

Retry Policy Design: Backoff, Jitter, Expiration, and Ordering

Retries should use exponential backoff with full jitter rather than fixed intervals that cause every failed event to return simultaneously. A formula such as random(0, min(cap, base × 2^attempt)) spreads load while keeping a practical upper bound. For a ten-attempt policy spanning 24 hours, a broad schedule might include target delays of 10 seconds, 1 minute, 5 minutes, 30 minutes, 2 hours, 6 hours, 12 hours, and 24 hours, with randomized jitter around each target. The chosen cap should be tested against the operational meaning of “late”; a compliance filing that must arrive in two hours needs a much shorter policy than a low-priority case digest.

Retry budgets also prevent one destination from monopolizing infrastructure. A sender might cap no more than 10% of its delivery capacity at 60% of the maximum attempts for one destination, while allowing lower-priority events to defer new work. This is especially important after a receiver recovers from an outage, because synchronized retries can create a self-inflicted denial of service. Token-bucket rate limits at both the event and destination level are generally more predictable than an application-wide queue that treats all webhooks alike.

Ordering cannot be guaranteed reliably over ordinary HTTP. Attempt 3 may reach a customer after attempt 4 if a timeout occurs after processing but before the sender receives a response. The payload should therefore carry an object version or monotonically increasing sequence, and the receiver should compare versions before applying an update. If strict ordering is essential, the sender can serialize events per case, but this trades throughput and failure isolation for order. Most B2B case-management integrations are better served by version-aware updates than by trying to force global ordering.

Idempotency: The Receiver’s Most Important Reliability Control

An HTTP network can acknowledge a request even when the response is lost. The sender will then retry, so the receiver will see the same event more than once. Idempotency handles that situation by assigning each event a stable key, commonly a UUID or sender-and-event-ID combination, and recording that key in the same database transaction as the business change. A second request with the same key should return the previously stored result or an equivalent 2xx response without repeating side effects.

A useful design retains the key for at least as long as the sender may retry, not merely for 24 hours. If delivery can continue for seven days, deduplication records may need the same seven-day window, or longer if audit rules require proof of prior processing. For a support case, this might mean not creating two identical cases, not assigning the same compliance task twice, and not sending a second public-affairs alert. A side effect such as sending an email should also be tied to the event-processing transaction through an internal job record; otherwise, the system can create the job twice before the idempotency transaction commits.

HTTP status alone cannot prove semantic success. A receiver may return 200 for a request it has deliberately stored for later processing, which is valid asynchronous behavior but shifts end-to-end responsibility to that receiver. Conversely, a 409 response may indicate that the event already exists and should not automatically be retried forever. The agreement between parties should define accepted status codes, response timeouts, payload schemas, signature verification, and the maximum event age. Without that contract, “webhook reliability” often means both sides assume the other handles failure correctly.

Practical Implementation Steps for an Issue-Ops Platform

First, define the delivery contract. Record a stable event ID, event type, schema version, resource ID, resource version, occurrence time, sender time, and correlation ID, and publish examples of valid and invalid payloads. For a case-house SaaS, events could include case.stage_changed, evidence.requested, approval.completed, and filing.submitted. Keep internal case identifiers separate from customer display references so integrations remain stable after records are migrated.

Second, create the durable outbox. In one database transaction, commit the case update and the pending outbound event; a dispatcher can then poll the outbox or consume a change stream. Use a short claim lease, heartbeat long requests where possible, and release claims after crashes. Every request needs an overall deadline, an idempotency header such as Idempotency-Key, and a body signature that includes the raw bytes and timestamp to prevent replay.

Third, implement classification and scheduling. Retry network timeouts, 408, 425, 429, and selected 5xx responses; quarantine 400, 401, 403, 404, 410, and 422 responses unless a documented retry condition applies. Parse Retry-After for 429 and 503 when supplied, but cap it so one endpoint cannot stop delivery for an unreasonable period. After the final attempt, mark the event as exhausted and move it to a review queue with replay, cancel, or corrected-endpoint actions.

Fourth, test the behavior. Include duplicate delivery, a 200 response followed by a lost connection, a slow endpoint, rate limiting, schema changes, clock skew, worker termination, and replay months later. Track how many attempts produce unique business processing. An SLO such as “99% of eligible notifications reach a receiver within 5 minutes and 99.9% are accepted within 24 hours” is more useful than a vague reliability target, but it must include the endpoint’s availability and the number of retries used.

Managed Services, Open Source, and Build-versus-Buy Alternatives

Teams can build the control plane on a database plus workers, adopt an outbound-webhook platform, or use a broad event-bus service and construct delivery semantics around it. AWS SNS can fan out events through SQS subscriptions, while Lambda, EventBridge, and container services can perform processing. Kafka is appropriate for high-throughput ordered event streams, but it does not remove the need for per-destination HTTP delivery, backoff, endpoint health, or receiver idempotency. A queue moves work; it does not by itself make a remote API call exactly once.

EventDock, advertised on Hacker News in the provided research as webhook reliability priced at $29 per month versus approximately $490 for alternatives, represents the managed outbound-webhook category. Nango, launched on Hacker News as a YC W23 source-available unified API platform, sits closer to integration infrastructure, where authentication and many API connectors may be the larger problem. Neither public launch description proves that a product meets every compliance, retention, regional, or data-residency requirement, so buyers should run a failure-injection trial rather than accept a successful demo as evidence.

Open-source outbound webhook infrastructure such as Outpost can offer control over deployment and data placement, but operations remain the customer’s responsibility. Oracle’s discussion of microservices with an agent registry illustrates that registries improve discovery and governance, not webhook delivery by themselves. For issues.house customers, the relevant choice is usually governed by connector count, delivery volume, compliance controls, and staffing rather than by peak requests per second alone.

FeatureDatabase plus workersManaged webhook or integration serviceEvent bus plus custom delivery layer
Initial design effortHigh, typically 2–8 engineer-weeks for a solid first versionLow to medium, depending on required controlsHigh
Operational ownershipFullUsually shared, but verify scopeFull
Scheduling and deduplicationBuilt and operated by the teamOften included in managed platformsMust be designed separately
Best fitStrict data residency or unusual routing rulesB2B teams needing reliability without a platform teamHigh-volume, many-destination event systems
Typical commercial pathInfrastructure and engineering laborRoughly $29/month for an entry managed offering in the cited EventDock example, versus about $490/month for named alternatives in that comparisonUsage-based queue, function, storage, and network charges plus engineering labor
Main weaknessReliability bugs and on-call burdenVendor limits, lock-in, and possible feature gapsCost and complexity can be disproportionate
## Common Failure Modes That Make Retry Systems Worse

The most common mistake is treating non-2xx responses as a single class. Repeatedly retrying a 401 or a schema validation error wastes capacity and may trigger account security controls. Another error is retaining the “at least once” model without defining duplicate behavior at the receiver. Short idempotency retention, non-transactional side effects, and replay tools that generate new event IDs can all defeat deduplication.

A second common error is using unbounded exponential backoff. A mathematical cap of 31 days may sound conservative, but it can conceal broken credentials for a month. A bounded 24-hour policy plus a visible dead-letter state is usually easier for support, compliance, and public-affairs teams to interpret. Teams should also avoid retry storms after recovery, ignore Retry-After, and let a single 90-second endpoint consume worker capacity. Per-destination concurrency, circuit breaking, jitter, and a global delivery budget address that operational risk.

Payload evolution is another source of incidents. Adding a required field without versioning can break consumers that worked for months. Additive optional fields are usually safer, while removals and semantic changes should use a new schema version and a published migration date. Event timestamps should not be treated as proof of delivery order, and signatures should be rejected outside a documented clock-skew window, often 5 minutes. Finally, sensitive data changes the threat model: secrets belong in managed secret stores, logs should redact tokens and regulated case content, and deletion requests may conflict with retention needed to prove delivery.

When to Act, and What It May Cost

An organization should implement a durable retry system before webhooks become cross-system dependencies for customer-visible actions. The trigger is not simply traffic volume; even 20 important notifications per day can matter if they represent filing deadlines, evidence requests, or approval decisions. Immediate action is warranted when failed requests are invisible, retries happen only in application memory, duplicate cases have already occurred, or no one can explain why a customer received a notification twice. A lighter implementation may be adequate when events are advisory, easily regenerated, and not tied to compliance or customer commitments.

The simplest inexpensive design uses an existing relational database, one outbox table, and a small worker pool, so direct cost can begin below $100 per month for low volume. At higher scale, compute, storage, observability, and egress become usage-dependent rather than following a clean per-webhook price. A managed service can reduce engineering time but introduces subscription and minimum-plan costs; the cited $29 per month EventDock proposition illustrates why buyers should compare functionality, not headline price alone. The stated $490 alternative is a vendor comparison point, not a universal market benchmark, and a fair evaluation should normalize event volume, destinations, retention, support, and included delivery attempts.

For issues.house, a sensible operating target is 99.9% acceptance within 24 hours for eligible events, with alerting after 5 minutes of queue age during business hours. Separate these SLOs by priority: payment or filing events may need escalation after one failed attempt, while digest notifications can wait. Review exhausted events daily, sample signatures and payloads quarterly, and test replay after each major receiver or queue change. The system is reliable when teams can explain every delay, duplicate, and final failure—not when the retry code merely runs without interruption.