TouricalDevelopers
Webhooks

Retries, auto-pause, redelivery

What happens when your endpoint times out, returns a 5xx, or drops to a steady stream of 4xx — and how to recover.

Webhook delivery is at-least-once. We do our best to deliver, with a generous retry cascade. You do your best to dedupe on the Tourical-Event-Id header.

The cascade

When a delivery attempt fails, we requeue it on a fixed schedule:

AttemptDelay after previous attempt
1(initial delivery)
21 minute
35 minutes
430 minutes
52 hours
612 hours
7+Dead-letter — no further automatic retries

So an endpoint that's been down since the initial event sees roughly: t=0, t+1m, t+6m, t+36m, t+2h36m, t+14h36m. Then no more — until you redeliver.

Each attempt is exactly one HTTP request. A delivery is claimed atomically by one worker; a worker that dies mid-request leaves the row in_flight and a sweeper requeues it after 20 minutes, so a crash can produce a duplicate attempt but never a lost one.

What counts as failure

  • 2xx (200–299) → success.
  • 3xx → followed up to 3 hops. Each redirect target is re-validated for SSRF safety. Loops or hop-limit hits are recorded as failure.
  • 4xx → failure. Counts toward auto-pause (see below). Not retried with a tighter schedule — we use the same cascade.
  • 5xx → failure. Retried.
  • Network timeout (15s per attempt) → failure. Retried.
  • DNS failure or private-IP resolution → failure. Recorded but the URL was rejected; if your DNS recovers we'll keep retrying.
  • Body size > 1 MB → never sent. Treated as a permanent failure (this shouldn't happen for events we emit; if it does it's a bug on our side).

Auto-pause

If a single subscription sees 24 consecutive 4xx responses and the streak has lasted at least 24 hours, we automatically:

  1. Set pausedAt and active = false on the subscription.
  2. Stop draining its deliveries — they stay in pending, nothing is dead-lettered by the pause itself.
  3. Notify every owner/admin of the tenant through the dashboard notification feed (kind webhook_subscription_paused).

Any successful delivery resets the streak. This keeps a forgotten or misconfigured endpoint from accumulating thousands of dead-letter rows.

To resume after fixing the consumer side: open the dashboard, Settings → Developer access, click into the subscription, hit Re-enable. The queued rows drain on the next sweep (within 5 minutes), oldest first, on the normal cascade.

Dead-letter rows

Deliveries that have exhausted the retry cascade move to status = "dead_letter". They stay in the database for audit but are never retried automatically. Two ways to recover:

  • Redeliver the row (below).
  • Reconcile via the API: list bookings, payments, etc. via the public REST API and reconcile state on your side. This is the right move when a long outage caused a large backlog.

Redelivery

POST /v1/webhooks/{id}/deliveries/{deliveryId}/redeliver (scope webhooks:write) — or the Redeliver button in the dashboard's Recent deliveries list — queues a fresh delivery:

  • A new delivery row is created (redeliveredFromId points at the original); the original row is never mutated.
  • The new row carries the same Tourical-Event-Id and the same body, with Tourical-Delivery-Attempt: 1 and a new Tourical-Request-Id. Consumers that dedupe on the event id — as this guide asks — see no duplicate side-effect.
  • Any status can be redelivered (success, dead_letter, pending) except in_flight, which returns 409 conflict with code: "in_flight".
  • Redelivery to a paused or deleted subscription returns 409 with code: "subscription_inactive"; re-enable it first.

What you should do on your side

  1. Dedupe on Tourical-Event-Id before doing the work. The same event can arrive twice — a retry after your 2xx was lost on the wire, a sweeper requeue after a worker crash, or a deliberate redelivery.
  2. Make handlers idempotent. Don't increment counters or re-send emails based on receipt alone — key the side-effect on the event id.
  3. Return 2xx fast. Anything you don't need to do synchronously should go into your own background queue. We give up at 15s.
  4. Log the request id. Tourical-Request-Id (<deliveryId>.<attempt>) is the correlation key for our logs and for the deliveries endpoint.
  5. Treat the Tourical-Signature header as authoritative. Don't trust the body until the signature verifies.

Inspecting delivery history

  • GET /v1/webhooks returns each subscription with its lastSuccessAt, lastFailureAt, and consecutiveFailures counters.
  • GET /v1/webhooks/{id}/deliveries (scope webhooks:read) lists delivery rows, newest first, cursor-paginated; filter with ?status=pending|in_flight|success|dead_letter. Each row carries eventId, eventType, status, attempts, lastResponseStatus, nextAttemptAt, deliveredAt, claimedAt, redeliveredFromId. The response body your endpoint returned is not stored — only its status code.
  • The operator dashboard's subscription detail shows the same rows with a Redeliver action.

For deeper forensics, contact support@tourical.com with the Tourical-Request-Id.

On this page