Skip to content

Retries & dead-letter

SchedStack accepts a delivery, then owns it until it reaches a terminal state. A delivery that fails on a retryable error is retried with exponential backoff. A delivery that can’t succeed is recorded (as dead_letter or expired), never silently dropped.

This page is the reliability model: the retry policy, what’s retryable vs terminal, the three ways a delivery dead-letters, and how TTL time-caps the whole process. For the operational how-to (listing, inspecting, and replaying failed deliveries), see Dead-letter & replay.

Every delivery ends in exactly one of these. All three are durably recorded and queryable.

State Meaning
succeeded The endpoint returned a 2xx.
dead_letter Gave up after a terminal response, exhausted attempts, or an exhausted retry budget.
expired The delivery’s TTL deadline passed before it could be (re)delivered.
Delivery lifecycleAn accepted delivery is scheduled, then claimed at fire time. A 2xx response ends it as succeeded. Non-2xx responses retry with backoff; exhausting retries parks it as dead_letter. If its TTL elapses first it becomes expired, and canceling the schedule makes it canceled. succeeded, dead_letter, expired, and canceled are terminal.acceptedscheduledclaimed2xxsucceededexhausteddead_letterttl elapsedexpiredretryingbackoffnon-2xxcanceledcancel

Retries use exponential backoff. The delay after a failed attempt is:

delay = min(base * factor^attempt, max)
Retry and backoff timelineA failed attempt is retried on an exponential backoff with jitter: attempt 1 fails, wait about 5 seconds, attempt 2 fails, wait about 10 seconds, attempt 3 succeeds. If attempts are exhausted without a 2xx, the delivery is parked in dead-letter.time →attempt 1failwait ~5sattempt 2failwait ~10s (jittered)attempt 32xx → succeededRetries exhausted with no 2xx → dead_letter

where attempt is 0-based (the delay after the first failure uses attempt = 0).

Defaults, applied when you don’t set retry_policy:

Field Default Meaning
max_attempts 8 Total attempts before dead-lettering (1–50).
base 5s Base delay.
factor 2 Growth factor per attempt (1–100).
max 1h Cap on any single backoff delay.

With the defaults, the backoff doubles each attempt. These are the base (un-jittered) delays; with jitter on by default, each actual wait is randomized within [delay/2, delay] (see below):

attempt 1 fails → wait 5s
attempt 2 fails → wait 10s
attempt 3 fails → wait 20s
attempt 4 fails → wait 40s
attempt 5 fails → wait 1m20s
attempt 6 fails → wait 2m40s
attempt 7 fails → wait 5m20s
attempt 8 fails → dead_letter (max_attempts reached)

Pass retry_policy when you create a schedule. Durations are strings (e.g. "30s", "1h"), consistent with delay and ttl.

Terminal window
curl -X POST https://api.schedstack.com/v1/schedules \
-H "Authorization: Bearer sk_test_…" \
-H "Content-Type: application/json" \
-d '{
"endpoint": "https://example.com/webhooks/orders",
"method": "POST",
"body": "{\"order_id\":\"o_123\"}",
"delay": "5m",
"retry_policy": {
"max_attempts": 12,
"base": "10s",
"factor": 2,
"max": "30m"
}
}'

Validation: max_attempts must be 1–50, factor must be 1–100, and base/max must be valid duration strings. Omitted fields fall back to the defaults above.

After each attempt, SchedStack classifies the outcome. Retryable outcomes back off and try again; terminal outcomes dead-letter immediately, without consuming the rest of your attempt budget.

Outcome Class
2xx Success: stop.
408, 429 Retryable.
5xx Retryable.
Transport fault (timeout, connection refused, DNS/TLS error) Retryable.
3xx Terminal: redirects are not followed; a 3xx is a misconfiguration.
4xx (other than 408/429) Terminal: retrying a client error won’t help.
Blocked address (SSRF guard) Terminal.
Scheme / header-injection rejected (request never sent) Terminal.

A delivery moves to dead_letter for exactly one of these reasons:

  1. Terminal response. The attempt returned a terminal class (3xx/4xx, blocked address, or a rejected request). Retrying can’t help, so SchedStack stops immediately.

  2. Attempts exhausted. A retryable failure occurred on the final attempt: the attempt count reached max_attempts.

  3. Retry budget exhausted. Each tenant has a per-endpoint retry budget that caps aggregate retry amplification. When it’s exhausted, further retries dead-letter instead of piling on. Only retries consume budget; first attempts don’t.

Whichever trigger fires, the final attempt (status code, timing, and error) is recorded on the delivery. Nothing disappears.

ttl time-caps the entire process, complementing the count-based max_attempts. When you set a ttl, the deadline is:

deadline = first_fire_time + ttl

A delivery becomes expired when:

  • it can’t be initiated before the deadline (e.g. a circuit-breaker deferral for an unhealthy endpoint would push the next attempt past it), or
  • a scheduled retry would land after the deadline. SchedStack expires it now rather than firing a request it knows is already too late.

expired is a distinct terminal state from dead_letter, but it’s recorded the same way: visible, queryable, and never a silent drop.

The core guarantee: every accepted delivery reaches succeeded, dead_letter, or expired, and each one is durable and inspectable. There is no fourth, silent outcome.

To find and act on failed deliveries (list the dead-letter queue, read the attempt history, and replay), continue to Dead-letter & replay.