Reliability Patterns
Reference for building resilient backend systems. Read when implementing workers, queues, external integrations, caching, or operational infrastructure.
Table of Contents
- Retries
- Idempotency
- Dead-Letter Queues
- Scheduled Jobs
- Worker Concurrency
- External API Timeouts
- Circuit Breakers
- Caching and Invalidation
- Graceful Shutdown
- Health Checks
Retries
- Exponential backoff with jitter:
delay = base * 2^attempt + random(0, base). - Max retry count: 3–5 for API calls, 5–10 for queue jobs.
- Only retry transient errors (5xx, timeouts, connection errors). Never retry 4xx.
- After max retries, route to dead-letter queue or alert.
Idempotency
- Workers must produce same result regardless of how many times message is processed.
- Use idempotency keys: store
(key, result)with unique constraint. - Use database upserts for naturally idempotent writes.
- Check "already processed" before side effects (emails, payments).
- Set TTL on idempotency records (24–72 hours).
Dead-Letter Queues
- Route messages exceeding max retry count to DLQ.
- Include: original message, error details, retry count, timestamp.
- Monitor DLQ size. Alert when messages appear.
- Build tooling to inspect, replay, or discard DLQ messages.
Scheduled Jobs
- Make scheduled jobs idempotent.
- Use distributed locks (Redis SETNX, database advisory locks) for multi-instance.
- Log job start, completion, duration, items processed.
- Set timeout to prevent runaway execution.
Worker Concurrency
- CPU-bound: match CPU cores. I/O-bound: 2–4x cores. Start at 2x, adjust.
- Separate queues for different priorities.
- Implement backpressure when downstream overloaded.
External API Timeouts
- Set connect timeout (3–5s) and read timeout (10–30s) on every HTTP client.
- Never use infinite timeouts.
- Tighter timeouts on user-facing paths (5–10s total).
Circuit Breakers
States: Closed (normal) → Open (fail fast) → Half-Open (test).
- Per-dependency circuit breakers. Not global.
- Return fallback when open (cached data, defaults).
- Config: 5 failures in 60s → open for 30s → half-open with 3 test requests.
Caching and Invalidation
| Strategy | Best for |
|---|---|
| TTL | API responses, config |
| Write-through | User profiles, settings |
| Event-based | Inventory, pricing |
- Set TTL on every cache entry. No unbounded caches.
- Cache-aside pattern for most cases.
- Handle cache stampede with locking or stale-while-revalidate.
- Never cache sensitive data in shared caches.
Graceful Shutdown
- Stop accepting new requests.
- Drain in-flight requests (timeout 30s).
- Close database/cache connections.
- Flush logs and metrics.
- Exit.
Handle SIGTERM/SIGINT. Drain timeout < orchestrator kill timeout.
Health Checks
Liveness (/health/live): process alive? Lightweight 200 response. No dependency checks.
Readiness (/health/ready): can handle requests? Check database, cache, queue connections. Timeout 2–5s.
Keep health check queries cheap (SELECT 1). No sensitive info in responses.