Advanced Error Handling
Advanced error handling patterns for durable functions, including timeout handling, circuit breakers, and conditional retry strategies.
Timeout Handling with Callbacks
Pattern: Wait for an external callback with a timeout, and implement fallback logic if the timeout is reached.
Implementation approach:
- Use
waitForCallback(TypeScript) orwait_for_callback(Python) with a timeout configuration set in the config argument - Wrap in try-catch to handle timeout errors
- Check if the error is a timeout
- Implement fallback logic in a step (e.g., escalate to manager, use default value, retry with different parameters)
- Return appropriate status indicating timeout occurred
Key considerations:
- Timeout errors are thrown when the callback doesn't complete within the specified duration
- Fallback logic should be in a step to ensure it's checkpointed
- Log timeout events for monitoring and debugging
Local Timeout with Promise.race in Typescript SDK
Pattern: Implement a timeout for a step operation within a single Lambda invocation.
Implementation approach:
- Use
Promise.race()to race the step operation against a timeout promise - The timeout promise rejects after the specified duration
- Catch the timeout error and implement fallback logic
- Execute fallback operation in a separate step
Important limitation: In TypeScript, native setTimeout (and patterns like Promise.race using it) will fail during execution replays. To create a reliable timeout that persists across execution (expands over multi invocations), always use the timeout parameter provided by waitForCallback
Conditional Retry Based on Error Type
Pattern: Retry operations selectively based on the type of error encountered.
Implementation approach:
- Define a custom retry strategy function that examines the error
- For client errors (4xx): Don't retry - these are permanent failures
- For server errors (5xx): Retry with exponential backoff
- For network errors: Retry with fixed delay
- For unknown errors: Don't retry by default
Key considerations:
- Client errors (400-499) typically indicate bad input and shouldn't be retried
- Server errors (500-599) are often transient and benefit from retry
- Network errors (connection refused, timeout) should retry with reasonable limits
- Use exponential backoff for server errors to avoid overwhelming the service
- Set maximum retry attempts to prevent infinite loops
Circuit Breaker Pattern
Pattern: Temporarily stop making requests to a failing external service to prevent cascading failures.
Implementation approach:
- Track failure count and last failure time (note: these reset on replay due to closure mutations)
- Check if circuit is "open" (too many recent failures)
- If open, throw a circuit breaker error and wait before retrying
- If closed, attempt the operation
- On success, reset failure count
- On failure, increment failure count and record timestamp
- Configure retry strategy to wait longer when circuit is open
Important caveat: The example implementations use closure variables (failureCount, lastFailureTime) which reset on replay. For production use, store circuit breaker state in:
- A step return value that persists across replays
- An external store like DynamoDB
- A durable variable pattern
Key considerations:
- Circuit breaker prevents cascading failures to downstream services
- The "open" duration should be long enough for the service to recover
- Reset the circuit on successful operations
- Log circuit state changes for monitoring
Error Handling Best Practices
- Timeout Handling: Always implement fallback logic for callback timeouts - don't let executions fail silently
- Conditional Retries: Classify errors as transient vs permanent, only retry transient errors
- Circuit Breakers: Protect against cascading failures to external services, especially for high-volume operations
- Structured Logging: Log error context (error type, attempt count, operation name) for debugging
- Graceful Degradation: Return partial results when possible rather than failing completely
- Error Classification: Distinguish between client errors (don't retry), server errors (retry with backoff), and network errors (retry with fixed delay)
Common Error Patterns
Transient Errors (Should Retry)
- Network timeouts
- Service unavailable (503)
- Rate limiting (429)
- Database connection failures
- Temporary infrastructure issues
Permanent Errors (Should Not Retry)
- Invalid input (400)
- Authentication failures (401, 403)
- Resource not found (404)
- Business logic violations
- Validation errors
Timeout Errors (Need Fallback)
- Callback timeouts - external system didn't respond in time
- External system delays - service is slow or unresponsive
- Long-running operations - operation exceeded expected duration