Error Handling
Original skill by ECC. Used with credit under its open-source license.
Use these rules to make errors safe, clear, and easy to fix.
When to use this skill
Use this skill when you:
- Add errors to a new module or service.
- Review an API for missing error checks.
- Call a network service, file system, or database.
- Add retries or a circuit breaker.
- Write error text for users.
- Fix hidden errors or chain failures.
Main rules
- Stop bad work as soon as you can.
- Use error types or codes, not text checks.
- Give users safe and useful text.
- Keep private details in server logs.
- Never hide an error without a clear reason.
- Keep the first error when adding more detail.
- List public error codes in the API docs.
- Test both success and failure paths.
Do not put passwords, keys, tokens, cookies, full request bodies, or private user data in an error or log.
Before writing code
Decide these points first:
- Which errors can the caller fix?
- Which errors are safe to retry?
- Which errors should reach the user?
- Which errors need a log?
- Which public code and HTTP status should each error use?
- Can a retry repeat a write or charge?
Use one public reply shape when possible:
{
"error": {
"code": "NOT_FOUND",
"message": "User not found",
"details": []
}
}Keep error codes stable. A client may use them in its own code.
TypeScript and JavaScript
Typed errors
export class AppError extends Error {
constructor(
message: string,
public readonly code: string,
public readonly statusCode = 500,
public readonly details?: unknown,
options?: ErrorOptions,
) {
super(message, options)
this.name = new.target.name
Object.setPrototypeOf(this, new.target.prototype)
}
}
export class NotFoundError extends AppError {
constructor(resource: string, id: string) {
super(`${resource} not found: ${id}`, 'NOT_FOUND', 404)
}
}
export class ValidationError extends AppError {
constructor(details: { field: string; message: string }[]) {
super('Request validation failed', 'VALIDATION_ERROR', 422, details)
}
}
export class UnauthorizedError extends AppError {
constructor() {
super('Authentication required', 'UNAUTHORIZED', 401)
}
}
export class RateLimitError extends AppError {
constructor(public readonly retryAfterMs: number) {
super('Too many requests', 'RATE_LIMITED', 429)
}
}Use cause when you turn one error into another:
try {
return await db.users.findUniqueOrThrow({ where: { id } })
} catch (cause) {
throw new AppError('User lookup failed', 'DB_ERROR', 500, undefined, {
cause,
})
}Do not show cause, stack traces, SQL, file paths, or service replies to users.
Result values
Use a result value when failure is common and expected, such as parsing input.
type Result<T, E = AppError> =
| { ok: true; value: T }
| { ok: false; error: E }
const ok = <T>(value: T): Result<T, never> => ({ ok: true, value })
const err = <E>(error: E): Result<never, E> => ({ ok: false, error })
function parseAge(value: string): Result<number, ValidationError> {
const age = Number(value)
if (!Number.isInteger(age) || age < 0) {
return err(
new ValidationError([
{ field: 'age', message: 'Enter a whole number that is 0 or more' },
]),
)
}
return ok(age)
}
const result = parseAge(input)
if (!result.ok) {
showFormErrors(result.error.details)
return
}
saveAge(result.value)Do not mix result values and thrown errors in the same layer without a clear rule.
API error handler
import { NextResponse } from 'next/server'
import { z } from 'zod'
export function handleApiError(error: unknown): NextResponse {
if (error instanceof z.ZodError) {
return NextResponse.json(
{
error: {
code: 'VALIDATION_ERROR',
message: 'Check the marked fields',
details: error.issues.map(issue => ({
field: issue.path.join('.'),
message: issue.message,
})),
},
},
{ status: 422 },
)
}
if (error instanceof AppError) {
return NextResponse.json(
{
error: {
code: error.code,
message: error.statusCode >= 500
? 'The request could not be completed'
: error.message,
...(error.details === undefined
? {}
: { details: error.details }),
},
},
{
status: error.statusCode,
headers:
error instanceof RateLimitError
? { 'Retry-After': String(Math.ceil(error.retryAfterMs / 1000)) }
: undefined,
},
)
}
console.error('Unexpected error', error)
return NextResponse.json(
{
error: {
code: 'INTERNAL_ERROR',
message: 'Something went wrong. Please try again.',
},
},
{ status: 500 },
)
}Log an unexpected error once at the service edge. Add a request ID if one exists. Do not log the same error at every layer.
React error boundary
import { Component, ErrorInfo, ReactNode } from 'react'
type Props = {
children: ReactNode
fallback: ReactNode
onError?: (error: Error, info: ErrorInfo) => void
}
type State = {
failed: boolean
}
export class ErrorBoundary extends Component<Props, State> {
state: State = { failed: false }
static getDerivedStateFromError(): State {
return { failed: true }
}
componentDidCatch(error: Error, info: ErrorInfo) {
this.props.onError?.(error, info)
}
render() {
return this.state.failed ? this.props.fallback : this.props.children
}
}An error boundary does not catch errors from event handlers, timers, server code, or most async work. Handle those errors where they run.
Python
Custom errors
class AppError(Exception):
def __init__(
self,
message: str,
code: str,
status_code: int = 500,
details: list[dict[str, str]] | None = None,
):
super().__init__(message)
self.code = code
self.status_code = status_code
self.details = details or []
class NotFoundError(AppError):
def __init__(self, resource: str, item_id: str):
super().__init__(
f"{resource} not found: {item_id}",
"NOT_FOUND",
404,
)
class ValidationError(AppError):
def __init__(self, details: list[dict[str, str]]):
super().__init__(
"Request validation failed",
"VALIDATION_ERROR",
422,
details,
)Keep the first error when adding detail:
try:
user = repository.find_user(user_id)
except DatabaseError as exc:
raise AppError("User lookup failed", "DB_ERROR") from excDo not use except Exception: pass. Catch a narrow error when you can.
FastAPI handlers
import logging
from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse
logger = logging.getLogger(__name__)
app = FastAPI()
@app.exception_handler(AppError)
async def handle_app_error(
request: Request,
exc: AppError,
) -> JSONResponse:
message = (
"The request could not be completed"
if exc.status_code >= 500
else str(exc)
)
return JSONResponse(
status_code=exc.status_code,
content={
"error": {
"code": exc.code,
"message": message,
"details": exc.details,
}
},
)
@app.exception_handler(Exception)
async def handle_unknown_error(
request: Request,
exc: Exception,
) -> JSONResponse:
logger.exception(
"Unexpected request error",
extra={"request_id": getattr(request.state, "request_id", None)},
)
return JSONResponse(
status_code=500,
content={
"error": {
"code": "INTERNAL_ERROR",
"message": "Something went wrong. Please try again.",
}
},
)Do not catch asyncio.CancelledError as a normal app error. Let canceled work stop.
Go
Wrap and check errors
package domain
import (
"context"
"database/sql"
"errors"
"fmt"
)
var (
ErrNotFound = errors.New("not found")
ErrUnauthorized = errors.New("unauthorized")
ErrConflict = errors.New("conflict")
)
func (r *UserRepository) FindByID(
ctx context.Context,
id string,
) (*User, error) {
user, err := r.queryUser(ctx, id)
if errors.Is(err, sql.ErrNoRows) {
return nil, fmt.Errorf("user %s: %w", id, ErrNotFound)
}
if err != nil {
return nil, fmt.Errorf("query user %s: %w", id, err)
}
return user, nil
}Check wrapped errors with errors.Is or errors.As. Do not compare error text.
func (h *Handler) GetUser(w http.ResponseWriter, r *http.Request) {
user, err := h.service.GetUser(
r.Context(),
chi.URLParam(r, "id"),
)
if err != nil {
switch {
case errors.Is(err, domain.ErrNotFound):
writeError(w, http.StatusNotFound, "NOT_FOUND", "User not found")
case errors.Is(err, domain.ErrUnauthorized):
writeError(w, http.StatusUnauthorized, "UNAUTHORIZED", "Sign in first")
case errors.Is(err, context.Canceled):
return
case errors.Is(err, context.DeadlineExceeded):
writeError(w, http.StatusGatewayTimeout, "TIMEOUT", "The request took too long")
default:
slog.Error("unexpected request error", "error", err)
writeError(
w,
http.StatusInternalServerError,
"INTERNAL_ERROR",
"Something went wrong. Please try again.",
)
}
return
}
writeJSON(w, http.StatusOK, user)
}Do not use panic for normal errors. If an HTTP server recovers from a panic, log it and return a safe 500 reply.
Retries
Retry only short faults that may clear soon:
- A network reset.
- A timeout.
- HTTP
408,429,502,503, or504. - A service error marked safe to retry.
Do not retry:
- Bad input.
- Sign-in or access errors.
- Missing data.
- Most
4xxreplies. - A canceled request.
- A write that may run twice, unless it has an idempotency key.
Use a max try count, a max wait, random jitter, and a time limit. Follow Retry-After when the service sends it.
type RetryOptions = {
maxAttempts?: number
baseDelayMs?: number
maxDelayMs?: number
signal?: AbortSignal
retryIf: (error: unknown) => boolean
}
const wait = (ms: number, signal?: AbortSignal) =>
new Promise<void>((resolve, reject) => {
const timer = setTimeout(resolve, ms)
signal?.addEventListener(
'abort',
() => {
clearTimeout(timer)
reject(signal.reason ?? new Error('Canceled'))
},
{ once: true },
)
})
export async function withRetry<T>(
work: () => Promise<T>,
options: RetryOptions,
): Promise<T> {
const {
maxAttempts = 3,
baseDelayMs = 250,
maxDelayMs = 5_000,
signal,
retryIf,
} = options
if (maxAttempts < 1) {
throw new RangeError('maxAttempts must be at least 1')
}
let lastError: unknown
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
signal?.throwIfAborted()
try {
return await work()
} catch (error) {
lastError = error
if (attempt === maxAttempts || !retryIf(error)) {
throw error
}
const cap = Math.min(
maxDelayMs,
baseDelayMs * 2 ** (attempt - 1),
)
await wait(Math.random() * cap, signal)
}
}
throw lastError
}Circuit breakers
Use a circuit breaker when a remote service keeps failing.
It has three states:
closed: Calls run as normal.open: Calls fail fast for a short time.half-open: One or a few test calls may run.
Count only service faults. Do not count bad user input or canceled work.
Set these values:
- How many faults open the circuit.
- How long it stays open.
- How many test calls may run.
- Which calls count as faults.
Keep breaker state per service or endpoint. Do not use one breaker for all remote calls. If the app has many copies, decide if each copy may keep its own state.
When the circuit is open, return a safe error such as SERVICE_UNAVAILABLE. Do not start a retry loop around an open circuit.
User messages
A user message should:
- Say what failed.
- Say what the user can do next.
- Avoid blame.
- Avoid private system facts.
- Avoid false claims that work was saved.
Good:
We could not save your changes. Check your connection and try again.Bad:
Prisma P2024 at db-prod-3: connection pool timed out.Show field errors near the field. Keep typed input when it is safe. Do not show the same error in many places.
Edge cases
Check these cases during review:
- Empty or invalid input.
- Timeouts and canceled work.
- A client that closes the link early.
- A service that sends bad JSON or HTML.
- A retry after part of a write has worked.
- Two requests that try the same write.
- A database rule or lock failure.
- An error while cleaning up.
- An error while handling another error.
- A full disk or closed file.
- A missing error body.
- A non-
Errorvalue thrown in JavaScript. - Too many errors causing too many logs.
- Private data inside error text.
Cleanup must still run. Use finally, defer, or a context manager. If cleanup also fails, keep the main error and record the cleanup error.
Concrete example
Task: Add an endpoint that creates an order and calls a payment service.
Use this plan:
- Check the request before any write.
- Return
VALIDATION_ERRORfor bad fields. - Create an idempotency key for the payment call.
- Set a short time limit.
- Retry only timeouts and safe
5xxreplies. - Stop retries if the request is canceled.
- Open the payment circuit after repeated service faults.
- Wrap the first error with order and request IDs.
- Log the full error once on the server.
- Return safe text to the user.
Example public reply:
{
"error": {
"code": "PAYMENT_UNAVAILABLE",
"message": "Payment is not ready right now. Please try again soon."
}
}Do not tell the user that payment failed if its final state is not known. Use a code such as PAYMENT_STATUS_UNKNOWN, then check the payment by its idempotency key before trying again.
Review checklist
Before finishing, confirm that:
- Every error is handled, returned, or thrown again.
- Error types or codes drive control flow.
- The first error is not lost.
- Public messages hide private details.
- Logs do not hold secrets or private user data.
- Expected errors do not fill logs.
- Retry rules are narrow and have limits.
- Writes cannot run twice by mistake.
- Timeouts and cancel signals reach all slow calls.
- HTTP status codes match the error.
- Public error codes are listed in the API docs.
- Tests cover success, known failure, unknown failure, timeout, cancel, and retry stop.