Observability Checklist
Reference for logging, metrics, tracing, and monitoring in backend systems. Read when adding observability, debugging production issues, or setting up monitoring.
Table of Contents
- Structured Logging
- Request IDs
- Metrics
- Traces
- Audit Logs
- Alert-Worthy Events
- Error Reporting
- Safe Log Redaction
- Health and Readiness Checks
Structured Logging
- Use JSON-formatted logs. Never unstructured
console.login production. - Include:
timestamp,level,message,requestId,userId,service,duration. - Log levels:
error(failures),warn(degradation),info(business events),debug(development only). - Log at request boundaries: request received, request completed (with duration and status).
- Log business events: order created, payment processed, user signed up.
{
"timestamp": "2024-01-15T10:30:00.123Z",
"level": "info",
"message": "Order created",
"requestId": "req_abc123",
"userId": "usr_456",
"orderId": "ord_789",
"duration": 142
}Request IDs
- Generate unique request ID at API gateway or first middleware.
- Propagate through all downstream service calls (via header
X-Request-Id). - Include in all log entries for the request.
- Return in response headers for client-side debugging.
- Use UUIDs or ULIDs for uniqueness.
Metrics
Key metrics to track:
| Metric | Type | What it tells you |
|---|---|---|
| Request rate | Counter | Traffic volume |
| Error rate | Counter | System health |
| Latency (p50/p95/p99) | Histogram | User experience |
| Queue depth | Gauge | Processing backlog |
| Database query time | Histogram | Database health |
| Cache hit rate | Gauge | Cache effectiveness |
| Active connections | Gauge | Resource usage |
- Use RED method for services: Rate, Errors, Duration.
- Use USE method for resources: Utilization, Saturation, Errors.
Traces
- Use distributed tracing for multi-service architectures.
- Trace context propagated via headers (W3C
traceparentor vendor-specific). - Create spans for: HTTP requests, database queries, cache operations, queue publish/consume, external API calls.
- Include span attributes: operation name, status, error message, relevant IDs.
Audit Logs
- Record security-sensitive actions: login, logout, permission changes, data access, configuration changes, admin actions.
- Include:
who(user/service),what(action),when(timestamp),where(IP, service),outcome(success/failure). - Store separately from application logs. Audit logs need longer retention.
- Make audit logs append-only. Never allow deletion by application code.
Alert-Worthy Events
Set alerts for:
- Error rate exceeds threshold (>1% of requests for 5 minutes)
- Latency p99 exceeds SLA
- Queue depth growing continuously
- Health check failures
- Authentication failure spikes
- Database connection pool exhaustion
- Disk space or memory approaching limits
- Certificate expiry approaching
- Zero traffic (service may be down)
Error Reporting
- Capture unhandled exceptions with stack traces.
- Group errors by type and source. Deduplicate repeated errors.
- Include context: request ID, user ID, request path, relevant input (redacted).
- Set severity levels. Page for critical errors, ticket for warnings.
- Track error trends over time. New errors after deploy = potential regression.
Safe Log Redaction
Never log:
- Passwords, tokens, API keys, secrets
- Credit card numbers, SSNs, government IDs
- Full request bodies containing sensitive fields
- Personal health information
Redaction strategies:
- Allowlist loggable fields (safer than blocklist)
- Mask sensitive values:
email: "d***@example.com",card: "****4242" - Use structured logging libraries with built-in redaction
- Audit log configuration periodically
Health and Readiness Checks
Liveness (/health/live): process running, not deadlocked. Return 200 with no dependency checks.
Readiness (/health/ready): process can serve traffic. Check database, cache, queue. Return 200 with dependency status.
Startup (/health/startup): process finished initialization. Used for slow-starting services.
- Timeout: 2–5 seconds. Slow health check = unhealthy.
- No sensitive info in responses.
- Cheap queries only (
SELECT 1,PING).