OpenTelemetry Best Practices & Distributed Tracing
Instrumentation strategy, semantic conventions, collector pipeline, sampling, tracing, telemetry correlation, GenAI observability, cost optimization
1. Instrumentation Strategy
| # | Practice | Description | Importance |
|---|---|---|---|
| OT-01 | Auto-first, Manual-second | Start with auto-instrumentation for immediate visibility, then add manual spans for business-critical paths | Required |
| OT-02 | SDK initialization first | Initialize OTel SDK before any application module imports | Required |
| OT-03 | Always close spans | Use async/await + finally blocks to ensure spans are closed | Required |
| OT-04 | Add business attributes | Enrich auto-generated spans with business context (customer_tier, order_value) | Recommended |
| OT-05 | Record errors and events | Log state transitions and errors as span events | Recommended |
Initialization order (critical):
// 1. Initialize OTel SDK first
const { NodeTracerProvider } = require('@opentelemetry/sdk-trace-node');
const provider = new NodeTracerProvider();
provider.register();
// 2. Then import application modules
const express = require('express');Manual Span Creation
@tracer.start_as_current_span("process_payment")
def process_payment(order_id: str, amount: float):
span = trace.get_current_span()
span.set_attribute("order.id", order_id)
span.set_attribute("payment.amount", amount)
span.set_attribute("payment.currency", "USD")
try:
result = charge_card(amount)
span.set_attribute("payment.status", "success")
span.set_status(StatusCode.OK)
return result
except PaymentError as e:
span.set_status(StatusCode.ERROR, str(e))
span.record_exception(e)
raise2. Semantic Conventions
Standard Attributes (must follow)
HTTP: http.method, http.status_code, http.route
DB: db.system, db.statement, db.name
Messaging: messaging.system, messaging.destination
RPC: rpc.system, rpc.service, rpc.methodApplication-Specific Attributes
- Prefix: app.* for custom attributes
- Naming: snake_case consistently
- Examples: app.customer_tier, app.order_value, app.feature_flagAttribute Anti-Patterns
x High-cardinality attributes (user_id) on all spans -> use traces not metrics
x Duplicate attributes on parent/child spans
x Mixing metric/log data into span attributes
x Abbreviations (svc -> service)3. Span Naming Conventions
| Layer | Format | Examples |
|---|---|---|
| HTTP server | HTTP {METHOD} {route} |
HTTP GET /api/users/:id |
| HTTP client | HTTP {METHOD} |
HTTP POST |
| Database | {db.system} {operation} {table} |
postgresql SELECT orders |
| Message publish | {queue} publish |
orders.created publish |
| Message consume | {queue} process |
orders.created process |
| Business logic | {verb}_{noun} |
validate_payment, calculate_tax |
| External service | {service}.{operation} |
stripe.create_charge |
4. Context Propagation
| Format | Header | Ecosystem |
|---|---|---|
| W3C Trace Context (recommended) | traceparent, tracestate |
Standard |
| B3 | X-B3-TraceId, X-B3-SpanId |
Zipkin |
| Jaeger | uber-trace-id |
Jaeger |
Propagation Checklist
- [ ] All HTTP clients inject trace context headers
- [ ] All HTTP servers extract trace context headers
- [ ] Message queues propagate trace context in headers/metadata
- [ ] Async workers link to parent span via context
- [ ] Batch jobs create new root spans with links to triggers
- [ ] Third-party API calls create client spansBaggage
- Use Baggage for cross-service context (customer_id, tenant_id)
- Caution: Baggage propagates to all downstream services
- Never include sensitive information in Baggage
5. Collector Deployment Patterns
| Pattern | Configuration | Pros | Cons | Scale |
|---|---|---|---|---|
| Agent | Sidecar per app | Network minimal, app isolation | Config management distributed | Small |
| Gateway | Central server | Centralized config, routing | SPOF risk | Medium |
| Hierarchical | Agent + Gateway | Optimal reliability/management balance | Complexity | Large (recommended) |
Collector Configuration
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
processors:
memory_limiter: # OOM prevention (MUST be first)
check_interval: 1s
limit_mib: 1000
batch: # Network efficiency
send_batch_size: 10000
timeout: 10s
exporters:
otlp:
endpoint: observability-backend:4317
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch] # memory_limiter always first
exporters: [otlp]Processor Ordering (Critical)
- memory_limiter (prevent crashes)
- enrichment (k8sattributes, resource)
- filter/transform (PII redaction, filtering)
- batch (efficient delivery, always last)
PII/PHI Filtering
processors:
filter:
spans:
include:
match_type: regexp
attributes:
- key: db.statement
value: "(?i)(?:password|passwd)\\s*=\\s*[^\\s,;]+"
actions:
- key: db.statement
action: update
value: "REDACTED"Operational Requirements
- Minimum 4GB node memory for graceful shutdown
- Version-lock Operator, Collector, and Target Allocator together
- Use
nodeAffinityto prevent deployment on small nodes - Monitor:
otelcol_receiver_refused_metric_points_total(non-zero = data loss)
6. Sampling Strategies
| Strategy | Decision Point | Pros | Cons | Use |
|---|---|---|---|---|
| Head Sampling | Trace start | Simple, low overhead | Misses error traces | Dev environments |
| Tail Sampling | Trace completion | Intelligent decisions | Requires buffering | Production (recommended) |
| Probabilistic | Random % | Predictable cost | Error miss risk | High traffic |
| Rate Limiting | Time-based cap | Spike control | Important trace loss risk | Burst protection |
Recommended: Composite Sampling Strategy
processors:
tail_sampling:
decision_wait: 10s
policies:
- name: errors
type: status_code
status_code: { status_codes: [ERROR] } # 100% error retention
- name: slow-requests
type: latency
latency: { threshold_ms: 2000 }
- name: critical-endpoints
type: string_attribute
string_attribute:
key: http.route
values: ["/api/payments", "/api/auth"]
- name: baseline
type: probabilistic
probabilistic: { sampling_percentage: 5 } # 5% normal traffic
decision_cache_size: 50000Metrics Accuracy Preservation
- Generate metrics BEFORE sampling (spanmetrics processor)
- Use
spanmetricsprocessor for automatic RED metric generation - Use
servicegraphprocessor for automatic dependency map generation
processors:
spanmetrics:
metrics_exporter: prometheus
dimensions:
- name: service.name
- name: http.method
- name: http.status_code7. Telemetry Correlation (Three Pillars)
Log-Trace correlation:
- Inject trace_id / span_id into logs automatically
- Use structured logging (JSON)
- ERROR/WARN logs must also be recorded as span events
Trace -> Metrics conversion:
- spanmetrics processor for RED metrics
- servicegraph processor for dependency maps
Performance tuning:
BatchSpanProcessor:
maxQueueSize: 2048
maxExportBatchSize: 512
scheduledDelayMillis: 5000
exportTimeoutMillis: 30000
- Enable gzip compression (bandwidth reduction)
- Circuit breaker (telemetry must not affect availability)8. Trace Analysis Patterns
| Pattern | What to Look For | Action |
|---|---|---|
| Long spans | Single span > SLO threshold | Optimize or decompose |
| Wide traces | Fan-out > 50 spans | Check N+1 queries |
| Deep traces | Depth > 10 levels | Simplify call chain |
| Orphan spans | Missing parent spans | Fix context propagation |
| Gap spans | Time gaps between child spans | Check queuing/scheduling |
9. Cardinality Management
Cardinality explosion example:
http_requests_total{method, path, status, user_id, client_ip}
-> method(5) x path(100) x status(10) x user_id(100K) x client_ip(50K)
-> 2.5 trillion unique time series -> system collapse
Detection:
1. Index size spikes (RAM/disk monitoring)
2. remote_write ingestion delays
3. Aggregation query (sum, avg) latency degradation
4. Observability platform cost spikes
Control strategy:
Tier 1: Per-service cardinality limits
- Business-critical metrics: higher thresholds
- Infrastructure metrics: strict limits
Tier 2: Tiered retention policies
- High resolution (raw data): 24-48 hours
- Medium resolution (1min aggregation): 30 days
- Low resolution (1hr aggregation): 13+ months
Tier 3: Adaptive downsampling
- Keep high-fidelity data at edge (local)
- Selective downsampling at central aggregation
- Prefer automation over manual recording rules10. Cost Optimization
Key cost levers (most to least effective):
1. Pipeline-level filtering BEFORE storage (Collector processors)
2. Intelligent sampling (tail-based, composite)
3. Tiered retention (raw -> aggregated -> archived)
4. Per-node Target Allocator (prevents 20-40x metric duplication)
5. Cardinality limits per service
Results benchmark (CNCF case study):
- 72% cost reduction vs previous vendor
- 100% APM trace coverage (was 5% sampling)
- Enabled by: OTel Collector + open-source backends (Loki, Weave, Mimir)
Observability budget framework:
- Set per-team telemetry budget (GB/day or cost/month)
- Monitor telemetry volume per service
- Alert on budget overruns
- Quarterly review: optimize top-5 cost contributors
Tool sprawl prevention:
- Standardize on OTel as the single collection layer
- Consolidate to single backend per signal type
- Avoid Prometheus + Datadog + New Relic + custom tools in parallel11. GenAI / Agent Observability
GenAI semantic conventions, agent span attributes, token cost tracking, and quality metrics → reference/llm-observability.md.
Instrumentation approaches:
Option 1: Baked-in (framework embeds OTel)
+ Simplified adoption, feature-release control
- Framework bloat, version lag
Option 2: External OTel libraries (recommended)
+ Decoupled, community-maintained
- Fragmentation risk if incompatible packagesDomain causality — spans alone cannot reconstruct an agent run
parent_span_id reconstructs who called whom. It does not answer the questions an agent incident actually
asks: which retrieved evidence produced this claim, which state version this tool call mutated, which
approval authorized this effect, and which earlier attempt this is a retry of. Time ordering is not causality
— two spans adjacent in the timeline may be unrelated, and the causal parent may be minutes earlier.
Carry domain IDs on the span alongside the span IDs:
state_before / state_after # state version the call read and wrote
produced_evidence # evidence IDs this step created
caused_state_mutation # mutation identity, not a boolean
capability_decision_id # which authorization decision allowed this
approval_ref # the approval this effect was bound to
retry_of # the prior attempt this supersedes
delegated_to # the child run this step handed off to
effect_id # external side-effect identity (see idempotency)The test: from a landed external effect, can you walk back to the approval, the plan, the state, the evidence, and the source? If any hop is missing, the trace records that something happened, not why.
Trace completeness is a measurable property, and worth a dashboard row each:
parentage_coverage (spans with a resolvable parent) · evidence_link_coverage (claims with an evidence ID)
· approval_binding_coverage (effects bound to an approval).
What not to record. Do not persist private chain-of-thought. Record instead: input/output hash and size, model + config, the structured route decision with its reason code, retrieval query template + parameters (not the full corpus), evidence IDs, tool name + schema version + status + effect identity, state version and mutation summary, approval decision ID, error class and retry relation, latency/token/cost. Storing full prompts and outputs requires a stated purpose, access rule, retention, and redaction — never as a debug default.
Sampling is a correctness concern here, not only a cost one. Dropping high-latency runs or keeping only successful traces produces a corpus in which the failures being investigated do not exist. Always retain error and side-effecting runs; sample the uneventful ones.
GenAI semantic conventions are still moving — the GenAI agent-span conventions are pre-stable. Keep a canonical internal event schema and map it to OTel at the exporter, recording the mapping version on the trace, so a convention change edits the exporter rather than the history.
12. Beacon Integration
Usage by mode:
1. DESIGN: OTel instrumentation strategy, collector pipeline design
2. SPECIFY: Collector pipeline specs, sampling configuration
3. MEASURE: Sampling strategy optimization, cardinality monitoring
4. Periodic review: Semantic Conventions compliance, cost optimization
Quality gates:
- OTel SDK initialization is first in app startup (OT-02)
- memory_limiter processor is first in pipeline
- Error traces retained at 100%
- Logs inject trace_id (correlation enabled)
- PII/PHI filtering in Collector
- Semantic Conventions compliance in attribute naming
- New metric addition requires cardinality estimate
- Telemetry budget per team/service definedSource: OTel Semantic Conventions v1.40 · Better Stack: OTel Best Practices · CNCF: Cost-Effective OTel · Dash0: OTel Collector Guide · OTel: AI Agent Observability · OTel GenAI SemConv · OTel Weaver
13. 2025 Ecosystem Updates
OTel eBPF Profiler (Public Alpha)
The OpenTelemetry eBPF Profiler enables zero-instrumentation continuous profiling at the kernel level.
Status: Public Alpha (2025)
SIG participants: Grafana, Splunk, Odigos, Elastic
Key capabilities:
- Language-agnostic: works with Go, Python, Java, Node.js, Rust, .NET without code changes
- Low overhead: < 1% CPU impact via eBPF
- Stack trace → OTel profiles → OTLP export
- Correlate profiles with traces (via trace_id on profile frames)
Profile data format:
- Follows OTel Profiles specification (experimental)
- pprof-compatible export for Grafana Pyroscope
When to use:
- CPU hotspot investigation without instrumentation
- Memory allocation profiling in production
- Correlate slow traces with profile dataOTel Logs Stability Status
Logs Bridge API: STABLE (as of OTel v1.x)
- Use for integrating existing logging frameworks (log4j, winston, etc.)
- Bridges log records into OTel pipeline with trace correlation
Event API: EXPERIMENTAL
- Use for structured event emission (e.g., user actions, state transitions)
- Not yet stable; API may change
Recommendation:
- Use Logs Bridge API in production for log → OTLP export
- Use span events for in-span structured data (stable)
- Avoid Event API in production until stableCollector Declarative Configuration Schema (RC3)
The OTel Collector is adopting a declarative configuration schema that replaces the current pipeline YAML.
# New declarative config format (RC3, 2025)
# Replaces: receivers/processors/exporters/service.pipelines
receivers:
otlp/grpc:
protocols:
grpc:
endpoint: 0.0.0.0:4317
processors:
# QueueBatcher replaces batch processor — combines queuing + batching
queuebatcher/traces:
max_size: 1000
timeout: 5s
memory_limiter:
limit_mib: 512
spike_limit_mib: 128
exporters:
otlphttp/backend:
endpoint: https://otel.backend.internal
# New: pipelines defined inline with connectors
pipelines:
traces:
receivers: [otlp/grpc]
processors: [memory_limiter, queuebatcher/traces]
exporters: [otlphttp/backend]Key change: QueueBatcher replaces the batch processor and adds built-in retry queuing, reducing common pipeline configuration complexity.
Adaptive Telemetry
Adaptive telemetry dynamically adjusts sampling rates based on observed error rates and latency SLOs.
Grafana Cloud Adaptive Metrics: 30-50% cost reduction observed in production
Strategy:
1. High-value signals: always-on (errors, SLO violations, critical paths)
2. Normal traffic: tail-based sampling (10-20%)
3. Healthy, low-latency traffic: head-based sampling (1-5%)
4. Metrics: adaptive scrape intervals (longer for stable metrics)
Grafana Adaptive Metrics rules example:
# Drop high-cardinality metrics with low query frequency
- match: {__name__=~"go_.*"}
keep_labels: [job, instance]
drop_if_unqueried_for: 7d14. 4-Layer Cost Reduction Framework
Layer 1: Generation (Reduce what you produce)
Techniques:
- Remove unused instrumentation (audit with OTel Weaver)
- Drop debug spans in production (environment-based filtering)
- Use exemplars instead of 100% trace sampling
- Instrument at service boundaries, not every function
Cardinality control:
- Never use user IDs or request IDs as metric labels
- Maximum 10 label combinations per metric
- Alert when cardinality exceeds threshold
Cardinality detection query (Prometheus):
# Find metrics with > 1000 unique label combinations
count by (__name__) (
count by (__name__, job, instance) ({__name__=~".+"})
) > 1000Layer 2: Transport (Reduce what you move)
Techniques:
- Enable OTLP gzip compression (50-70% size reduction)
- Batch spans (QueueBatcher: timeout=5s, max_size=1000)
- Filter at Collector, not at backend (cheaper CPU)
- Use tail-based sampling to drop healthy traces before export
Collector filter example:
processors:
filter/drop_healthy:
error_mode: ignore
traces:
span:
# Drop spans where no error AND duration < 100ms
- 'status.code == STATUS_CODE_OK and duration < 100ms and not IsRootSpan()'Layer 3: Storage (Reduce what you keep)
Retention tiers (recommended):
| Signal | Hot (query-ready) | Warm (compressed) | Cold (archive) |
|---------|-------------------|-------------------|----------------|
| Metrics | 15 days | 90 days | 1 year |
| Traces | 3 days | 14 days | 90 days |
| Logs | 7 days | 30 days | 1 year |
Aggregation:
- Pre-aggregate high-cardinality metrics with recording rules
- Store raw traces only for errors and SLO violations after hot tier
- Use log sampling for INFO-level logs after hot tierLayer 4: Query (Reduce what you read)
Techniques:
- Create recording rules for frequent, expensive queries
- Use metric resolution (5m avg) for long-range dashboards
- Avoid full-table log scans (use structured log fields)
- Cache dashboard queries (Grafana: 30s-5m depending on panel)
Scrape interval optimization:
| Metric type | Recommended interval |
|--------------------------|----------------------|
| SLO error budget | 15s |
| Service RED metrics | 15s |
| Infrastructure (CPU/mem) | 30s |
| Capacity planning | 60s |
| Business metrics | 60s |
| Build/deploy metrics | 300s |OTel and Profiling Long Form (SKILL.md excerpt)
For brownfield services, evaluate OTel eBPF Instrumentation (OBI) for zero-code observability before committing to SDK integration. OBI captures HTTP/gRPC traces and RED metrics without code changes, suitable for initial visibility; add SDK instrumentation selectively for business-critical spans. OBI is in beta (2026), targeting a stable 1.0 release; expanding protocol coverage to messaging (MQTT, AMQP, NATS) and NoSQL (MongoDB). Evaluate for initial rollout in Kubernetes environments.
Mandate OTel semantic conventions (stable core since 1.28; track latest release, currently 1.40+) for all instrumentation — non-negotiable for cross-service correlation and vendor portability. For GenAI workloads, adopt
gen_ai.*namespace conventions including agent spans (create_agent,invoke_agentoperations); these remain experimental as of 2026 — setOTEL_SEMCONV_STABILITY_OPT_IN=http/dupfor dual-emission during version transitions to avoid breaking changes on stabilization.Prefer OTel Declarative Configuration (YAML-based SDK config) over code-based setup — stable since 1.0.0 (JSON schema, YAML data model,
OTEL_CONFIG_FILEenv var). Implementations available in Java, Go, PHP, JS, and C++; .NET and Python in development. Reduces instrumentation drift across services and enables configuration-as-code alongside SLOs-as-code.For environments with 10+ Collectors, adopt OpAMP (Open Agent Management Protocol) with supervisor-based orchestration for fleet management — enables remote configuration reload, health reporting, version discovery, and dynamic pipeline reconfiguration without redeployment. OpAMP Gateway Extension addresses WebSocket connection scaling limits for large fleets.
Evaluate OTel Profiles (continuous profiling) as the 4th observability pillar during the DESIGN phase. Profiles entered public Alpha in March 2026 with eBPF-based whole-system profiling (donated by Elastic); include profiling assessment for latency-sensitive services but mark as experimental in implementation specs until the signal reaches stable status.
Standardise continuous profiling on Pyroscope 2.0 / Parca for production-scale. Pyroscope 2.0 ingests 19.5 PB/year at Grafana with 95% symbol-storage reduction via write-once symbols; Parca offers the same continuous-profiling primitives under a CNCF-incubating posture. Add continuous profiling as the third pillar alongside metrics (Prometheus / Mimir) and traces (Weave / Jaeger) — flame graphs over time make the "slow in production only" class of bugs observable. Coordinate with
siege(concurrency recipe) for memory-leak handoffs (temporal flame graphs) and withboltfor CPU hotspot remediation. [Source: grafana.com/blog/pyroscope-2-0-release/; parca.dev]Wire flame-graph temporal-window analysis into the leak-detection runbook.
memray(Python) emits temporal flame graphs that isolate "allocations made inside a window that remain unfreed at the window's end" — the canonical leak signature, not "high allocation rate". Same primitive injemalloc heap profiling, Pyroscope 2.0, and Parca. Surface continuous-profiling burn-rate alerts (allocation rate × retention rate) alongside latency / error burn rates. [Source: bloomberg.github.io/memray/temporal-flame-graphs.html]