All skills
simota avatar

/beacon

@35ffd55
by shingo imotasimota/agent-skills85 stars
15

Engineering observability and reliability: SLO/SLI design, distributed tracing, alerting, dashboards, capacity planning, toil automation, reliability review. Use for instrumentation or SLO definition.

Use this Skill: https://skilld.dev/gh/simota/agent-skills/beacon

This session only. Nothing lands on disk.

referenceopentelemetry-best-practices.md

≈6k tokens on demand. Your agent reads this file only when SKILL.md points to it.

OpenTelemetry Best Practices & Distributed Tracing

Instrumentation strategy, semantic conventions, collector pipeline, sampling, tracing, telemetry correlation, GenAI observability, cost optimization


1. Instrumentation Strategy

# Practice Description Importance
OT-01 Auto-first, Manual-second Start with auto-instrumentation for immediate visibility, then add manual spans for business-critical paths Required
OT-02 SDK initialization first Initialize OTel SDK before any application module imports Required
OT-03 Always close spans Use async/await + finally blocks to ensure spans are closed Required
OT-04 Add business attributes Enrich auto-generated spans with business context (customer_tier, order_value) Recommended
OT-05 Record errors and events Log state transitions and errors as span events Recommended
Initialization order (critical):
  // 1. Initialize OTel SDK first
  const { NodeTracerProvider } = require('@opentelemetry/sdk-trace-node');
  const provider = new NodeTracerProvider();
  provider.register();

  // 2. Then import application modules
  const express = require('express');

Manual Span Creation

@tracer.start_as_current_span("process_payment")
def process_payment(order_id: str, amount: float):
    span = trace.get_current_span()
    span.set_attribute("order.id", order_id)
    span.set_attribute("payment.amount", amount)
    span.set_attribute("payment.currency", "USD")

    try:
        result = charge_card(amount)
        span.set_attribute("payment.status", "success")
        span.set_status(StatusCode.OK)
        return result
    except PaymentError as e:
        span.set_status(StatusCode.ERROR, str(e))
        span.record_exception(e)
        raise

2. Semantic Conventions

Standard Attributes (must follow)

HTTP:       http.method, http.status_code, http.route
DB:         db.system, db.statement, db.name
Messaging:  messaging.system, messaging.destination
RPC:        rpc.system, rpc.service, rpc.method

Application-Specific Attributes

- Prefix: app.* for custom attributes
- Naming: snake_case consistently
- Examples: app.customer_tier, app.order_value, app.feature_flag

Attribute Anti-Patterns

x  High-cardinality attributes (user_id) on all spans -> use traces not metrics
x  Duplicate attributes on parent/child spans
x  Mixing metric/log data into span attributes
x  Abbreviations (svc -> service)

3. Span Naming Conventions

Layer Format Examples
HTTP server HTTP {METHOD} {route} HTTP GET /api/users/:id
HTTP client HTTP {METHOD} HTTP POST
Database {db.system} {operation} {table} postgresql SELECT orders
Message publish {queue} publish orders.created publish
Message consume {queue} process orders.created process
Business logic {verb}_{noun} validate_payment, calculate_tax
External service {service}.{operation} stripe.create_charge

4. Context Propagation

Format Header Ecosystem
W3C Trace Context (recommended) traceparent, tracestate Standard
B3 X-B3-TraceId, X-B3-SpanId Zipkin
Jaeger uber-trace-id Jaeger

Propagation Checklist

- [ ] All HTTP clients inject trace context headers
- [ ] All HTTP servers extract trace context headers
- [ ] Message queues propagate trace context in headers/metadata
- [ ] Async workers link to parent span via context
- [ ] Batch jobs create new root spans with links to triggers
- [ ] Third-party API calls create client spans

Baggage

  • Use Baggage for cross-service context (customer_id, tenant_id)
  • Caution: Baggage propagates to all downstream services
  • Never include sensitive information in Baggage

5. Collector Deployment Patterns

Pattern Configuration Pros Cons Scale
Agent Sidecar per app Network minimal, app isolation Config management distributed Small
Gateway Central server Centralized config, routing SPOF risk Medium
Hierarchical Agent + Gateway Optimal reliability/management balance Complexity Large (recommended)

Collector Configuration

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317

processors:
  memory_limiter:        # OOM prevention (MUST be first)
    check_interval: 1s
    limit_mib: 1000
  batch:                 # Network efficiency
    send_batch_size: 10000
    timeout: 10s

exporters:
  otlp:
    endpoint: observability-backend:4317

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]  # memory_limiter always first
      exporters: [otlp]

Processor Ordering (Critical)

  1. memory_limiter (prevent crashes)
  2. enrichment (k8sattributes, resource)
  3. filter/transform (PII redaction, filtering)
  4. batch (efficient delivery, always last)

PII/PHI Filtering

processors:
  filter:
    spans:
      include:
        match_type: regexp
        attributes:
          - key: db.statement
            value: "(?i)(?:password|passwd)\\s*=\\s*[^\\s,;]+"
      actions:
        - key: db.statement
          action: update
          value: "REDACTED"

Operational Requirements

  • Minimum 4GB node memory for graceful shutdown
  • Version-lock Operator, Collector, and Target Allocator together
  • Use nodeAffinity to prevent deployment on small nodes
  • Monitor: otelcol_receiver_refused_metric_points_total (non-zero = data loss)

6. Sampling Strategies

Strategy Decision Point Pros Cons Use
Head Sampling Trace start Simple, low overhead Misses error traces Dev environments
Tail Sampling Trace completion Intelligent decisions Requires buffering Production (recommended)
Probabilistic Random % Predictable cost Error miss risk High traffic
Rate Limiting Time-based cap Spike control Important trace loss risk Burst protection

Recommended: Composite Sampling Strategy

processors:
  tail_sampling:
    decision_wait: 10s
    policies:
      - name: errors
        type: status_code
        status_code: { status_codes: [ERROR] }     # 100% error retention
      - name: slow-requests
        type: latency
        latency: { threshold_ms: 2000 }
      - name: critical-endpoints
        type: string_attribute
        string_attribute:
          key: http.route
          values: ["/api/payments", "/api/auth"]
      - name: baseline
        type: probabilistic
        probabilistic: { sampling_percentage: 5 }   # 5% normal traffic
    decision_cache_size: 50000

Metrics Accuracy Preservation

  • Generate metrics BEFORE sampling (spanmetrics processor)
  • Use spanmetrics processor for automatic RED metric generation
  • Use servicegraph processor for automatic dependency map generation
processors:
  spanmetrics:
    metrics_exporter: prometheus
    dimensions:
      - name: service.name
      - name: http.method
      - name: http.status_code

7. Telemetry Correlation (Three Pillars)

Log-Trace correlation:
  - Inject trace_id / span_id into logs automatically
  - Use structured logging (JSON)
  - ERROR/WARN logs must also be recorded as span events

Trace -> Metrics conversion:
  - spanmetrics processor for RED metrics
  - servicegraph processor for dependency maps

Performance tuning:
  BatchSpanProcessor:
    maxQueueSize: 2048
    maxExportBatchSize: 512
    scheduledDelayMillis: 5000
    exportTimeoutMillis: 30000
  - Enable gzip compression (bandwidth reduction)
  - Circuit breaker (telemetry must not affect availability)

8. Trace Analysis Patterns

Pattern What to Look For Action
Long spans Single span > SLO threshold Optimize or decompose
Wide traces Fan-out > 50 spans Check N+1 queries
Deep traces Depth > 10 levels Simplify call chain
Orphan spans Missing parent spans Fix context propagation
Gap spans Time gaps between child spans Check queuing/scheduling

9. Cardinality Management

Cardinality explosion example:
  http_requests_total{method, path, status, user_id, client_ip}
  -> method(5) x path(100) x status(10) x user_id(100K) x client_ip(50K)
  -> 2.5 trillion unique time series -> system collapse

Detection:
  1. Index size spikes (RAM/disk monitoring)
  2. remote_write ingestion delays
  3. Aggregation query (sum, avg) latency degradation
  4. Observability platform cost spikes

Control strategy:
  Tier 1: Per-service cardinality limits
    - Business-critical metrics: higher thresholds
    - Infrastructure metrics: strict limits

  Tier 2: Tiered retention policies
    - High resolution (raw data): 24-48 hours
    - Medium resolution (1min aggregation): 30 days
    - Low resolution (1hr aggregation): 13+ months

  Tier 3: Adaptive downsampling
    - Keep high-fidelity data at edge (local)
    - Selective downsampling at central aggregation
    - Prefer automation over manual recording rules

10. Cost Optimization

Key cost levers (most to least effective):
  1. Pipeline-level filtering BEFORE storage (Collector processors)
  2. Intelligent sampling (tail-based, composite)
  3. Tiered retention (raw -> aggregated -> archived)
  4. Per-node Target Allocator (prevents 20-40x metric duplication)
  5. Cardinality limits per service

Results benchmark (CNCF case study):
  - 72% cost reduction vs previous vendor
  - 100% APM trace coverage (was 5% sampling)
  - Enabled by: OTel Collector + open-source backends (Loki, Weave, Mimir)

Observability budget framework:
  - Set per-team telemetry budget (GB/day or cost/month)
  - Monitor telemetry volume per service
  - Alert on budget overruns
  - Quarterly review: optimize top-5 cost contributors

Tool sprawl prevention:
  - Standardize on OTel as the single collection layer
  - Consolidate to single backend per signal type
  - Avoid Prometheus + Datadog + New Relic + custom tools in parallel

11. GenAI / Agent Observability

GenAI semantic conventions, agent span attributes, token cost tracking, and quality metrics → reference/llm-observability.md.

Instrumentation approaches:
  Option 1: Baked-in (framework embeds OTel)
    + Simplified adoption, feature-release control
    - Framework bloat, version lag

  Option 2: External OTel libraries (recommended)
    + Decoupled, community-maintained
    - Fragmentation risk if incompatible packages

Domain causality — spans alone cannot reconstruct an agent run

parent_span_id reconstructs who called whom. It does not answer the questions an agent incident actually asks: which retrieved evidence produced this claim, which state version this tool call mutated, which approval authorized this effect, and which earlier attempt this is a retry of. Time ordering is not causality — two spans adjacent in the timeline may be unrelated, and the causal parent may be minutes earlier.

Carry domain IDs on the span alongside the span IDs:

state_before / state_after      # state version the call read and wrote
produced_evidence               # evidence IDs this step created
caused_state_mutation           # mutation identity, not a boolean
capability_decision_id          # which authorization decision allowed this
approval_ref                    # the approval this effect was bound to
retry_of                        # the prior attempt this supersedes
delegated_to                    # the child run this step handed off to
effect_id                       # external side-effect identity (see idempotency)

The test: from a landed external effect, can you walk back to the approval, the plan, the state, the evidence, and the source? If any hop is missing, the trace records that something happened, not why.

Trace completeness is a measurable property, and worth a dashboard row each: parentage_coverage (spans with a resolvable parent) · evidence_link_coverage (claims with an evidence ID) · approval_binding_coverage (effects bound to an approval).

What not to record. Do not persist private chain-of-thought. Record instead: input/output hash and size, model + config, the structured route decision with its reason code, retrieval query template + parameters (not the full corpus), evidence IDs, tool name + schema version + status + effect identity, state version and mutation summary, approval decision ID, error class and retry relation, latency/token/cost. Storing full prompts and outputs requires a stated purpose, access rule, retention, and redaction — never as a debug default.

Sampling is a correctness concern here, not only a cost one. Dropping high-latency runs or keeping only successful traces produces a corpus in which the failures being investigated do not exist. Always retain error and side-effecting runs; sample the uneventful ones.

GenAI semantic conventions are still moving — the GenAI agent-span conventions are pre-stable. Keep a canonical internal event schema and map it to OTel at the exporter, recording the mapping version on the trace, so a convention change edits the exporter rather than the history.


12. Beacon Integration

Usage by mode:
  1. DESIGN: OTel instrumentation strategy, collector pipeline design
  2. SPECIFY: Collector pipeline specs, sampling configuration
  3. MEASURE: Sampling strategy optimization, cardinality monitoring
  4. Periodic review: Semantic Conventions compliance, cost optimization

Quality gates:
  - OTel SDK initialization is first in app startup (OT-02)
  - memory_limiter processor is first in pipeline
  - Error traces retained at 100%
  - Logs inject trace_id (correlation enabled)
  - PII/PHI filtering in Collector
  - Semantic Conventions compliance in attribute naming
  - New metric addition requires cardinality estimate
  - Telemetry budget per team/service defined

Source: OTel Semantic Conventions v1.40 · Better Stack: OTel Best Practices · CNCF: Cost-Effective OTel · Dash0: OTel Collector Guide · OTel: AI Agent Observability · OTel GenAI SemConv · OTel Weaver


13. 2025 Ecosystem Updates

OTel eBPF Profiler (Public Alpha)

The OpenTelemetry eBPF Profiler enables zero-instrumentation continuous profiling at the kernel level.

Status: Public Alpha (2025)
SIG participants: Grafana, Splunk, Odigos, Elastic

Key capabilities:
  - Language-agnostic: works with Go, Python, Java, Node.js, Rust, .NET without code changes
  - Low overhead: < 1% CPU impact via eBPF
  - Stack trace → OTel profiles → OTLP export
  - Correlate profiles with traces (via trace_id on profile frames)

Profile data format:
  - Follows OTel Profiles specification (experimental)
  - pprof-compatible export for Grafana Pyroscope

When to use:
  - CPU hotspot investigation without instrumentation
  - Memory allocation profiling in production
  - Correlate slow traces with profile data

OTel Logs Stability Status

Logs Bridge API: STABLE (as of OTel v1.x)
  - Use for integrating existing logging frameworks (log4j, winston, etc.)
  - Bridges log records into OTel pipeline with trace correlation

Event API: EXPERIMENTAL
  - Use for structured event emission (e.g., user actions, state transitions)
  - Not yet stable; API may change

Recommendation:
  - Use Logs Bridge API in production for log → OTLP export
  - Use span events for in-span structured data (stable)
  - Avoid Event API in production until stable

Collector Declarative Configuration Schema (RC3)

The OTel Collector is adopting a declarative configuration schema that replaces the current pipeline YAML.

# New declarative config format (RC3, 2025)
# Replaces: receivers/processors/exporters/service.pipelines
receivers:
  otlp/grpc:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317

processors:
  # QueueBatcher replaces batch processor — combines queuing + batching
  queuebatcher/traces:
    max_size: 1000
    timeout: 5s

  memory_limiter:
    limit_mib: 512
    spike_limit_mib: 128

exporters:
  otlphttp/backend:
    endpoint: https://otel.backend.internal

# New: pipelines defined inline with connectors
pipelines:
  traces:
    receivers: [otlp/grpc]
    processors: [memory_limiter, queuebatcher/traces]
    exporters: [otlphttp/backend]

Key change: QueueBatcher replaces the batch processor and adds built-in retry queuing, reducing common pipeline configuration complexity.

Adaptive Telemetry

Adaptive telemetry dynamically adjusts sampling rates based on observed error rates and latency SLOs.

Grafana Cloud Adaptive Metrics: 30-50% cost reduction observed in production

Strategy:
  1. High-value signals: always-on (errors, SLO violations, critical paths)
  2. Normal traffic: tail-based sampling (10-20%)
  3. Healthy, low-latency traffic: head-based sampling (1-5%)
  4. Metrics: adaptive scrape intervals (longer for stable metrics)

Grafana Adaptive Metrics rules example:
  # Drop high-cardinality metrics with low query frequency
  - match: {__name__=~"go_.*"}
    keep_labels: [job, instance]
    drop_if_unqueried_for: 7d

14. 4-Layer Cost Reduction Framework

Layer 1: Generation (Reduce what you produce)

Techniques:
  - Remove unused instrumentation (audit with OTel Weaver)
  - Drop debug spans in production (environment-based filtering)
  - Use exemplars instead of 100% trace sampling
  - Instrument at service boundaries, not every function

Cardinality control:
  - Never use user IDs or request IDs as metric labels
  - Maximum 10 label combinations per metric
  - Alert when cardinality exceeds threshold

Cardinality detection query (Prometheus):
  # Find metrics with > 1000 unique label combinations
  count by (__name__) (
    count by (__name__, job, instance) ({__name__=~".+"})
  ) > 1000

Layer 2: Transport (Reduce what you move)

Techniques:
  - Enable OTLP gzip compression (50-70% size reduction)
  - Batch spans (QueueBatcher: timeout=5s, max_size=1000)
  - Filter at Collector, not at backend (cheaper CPU)
  - Use tail-based sampling to drop healthy traces before export

Collector filter example:
  processors:
    filter/drop_healthy:
      error_mode: ignore
      traces:
        span:
          # Drop spans where no error AND duration < 100ms
          - 'status.code == STATUS_CODE_OK and duration < 100ms and not IsRootSpan()'

Layer 3: Storage (Reduce what you keep)

Retention tiers (recommended):
  | Signal  | Hot (query-ready) | Warm (compressed) | Cold (archive) |
  |---------|-------------------|-------------------|----------------|
  | Metrics | 15 days           | 90 days           | 1 year         |
  | Traces  | 3 days            | 14 days           | 90 days        |
  | Logs    | 7 days            | 30 days           | 1 year         |

Aggregation:
  - Pre-aggregate high-cardinality metrics with recording rules
  - Store raw traces only for errors and SLO violations after hot tier
  - Use log sampling for INFO-level logs after hot tier

Layer 4: Query (Reduce what you read)

Techniques:
  - Create recording rules for frequent, expensive queries
  - Use metric resolution (5m avg) for long-range dashboards
  - Avoid full-table log scans (use structured log fields)
  - Cache dashboard queries (Grafana: 30s-5m depending on panel)

Scrape interval optimization:
  | Metric type              | Recommended interval |
  |--------------------------|----------------------|
  | SLO error budget         | 15s                  |
  | Service RED metrics      | 15s                  |
  | Infrastructure (CPU/mem) | 30s                  |
  | Capacity planning        | 60s                  |
  | Business metrics         | 60s                  |
  | Build/deploy metrics     | 300s                 |

OTel and Profiling Long Form (SKILL.md excerpt)

  • For brownfield services, evaluate OTel eBPF Instrumentation (OBI) for zero-code observability before committing to SDK integration. OBI captures HTTP/gRPC traces and RED metrics without code changes, suitable for initial visibility; add SDK instrumentation selectively for business-critical spans. OBI is in beta (2026), targeting a stable 1.0 release; expanding protocol coverage to messaging (MQTT, AMQP, NATS) and NoSQL (MongoDB). Evaluate for initial rollout in Kubernetes environments.

  • Mandate OTel semantic conventions (stable core since 1.28; track latest release, currently 1.40+) for all instrumentation — non-negotiable for cross-service correlation and vendor portability. For GenAI workloads, adopt gen_ai.* namespace conventions including agent spans (create_agent, invoke_agent operations); these remain experimental as of 2026 — set OTEL_SEMCONV_STABILITY_OPT_IN=http/dup for dual-emission during version transitions to avoid breaking changes on stabilization.

  • Prefer OTel Declarative Configuration (YAML-based SDK config) over code-based setup — stable since 1.0.0 (JSON schema, YAML data model, OTEL_CONFIG_FILE env var). Implementations available in Java, Go, PHP, JS, and C++; .NET and Python in development. Reduces instrumentation drift across services and enables configuration-as-code alongside SLOs-as-code.

  • For environments with 10+ Collectors, adopt OpAMP (Open Agent Management Protocol) with supervisor-based orchestration for fleet management — enables remote configuration reload, health reporting, version discovery, and dynamic pipeline reconfiguration without redeployment. OpAMP Gateway Extension addresses WebSocket connection scaling limits for large fleets.

  • Evaluate OTel Profiles (continuous profiling) as the 4th observability pillar during the DESIGN phase. Profiles entered public Alpha in March 2026 with eBPF-based whole-system profiling (donated by Elastic); include profiling assessment for latency-sensitive services but mark as experimental in implementation specs until the signal reaches stable status.

  • Standardise continuous profiling on Pyroscope 2.0 / Parca for production-scale. Pyroscope 2.0 ingests 19.5 PB/year at Grafana with 95% symbol-storage reduction via write-once symbols; Parca offers the same continuous-profiling primitives under a CNCF-incubating posture. Add continuous profiling as the third pillar alongside metrics (Prometheus / Mimir) and traces (Weave / Jaeger) — flame graphs over time make the "slow in production only" class of bugs observable. Coordinate with siege (concurrency recipe) for memory-leak handoffs (temporal flame graphs) and with bolt for CPU hotspot remediation. [Source: grafana.com/blog/pyroscope-2-0-release/; parca.dev]

  • Wire flame-graph temporal-window analysis into the leak-detection runbook. memray (Python) emits temporal flame graphs that isolate "allocations made inside a window that remain unfreed at the window's end" — the canonical leak signature, not "high allocation rate". Same primitive in jemalloc heap profiling, Pyroscope 2.0, and Parca. Surface continuous-profiling burn-rate alerts (allocation rate × retention rate) alongside latency / error burn rates. [Source: bloomberg.github.io/memray/temporal-flame-graphs.html]

Source: SKILL.md on GitHub

1 warning13d5 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The Beacon skill is a specialized observability and reliability engineering assistant that provides robust guidance for designing SLOs, alerting strategies, and distributed tracing. It adheres to security best practices by emphasizing PII redaction, structured logging, and a separation of duties between design and implementation. No security threats were identified.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

  • Runlayer6mo

    3/9 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 35ffd55. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 2 weeks ago

README badge

README badge for simota/agent-skills/beacon