All skills
wshobson avatar

/distributed-tracing

@be57c0b
by Seth Hobsonwshobson/agents40k stars
4,281

Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks. Use when debugging microservices, analyzing request flows, or implementing observability for distributed systems.

Use this Skill: https://skilld.dev/gh/wshobson/agents/distributed-tracing

This session only. Nothing lands on disk.

SKILL.md

≈66 tokens always: the name and description. ≈454 when used: this file. ≈2k more on demand in 1 file.

Distributed Tracing

Implement distributed tracing with Jaeger and Tempo for request flow visibility across microservices.

Purpose

Track requests across distributed systems to understand latency, dependencies, and failure points.

When to Use

  • Debug latency issues
  • Understand service dependencies
  • Identify bottlenecks
  • Trace error propagation
  • Analyze request paths

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Best Practices

  1. Sample appropriately (1-10% in production)
  2. Add meaningful tags (user_id, request_id)
  3. Propagate context across all service boundaries
  4. Log exceptions in spans
  5. Use consistent naming for operations
  6. Monitor tracing overhead (<1% CPU impact)
  7. Set up alerts for trace errors
  8. Implement distributed context (baggage)
  9. Use span events for important milestones
  10. Document instrumentation standards

Integration with Logging

Correlated Logs

import logging
from opentelemetry import trace

logger = logging.getLogger(__name__)

def process_request():
    span = trace.get_current_span()
    trace_id = span.get_span_context().trace_id

    logger.info(
        "Processing request",
        extra={"trace_id": format(trace_id, '032x')}
    )

Troubleshooting

No traces appearing:

  • Check collector endpoint
  • Verify network connectivity
  • Check sampling configuration
  • Review application logs

High latency overhead:

  • Reduce sampling rate
  • Use batch span processor
  • Check exporter configuration

Related Skills

  • prometheus-configuration - For metrics
  • grafana-dashboards - For visualization
  • slo-implementation - For latency SLOs

Source: SKILL.md on GitHub

No alerts1d5 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    This skill provides standard patterns and configurations for implementing distributed tracing using Jaeger and Tempo. It includes Kubernetes deployment manifests and application instrumentation examples for Python, Node.js, and Go using official OpenTelemetry libraries. No security issues were detected.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: LOW · No issues

  • Runlayer6mo

    1/1 file flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at be57c0b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 4 months ago
  • distributed-tracing
  • jaeger
  • tempo
  • microservices
  • observability
  • opentelemetry
  • instrumentation
  • performance-debugging

README badge

README badge for wshobson/agents/distributed-tracing

Implement distributed tracing with Jaeger and Tempo to track requests across microservices, identify latency issues, and understand service dependencies. The skill covers span instrumentation, context propagation, sampling strategies, and integration with logging to correlate traces with application logs.

Generated from the current SKILL.md.

Does this skill cover both Jaeger and Tempo?
Yes. The skill provides patterns for implementing distributed tracing with either Jaeger or Tempo to track requests across microservices.
What sampling rate should I use in production?
The skill recommends sampling 1-10% of traces in production to balance observability with overhead.
How do I correlate traces with logs?
The skill includes a Python example showing how to extract the trace_id from the current span and add it to log output for correlation.
What should I do if traces aren't appearing?
Check the collector endpoint, verify network connectivity, review sampling configuration, and examine application logs for errors.
How much CPU overhead should distributed tracing have?
The skill recommends monitoring to keep tracing overhead below 1% CPU impact in production.

Generated from the current SKILL.md. These answers refresh after source changes.