All skills
wshobson avatar

/service-mesh-observability

@be57c0b
by Seth Hobsonwshobson/agents40k stars
4,281

Implement comprehensive observability for service meshes including distributed tracing, metrics, and visualization. Use when setting up mesh monitoring, debugging latency issues, or implementing SLOs for service communication.

Use this Skill: https://skilld.dev/gh/wshobson/agents/service-mesh-observability

This session only. Nothing lands on disk.

SKILL.md

β‰ˆ64 tokens always: the name and description. β‰ˆ638 when used: this file. β‰ˆ1.9k more on demand in 1 file.

Service Mesh Observability

Complete guide to observability patterns for Istio, Linkerd, and service mesh deployments.

When to Use This Skill

  • Setting up distributed tracing across services
  • Implementing service mesh metrics and dashboards
  • Debugging latency and error issues
  • Defining SLOs for service communication
  • Visualizing service dependencies
  • Troubleshooting mesh connectivity

Core Concepts

1. Three Pillars of Observability

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                  Observability                       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚     Metrics     β”‚     Traces      β”‚      Logs       β”‚
β”‚                 β”‚                 β”‚                 β”‚
β”‚ β€’ Request rate  β”‚ β€’ Span context  β”‚ β€’ Access logs   β”‚
β”‚ β€’ Error rate    β”‚ β€’ Latency       β”‚ β€’ Error details β”‚
β”‚ β€’ Latency P50   β”‚ β€’ Dependencies  β”‚ β€’ Debug info    β”‚
β”‚ β€’ Saturation    β”‚ β€’ Bottlenecks   β”‚ β€’ Audit trail   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

2. Golden Signals for Mesh

Signal Description Alert Threshold
Latency Request duration P50, P99 P99 > 500ms
Traffic Requests per second Anomaly detection
Errors 5xx error rate > 1%
Saturation Resource utilization > 80%

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's

  • Sample appropriately - 100% in dev, 1-10% in prod
  • Use trace context - Propagate headers consistently
  • Set up alerts - For golden signals
  • Correlate metrics/traces - Use exemplars
  • Retain strategically - Hot/cold storage tiers

Don'ts

  • Don't over-sample - Storage costs add up
  • Don't ignore cardinality - Limit label values
  • Don't skip dashboards - Visualize dependencies
  • Don't forget costs - Monitor observability costs

Source: SKILL.md on GitHub

1 warning16d5 checks Β· Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides safe templates, configurations, and commands for implementing observability in service meshes such as Istio and Linkerd. No malicious patterns or security risks were detected.

  • Socket16d

    1 alert: gptAnomaly

  • Snyk16d

    Risk: LOW Β· No issues

  • Runlayer6mo

    1 file scanned Β· No issues

  • ZeroLeaks5mo

    Score: 93/100 Β· 2 sections analyzed

Signed by skilld at be57c0b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 4 months ago
  • service-mesh
  • istio
  • linkerd
  • distributed-tracing
  • metrics
  • observability
  • slo
  • latency
  • monitoring

README badge

README badge for wshobson/agents/service-mesh-observability

Implements distributed tracing, metrics collection, and visualization for Istio and Linkerd service meshes, with templates for golden signals monitoring and SLO definition. Use when debugging latency issues, setting up mesh dashboards, or establishing observability across service-to-service communication.

Generated from the current SKILL.md.

Does this skill cover both Istio and Linkerd?
Yes. The skill provides observability patterns for both Istio and Linkerd service mesh deployments, including distributed tracing, metrics, and dashboards.
What are the golden signals this skill teaches?
Latency (P50, P99), traffic (requests per second), errors (5xx rate), and saturation (resource utilization). The skill includes recommended alert thresholds for each.
Does this skill include template code or just theory?
The skill includes concrete templates and worked examples in a references/details.md file that covers sampling rates, trace propagation, alerting setup, and retention strategies.
Can this skill help with debugging latency issues?
Yes. The skill covers distributed tracing across services, correlating metrics with traces using exemplars, and visualizing service dependencies to identify bottlenecks.

Generated from the current SKILL.md. These answers refresh after source changes.