All skills
wshobson avatar

/slo-implementation

@156b7a5
by Seth Hobsonwshobson/agents40k stars
4,281

Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting. Use when establishing reliability targets, implementing SRE practices, or measuring service performance.

Use this Skill: https://skilld.dev/gh/wshobson/agents/slo-implementation

This session only. Nothing lands on disk.

referencesdetails.md

≈345 tokens on demand. Your agent reads this file only when SKILL.md points to it.

slo-implementation — additional patterns and templates

Multi-Window Burn Rate Alerts

# Combination of short and long windows reduces false positives
rules:
  - alert: SLOBurnRateHigh
    expr: |
      (
        slo:http_availability:burn_rate_1h > 14.4
        and
        slo:http_availability:burn_rate_5m > 14.4
      )
      or
      (
        slo:http_availability:burn_rate_6h > 6
        and
        slo:http_availability:burn_rate_30m > 6
      )
    labels:
      severity: critical

SLO Review Process

Weekly Review

  • Current SLO compliance
  • Error budget status
  • Trend analysis
  • Incident impact

Monthly Review

  • SLO achievement
  • Error budget usage
  • Incident postmortems
  • SLO adjustments

Quarterly Review

  • SLO relevance
  • Target adjustments
  • Process improvements
  • Tooling enhancements

Best Practices

  1. Start with user-facing services
  2. Use multiple SLIs (availability, latency, etc.)
  3. Set achievable SLOs (don't aim for 100%)
  4. Implement multi-window alerts to reduce noise
  5. Track error budget consistently
  6. Review SLOs regularly
  7. Document SLO decisions
  8. Align with business goals
  9. Automate SLO reporting
  10. Use SLOs for prioritization

Related Skills

  • prometheus-configuration - For metric collection
  • grafana-dashboards - For SLO visualization

Source: SKILL.md on GitHub

No alerts1d5 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    The skill provides templates and guidelines for implementing Service Level Objectives (SLOs) and indicators. No security issues were detected.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: LOW · No issues

  • Runlayer6mo

    1 file scanned · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 156b7a5. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 3 days ago
  • slo
  • sre
  • monitoring
  • prometheus
  • alerting
  • reliability
  • error-budget
  • grafana

README badge

README badge for wshobson/agents/slo-implementation

Defines Service Level Indicators, Service Level Objectives, and error budgets using Prometheus recording and alerting rules to track reliability targets. Includes templates for availability and latency SLIs, error budget policies, burn rate alerts, and Grafana dashboard queries for SRE practices.

Generated from the current SKILL.md.

Does this skill work with Prometheus and Grafana?
Yes. The skill provides Prometheus recording rules and alerting rules for SLO compliance and error budget tracking, plus example Grafana dashboard queries.
What SLI types does this skill cover?
The skill includes templates for availability, latency, and durability SLIs, with example PromQL queries for each.
Can I use this skill if I'm not familiar with SRE practices?
Yes. The skill includes the SLI/SLO/SLA hierarchy, a table of downtime equivalents for different SLO percentages, and guidance on choosing appropriate SLO targets based on user expectations and business requirements.
Does this skill provide alerting rules?
Yes. It includes Prometheus alerting rules for fast burn (14.4x rate), slow burn (6x rate), and error budget exhaustion, each with different severity levels and time windows.

Generated from the current SKILL.md. These answers refresh after source changes.