Topics
- slo
- sre
- monitoring
- prometheus
- alerting
- reliability
- error-budget
- grafana
What it does
Defines Service Level Indicators, Service Level Objectives, and error budgets using Prometheus recording and alerting rules to track reliability targets. Includes templates for availability and latency SLIs, error budget policies, burn rate alerts, and Grafana dashboard queries for SRE practices.
Generated from the current SKILL.md.
Frequently asked
Does this skill work with Prometheus and Grafana?
Yes. The skill provides Prometheus recording rules and alerting rules for SLO compliance and error budget tracking, plus example Grafana dashboard queries.
What SLI types does this skill cover?
The skill includes templates for availability, latency, and durability SLIs, with example PromQL queries for each.
Can I use this skill if I'm not familiar with SRE practices?
Yes. The skill includes the SLI/SLO/SLA hierarchy, a table of downtime equivalents for different SLO percentages, and guidance on choosing appropriate SLO targets based on user expectations and business requirements.
Does this skill provide alerting rules?
Yes. It includes Prometheus alerting rules for fast burn (14.4x rate), slow burn (6x rate), and error budget exhaustion, each with different severity levels and time windows.
Generated from the current SKILL.md. These answers refresh after source changes.