Knowledge Lifecycle Reference
Use knowledge capture to make future investigations faster without creating stale runbook sprawl.
Capture Only Reusable Knowledge
Capture confirmed patterns, detection signals, safe diagnostic commands, mitigation decision criteria, and prevention actions. Avoid copying long incident transcripts or speculative hypotheses.
Lifecycle
- After useful incidents, draft a known-error record or focused runbook update.
- Review knowledge quality before publishing.
- Set owner and review date.
- Run scheduled stale-runbook reviews.
- Retire or rewrite outdated content.
Evaluation Data Quality Tiers
Use quality tiers to gate autonomy promotion. An agent workflow is only ready for L3 (Autonomous) when its evaluation data includes Gold-quality entries validated against real incidents.
| Tier | Source | Use for |
|---|---|---|
| Bronze | Auto-generated from incident metadata, heuristic labels, or agent-suggested mitigations | Initial training data, broad coverage, cold-start |
| Silver | Programmatically generated but calibrated against Gold data; minimum confidence threshold enforced | Nightly evaluations, regression gates, release readiness |
| Gold | Human-verified mitigation labels; exact action, parameters, and outcome confirmed | Autonomy promotion gates, precision/recall measurement |
Generating Gold data without overhead: When an oncaller declares an incident mitigated, auto-suggest the exact mitigation applied (action, target, parameters). The SRE accepts, modifies, or rejects during their standard workflow, this feeds Gold labels back into the evaluation pipeline with zero extra steps.
Calibration: Silver data must be mathematically calibrated against Gold to measure True Precision (not Observed Precision). Use stratified sampling to surface diverse incidents for manual Gold review, this catches edge cases that heuristic Bronze labels miss.
When to use each tier:
- New workflow in
Review→ Bronze is sufficient for initial training. - Promoting to L3 Autonomous → must have Silver-calibrated evaluation data.
- High-risk or write-heavy workflows → require Gold-verified entries before any autonomous execution.
KT Fit
- P1/P2: include full KT summary where useful.
- P3/P4: capture concise symptom, cause, action, and validation.
- Read-only health checks: only capture knowledge when a repeated pattern is found.