Drasi on AKS Playbook Template
Use this template when creating an SRE Agent for Drasi workloads on AKS. Adapt resource names, namespaces, and SLO thresholds to your environment.
Scope
This playbook targets:
- Drasi control and processing components on AKS
- event ingestion and processing health
- query/result freshness and pipeline lag
- AKS platform dependencies impacting Drasi behavior
Custom Agent Set
drasi-incident-triage- classify incident and determine likely failure domain
drasi-runtime-diagnostics- inspect Drasi workloads, logs, queue/lag indicators
aks-platform-diagnostics- inspect node/pod/network/storage/control-plane symptoms
drasi-remediation-review- propose and execute approved remediation steps only
Trigger Design
Incident plans
drasi-processing-lag- severity: P1/P2
- keywords: "lag", "stale", "backlog"
- handler:
drasi-runtime-diagnostics
drasi-query-staleness- severity: P2/P3
- keywords: "stale data", "freshness"
- handler:
drasi-runtime-diagnostics
drasi-platform-fault- severity: P1/P2
- keywords: "pod crash", "node not ready", "dns", "ingress"
- handler:
aks-platform-diagnostics
Scheduled tasks
- 15-minute health probe:
- backlog/lag trend
- failed processing rate
- pod restart anomalies
- daily resilience report:
- top failure signatures
- recurring failure windows
- capacity pressure indicators
Diagnostic Workflow
- Search memory for similar Drasi incidents.
- Determine failure domain first:
- application processing logic
- data/source connector path
- AKS platform/runtime path
- Gather evidence:
- pod status and recent events
- workload logs and error signatures
- latency/throughput/freshness trend
- Correlate with recent changes:
- deployments/revisions
- config changes
- cluster events
- Produce action recommendation with risk + rollback.
Evidence Contract
Every incident output must include:
- impact statement
- UTC timeline
- confidence level
- failing component boundary
- remediation options (safe-first ordering)
Prompt Starter: Drasi Runtime Check
Investigate Drasi processing health in AKS namespace @@DRASI_NAMESPACE@@:
1. Identify unhealthy pods and restart patterns.
2. Summarize top runtime errors in the last 60 minutes.
3. Estimate processing lag/freshness impact from available telemetry.
4. Correlate with recent deployment or config changes.
5. Classify likely failure domain and propose next best action.Prompt Starter: Drasi Post-Mitigation Validation
Validate Drasi recovery after mitigation for 30 minutes:
1. Run checks every 1 minute for up to 30 executions.
2. Detect regressions in lag, error rate, and pod health.
3. Escalate immediately on regression with evidence.
4. On success, output final pass/fail summary with supporting data.Governance and Safety
- Keep write actions in Review mode by default.
- Add Stop hook approval gate for disruptive actions.
- Require rollback path before execution.
- Audit all write-capable tool usage.
- For P1/P2 incidents, require full KT sections (
SA,PA,DA,PPA) in outputs.
Integration Notes
Use this with:
- aks-containerapps-production.md
- hooks-governance.md
- deployment-patterns.md
- kt-methodology.md
- kt-templates.md
This template is intentionally environment-neutral and avoids hardcoded IDs.
Bundle Mapping
Recommended bundle composition:
../bundles/base-core../bundles/aks-production../bundles/drasi-aks-production../bundles/governance-kt../bundles/connectors-observability(if external integrations are needed)