All skills

USE FOR: Drasi continuous-query solutions - real-time queries, change detection, reactive events, data-trigger pipelines on Drasi Server, Drasi for Kubernetes, or drasi-lib. Router: load bundle guides as needed. DO NOT USE for non-Drasi messaging (event-driven-messaging) or pure AKS/ACA hosting (aks-cluster-architecture, azure-container-apps).

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/drasi

This session only. Nothing lands on disk.

bundlesobservabilityguide.md

≈2.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Observability bundle

Use this bundle for logs, metrics, alerts, dashboards, SLOs, telemetry, operational evidence, and proactive monitoring for Drasi solutions.

Observability principle

A Drasi project should show whether change detection, query evaluation, and reaction delivery are working without requiring a user to manually inspect every component.

New-project telemetry baseline

Capture signals for:

  • Runtime health.
  • Source availability.
  • Source lag, source event lag, or stale source indicators where available.
  • Query status and result-change activity.
  • Reaction delivery success, failure, retry, and latency.
  • Downstream rejection or throttling.
  • Container restart count and revision health.
  • Authentication and authorization failures.
  • Secret or identity configuration errors.
  • Version and deployment metadata.

Recommended dashboard sections

  1. Runtime health.
  2. Source status.
  3. Query status.
  4. Reaction status.
  5. End-to-end latency from source change to downstream effect.
  6. Error budget or failure count.
  7. Recent deployments and version changes.
  8. Top failing source/query/reaction names.
  9. Security-relevant events, including unauthorized API access.

Azure hosting telemetry

For Azure Container Apps-hosted Drasi Server, prefer:

  • Log Analytics workspace connected to the Container App environment.
  • Container App revision health and restart alerts.
  • HTTP 5xx and latency alerts for the protected ingress path.
  • Auth failure metrics at APIM, front door, or proxy layer.
  • Secret reference and managed identity error visibility.
  • Deployment metadata in logs or release annotations.

Kubernetes telemetry

For Drasi for Kubernetes, prefer:

  • Kubernetes events for the Drasi namespace.
  • Pod restart and readiness alerts for Drasi components.
  • Resource status collection for Sources, ContinuousQueries, and Reactions.
  • Logs for provider pods, query hosts, and reaction pods.
  • Namespace-level RBAC and Secret access monitoring where available.

Synthetic canary

For production or customer-facing solutions, prefer a small labelled synthetic change that periodically verifies the full path: Source ingestion, ContinuousQuery result update, Reaction delivery, and downstream acknowledgement. Exclude canary records from business metrics and make cleanup/reconciliation explicit.

Alert design

Create practical alerts only. Each alert should include:

  • Symptom.
  • Impact.
  • Likely owner.
  • First diagnostic command or dashboard link.
  • Runbook link.
  • Severity.
  • Auto-remediation status, if any.

Avoid noisy alerts for transient startup states unless they exceed the expected bootstrapping window.

SLO candidates

For production systems, consider SLOs for:

  • Source available percentage.
  • Query running percentage.
  • Reaction delivery success rate.
  • End-to-end change delivery latency.
  • Time to detect event-flow failure.
  • Time to recover reaction delivery.

Drasi-specific SLO baselines

These are pragmatic starting targets for a production Drasi solution. Tune them against your own measured baseline; record the actual target in the architecture decision record (templates/drasi-architecture-decision-record.md). All targets are evaluated on a 30-day rolling window unless noted.

Signal Starting SLO Critical alert threshold Notes
Source availability (running) ≥ 99.5% < 99.0% over 1 h Per Source; weight critical sources higher.
Source CDC lag (event time → ingest) p95 ≤ 5 s; p99 ≤ 30 s p95 > 30 s for 5 min Use source connector lag metric or replication slot lag.
ContinuousQuery freshness (source change → query result) p95 ≤ 10 s; p99 ≤ 60 s p95 > 60 s for 5 min Measure end-to-end via a synthetic source change.
Query active state ≥ 99.5% inactive > 10 min while source is active Inactive query while source is healthy is a control-plane defect.
Reaction delivery success ≥ 99% within 5 min; ≥ 99.9% within 60 min < 95% over 5 min Measure after configured retries; pair with DLQ depth metric.
Reaction delivery latency (query result → side effect) p95 ≤ 30 s p95 > 2 min Use the reaction-specific delivery latency metric.
Container restart rate < 1 restart / pod / day > 3 restarts in 15 min Apply per-component, not aggregate.
Auth / identity errors 0 sustained any error sustained > 5 min Auth errors must page immediately.

Alert design rules

  • Each alert must link to a runbook entry that names the bundle and exact command to run (e.g., "see operations/guide.md symptom table for Query active but stale").
  • Avoid CPU/memory-only alerts. Pair them with a Drasi-level effect (e.g., query lag) to prevent paging for harmless saturation.
  • Tag every alert with runtime_form=(drasi-server|drasi-k8s|drasi-lib) so on-call can route correctly.

Validation that observability actually works

Run this once per release:

  1. Inject a synthetic source change.
  2. Confirm: source-lag metric reports a brief spike, then recovers; query freshness emits a measurement; reaction delivery metric increments.
  3. Confirm an alert fires when you deliberately violate one SLO (e.g., pause a Reaction's downstream for 10 min).
  4. Confirm the alert links resolve to a current runbook.

These baselines are also enforced as release gates in bundles/delivery/guide.md - p95 reaction latency, reaction success rate, and source lag are required PR-time evidence for production-bound merges. If a baseline is missing for your environment, set one before promoting to production; an unset baseline blocks merge rather than allowing it.

Telemetry cost controls

Drasi observability scales with event rate, query cardinality, and trace fan-out. Telemetry is one of the largest cost surfaces of a Drasi deployment - Log Analytics ingestion, metric cardinality explosions, and unsampled traces can each dominate the compute bill. Budget telemetry the same way you budget compute.

Cost-control rules per signal

Signal Cost driver Control
Container logs (Log Analytics) Verbose log level × event rate × payload size Cap production log level at info. Reserve debug for short, time-bounded triage windows with an automatic revert. Drop or sample noisy log streams at the agent (DCR-based transform on Azure Monitor Agent).
Custom metrics High-cardinality labels (per-source-id, per-query-id, per-reaction-id × per-tenant × per-status) Allow per-source-id and per-query-id. Reject per-tenant, per-user, per-request-id labels. Pre-aggregate at the SDK or use a recording rule before metrics reach the workspace.
Distributed traces Trace volume × span fan-out × retention Sample at 1% for steady state, 100% for error traces (use OpenTelemetry tail-based sampling). Drop spans for hot loops (per-event source-connector reads) entirely; keep request-shaped spans only.
Synthetic canary Canary event rate × telemetry per event Run canaries at 1/min in production, 1/15-min in dev. Tag canary events with drasi.synthetic=true and exclude them from business metrics.
Health-probe logs Probe rate (typically 1/sec per pod) × log size Suppress 200 OK probe responses at the ingress logger; log probe failures only.

Budget enforcement

  1. Set a daily Log Analytics ingestion cap per workspace via the Azure Monitor "Daily Cap" setting. Treat hitting the cap as a sev-3 alert - telemetry blackout is operationally worse than the cost overrun.
  2. Tag every dashboard, alert, and ingestion pipeline with a cost-centre tag (cost-centre: <project>) so a FinOps query can split Drasi telemetry from sibling workloads.
  3. Review the top-10 highest-volume log streams and top-10 highest-cardinality metrics monthly; remove or sample any stream that does not have a named alert or runbook consuming it.
  4. Confirm at the pinned release whether Drasi's built-in logger has a configurable level and whether sensitive fields are redacted before they reach the workspace - see bundles/security/guide.md "Logging hygiene".

A common pattern: a Drasi project goes to production with a healthy SLO baseline, then telemetry cost grows linearly with traffic because per-event traces and per-tenant labels were left on. Cost controls MUST be set at first production cutover, not retrofitted after the first invoice.

Cross-reference: bundles/azure-hosting/guide.md "ACA plan choice for Drasi Server" for cost-centre tagging on the Container App; bundles/scaling-and-capacity/guide.md "Storage budgeting" for log-volume budgeting.

Cross-runtime end-to-end signals

A Drasi event that crosses a runtime boundary (drasi-lib edge → Drasi Server hub, or Drasi for Kubernetes core → drasi-lib worker) needs a single end-to-end view; per-runtime SLOs are not additive. This section names what to add when more than one runtime is in scope; single-runtime projects use the SLO baselines above unchanged.

Trace context propagation

  • Propagate W3C traceparent across runtime boundaries on every event-bearing call. drasi-lib emits a span when it produces an event; Drasi Server / for-Kubernetes consumes the same traceparent on ingest and joins the trace. Without propagation, the cross-runtime hop is invisible and root-cause analysis stalls at the boundary.
  • Tag every span with drasi.runtime_form=(drasi-server|drasi-k8s|drasi-lib) so on-call can filter end-to-end traces by runtime hop.
  • If the boundary is asynchronous (e.g. drasi-lib emits to a queue that Drasi Server consumes), include the queue's enqueue/dequeue spans in the trace. The hidden queue latency is the most common cross-runtime SLO defect.

End-to-end freshness budget

Per-runtime SLOs are not additive - a 5s drasi-lib SLO and a 10s Drasi Server SLO do NOT imply 15s end-to-end. The end-to-end budget must be set explicitly and decomposed:

Hop Budget (example; tune to project) Owning bundle
Source change → drasi-lib emit p95 ≤ 2 s bundles/drasi-lib/guide.md
Cross-runtime transport (queue / HTTP / gRPC) p95 ≤ 1 s This bundle (instrumented at the SDK)
Drasi Server / Kubernetes ingest → reaction effect p95 ≤ 7 s SLO baselines above
End-to-end (source → downstream effect) p95 ≤ 10 s Cross-runtime SLO

Set the end-to-end budget once and decompose; do not let per-runtime SLOs drift independently of the end-to-end commitment.

Cross-runtime alerts

  • Alert on end-to-end p95 latency, not on per-runtime latency in isolation. Per-runtime alerts page twice for one user-facing incident.
  • Tag every cross-runtime alert with the boundary identifier so on-call can route to the correct owning team (drasi-lib worker team vs Drasi Server hub team).
  • Pair end-to-end latency alerts with cross-runtime recovery sequencing in bundles/recovery/guide.md - when the end-to-end budget breaks, the recovery action depends on which leg owns the breach.

Evidence for releases

Every release should capture:

  • Before and after runtime health.
  • Current Drasi versions.
  • Source/query/reaction status.
  • Validation test result.
  • Any alert suppressions or expected transient failures.
  • Links to logs or dashboards.

Source: SKILL.md on GitHub

1 warning8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Drasi skill package is a highly structured and security-conscious set of instructions for managing data change detection pipelines. It includes extensive documentation on threat modeling, workload identity setup on AKS, and specific guidance for preventing prompt injection when source data is fed into AI agents. All documented commands and scripts are legitimate operational tools for the Drasi platform, and no malicious patterns such as obfuscation, persistence, or data exfiltration were found.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: MEDIUM · 1 issue

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
metadata
{
  "last_verified": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/drasi