Observability bundle
Use this bundle for logs, metrics, alerts, dashboards, SLOs, telemetry, operational evidence, and proactive monitoring for Drasi solutions.
Observability principle
A Drasi project should show whether change detection, query evaluation, and reaction delivery are working without requiring a user to manually inspect every component.
New-project telemetry baseline
Capture signals for:
- Runtime health.
- Source availability.
- Source lag, source event lag, or stale source indicators where available.
- Query status and result-change activity.
- Reaction delivery success, failure, retry, and latency.
- Downstream rejection or throttling.
- Container restart count and revision health.
- Authentication and authorization failures.
- Secret or identity configuration errors.
- Version and deployment metadata.
Recommended dashboard sections
- Runtime health.
- Source status.
- Query status.
- Reaction status.
- End-to-end latency from source change to downstream effect.
- Error budget or failure count.
- Recent deployments and version changes.
- Top failing source/query/reaction names.
- Security-relevant events, including unauthorized API access.
Azure hosting telemetry
For Azure Container Apps-hosted Drasi Server, prefer:
- Log Analytics workspace connected to the Container App environment.
- Container App revision health and restart alerts.
- HTTP 5xx and latency alerts for the protected ingress path.
- Auth failure metrics at APIM, front door, or proxy layer.
- Secret reference and managed identity error visibility.
- Deployment metadata in logs or release annotations.
Kubernetes telemetry
For Drasi for Kubernetes, prefer:
- Kubernetes events for the Drasi namespace.
- Pod restart and readiness alerts for Drasi components.
- Resource status collection for Sources, ContinuousQueries, and Reactions.
- Logs for provider pods, query hosts, and reaction pods.
- Namespace-level RBAC and Secret access monitoring where available.
Synthetic canary
For production or customer-facing solutions, prefer a small labelled synthetic change that periodically verifies the full path: Source ingestion, ContinuousQuery result update, Reaction delivery, and downstream acknowledgement. Exclude canary records from business metrics and make cleanup/reconciliation explicit.
Alert design
Create practical alerts only. Each alert should include:
- Symptom.
- Impact.
- Likely owner.
- First diagnostic command or dashboard link.
- Runbook link.
- Severity.
- Auto-remediation status, if any.
Avoid noisy alerts for transient startup states unless they exceed the expected bootstrapping window.
SLO candidates
For production systems, consider SLOs for:
- Source available percentage.
- Query running percentage.
- Reaction delivery success rate.
- End-to-end change delivery latency.
- Time to detect event-flow failure.
- Time to recover reaction delivery.
Drasi-specific SLO baselines
These are pragmatic starting targets for a production Drasi solution. Tune them against your own measured baseline; record the actual target in the architecture decision record (templates/drasi-architecture-decision-record.md). All targets are evaluated on a 30-day rolling window unless noted.
| Signal | Starting SLO | Critical alert threshold | Notes |
|---|---|---|---|
| Source availability (running) | ≥ 99.5% | < 99.0% over 1 h | Per Source; weight critical sources higher. |
| Source CDC lag (event time → ingest) | p95 ≤ 5 s; p99 ≤ 30 s | p95 > 30 s for 5 min | Use source connector lag metric or replication slot lag. |
| ContinuousQuery freshness (source change → query result) | p95 ≤ 10 s; p99 ≤ 60 s | p95 > 60 s for 5 min | Measure end-to-end via a synthetic source change. |
| Query active state | ≥ 99.5% | inactive > 10 min while source is active | Inactive query while source is healthy is a control-plane defect. |
| Reaction delivery success | ≥ 99% within 5 min; ≥ 99.9% within 60 min | < 95% over 5 min | Measure after configured retries; pair with DLQ depth metric. |
| Reaction delivery latency (query result → side effect) | p95 ≤ 30 s | p95 > 2 min | Use the reaction-specific delivery latency metric. |
| Container restart rate | < 1 restart / pod / day | > 3 restarts in 15 min | Apply per-component, not aggregate. |
| Auth / identity errors | 0 sustained | any error sustained > 5 min | Auth errors must page immediately. |
Alert design rules
- Each alert must link to a runbook entry that names the bundle and exact command to run (e.g., "see
operations/guide.mdsymptom table forQuery active but stale"). - Avoid CPU/memory-only alerts. Pair them with a Drasi-level effect (e.g., query lag) to prevent paging for harmless saturation.
- Tag every alert with
runtime_form=(drasi-server|drasi-k8s|drasi-lib)so on-call can route correctly.
Validation that observability actually works
Run this once per release:
- Inject a synthetic source change.
- Confirm: source-lag metric reports a brief spike, then recovers; query freshness emits a measurement; reaction delivery metric increments.
- Confirm an alert fires when you deliberately violate one SLO (e.g., pause a Reaction's downstream for 10 min).
- Confirm the alert links resolve to a current runbook.
These baselines are also enforced as release gates in bundles/delivery/guide.md - p95 reaction latency, reaction success rate, and source lag are required PR-time evidence for production-bound merges. If a baseline is missing for your environment, set one before promoting to production; an unset baseline blocks merge rather than allowing it.
Telemetry cost controls
Drasi observability scales with event rate, query cardinality, and trace fan-out. Telemetry is one of the largest cost surfaces of a Drasi deployment - Log Analytics ingestion, metric cardinality explosions, and unsampled traces can each dominate the compute bill. Budget telemetry the same way you budget compute.
Cost-control rules per signal
| Signal | Cost driver | Control |
|---|---|---|
| Container logs (Log Analytics) | Verbose log level × event rate × payload size | Cap production log level at info. Reserve debug for short, time-bounded triage windows with an automatic revert. Drop or sample noisy log streams at the agent (DCR-based transform on Azure Monitor Agent). |
| Custom metrics | High-cardinality labels (per-source-id, per-query-id, per-reaction-id × per-tenant × per-status) | Allow per-source-id and per-query-id. Reject per-tenant, per-user, per-request-id labels. Pre-aggregate at the SDK or use a recording rule before metrics reach the workspace. |
| Distributed traces | Trace volume × span fan-out × retention | Sample at 1% for steady state, 100% for error traces (use OpenTelemetry tail-based sampling). Drop spans for hot loops (per-event source-connector reads) entirely; keep request-shaped spans only. |
| Synthetic canary | Canary event rate × telemetry per event | Run canaries at 1/min in production, 1/15-min in dev. Tag canary events with drasi.synthetic=true and exclude them from business metrics. |
| Health-probe logs | Probe rate (typically 1/sec per pod) × log size | Suppress 200 OK probe responses at the ingress logger; log probe failures only. |
Budget enforcement
- Set a daily Log Analytics ingestion cap per workspace via the Azure Monitor "Daily Cap" setting. Treat hitting the cap as a sev-3 alert - telemetry blackout is operationally worse than the cost overrun.
- Tag every dashboard, alert, and ingestion pipeline with a cost-centre tag (
cost-centre: <project>) so a FinOps query can split Drasi telemetry from sibling workloads. - Review the top-10 highest-volume log streams and top-10 highest-cardinality metrics monthly; remove or sample any stream that does not have a named alert or runbook consuming it.
- Confirm at the pinned release whether Drasi's built-in logger has a configurable level and whether sensitive fields are redacted before they reach the workspace - see
bundles/security/guide.md"Logging hygiene".
A common pattern: a Drasi project goes to production with a healthy SLO baseline, then telemetry cost grows linearly with traffic because per-event traces and per-tenant labels were left on. Cost controls MUST be set at first production cutover, not retrofitted after the first invoice.
Cross-reference: bundles/azure-hosting/guide.md "ACA plan choice for Drasi Server" for cost-centre tagging on the Container App; bundles/scaling-and-capacity/guide.md "Storage budgeting" for log-volume budgeting.
Cross-runtime end-to-end signals
A Drasi event that crosses a runtime boundary (drasi-lib edge → Drasi Server hub, or Drasi for Kubernetes core → drasi-lib worker) needs a single end-to-end view; per-runtime SLOs are not additive. This section names what to add when more than one runtime is in scope; single-runtime projects use the SLO baselines above unchanged.
Trace context propagation
- Propagate W3C
traceparentacross runtime boundaries on every event-bearing call. drasi-lib emits a span when it produces an event; Drasi Server / for-Kubernetes consumes the sametraceparenton ingest and joins the trace. Without propagation, the cross-runtime hop is invisible and root-cause analysis stalls at the boundary. - Tag every span with
drasi.runtime_form=(drasi-server|drasi-k8s|drasi-lib)so on-call can filter end-to-end traces by runtime hop. - If the boundary is asynchronous (e.g. drasi-lib emits to a queue that Drasi Server consumes), include the queue's enqueue/dequeue spans in the trace. The hidden queue latency is the most common cross-runtime SLO defect.
End-to-end freshness budget
Per-runtime SLOs are not additive - a 5s drasi-lib SLO and a 10s Drasi Server SLO do NOT imply 15s end-to-end. The end-to-end budget must be set explicitly and decomposed:
| Hop | Budget (example; tune to project) | Owning bundle |
|---|---|---|
| Source change → drasi-lib emit | p95 ≤ 2 s | bundles/drasi-lib/guide.md |
| Cross-runtime transport (queue / HTTP / gRPC) | p95 ≤ 1 s | This bundle (instrumented at the SDK) |
| Drasi Server / Kubernetes ingest → reaction effect | p95 ≤ 7 s | SLO baselines above |
| End-to-end (source → downstream effect) | p95 ≤ 10 s | Cross-runtime SLO |
Set the end-to-end budget once and decompose; do not let per-runtime SLOs drift independently of the end-to-end commitment.
Cross-runtime alerts
- Alert on end-to-end p95 latency, not on per-runtime latency in isolation. Per-runtime alerts page twice for one user-facing incident.
- Tag every cross-runtime alert with the boundary identifier so on-call can route to the correct owning team (drasi-lib worker team vs Drasi Server hub team).
- Pair end-to-end latency alerts with cross-runtime recovery sequencing in
bundles/recovery/guide.md- when the end-to-end budget breaks, the recovery action depends on which leg owns the breach.
Evidence for releases
Every release should capture:
- Before and after runtime health.
- Current Drasi versions.
- Source/query/reaction status.
- Validation test result.
- Any alert suppressions or expected transient failures.
- Links to logs or dashboards.