Operations bundle
Use this bundle for runtime triage, inactive resources, missing events, unhealthy components, logs, drift, and API/CLI mismatch.
Triage principle
Diagnose from the narrowest observable failure and preserve evidence before changing state.
Dependency order for investigation:
- Runtime form and version.
- Hosting/runtime health.
- Source availability.
- Query status and results.
- Reaction status and downstream delivery.
- External dependency health.
First evidence to collect
For any issue, capture:
- Runtime form: Drasi Server, Drasi for Kubernetes, or drasi-lib.
- Pinned version, CLI version, and container image if applicable.
- Namespace, resource group, or environment.
- Source/query/reaction names.
- Exact error text.
- Recent changes.
- Logs around first failure time.
- Whether the issue is create-time, runtime, delivery, or recovery.
Drasi for Kubernetes checks
kubectl config current-context
kubectl get namespace "$DRASI_NAMESPACE"
drasi list source -n "$DRASI_NAMESPACE"
drasi list query -n "$DRASI_NAMESPACE"
drasi list reaction -n "$DRASI_NAMESPACE"
drasi describe source "$SOURCE_NAME" -n "$DRASI_NAMESPACE"
drasi describe query "$QUERY_NAME" -n "$DRASI_NAMESPACE"
drasi describe reaction "$REACTION_NAME" -n "$DRASI_NAMESPACE"Use Kubernetes logs only after Drasi resource status points to a component-level issue.
Exposing reactions/sources over HTTP on Kubernetes: use drasi ingress init, which installs the Contour ingress controller by default (namespace projectcontour) or integrates an existing controller via --use-existing --ingress-service-name <name> --ingress-namespace <ns> --ingress-class-name <class>. This is the supported path for reaching reaction endpoints (e.g. the MCP Reaction on port 3000) from outside the cluster ; see https://drasi.io/drasi-kubernetes/reference/ingress/.
Drasi Server checks
curl -fsS "$DRASI_BASE_URL/health"
curl -fsS "$DRASI_BASE_URL/api/v1/openapi.json" | jq '.openapi // .swagger'
curl -fsS "$DRASI_BASE_URL/api/v1/sources" | jq '.'
curl -fsS "$DRASI_BASE_URL/api/v1/queries" | jq '.'
curl -fsS "$DRASI_BASE_URL/api/v1/reactions" | jq '.'If the OpenAPI or route shape differs, trust the live OpenAPI from the pinned server version and update examples accordingly.
Symptom routing
| Symptom | Likely layer | Next bundle |
|---|---|---|
| Container App unhealthy | Azure hosting/runtime | azure-hosting |
| Source unavailable | Provider credentials, CDC, network, permissions | sources |
| Source flaps between available and terminal | PostgreSQL prereq or network instability | sources |
| Source reactivator fails; proxy logs "access token has invalid format" | Azure Workload ID injection interfering with password auth JDBC | sources (see Azure PostgreSQL prerequisite gate) |
| Source reactivator fails with "must be owner of table/publication" | Replication user does not own CDC tables or publication | sources (see publication ownership runbook) |
| Query TerminalError | Syntax, unsupported function, source unavailable | continuous-queries |
| Query active but stale | Source event flow, CDC, source permissions, query filters | sources, then continuous-queries |
Source running but query says Table <x> not found |
Wrong database/schema or schema not initialized | sources |
| Reaction healthy but no downstream effect | Template, auth, network, downstream rejection | reactions |
| Management API reachable publicly | Exposure/auth issue | security |
| Repeated crashes, corrupted state, or failed reset | Potential destructive recovery | recovery |
Defining "stale" and proving event flow
"Query active but stale" is a common but ambiguous symptom. Pin it down before triage.
Definition
A query is stale when no result update has been observed for 5 × p95-source-change-latency while the upstream source is producing changes. For a source running at p95 = 5 s, stale = no update for 25 s. Tune per query.
Proof procedure
Trigger a known source change (see
validation/guide.mdper-provider commands).Watch the query result stream:
# Drasi Server curl -N http://localhost:8080/api/v1/events | jq 'select(.queryId=="<id>")' # Drasi for Kubernetes drasi watch query <id>Measure: t1 = trigger time, t2 = first result update visible.
t2 - t1is the end-to-end freshness.If
t2never arrives, the chain is broken. Walk the dependency order:
- Source state:
drasi describe source <name>orcurl /api/v1/sources/<id>- is itrunning? Is the CDC slot/connector healthy? - Source-to-query wiring: does the query's
sources[].sourceIdmatch the sourceidexactly? - Query state:
drasi describe query <id>- is itrunning? Any compile errors in logs? - Reaction subscription: does the reaction's
querieslist contain the query ID?
Common root causes
| Observation | Likely cause |
|---|---|
| Source healthy, no query update, query CPU idle | Query predicate filters out the test change. Adjust the trigger to match the predicate. |
| Source healthy, query update arrives, no reaction effect | Reaction misconfigured, downstream rejecting, retry budget exhausted. Check reaction logs and DLQ. |
Source running but ingest lag growing |
Replication slot/connector backpressure. See sources/guide.md PG slot runbook. |
Query alternates between running and inactive |
Control-plane flap; check operator logs and resource pressure on the query container. |
Corrective action order
- Prefer configuration fix over restart.
- Prefer component restart over resource delete/recreate.
- Prefer affected resource rebuild over full runtime reset.
- Escalate to
recoverybefore any destructive action.
Runbook - Source says running, query says table not found
- Verify the query's source labels map to table names expected in the source definition.
- Run SQL preflight from
bundles/sources/guide.mdto confirm target schema tables actually exist. - If tables are missing, run database initialization/migrations first. Do not continue query parser triage until schema exists.
- Re-apply source, then queries, then reactions in dependency order.
Runbook - PostgreSQL source flaps or auth oscillates
Symptoms seen together:
drasi list sourcealternates between available and unavailable.- Query status flips to
TerminalErrorwith source fetch/auth errors. - Reactivator logs include one or more of:
wal_level property must be 'logical' but is: 'replica'password authentication failedThe connection attempt failed. Connect timed out
Remediation order:
- Pause query rollout. Do not keep applying new queries while source is unstable.
- Validate PostgreSQL prereqs: effective
wal_level=logical, replication role, publication/slot prerequisites. - Validate network posture consistency (public or private). Avoid half-configured transitions.
- Restart source reactivator and re-check logs.
- Run the source stability gate (10+ minutes continuous availability) before re-applying affected queries.
If stability cannot be proven, keep queries disabled and preserve baseline running queries only.
Additional diagnosis for "access token has invalid format" (Azure PostgreSQL + Workload ID)
Symptom: Source reactivator logs show password authentication failed OR the query-api/proxy logs show FATAL: The access token has invalid format. The source pod has AZURE_CLIENT_ID in its environment.
Root cause: The Azure Workload Identity webhook injects Entra token env vars into the pod even when the source is configured for password auth. The PostgreSQL JDBC driver prefers token auth when these env vars are present.
Fix:
# 1. Remove workload identity annotation from the default service account
kubectl annotate serviceaccount default --namespace <drasi-namespace> azure.workload.identity/client-id-
# 2. Delete all PostgreSQL source pods to force recreation without env vars
kubectl delete pods -n <drasi-namespace> -l drasi/resource=<postgres-source-name>
# 3. Wait for full recovery and re-check source status
drasi wait source <postgres-source-name> --namespace <drasi-namespace> --timeout 300Verification: Confirm AZURE_CLIENT_ID is no longer present in the proxy pod environment:
kubectl get pods -n <drasi-namespace> -l drasi/service=proxy,drasi/resource=<source-name> \
-o jsonpath='{.items[0].spec.containers[?(@.name=="proxy")].env[?(@.name=="AZURE_CLIENT_ID")]}'
# Should return emptyRunbook - Publication ownership failure after switching replication users
Symptom: Reactivator logs show must be owner of publication <name> or must be owner of table <name> after changing the replication user or auth mode.
Root cause: The publication and tables are owned by the previous replication user. The new user cannot create or update filtered publications.
Fix:
Drop the stale publication and replication slot (from the
postgresdatabase or as the server admin):DROP PUBLICATION IF EXISTS rg_<publication_name>; SELECT pg_drop_replication_slot('<slot_name>');Transfer table ownership to the new replication user:
ALTER TABLE <schema>.<table> OWNER TO <new_replication_user>;
Repeat for every table in the Drasi source's table list.
If the new user needs inherited privileges (e.g.,
azure_pg_admin), fix the role membership chain:GRANT <server_admin_role> TO <new_replication_user>;Restart the reactivator pod:
kubectl delete pods -n <drasi-namespace> -l drasi/service=reactivator,drasi/resource=<source-name>Verify the source comes online and all queries return to
Runningstatus.
Drift and version mismatch
Drasi is moving quickly. If commands or resource shapes fail unexpectedly:
- Compare local CLI version with target platform/server version.
- Check live OpenAPI for Drasi Server.
- Check current provider docs for Drasi for Kubernetes.
- Avoid mixing examples from older tutorials with current release schemas.
- Pin the working version in repo documentation and automation.
Incident summary requirements
When reporting an operational fix, include:
- Root cause or most likely cause.
- Evidence used.
- Scope of impact.
- Change made.
- Validation result.
- Follow-up hardening item.
Failure-mode taxonomy
Single consolidated index of Drasi failure classes. Each row names the severity, the bundle that owns the runbook, and the operator's first action. When triaging, find the row first, then jump to the owning bundle.
| Failure class | Severity | Owning bundle | First action |
|---|---|---|---|
| Source disconnected | High | sources |
Confirm source state via drasi describe source / /api/v1/sources/<id>; do not restart until reason captured. |
| Replication-slot bloat | High | sources |
Capture current slot lag and disk headroom on the upstream; freeze schema changes until drained. |
| Source-credential expiry / rotation half-state | High | sources + security |
Verify which credential version Drasi is presenting; complete or roll back the rotation rather than guessing. |
| Query inactive | Medium | continuous-queries |
Read query status reason; do not delete/recreate until the reason is captured. |
| Query container OOM | High | operations (see OOM triage below) + scaling-and-capacity |
kubectl describe pod for OOMKilled; preserve last metrics before restart. |
| Query alternating active/inactive (control-plane flap) | High | operations (see deep dive below) |
Capture operator/control-plane logs filtered by query name before any intervention. |
| Reaction queue full | High | reactions + scaling-and-capacity |
Inspect DLQ depth and downstream latency; do not scale replicas blindly. |
| Reaction endpoint 401/403 | High | reactions + security (see playbook below) |
Confirm the identity Drasi is presenting before touching the downstream. |
| Reaction endpoint 5xx | Medium | reactions |
Check downstream health and retry budget; preserve a sample failing payload. |
| Secret rotation mid-flight | High | security |
Identify which side (Drasi or downstream) holds the stale value; do not abort, complete the overlap. |
| Identity (Workload ID) token expiry | High | security + operations (runbook below) |
Check federated credential subject and AKS OIDC issuer before rotating anything. |
| MCP transport reconnect | Low | reactions (MCP failure handling) |
Confirm clients reconnect and re-read; check subscription count for leaks. |
| Network partition source-side | High | sources + recovery |
Capture last good CDC position; do not reset state until partition is confirmed healed. |
| Network partition reaction-side | High | reactions |
Inspect egress allowlist/NetworkPolicy first; the partition may be policy, not infrastructure. |
| Control-plane API unauthorized | High | security |
Confirm caller identity and route; do not relax auth as a workaround. |
| Deployed but nothing happens (apply succeeded, no events, no logs) | P3 | operations |
Verify the query's sources: reference matches the source name exactly (case-sensitive) and that the source is running, not provisioning; see bundles/sources/guide.md. |
| Persistent-index inconsistency after crash | High | recovery |
Confirm RocksDB/Garnet index store consistency; if a crash occurred mid-source-change on a deployment pinned to drasi-platform <= 0.9.x (i.e. before the drasi-core Session-Scoped-Transactions Phase 3 fix from issue #290 propagated through to a platform release), do not trust the persistent indexes - re-bootstrap from source. See https://github.com/drasi-project/drasi-core/issues/290. |
| Aggregation phantom rows after groups empty out | Medium | continuous-queries |
Drop and recreate the affected ContinuousQuery; watch for resolution in drasi-core #384. Symptom: aggregation result shows stale identity-valued rows for a GROUP BY group long after all source elements for that group are gone. See https://github.com/drasi-project/drasi-core/issues/384. |
| Plugin directory listing pagination cap (~50 plugins) | Medium | sources (also relevant to reactions) |
Limit deployed plugin count under 50, or wait for drasi-core #414 to add pagination. Symptom: CLI/API plugin enumeration silently truncates at the cap. See https://github.com/drasi-project/drasi-core/issues/414. |
drasi init / drasi list targets the wrong cluster |
High | operations (see CLI env trap in SKILL.md) |
The drasi CLI persists its own environment config across sessions. After switching kubectl contexts, re-run drasi env kube before any cluster-scoped command. Symptom: drasi init reports success but resources appear in the wrong/previous cluster. |
Reaction/Source pod ImagePullBackOff after drasi apply |
High | operations + reactions/sources |
The provider is registered in the catalog but its image was not built for the platform version. Verify the image exists on GHCR at the platform tag before applying. Symptom: drasi apply succeeds, the deployment is created, but the pod never starts because the image tag returns 404. See bundles/reactions/guide.md ## Reaction provider image verification gate. |
Recovery / triage stubs for the new rows above:
- Persistent-index inconsistency after crash. If a crash interrupted a source change on
drasi-platform <= 0.9.x(pre-#290-fix-propagation), assume the RocksDB/Garnet persistent index may diverge from source-of-truth. Capture the index store state for forensics, then jump to therecoverybundle and re-bootstrap the affected source rather than restarting in place - replays from a divergent index will compound the inconsistency. - Aggregation phantom rows after groups empty out. When a GROUP BY aggregation continues to emit identity-valued rows for a group whose source elements have all been deleted (drasi-core #384), the only safe remediation today is to drop and recreate the ContinuousQuery; do not attempt to "delete" the phantom row at the reaction layer because the next aggregation update will reintroduce it. Track #384 for a fix in drasi-core.
- Plugin directory listing pagination cap (~50 plugins). If
drasi list source/drasi list reactionor the equivalent API plugin enumeration returns suspiciously short or stable counts near 50, treat the listing as truncated, not authoritative. Until drasi-core #414 adds pagination, cap deployed plugin counts under 50 and inventory plugins from your manifests in source control, not from the live listing. - Reaction/Source pod
ImagePullBackOffafterdrasi apply. The provider catalog and image builds are decoupled,drasi list reactionprovider/sourceproviderlists definitions bundled into the release, but the image may not exist at the platform version tag. When a pod entersImagePullBackOffimmediately after apply: (a) rundrasi describe reactionprovider <name>(orsourceprovider) to find the image name fromspec.services.<service>.image, (b) verifyghcr.io/drasi-project/<image>:<platform-version>exists using the manifest-list Accept header (seebundles/reactions/guide.md), (c) if missing, either choose a different provider at the current version or verify ALL components exist at an older version before downgrading (cross-minor downgrade is a full migration, seebundles/recovery/guide.md).
For MCP-specific runbooks (503 retry-hint, 30-second cleanup, cursor-resume reconnect), see bundles/reactions/guide.md MCP failure handling section.
Numbered playbook - "Reaction healthy but no downstream effect"
Run in order; do not skip steps even if the answer "feels" obvious.
- Confirm the Reaction subscription is active. For Drasi Server,
curl -fsS "$DRASI_BASE_URL/api/v1/reactions/<name>" | jq '.status, .queries'; for Drasi for Kubernetes,drasi describe reaction <name>. Thequeriesarray MUST contain the expected query ID exactly. - Tail Reaction logs for outbound attempts.
kubectl logs -n "$DRASI_NAMESPACE" -l drasi/reaction=<name> --since=15m(or the equivalentdrasi logs reaction <name>). Confirm that for a known source change you see an outbound attempt log line with the query ID, change kind, and a target URL/identity. - Verify the downstream actually receives. Inspect downstream access logs, ingestion metric, or DLQ. "No log line on the downstream" is the failure shape; "log line shows 4xx/5xx" routes you to a different playbook.
- Check auth in isolation. Send a synthetic POST from inside the Drasi network namespace (
kubectl run curl --rm -it --image=curlimages/curl -- curl -sS -X POST <downstream-url> ...) using the same identity/token. Compare its outcome to the Reaction's outbound. A divergence points at Reaction-pod identity drift. - Confirm idempotency. Re-send a known payload (same idempotency key) and confirm the downstream suppresses the duplicate rather than processing twice. If duplicates are processed, the downstream contract is broken regardless of Drasi behaviour.
- Inspect DLQ depth. If the Reaction is configured with a dead-letter destination, depth > 0 means messages are being emitted but rejected - re-route to the 5xx playbook or downstream owner.
- Escalate to the
recoverybundle if Reaction state is suspected corrupted (e.g., subscription appears active in API but no outbound attempts logged for any change in a confirmed-fresh query).
Numbered playbook - "Reaction endpoint 401/403"
- Confirm the Reaction identity is the one expected. For Workload Identity:
kubectl get sa -n "$DRASI_NAMESPACE" reaction.<name> -o yamland confirm theazure.workload.identity/client-idannotation matches the managed identity recorded in the security evidence for this reaction. - Check the Workload Identity federated-credential subject matches the actual pod. Subject should be
system:serviceaccount:<namespace>:reaction.<name>(or the documented Drasi pattern). Subject drift typically follows a namespace rename, SA rename, or AKS cluster recreate. - Validate the token at the downstream. Capture the bearer token the pod is presenting (
kubectl execinto the reaction pod and read the projected token file) and decode it offline. Confirmaud,iss, andsubmatch the downstream's expected values. - Confirm a secret rotation overlap is in progress and was not abandoned. If the security bundle's rotation log shows a rotation started but not completed, the downstream may have already revoked the old credential while Drasi is still presenting it.
- If HMAC is in use: confirm the signing-key version (
X-Drasi-Signature-KeyId) matches a key the receiver currently accepts. Receiver-side log usually says "unknown key id" vs. "bad signature" - the two have different fixes. - If the incident's start time aligns with a change-control window, roll back the last identity/secret change first and re-test before further diagnosis. Forward fixes during an active outage tend to compound.
Query container OOM triage
- Confirm OOMKilled via
kubectl describe pod -n "$DRASI_NAMESPACE" <query-pod>- look forLast State: Terminated, Reason: OOMKilledand the exit code 137. Without this confirmation, do not treat the symptom as an OOM. - Capture last metrics and result-set size before restart:
kubectl logs --previousfor the query pod, plus any cached result size fromdrasi describe queryor the management API. Preserve these numbers in the incident record. - Confirm slot state was preserved on the source side. A query OOM does not by itself corrupt the source replication slot, but a forced delete/recreate of the source will. Point at
bundles/sources/guide.mdslot runbook before touching source state. - Choose remediation in this order: (a) raise the query container memory limit to a value supported by capacity headroom - see
bundles/scaling-and-capacity/guide.md; (b) shard the query (split by tenant, time window, or partition key) - seebundles/continuous-queries/guide.md; (c) reduce join cardinality or add a more selective predicate - alsobundles/continuous-queries/guide.md. Skip straight to (b) or (c) if the working set is unbounded; raising memory only delays the next OOM. - Restart the query and verify it becomes active again. Use
drasi describe query/ API status, and assert a known source change produces a result update within the freshness budget defined in## Defining "stale" and proving event flow. - Record the incident summary with concrete numbers: peak RSS observed, configured limit before and after, result-set cardinality, time-to-OOM, and the chosen remediation lever. Vague summaries ("increased memory") are not acceptable.
Identity-token-expiry runbook (Entra Workload ID)
Symptom
Sudden 401s from a Reaction or Source endpoint after a previously stable period - often "everything was fine last night, broke this morning". Affects whichever Drasi component federates with Entra (Reactions calling Azure Event Grid / SignalR, Sources reading Azure-hosted databases with managed identity, MCP reactions writing telemetry).
Diagnostic
- Check token TTL. Exec into the affected pod and inspect the projected service-account token file; decode and read the
expclaim. A normal token has minutes-to-an-hour of life; a token already pastexpmeans the workload-identity webhook isn't refreshing. - Check for federated-credential subject drift. If the namespace or service account was renamed (e.g., during a chart upgrade or a cluster rebuild), the federated credential on the Entra app/managed identity may still be pinned to the old
system:serviceaccount:<old-ns>:<old-sa>subject. Compare the credential'ssubjectagainstkubectl get sa -n "$DRASI_NAMESPACE" <sa-name>. - Check AKS OIDC issuer rotation.
az aks show -g <rg> -n <cluster> --query "oidcIssuerProfile.issuerUrl"- if this URL has changed (cluster recreate, region failover, manual rotate), every federated credential referencing the old issuer is now invalid.
Recovery
- Refresh / re-create the federated credential with the correct subject and current OIDC issuer URL.
- Restart the affected pods to force a new token acquisition:
kubectl rollout restart deployment/<reaction-or-source-deployment> -n "$DRASI_NAMESPACE". The workload-identity webhook injects a fresh token on pod start. - Re-validate end-to-end: trigger a known source change and assert the downstream side effect lands within the freshness budget, per
## Defining "stale" and proving event flow. - Record in the security bundle's rotation/identity log: which credential was repaired, the prior failing subject, the new subject, and the issuer URL in effect.
Distroless DNS resolution failures
Drasi containers built on gcr.io/distroless/cc (notably publish-api and query-host) can fail to resolve short Kubernetes service names like drasi-redis. The symptom is a startup crash with:
Error connecting to redis: failed to lookup address information: Name or service not knownor in the query-host:
thread panicked at ... query_worker.rs:157: called `Result::unwrap()` on an `Err` value: ConnectionError("Error connecting to redis: failed to lookup address information: Name or service not known")Root cause: Distroless base images ship a minimal libc that doesn't include the full DNS resolver chain (nsswitch, resolv.conf parsing) that standard glibc provides. Short DNS names (single-label) fail to resolve; fully qualified domain names (FQDN) work.
Fix: Set the Redis connection env var to the FQDN of the Drasi Redis service:
kubectl set env deployment/default-publish-api -n drasi-system \
REDIS_BROKER="redis://drasi-redis.drasi-system.svc.cluster.local:6379"For the query-host, patch the STORE_0_CONNECTION_STRING env var similarly:
kubectl set env deployment/default-query-host -n drasi-system \
STORE_0_CONNECTION_STRING="redis://drasi-redis.drasi-system.svc.cluster.local:6379"Note: The resource provider may overwrite these env vars on reconciliation. If the fix doesn't persist, add a post-install hook that re-applies the env var after every drasi apply cycle.
Client-only gRPC app-port deadlock
The Drasi source reactivator (Debezium-based CDC) is a gRPC client only, it connects to the Dapr sidecar to save CDC offsets but never listens on any port itself. If dapr.io/app-port is set (e.g. 80), the Dapr sidecar waits for the app to be ready on that port before completing startup. The app meanwhile waits for Dapr's gRPC endpoint (localhost:50001). This creates a deadlock:
Reactizator: connect to Dapr gRPC → Connection refused on :50001
Dapr sidecar: wait for app ready on :80 → blockedSymptom: The reactivator pod shows 1/2 Ready (only Dapr sidecar is up), the main container repeatedly crashes with:
Caused by: java.net.ConnectException: finishConnect(..) failed: Connection refused
... at /127.0.0.1:50001Fix: Remove dapr.io/app-port from the deployment annotation entirely:
kubectl patch deployment stepup-reactivator -n drasi-system --type json \
-p '[{"op":"remove","path":"/spec/template/metadata/annotations/dapr.io~1app-port"}]'Also set dapr.io/app-protocol: grpc so Dapr knows the app communicates via gRPC:
kubectl patch deployment stepup-reactivator -n drasi-system --type json \
-p '[{"op":"add","path":"/spec/template/metadata/annotations/dapr.io~1app-protocol","value":"grpc"}]'Verify: After the fix, the pod should reach 2/2 Ready and the reactivator logs should show Processing messages (indicating CDC streaming is active).
"Query alternating running/inactive" close look
Augments the one-line entry in ## Defining "stale" and proving event flow. A query that flips between running and inactive is rarely a query bug - it is almost always a symptom of an unstable dependency or undersized resources.
Signals to inspect
- Operator / control-plane pod logs filtered by the query name.
kubectl logs -n "$DRASI_NAMESPACE" -l app=drasi-operator --since=30m | grep "<query-id>"(or equivalent for the runtime form). Look for repeated state-transition lines, reconciler errors, and any "source unavailable" / "reconnect" mentions tied to the query's sources. - Query container restart count.
kubectl get pod -n "$DRASI_NAMESPACE" -l drasi/query=<id> -o jsonpath='{.items[*].status.containerStatuses[*].restartCount}'. A non-zero and growing count alongside the active/inactive flap points at OOM, liveness probe failure, or crash-loop - route to the OOM triage above or tobundles/scaling-and-capacity/guide.md. - Source-event arrival rate. If the upstream source is itself bouncing (CDC reconnects, replication slot churn), the query's "inactive" windows will line up with source gaps. Cross-reference with the source's connection log.
- Downstream-reaction error rate. A reaction that is rejecting deliveries can apply backpressure that surfaces upstream as query instability in some Drasi configurations. Check whether reaction error spikes precede query state flips.
Remediation order
- Stabilize the source first. A flapping source will produce a flapping query no matter what you do to the query container. Use
bundles/sources/guide.mdto fix CDC, slot, or credential issues before touching the query. - Stabilize the query container resources second. Address OOM, CPU starvation, or liveness-probe tuning per the OOM triage above and
bundles/scaling-and-capacity/guide.md. Do not raise resources blindly; capture peak usage first. - Re-validate end-to-end last. Trigger a known source change, assert the query stays
runningfor at least one full freshness window with no state transitions, and confirm the downstream reaction observes the change. Only then close the incident.
Related bundles
references/schema-evolution-and-failure-patterns.md, handling source schema changes, query sprawl, zombie queries, reaction stormsreferences/schema-evolution-and-failure-patterns.md, handling source schema changes, query sprawl, zombie queries, reaction stormsreferences/schema-evolution-and-failure-patterns.md, handling source schema changes, query sprawl, zombie queries, reaction stormsreferences/schema-evolution-and-failure-patterns.md, handling source schema changes, query sprawl, zombie queries, reaction storms