All skills

USE FOR: Drasi continuous-query solutions - real-time queries, change detection, reactive events, data-trigger pipelines on Drasi Server, Drasi for Kubernetes, or drasi-lib. Router: load bundle guides as needed. DO NOT USE for non-Drasi messaging (event-driven-messaging) or pure AKS/ACA hosting (aks-cluster-architecture, azure-container-apps).

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/drasi

This session only. Nothing lands on disk.

bundlesreactionsguide.md

≈12k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Reactions bundle

Use this bundle for Drasi Reaction setup, delivery contracts, webhooks, Debug, Result, Event Grid, SignalR, MCP, custom reactions, and downstream side-effect validation.

Reaction authoring rules

  1. Verify the selected Reaction schema against current Drasi docs.
  2. Verify the reaction provider image exists at the platform version tag on GHCR BEFORE writing YAML or committing to a provider. drasi list reactionprovider returns every provider definition bundled into the drasi init release, but the matching container image may not have been built for that version. A provider appearing in the catalog and being fully documented at drasi.io is NOT evidence its image exists (see ## Reaction provider image verification gate below).
  3. Subscribe only to queries that exist and have a documented result contract.
  4. Define added, updated, and deleted behavior explicitly.
  5. Make downstream effects idempotent where possible.
  6. Protect reaction endpoints with auth and network controls.
  7. Redact secrets and sensitive payload fields from logs.
  8. Validate the downstream side effect, not only the reaction status.
  9. Design retry, backoff, and failure handling based on downstream system behavior.
  10. For kind: Http, use properties.baseUrl and per-query added|updated|deleted blocks under queries. Do not place url/method directly under properties.
  11. For Azure-hosted reactions using Workload Identity, reactions use a per-reaction service account in the Drasi namespace: reaction.<reaction-name>. The federated credential subject is system:serviceaccount:<drasi-namespace>:reaction.<reaction-name>. You must also patch the reaction deployment's pod template with the azure.workload.identity/use: "true" label. SA annotation alone is insufficient. See ## Reaction workload identity on AKS below.
  12. Prefer native Dapr reactions over custom bridge services. Before writing a custom HTTP endpoint to receive Drasi reaction output and forward it to a message broker, check whether a native reaction exists. PostDaprPubSub, SyncDaprStateStore, EventGrid, and SignalR reactions publish directly to downstream systems, no intermediate service needed. See ## PostDaprPubSub reaction (direct-to-Service-Bus / Event-Hub) below.

Reaction provider image verification gate

The provider catalog and image builds are decoupled in Drasi's release process. A provider definition ships with the drasi init release, but the CI pipeline for that provider's image may not have run for every platform release tag.

Production-verified example (2026-06-23; GHCR gap re-verified against the live tag list 2026-08-16): The PostDaprPubSub reaction provider IS registered under Drasi for Kubernetes 0.10.0 (drasi list reactionprovider shows it) and fully documented in the drasi.io how-to guides (https://drasi.io/drasi-kubernetes/how-to-guides/configure-reactions/configure-post-pubsub-reaction/). However, ghcr.io/drasi-project/reaction-post-dapr-pubsub:0.10.0 does NOT exist on GHCR, the image tag stops at 0.9.2. drasi apply kind: PostDaprPubSub against a 0.10.0 cluster would fail with ImagePullBackOff.

The fix: externalImage: true (verified against 0.10.0 source code). Drasi resolves provider images two ways:

  • externalImage: true: the image field is used verbatim as a fully-qualified image reference, no prefix/tag resolution applied.
  • externalImage: true: the image field is used verbatim as a fully-qualified image reference, no prefix/tag resolution applied.

Register a custom provider with the version tag that does exist:

apiVersion: v1
kind: ReactionProvider
name: PostDaprPubSub
spec:
  services:
    reaction:
      image: ghcr.io/drasi-project/reaction-post-dapr-pubsub:0.9.2
      externalImage: true

Apply this with drasi apply -f <file>, then apply your reaction normally. No platform downgrade required.

Before committing to any provider, verify the image exists:

$accept = "application/vnd.docker.distribution.manifest.list.v2+json, application/vnd.docker.distribution.manifest.v2+json, application/vnd.oci.image.index.v1+json, application/vnd.oci.image.manifest.v1+json"
$t = (Invoke-RestMethod -Uri "https://ghcr.io/token?scope=repository:drasi-project/<image-name>:pull&service=ghcr.io").token
try {
    Invoke-WebRequest -Uri "https://ghcr.io/v2/drasi-project/<image-name>/manifests/<version-tag>" -Headers @{Authorization="Bearer $t"; Accept=$accept} -UseBasicParsing | Out-Null
    "EXISTS"
} catch {
    "MISSING - use externalImage: true with an older tag that exists"
}

Use the manifest-list media types (the $accept header above). GHCR returns 404 for multi-arch images when queried with only application/vnd.docker.distribution.manifest.v2+json.

To find the image name for a provider, run drasi describe reactionprovider <ProviderName> and read the spec.services.<service>.image field, then prepend ghcr.io/drasi-project/ and append :<platform-version>.

Reaction workload identity on AKS

Verified against Drasi for Kubernetes 0.10.0 with PostDaprPubSub reaction.

Two layers of identity, don't conflate them

  1. Reaction spec.identity — the identity declared on the Reaction resource for the reaction's own source/target (e.g. when the reaction needs to authenticate to a downstream Azure resource directly). This is a top-level spec field, verified against Drasi 0.10.0 openapi ServiceIdentityDto.

  2. Dapr sidecar workload identity — when a reaction like PostDaprPubSub uses a Dapr sidecar to publish to Azure Service Bus / Event Hub / Storage, it is the Dapr sidecar that authenticates to Azure, not the reaction container. The sidecar needs workload identity env vars (AZURE_CLIENT_ID, AZURE_TENANT_ID, AZURE_FEDERATED_TOKEN_FILE) injected by the Azure Workload Identity mutating webhook.

Critical: the pod template label, not just the SA annotation

The Azure Workload Identity webhook injects env vars only when the pod template has the azure.workload.identity/use: "true" label. Annotating the service account alone is not sufficient, the webhook keys off the pod template label.

Drasi manages reaction deployments via its own controller. You must patch the deployment's pod template to add the label, because drasi apply does not add it automatically:

# 1. Annotate the per-reaction SA (verify the exact SA name for your reaction)
kubectl annotate sa reaction.<reaction-name> -n <drasi-namespace> azure.workload.identity/client-id=<mi-client-id>
kubectl label sa reaction.<reaction-name> -n <drasi-namespace> azure.workload.identity/use=true

# 2. Patch the deployment pod template with the WI label (THIS IS THE MISSING STEP)
kubectl patch deploy <reaction-name>-reaction -n <drasi-namespace> \
  --type=json -p='[{"op":"add","path":"/spec/template/metadata/labels/azure.workload.identity~1use","value":"true"}]'

# 3. Wait for rollout to complete
kubectl rollout status deploy <reaction-name>-reaction -n <drasi-namespace>

# 4. Verify env vars were injected
kubectl get pod -n <drasi-namespace> -l drasi/resource=<reaction-name> -o yaml | grep AZURE_CLIENT_ID

If you skip step 2, the reaction container starts but the Dapr sidecar falls back to IMDS and fails with "Identity not found" / 400 Bad Request from 169.254.169.254.

Which service account do reactions use?

Verified against Drasi for Kubernetes 0.10.0: reactions use a per-reaction service account in the Drasi namespace, named reaction.<reaction-name>. This differs from the per-source SA pattern (source.<source-name>), but both are per-resource. The federated credential subject is system:serviceaccount:<drasi-namespace>:reaction.<reaction-name>.

This may be version-specific, always verify with:

kubectl get pod -n <drasi-namespace> -l drasi/resource=<reaction-name> -o jsonpath='{.items[0].spec.serviceAccountName}'

PostDaprPubSub + Service Bus full checklist

For a PostDaprPubSub reaction publishing to Azure Service Bus via Dapr:

  1. Register the PostDaprPubSub provider with externalImage: true (see ## Reaction provider image verification gate).
  2. Create a user-assigned managed identity and grant it Azure Service Bus Data Sender on the Service Bus namespace.
  3. Create a federated credential for subject system:serviceaccount:<drasi-namespace>:reaction.<reaction-name>.
  4. Annotate and label the reaction.<reaction-name> SA with workload identity.
  5. Apply a Dapr Component (pubsub.azure.servicebus) in the Drasi namespace with azureClientId set to the managed identity client ID.
  6. Apply the PostDaprPubSub reaction referencing that Dapr component (pubsubName must match the component metadata.name).
  7. Patch the reaction deployment's pod template with the azure.workload.identity/use: "true" label.
  8. Restart the reaction pods and verify the Dapr sidecar logs show no "Identity not found" errors.

PostDaprPubSub reaction (direct-to-Service-Bus / Event-Hub)

The PostDaprPubSub reaction is the native Dapr-based message broker integration for Drasi. It eliminates the need for custom bridge/forwarder services by publishing query results directly to Azure Service Bus, Event Hubs, Kafka, RabbitMQ, or any Dapr-supported pub/sub component via a Dapr sidecar.

When to use it (and when it replaces custom code)

Instead of writing Use PostDaprPubSub with
Custom HTTP endpoint receiving Drasi reaction POSTs → forwarding to Service Bus Dapr pubsub.azure.servicebus component
Custom HTTP endpoint → forwarding to Event Hubs Dapr pubsub.azure.eventhubs component
Custom HTTP endpoint → forwarding to Kafka Dapr pubsub.kafka component
Custom HTTP endpoint → forwarding to RabbitMQ Dapr pubsub.rabbitmq component

Architecture it replaces:

Before (custom bridge — more moving parts):
  Drasi query → HTTP Reaction → custom HTTP service → SDK → Service Bus
  (requires: building/pushing a Docker image, managing a deployment,
   managing a managed identity, health checks, scaling, debugging)

After (native Dapr — Drasi-managed):
  Drasi query → PostDaprPubSub Reaction → Dapr sidecar → Service Bus
  (requires: Dapr component YAML + workload identity setup)

How it works internally

Drasi ContinuousQuery detects a change
  → Query result published to internal Redis pubsub (drasi-pubsub-<reaction-name>)
  → Dapr sidecar subscribes to that Redis pubsub
  → Dapr sidecar calls the reaction's HTTP endpoint (/eventhandler)
  → Reaction container calls Dapr Publish API on the configured pubsubName
  → Dapr sidecar authenticates to Azure (Workload Identity) and publishes

The pubsubName in the reaction's per-query config must match a Dapr Component metadata.name in the Drasi namespace. The topicName maps to the Service Bus queue/topic name or Event Hub name.

Per-query configuration

Each query gets its own config block specifying the Dapr component and target:

apiVersion: v1
kind: Reaction
name: forward-to-servicebus
spec:
  kind: PostDaprPubSub
  queries:
    order-shipped: >
      {
        "pubsubName": "drasi-servicebus",
        "topicName": "order-events",
        "format": "Unpacked",
        "skipControlSignals": true
      }
    order-status-change: >
      {
        "pubsubName": "drasi-servicebus",
        "topicName": "order-events",
        "format": "Unpacked",
        "skipControlSignals": true
      }

Format options:

  • Packed, sends all results from a single change event as one message array.
  • Packed, sends all results from a single change event as one message array.

skipControlSignals: true suppresses Drasi control messages (bootstrap, query restart) so only actual data changes are published.

Message format caveat for downstream consumers

The Unpacked format sends the query's RETURN fields directly (e.g. {"matchId":"M001","homeScore":3,...}). Downstream consumers expecting a different envelope format (e.g. a wrapper with Envelope, Payload, BufferedAtUtc) will need to either:

  • Be updated to accept the raw query result shape, or
  • Use Packed format and add a mapping layer, or
  • Accept the impedance mismatch and handle deserialization accordingly

This is a consumer-side concern, not a Drasi issue, but plan for it when migrating from a custom bridge that transformed the payload.

Prerequisites

  1. Dapr must be installed in the cluster (dapr init -k or via Drasi's Dapr dependency). Drasi itself installs Dapr as a dependency for source/reaction sidecars.
  2. Dapr Component YAML for the target pub/sub (Service Bus, Event Hubs, etc.) must be applied to the Drasi namespace, not the application namespace.
  3. Workload identity must be configured on the reaction pod (see ## Reaction workload identity on AKS above and the checklist).

Verified Dapr pub/sub components

Dapr component type Azure target Verified
pubsub.azure.servicebus Azure Service Bus queues and topics ✅ Production-verified 2026-06-23
pubsub.azure.eventhubs Azure Event Hubs [VERIFY] — same workload identity pattern
pubsub.kafka Apache Kafka / Confluent [VERIFY] — SASL/SCRAM auth
pubsub.rabbitmq RabbitMQ [VERIFY] — connection string auth

Reaction selection

Need Prefer
Local or development visibility Debug Reaction or result polling
Push to HTTP service HTTP Reaction with auth, idempotency, timeout, and retry strategy
Azure event distribution Azure Event Grid Reaction with documented identity support
AWS event distribution AWS EventBridge Reaction with IRSA or access key auth
Real-time UI updates SignalR Reaction with documented identity support
Agent/MCP integration MCP Reaction, plus agent-integration bundle
Vector store / RAG sync SyncVectorStore Reaction with Qdrant, Azure AI Search, or InMemory
State store sync (Dapr) SyncDaprStateStore Reaction
Pub/sub messaging (Dapr) PostDaprPubSub Reaction
Kafka / Debezium ecosystem Debezium Reaction after schema review
Graph database sync Gremlin Reaction
Database side effects StoredProc Reaction (MSSQL, MySQL, PostgreSQL)
Azure queue storage Azure StorageQueue Reaction

SignalR Reaction: in-process vs Azure SignalR Service modes

The SignalR Reaction supports two hosting modes. Choose based on your deployment and scale requirements.

Mode When to use Pros Cons
In-process (omit connectionString) Local development, small-scale deployments, or when you want to avoid an Azure SignalR Service dependency No external dependency; lower latency; simpler configuration Connections are bound to the reaction pod's memory; scaling requires sticky sessions or a Redis backplane
Azure SignalR Service (set connectionString) Production, multi-replica, or when clients need to reconnect across pod restarts Scales independently; serverless-friendly; handles reconnection and broadcast Azure dependency and cost; additional latency hop

In-process configuration, simply omit properties.connectionString:

kind: Reaction
apiVersion: v1
name: my-signalr-reaction
spec:
  kind: SignalR
  queries:
    my-query:

Azure SignalR Service configuration:

kind: Reaction
apiVersion: v1
name: my-signalr-reaction
spec:
  kind: SignalR
  properties:
    connectionString: "Endpoint=https://<resource>.service.signalr.net;AccessKey=<key>;Version=1.0;"
  queries:
    my-query:

For managed-identity-based auth to Azure SignalR Service, use:

properties:
  connectionString: "Endpoint=https://<resource>.service.signalr.net;AuthType=azure"

And configure the identity block per the Azure identity guidance.

Client connection: In both modes, clients connect via drasi tunnel reaction <name> <port> (development) or directly to the service endpoint (production). The reaction exposes a SignalR hub per subscribed query, and the client library (@microsoft/signalr for JS, Microsoft.AspNetCore.SignalR.Client for .NET) connects to /query/<query-id>.

SignalR reaction actor limitation (pre-1.0)

The SignalR reaction requires the query host to support a ReactionResource actor type that handles the configure method. When the Drasi API applies a SignalR reaction, the resource provider calls PUT /actors/ReactionResource/<name>/method/configure on the query host to set up the reaction. If the query host binary doesn't include SignalR actor support (e.g. main-branch builds where the actor plugin isn't compiled in), this call fails with:

Error configuring resource: GrpcError { ... "error invoke actor method: Put \"http://127.0.0.1:8080/actors/ReactionResource/dashboard/method/configure\": EOF" }

This is a compile-time limitation, not a configuration error. The query-host binary either includes the SignalR reaction actor or it doesn't, there is no runtime flag or env var to enable it. If the query-host image was built without SignalR support, no configuration change can fix it.

Workarounds (none guaranteed):

  • Use a different query-host image tag that includes SignalR support. In the drasi-platform repo, SignalR reaction support is compiled in when the signalr feature flag is enabled in query-container/query-host/Cargo.toml.
  • Replace the SignalR reaction with an HTTP reaction connected to a standalone SignalR hub, or with the PostDaprPubSub reaction feeding a custom bridge service.

Verify: Run drasi apply -f signalr-reaction.yaml. If it succeeds (HTTP 200), the query host supports SignalR. If it returns HTTP 500 with the actor EOF error, the query host was built without SignalR support.

SignalR consumer reference implementation

The drasi-project/grafana-signalr repository provides a real-world SignalR consumer as a Grafana data source plugin. It validates:

  • Operation types: Drasi SignalR reactions use single-letter operation codes, i (insert/added), u (update), d (delete), x (control). These differ from the MCP reaction's added/updated/deleted strings. Ensure your SignalR client handles all four.
  • Snapshot loading: The plugin calls a snapshot/query endpoint on initial connect to recover current state before streaming incremental changes. This is the SignalR equivalent of MCP's resources/read.
  • Error handling: The plugin documents connection failure, missing query ID, and data processing error paths.

Cross-reference: github.com/drasi-project/grafana-signalr for the full implementation and operational experience.

SyncDaprStateStore reaction

Synchronises the current result set of a Continuous Query to a Dapr state store component. Each item in the result set is stored as a separate key-value entry, keyed by the query result's identity column.

Per-query configuration

kind: Reaction
apiVersion: v1
name: my-dapr-state-reaction
spec:
  kind: SyncDaprStateStore
  properties:
    stateStoreName: my-statestore
  queries:
    my-query-id:

Properties

Property Type Required Default Description
stateStoreName string Yes — The metadata.name of the Dapr state store component. The component MUST be defined in the Drasi namespace (drasi-system by default) and point to the same state store used by your application.

State store behaviour

  • Added results → new key-value entry created.
  • Updated results → existing entry overwritten with upsert semantics.
  • Deleted results → existing entry deleted from the state store.
  • The key is derived from the query result's identity, typically the first RETURN column. Verify the identity column is stable and unique before using this reaction.

Prerequisites

  1. Dapr must be installed in the cluster.
  2. Dapr Component YAML for the target state store (e.g. state.azure.cosmosdb, state.redis) must be applied to the Drasi namespace.
  3. Workload identity must be configured on the reaction pod if the state store uses Azure AD auth (see ## Reaction workload identity on AKS).

Validation

Validation Method
Reaction is healthy drasi list reaction shows status Available
State store contains expected keys Direct read from the Dapr state store API (localhost:<dapr-port>/v1.0/state/<stateStoreName>/<key>) returns the current query result for each identity
Deleted results are removed Read the key for a deleted result — the store returns a null or 404
Upsert idempotency Repeated application of the same query result produces the same store entry without errors

SyncVectorStore reaction

The SyncVectorStore reaction continuously synchronizes ContinuousQuery results to a vector database, generating embeddings for each result row and creating searchable vector collections. This is the primary reaction type for AI/RAG integration patterns.

Supported backends

Backend Type Use case
Qdrant Persistent vector DB Production, high-scale semantic search
Azure AI Search Cloud search service Azure-native, hybrid (vector + full-text) search
InMemory In-process memory Development, testing, demos (data lost on restart)

Embedding service configuration

Setting Required Default Notes
embeddingServiceType Yes — Currently only AzureOpenAI and OpenAI supported
embeddingEndpoint Yes — Full URL of the embeddings API endpoint
embeddingApiKey Yes — API key; supports inline value or kind: Secret reference
embeddingModel Yes text-embedding-3-large Model or deployment name
embeddingDimensions Yes 3072 Vector dimension count (max 20000)

Vector store configuration

Setting Required Default Notes
vectorStoreType Yes — Qdrant, AzureAISearch, or InMemory
connectionString Yes — Format varies by backend (see examples)
distanceFunction No CosineSimilarity Also: CosineDistance, EuclideanDistance, DotProductSimilarity, ManhattanDistance
indexKind No Hnsw Also: Flat, IvfFlat, DiskAnn
isFilterable No true Enables field-based filtering on vector search
isFullTextSearchable No false Full-text search on content fields (Azure AI Search only)

Per-query configuration

Each query maps to one vector collection with a Handlebars template for document formatting:

queries:
  my-query-id: |
    {
      "collectionName": "products",
      "keyField": "product_id",
      "documentTemplate": "Product: {{name}}\nDescription: {{description}}",
      "titleTemplate": "{{name}}",
      "vectorField": "content_vector",
      "createCollection": true
    }
Setting Required Notes
collectionName Yes Name of the collection in the vector store
keyField Yes Field from the query result used as the document key
documentTemplate Yes Handlebars template for the text to embed
titleTemplate No Optional title for search result display
vectorField No Default: embedding or backend-specific default
createCollection No true to auto-create on first sync

Configuration example (Qdrant)

apiVersion: v1
kind: Reaction
name: product-catalog-qdrant
spec:
  kind: SyncVectorStore
  properties:
    vectorStoreType: "Qdrant"
    connectionString: "Endpoint=qdrant-service:6334"
    embeddingServiceType: "AzureOpenAI"
    embeddingEndpoint: "https://your-resource.openai.azure.com/"
    embeddingApiKey:
      kind: Secret
      name: azure-openai-creds
      key: api-key
    embeddingModel: "text-embedding-3-large"
    embeddingDimensions: 3072
    distanceFunction: "CosineSimilarity"
    indexKind: "Hnsw"
    isFilterable: true
  queries:
    product-catalog: |
      {
        "collectionName": "products",
        "keyField": "id",
        "documentTemplate": "{{name}}: {{description}}",
        "titleTemplate": "{{name}}",
        "vectorField": "embedding",
        "createCollection": true
      }

Configuration example (Azure AI Search)

spec:
  kind: SyncVectorStore
  properties:
    vectorStoreType: "AzureAISearch"
    connectionString: "Endpoint=https://your-search.search.windows.net;ApiKey=your-key"
    embeddingServiceType: "AzureOpenAI"
    embeddingEndpoint: "https://your-resource.openai.azure.com/"
    embeddingApiKey:
      kind: Secret
      name: azure-openai-creds
      key: api-key
    embeddingModel: "text-embedding-3-large"
    embeddingDimensions: 3072
    isFullTextSearchable: true       # Azure AI Search supports hybrid search
  queries:
    products: |
      {
        "collectionName": "products",
        "keyField": "product_id",
        "documentTemplate": "Product: {{name}}\n{{description}}",
        "vectorField": "content_vector",
        "createCollection": true
      }

Validation

Check Method
Reaction is healthy drasi list reaction shows Ready
Vector store has documents Query the vector store collection count (Qdrant /collections/{name}/points/count, Azure AI Search count parameter)
Document content matches query result Read a document from the vector store and verify fields match the Handlebars template output
Embedding dimensions match config Verify vector dimension matches embeddingDimensions setting
All three change types propagate Insert, update, and delete a source row; assert added/updated/deleted in the vector store

Reliability defaults

Property Baseline
Embedding retry Bounded retries (3 attempts) on embedding service 5xx and 429 with exponential backoff
Vector store retry Bounded retries (3 attempts) on write failures
Delivery guarantee At-least-once per result change; duplicate upserts to the vector store MUST be idempotent (use keyField as the upsert key)
Backpressure Reaction syncs asynchronously; vector store writes do not block the source change pipeline
Error handling Failed embeddings are logged with query ID and result key; a failed vector write is retried then logged; poison documents are not automatically dropped

See bundles/agent-integration/guide.md for the consumer-side architecture that reads from the vector store.

Output format patterns

Multiple Drasi reactions (EventGrid, EventBridge, StorageQueue, and custom reactions using the same pattern) support three output formats that control how query results are packaged into events.

Format comparison

Format Behaviour Event count per change Use when
packed All query result rows are bundled into a single event on any change 1 event per source change Downstream wants a complete snapshot; fan-out is acceptable; latency is more important than per-row granularity
unpacked Each added/updated/deleted row produces a separate event N events per source change (one per affected row) Downstream processes each row individually; per-row idempotency is needed; event-driven consumers that map one event to one action
template A Handlebars template produces a custom-shaped event from the query result 1 event per source change Downstream requires a specific schema; field renaming, filtering, or enrichment is needed before delivery

When to use which

Use packed when:

  • The downstream is a cache or state store that replaces its entire dataset on each change.
  • The downstream can efficiently diff the full payload against its current state.
  • Event volume must be minimized (each source change = 1 event, not N).

Use unpacked when:

  • Each result row maps to an independent downstream action (email per new order, alert per new sensor reading).
  • The downstream needs per-row idempotency keys.
  • A single source change affects a small, predictable number of rows.

Use template when:

  • The downstream schema does not match the query result contract.
  • You need to rename fields, combine fields, or add computed fields before delivery.
  • You need to filter out fields from the result contract before they reach the downstream.
  • The same query result needs different shapes for different reactions.

Template availability

Reaction kind packed unpacked template
EventGrid Yes (default) Yes Yes
EventBridge Yes (default) Yes Yes
StorageQueue Yes (default) Yes Yes
Http Per-query body template Per-query added/updated/deleted blocks Inherits query-level Handlebars
SyncVectorStore Not applicable (always per-document) Not applicable Per-query documentTemplate

Per-query template config on EventGrid and EventBridge

EventGrid and EventBridge reactions accept Handlebars templates per query with per-change-type blocks (verified against drasi-platform PR #343/#344, merged 2025-12/2026-01; shape confirmed from PR body 2026-08-24):

kind: Reaction
name: my-eventgrid-reaction
spec:
  kind: EventGrid
  # ... connection fields ...
  format: template
  queries:
    my-query: |
      added:
        template: |
          { "eventType": "ItemAdded", "data": {{json after}} }
        metadata:            # becomes CloudEvent extension attributes
          changeType: added
      updated:
        template: |
          { "before": {{json before}}, "after": {{json after}} }
      deleted:
        template: |
          { "id": "{{before.id}}" }

Rules: each change-type block has template (Handlebars string) and optional metadata (key-value pairs applied as CloudEvent extension attributes); templates receive after (result data for added/updated) and before (previous result for updated/deleted); {{json <expr>}} emits JSON-safe output. Verify the exact field names against the pinned platform release docs before writing production YAML; this shape ships in the platform source, not yet in a tagged release newer than 0.10.0.

Per-reaction state store (direction of travel)

Platform PR #442 (merged 2026-08-15, not in any tagged release as of 2026-08-24) adds per-reaction Dapr state store support: reaction providers can declare state_store: true, the platform provisions and cleans up a per-reaction Dapr state store component, and injects a StateStoreName into reaction services. Do not write YAML against this surface until it lands in a release you pin; re-check the release notes when upgrading past 0.10.0.

Reliability contract

Every production reaction should document:

  • Query subscriptions.
  • Payload template and schema.
  • Auth mechanism.
  • Retry behavior.
  • Timeout behavior.
  • Idempotency key or duplicate handling.
  • Dead-letter or manual replay process, where supported.
  • Data classification and redaction rules.
  • Downstream ownership and operational contact.

Restart durability: since the 2026-08-20 drasi-core change (PR #735, in the Server 0.2.2 line), the query-result sequence clock survives restarts via the configured state store; duplicate suppression no longer silently resets on pod restart. Details and re-verification guidance: bundles/recovery/guide.md "Restart semantics per runtime form".

Token and credential rotation

Use overlap rotation for webhook tokens and downstream credentials:

  1. Add new credential to downstream service while old credential remains valid.
  2. Update Reaction configuration to send the new credential.
  3. Validate successful delivery.
  4. Remove old credential.
  5. Record rotation evidence.

Do not perform a single cutover if failed delivery would be user-visible or hard to replay.

Azure identity guidance

For Azure Event Grid, SignalR, and other reactions using Microsoft Entra Workload ID, current Drasi docs document Workload ID flows. Verify current docs and ensure:

  • AKS Workload Identity is enabled.
  • Federated credential subject matches system:serviceaccount:drasi-system:reaction.<reaction-name> (see ## Reaction workload identity on AKS above for the full apply-annotate-restart sequence).
  • The managed identity has the minimum role, for example EventGrid Data Contributor, SignalR App Server, or Azure Service Bus Data Sender where applicable.
  • Role assignment propagation is accounted for in validation.

Logging and data minimization

Never log:

  • Raw Authorization headers.
  • Webhook tokens.
  • Database passwords.
  • Full payloads containing PII or sensitive operational data unless explicitly approved and redacted.

Prefer structured logs with correlation ids, query name, operation type, delivery status, latency, and redacted error details.

Validation

A reaction is only complete when all are true:

  • Reaction resource is healthy.
  • It subscribes to the expected query.
  • A synthetic or real source change produces an expected reaction payload.
  • The downstream system accepts and processes the payload.
  • Failure behavior is tested or documented.

Per-reaction validation patterns

"Reaction delivered" must be defined per reaction type. Asserting only that Drasi sent the event is not sufficient - verify the downstream side effect.

Reaction kind Validation method Sample assertion
Debug / Log Inspect Drasi logs after a trigger drasi logs reaction <id> --since 2m contains the expected change payload.
HTTP / Webhook Capture the receiver call Receiver returns 2xx; log line on the receiver shows the query ID, change kind, and a stable payload hash matching the query result.
gRPC Receiver-side metric or log gRPC server records a unary/stream call with the expected method and payload.
Server-Sent Events (SSE) Open the SSE stream and trigger a change curl -N http://<host>/api/v1/queries/<id>/stream emits a message event matching the query result within p95 ≤ 30 s.
Azure Event Grid Event Grid subscription delivery Event arrives on the subscription endpoint (e.g., Service Bus, Function) within the delivery SLA; eventType matches; data matches the query result contract.
Azure SignalR Connected client receives message A test SignalR client connected to the hub receives the expected message group/event within p95 ≤ 30 s.
MCP Reaction MCP client receives notification The MCP client gets notifications/resources/updated for drasi://query/<id> and resources/read returns the new result.
SyncVectorStore Vector store has documents matching query Query vector store collection count; assert document fields match Handlebars template output; assert all three change types propagate
AWS EventBridge CloudEvent arrives on EventBus EventBridge target (e.g., Step Function, Lambda) receives the expected CloudEvent with matching payload
Dapr StateStore Dapr state store has expected key Read state from Dapr state store component; key matches keyField and value matches query result
Dapr PubSub Message arrives on Dapr pub/sub topic Dapr subscriber receives the expected CloudEvent with matching topic and payload
Debezium Kafka topic receives change event Kafka consumer on the configured topic receives a Debezium-compatible payload event matching the query result
StorageQueue Message appears in Azure Storage Queue az storage message peek on the target queue shows a message matching the expected format
StoredProc Target database shows side effect Query the target stored procedure's effect table (e.g., audit log, summary row) and assert the expected state change

Http Reaction manifest shape guardrail

When authoring Drasi for Kubernetes kind: Http reactions, this shape is required:

kind: Reaction
apiVersion: v1
name: my-http-reaction
spec:
  kind: Http
  properties:
    baseUrl: "https://api.example.com"
    token: "<optional-bearer-token>"
  queries:
    my-query-id: >
      added:
        url: "/api/path"
        method: "POST"
        body: |
          { "queryId": "my-query-id", "changeKind": "added", "result": {{json after}} }
        headers:
          Content-Type: "application/json"

Failure pattern from field incidents:

  • If drasi apply reaction returns 400 ... missing field kind, check that spec.kind exists and that queries values are structured mappings (not empty/null entries).

Required evidence per release

For every reaction, record in the acceptance evidence (templates/drasi-acceptance-evidence.md):

  • The reaction's idempotency strategy (e.g., dedup key, target-side upsert).
  • The configured retry policy and timeout.
  • A test trigger timestamp and the downstream observation timestamp.
  • The verification that retried/duplicate deliveries are absorbed by the downstream system without double effects.

Common reaction failure modes

Symptom Likely cause First action
Reaction state shows running, no downstream effect Reaction not subscribed to the right query Confirm reactions[].queries lists the query ID exactly.
Sporadic deliveries with same payload hash Source emitting duplicate change notifications Verify dedup on the reaction's downstream consumer.
Reaction errors only under load Downstream rate limit or auth-token expiry race Lower maxReplicas, rotate identity, confirm token lifetime.
Reaction stops emitting after secret rotation Reaction container did not pick up the new secret Confirm secret reload strategy (pod restart vs. env reload) for the runtime form in use.

Per-reaction reliability defaults (baseline contract)

Treat the values below as a required contract the reaction author MUST document and validate before the reaction goes to production. The reaction author MUST establish concrete numbers per row; the table names what each row SHOULD specify, not what Drasi necessarily ships. If the upstream Drasi release publishes contradicting defaults, the published value wins - record both in the validation evidence.

Reaction kind Default retry policy expected Backoff curve expected Idempotency requirement DLQ destination
HTTP webhook Bounded retries (e.g., 5 attempts) on 5xx and 408/429; no retry on 4xx other than 408/429 Exponential with full jitter, capped (e.g., base 1 s, max 30 s) Receiver MUST dedupe on a stable key derived from query ID + change kind + payload hash Operator-owned durable store (e.g., Service Bus dead-letter, S3 prefix); poison messages preserved with original headers
Event Grid Honor Event Grid retry schedule; do not retry locally on top of it Event Grid published curve (publisher MUST NOT add its own backoff) Event ID + eventTime treated as unique by subscriber Event Grid dead-letter container in storage
SignalR Best-effort, no retry for transient client disconnects; reconnect handled by client None at publisher; client SHOULD reconnect with exponential backoff Client MUST tolerate redelivery of the last N messages on reconnect Not applicable for fan-out; missed messages require client re-query against the result endpoint
MCP No silent retry; surface upstream errors to the MCP client Client-driven; server returns 503 with retry hint when query briefly inactive resources/read MUST be safe to call repeatedly; result is the current snapshot None; the MCP client recovers by re-reading after notifications/resources/updated
SSE No retry at server; client SHOULD reconnect using Last-Event-ID Client-driven exponential backoff Stable event IDs so clients can resume without duplicates None; gaps recovered by client re-read on reconnect
gRPC Bounded retries on UNAVAILABLE, DEADLINE_EXCEEDED, RESOURCE_EXHAUSTED; no retry on INVALID_ARGUMENT, PERMISSION_DENIED Exponential with jitter, capped; honor server RetryInfo if returned Server MUST dedupe by request ID metadata Operator-owned durable store; poison RPC payloads preserved
Debug None (development only) None Not required Not applicable
Log None None Not required Not applicable (logs themselves are the record)

Record the actual configured values per row in templates/drasi-acceptance-evidence.md. If a row is left at "operator MUST establish" with no concrete value, the reaction is not production-ready.

Webhook reaction outbound integrity (HMAC, mTLS, allowlist)

Without integrity controls, a tampered Reaction (compromised container, hijacked egress, MITM on the path to the downstream) can exfiltrate row-level data from the query result. HTTP/webhook Reactions MUST enforce all of the following:

HMAC-SHA256 signature

  • Header: X-Drasi-Signature: sha256=<hex-digest>.
  • Computed over the raw request body using a shared secret.
  • Secret rotated per the rotation runbook in bundles/security/guide.md. Use overlap rotation (current + next valid simultaneously) so a rolling restart does not drop deliveries.
  • Include a key version header (e.g., X-Drasi-Signature-KeyId: 2) so the receiver can pick the correct secret during overlap.

Replay protection

  • Header: X-Drasi-Timestamp: <unix-seconds>.
  • Receiver rejects requests outside a 5-minute acceptance window (abs(now - timestamp) > 300).
  • Timestamp MUST be inside the signed payload prefix (e.g., sign "<timestamp>.<body>") so it cannot be rewritten in flight.

Receiver verification example (bash + openssl)

# $SECRET, $BODY_FILE, $TS, $SIG come from the incoming request
EXPECTED=$(printf '%s.' "$TS" | cat - "$BODY_FILE" \
  | openssl dgst -sha256 -hmac "$SECRET" -binary | xxd -p -c 256)
# Constant-time compare
[ "$SIG" = "sha256=$EXPECTED" ] || { echo "bad signature"; exit 1; }
NOW=$(date +%s); [ $((NOW - TS)) -le 300 ] || { echo "stale"; exit 1; }

mTLS for high-sensitivity downstreams

  • Mutual TLS where the Reaction presents a client certificate issued from a private CA the downstream trusts.
  • Certificate rotation tracked in the same evidence record as the HMAC secret.
  • Required whenever payloads carry regulated data classes (PII, PHI, financial, secrets).

Outbound URL allowlist

  • The Reaction pod's egress MUST be constrained to the explicit downstream FQDN(s).
  • On AKS: enforce via Kubernetes NetworkPolicy with an egress to block listing only the required FQDN/IP and port; combine with the cluster's egress firewall.
  • On Azure Container Apps: configure ACA egress rules (or place the ACA environment behind a controlled VNet egress) restricting outbound to the allowlisted endpoints.
  • Point at bundles/security/guide.md egress-controls section for the canonical policy and review cadence. Wildcards are not acceptable - list FQDNs explicitly.

Validation evidence MUST record: HMAC key id in use, last rotation date, mTLS cert thumbprint and expiry, the exact allowlist set, and a synthetic negative test showing a tampered body or out-of-window timestamp is rejected.

Shadow / dry-run reaction rollout

Roll a new (or materially changed) reaction out in stages rather than pointing it at the production downstream on first deploy.

  1. Stage A - Debug shadow. Deploy the reaction subscribed to the production query but configured as a Debug or Log reaction (no external side effect). Verify it observes the same change stream the production reaction would.
  2. Stage B - Staging downstream. Re-point to a non-production downstream (test webhook receiver, dev Event Grid topic, dev SignalR hub, sandbox MCP client). Capture every emitted payload for a representative interval that covers normal traffic and at least one high-cardinality change burst.
  3. Stage C - Diff against expected. For each captured payload assert: query ID matches, change kind is one of added|updated|deleted, payload hash matches the query result contract, idempotency key is present and stable, no PII outside the declared classification.
  4. Stage D - Production promotion. Re-point to the production downstream. Keep Stage B running in parallel for one validation window so any drift is visible.
  5. Stage E - Decommission shadow. Remove the shadow reaction only after the production reaction has passed the freshness and delivery checks in ## Validation.

Validation evidence required

  • Stage B capture file or log query showing N payloads observed.
  • Stage C diff report showing zero schema or classification violations.
  • Stage D timestamp of cutover and the first successful production downstream observation.
  • Sign-off from the downstream owner that they expect and can absorb the traffic shape.

MCP-specific failure handling

The MCP Reaction sits on top of an active continuous query and a long-lived client transport. Three failure shapes MUST be handled explicitly.

Upstream query briefly inactive during resources/read

  • The server SHOULD respond with HTTP 503 (or the JSON-RPC equivalent error) and include a retry hint (e.g., Retry-After: 2).
  • The error body MUST identify the query ID and the inactive reason (compiling, restarting, source disconnected) without leaking source credentials.
  • Clients MUST treat 503 as transient and retry with bounded exponential backoff rather than re-subscribing.

Client disconnect during streaming

  • On transport close the server MUST clean up the subscription state for that client within a bounded interval - recommend a 30-second budget - to avoid leaking subscription handles.
  • Observable signal: a metric or log line on the MCP reaction showing subscription_closed{reason="client_disconnect"} (or equivalent) within the budget. Validation evidence MUST cite the exact signal name and the measured cleanup latency.
  • A leaked subscription is recognizable as a monotonically growing subscription count on the reaction with no matching active clients.

Transport reconnect behaviour

  • Clients SHOULD resume from the last observed notifications/resources/updated cursor by re-reading the resource and reconciling against locally cached state.
  • If the cursor was not persisted across the disconnect, the client MUST perform a full resources/read and treat all results as potentially new - partial replay is not safe.
  • The server MUST NOT replay missed notifications on reconnect; the recovery path is always "client re-reads current state".

Source: SKILL.md on GitHub

1 warning8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Drasi skill package is a highly structured and security-conscious set of instructions for managing data change detection pipelines. It includes extensive documentation on threat modeling, workload identity setup on AKS, and specific guidance for preventing prompt injection when source data is fed into AI agents. All documented commands and scripts are legitimate operational tools for the Drasi platform, and no malicious patterns such as obfuscation, persistence, or data exfiltration were found.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: MEDIUM · 1 issue

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
metadata
{
  "last_verified": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/drasi