Reactions bundle
Use this bundle for Drasi Reaction setup, delivery contracts, webhooks, Debug, Result, Event Grid, SignalR, MCP, custom reactions, and downstream side-effect validation.
Reaction authoring rules
- Verify the selected Reaction schema against current Drasi docs.
- Verify the reaction provider image exists at the platform version tag on GHCR BEFORE writing YAML or committing to a provider.
drasi list reactionproviderreturns every provider definition bundled into thedrasi initrelease, but the matching container image may not have been built for that version. A provider appearing in the catalog and being fully documented at drasi.io is NOT evidence its image exists (see## Reaction provider image verification gatebelow). - Subscribe only to queries that exist and have a documented result contract.
- Define added, updated, and deleted behavior explicitly.
- Make downstream effects idempotent where possible.
- Protect reaction endpoints with auth and network controls.
- Redact secrets and sensitive payload fields from logs.
- Validate the downstream side effect, not only the reaction status.
- Design retry, backoff, and failure handling based on downstream system behavior.
- For
kind: Http, useproperties.baseUrland per-queryadded|updated|deletedblocks underqueries. Do not placeurl/methoddirectly underproperties. - For Azure-hosted reactions using Workload Identity, reactions use a per-reaction service account in the Drasi namespace:
reaction.<reaction-name>. The federated credential subject issystem:serviceaccount:<drasi-namespace>:reaction.<reaction-name>. You must also patch the reaction deployment's pod template with theazure.workload.identity/use: "true"label. SA annotation alone is insufficient. See## Reaction workload identity on AKSbelow. - Prefer native Dapr reactions over custom bridge services. Before writing a custom HTTP endpoint to receive Drasi reaction output and forward it to a message broker, check whether a native reaction exists.
PostDaprPubSub,SyncDaprStateStore,EventGrid, andSignalRreactions publish directly to downstream systems, no intermediate service needed. See## PostDaprPubSub reaction (direct-to-Service-Bus / Event-Hub)below.
Reaction provider image verification gate
The provider catalog and image builds are decoupled in Drasi's release process. A provider definition ships with the drasi init release, but the CI pipeline for that provider's image may not have run for every platform release tag.
Production-verified example (2026-06-23; GHCR gap re-verified against the live tag list 2026-08-16): The PostDaprPubSub reaction provider IS registered under Drasi for Kubernetes 0.10.0 (drasi list reactionprovider shows it) and fully documented in the drasi.io how-to guides (https://drasi.io/drasi-kubernetes/how-to-guides/configure-reactions/configure-post-pubsub-reaction/). However, ghcr.io/drasi-project/reaction-post-dapr-pubsub:0.10.0 does NOT exist on GHCR, the image tag stops at 0.9.2. drasi apply kind: PostDaprPubSub against a 0.10.0 cluster would fail with ImagePullBackOff.
The fix: externalImage: true (verified against 0.10.0 source code). Drasi resolves provider images two ways:
externalImage: true: theimagefield is used verbatim as a fully-qualified image reference, no prefix/tag resolution applied.externalImage: true: theimagefield is used verbatim as a fully-qualified image reference, no prefix/tag resolution applied.
Register a custom provider with the version tag that does exist:
apiVersion: v1
kind: ReactionProvider
name: PostDaprPubSub
spec:
services:
reaction:
image: ghcr.io/drasi-project/reaction-post-dapr-pubsub:0.9.2
externalImage: trueApply this with drasi apply -f <file>, then apply your reaction normally. No platform downgrade required.
Before committing to any provider, verify the image exists:
$accept = "application/vnd.docker.distribution.manifest.list.v2+json, application/vnd.docker.distribution.manifest.v2+json, application/vnd.oci.image.index.v1+json, application/vnd.oci.image.manifest.v1+json"
$t = (Invoke-RestMethod -Uri "https://ghcr.io/token?scope=repository:drasi-project/<image-name>:pull&service=ghcr.io").token
try {
Invoke-WebRequest -Uri "https://ghcr.io/v2/drasi-project/<image-name>/manifests/<version-tag>" -Headers @{Authorization="Bearer $t"; Accept=$accept} -UseBasicParsing | Out-Null
"EXISTS"
} catch {
"MISSING - use externalImage: true with an older tag that exists"
}Use the manifest-list media types (the $accept header above). GHCR returns 404 for multi-arch images when queried with only application/vnd.docker.distribution.manifest.v2+json.
To find the image name for a provider, run drasi describe reactionprovider <ProviderName> and read the spec.services.<service>.image field, then prepend ghcr.io/drasi-project/ and append :<platform-version>.
Reaction workload identity on AKS
Verified against Drasi for Kubernetes 0.10.0 with PostDaprPubSub reaction.
Two layers of identity, don't conflate them
Reaction
spec.identity— the identity declared on the Reaction resource for the reaction's own source/target (e.g. when the reaction needs to authenticate to a downstream Azure resource directly). This is a top-levelspecfield, verified against Drasi 0.10.0 openapiServiceIdentityDto.Dapr sidecar workload identity — when a reaction like
PostDaprPubSubuses a Dapr sidecar to publish to Azure Service Bus / Event Hub / Storage, it is the Dapr sidecar that authenticates to Azure, not the reaction container. The sidecar needs workload identity env vars (AZURE_CLIENT_ID,AZURE_TENANT_ID,AZURE_FEDERATED_TOKEN_FILE) injected by the Azure Workload Identity mutating webhook.
Critical: the pod template label, not just the SA annotation
The Azure Workload Identity webhook injects env vars only when the pod template has the azure.workload.identity/use: "true" label. Annotating the service account alone is not sufficient, the webhook keys off the pod template label.
Drasi manages reaction deployments via its own controller. You must patch the deployment's pod template to add the label, because drasi apply does not add it automatically:
# 1. Annotate the per-reaction SA (verify the exact SA name for your reaction)
kubectl annotate sa reaction.<reaction-name> -n <drasi-namespace> azure.workload.identity/client-id=<mi-client-id>
kubectl label sa reaction.<reaction-name> -n <drasi-namespace> azure.workload.identity/use=true
# 2. Patch the deployment pod template with the WI label (THIS IS THE MISSING STEP)
kubectl patch deploy <reaction-name>-reaction -n <drasi-namespace> \
--type=json -p='[{"op":"add","path":"/spec/template/metadata/labels/azure.workload.identity~1use","value":"true"}]'
# 3. Wait for rollout to complete
kubectl rollout status deploy <reaction-name>-reaction -n <drasi-namespace>
# 4. Verify env vars were injected
kubectl get pod -n <drasi-namespace> -l drasi/resource=<reaction-name> -o yaml | grep AZURE_CLIENT_IDIf you skip step 2, the reaction container starts but the Dapr sidecar falls back to IMDS and fails with "Identity not found" / 400 Bad Request from 169.254.169.254.
Which service account do reactions use?
Verified against Drasi for Kubernetes 0.10.0: reactions use a per-reaction service account in the Drasi namespace, named reaction.<reaction-name>. This differs from the per-source SA pattern (source.<source-name>), but both are per-resource. The federated credential subject is system:serviceaccount:<drasi-namespace>:reaction.<reaction-name>.
This may be version-specific, always verify with:
kubectl get pod -n <drasi-namespace> -l drasi/resource=<reaction-name> -o jsonpath='{.items[0].spec.serviceAccountName}'PostDaprPubSub + Service Bus full checklist
For a PostDaprPubSub reaction publishing to Azure Service Bus via Dapr:
- Register the PostDaprPubSub provider with
externalImage: true(see## Reaction provider image verification gate). - Create a user-assigned managed identity and grant it Azure Service Bus Data Sender on the Service Bus namespace.
- Create a federated credential for subject
system:serviceaccount:<drasi-namespace>:reaction.<reaction-name>. - Annotate and label the
reaction.<reaction-name>SA with workload identity. - Apply a Dapr Component (
pubsub.azure.servicebus) in the Drasi namespace withazureClientIdset to the managed identity client ID. - Apply the PostDaprPubSub reaction referencing that Dapr component (
pubsubNamemust match the componentmetadata.name). - Patch the reaction deployment's pod template with the
azure.workload.identity/use: "true"label. - Restart the reaction pods and verify the Dapr sidecar logs show no
"Identity not found"errors.
PostDaprPubSub reaction (direct-to-Service-Bus / Event-Hub)
The PostDaprPubSub reaction is the native Dapr-based message broker integration for Drasi. It eliminates the need for custom bridge/forwarder services by publishing query results directly to Azure Service Bus, Event Hubs, Kafka, RabbitMQ, or any Dapr-supported pub/sub component via a Dapr sidecar.
When to use it (and when it replaces custom code)
| Instead of writing | Use PostDaprPubSub with |
|---|---|
| Custom HTTP endpoint receiving Drasi reaction POSTs → forwarding to Service Bus | Dapr pubsub.azure.servicebus component |
| Custom HTTP endpoint → forwarding to Event Hubs | Dapr pubsub.azure.eventhubs component |
| Custom HTTP endpoint → forwarding to Kafka | Dapr pubsub.kafka component |
| Custom HTTP endpoint → forwarding to RabbitMQ | Dapr pubsub.rabbitmq component |
Architecture it replaces:
Before (custom bridge — more moving parts):
Drasi query → HTTP Reaction → custom HTTP service → SDK → Service Bus
(requires: building/pushing a Docker image, managing a deployment,
managing a managed identity, health checks, scaling, debugging)
After (native Dapr — Drasi-managed):
Drasi query → PostDaprPubSub Reaction → Dapr sidecar → Service Bus
(requires: Dapr component YAML + workload identity setup)How it works internally
Drasi ContinuousQuery detects a change
→ Query result published to internal Redis pubsub (drasi-pubsub-<reaction-name>)
→ Dapr sidecar subscribes to that Redis pubsub
→ Dapr sidecar calls the reaction's HTTP endpoint (/eventhandler)
→ Reaction container calls Dapr Publish API on the configured pubsubName
→ Dapr sidecar authenticates to Azure (Workload Identity) and publishesThe pubsubName in the reaction's per-query config must match a Dapr Component metadata.name in the Drasi namespace. The topicName maps to the Service Bus queue/topic name or Event Hub name.
Per-query configuration
Each query gets its own config block specifying the Dapr component and target:
apiVersion: v1
kind: Reaction
name: forward-to-servicebus
spec:
kind: PostDaprPubSub
queries:
order-shipped: >
{
"pubsubName": "drasi-servicebus",
"topicName": "order-events",
"format": "Unpacked",
"skipControlSignals": true
}
order-status-change: >
{
"pubsubName": "drasi-servicebus",
"topicName": "order-events",
"format": "Unpacked",
"skipControlSignals": true
}Format options:
Packed, sends all results from a single change event as one message array.Packed, sends all results from a single change event as one message array.
skipControlSignals: true suppresses Drasi control messages (bootstrap, query restart) so only actual data changes are published.
Message format caveat for downstream consumers
The Unpacked format sends the query's RETURN fields directly (e.g. {"matchId":"M001","homeScore":3,...}). Downstream consumers expecting a different envelope format (e.g. a wrapper with Envelope, Payload, BufferedAtUtc) will need to either:
- Be updated to accept the raw query result shape, or
- Use
Packedformat and add a mapping layer, or - Accept the impedance mismatch and handle deserialization accordingly
This is a consumer-side concern, not a Drasi issue, but plan for it when migrating from a custom bridge that transformed the payload.
Prerequisites
- Dapr must be installed in the cluster (
dapr init -kor via Drasi's Dapr dependency). Drasi itself installs Dapr as a dependency for source/reaction sidecars. - Dapr Component YAML for the target pub/sub (Service Bus, Event Hubs, etc.) must be applied to the Drasi namespace, not the application namespace.
- Workload identity must be configured on the reaction pod (see
## Reaction workload identity on AKSabove and the checklist).
Verified Dapr pub/sub components
| Dapr component type | Azure target | Verified |
|---|---|---|
pubsub.azure.servicebus |
Azure Service Bus queues and topics | ✅ Production-verified 2026-06-23 |
pubsub.azure.eventhubs |
Azure Event Hubs | [VERIFY] — same workload identity pattern |
pubsub.kafka |
Apache Kafka / Confluent | [VERIFY] — SASL/SCRAM auth |
pubsub.rabbitmq |
RabbitMQ | [VERIFY] — connection string auth |
Reaction selection
| Need | Prefer |
|---|---|
| Local or development visibility | Debug Reaction or result polling |
| Push to HTTP service | HTTP Reaction with auth, idempotency, timeout, and retry strategy |
| Azure event distribution | Azure Event Grid Reaction with documented identity support |
| AWS event distribution | AWS EventBridge Reaction with IRSA or access key auth |
| Real-time UI updates | SignalR Reaction with documented identity support |
| Agent/MCP integration | MCP Reaction, plus agent-integration bundle |
| Vector store / RAG sync | SyncVectorStore Reaction with Qdrant, Azure AI Search, or InMemory |
| State store sync (Dapr) | SyncDaprStateStore Reaction |
| Pub/sub messaging (Dapr) | PostDaprPubSub Reaction |
| Kafka / Debezium ecosystem | Debezium Reaction after schema review |
| Graph database sync | Gremlin Reaction |
| Database side effects | StoredProc Reaction (MSSQL, MySQL, PostgreSQL) |
| Azure queue storage | Azure StorageQueue Reaction |
SignalR Reaction: in-process vs Azure SignalR Service modes
The SignalR Reaction supports two hosting modes. Choose based on your deployment and scale requirements.
| Mode | When to use | Pros | Cons |
|---|---|---|---|
In-process (omit connectionString) |
Local development, small-scale deployments, or when you want to avoid an Azure SignalR Service dependency | No external dependency; lower latency; simpler configuration | Connections are bound to the reaction pod's memory; scaling requires sticky sessions or a Redis backplane |
Azure SignalR Service (set connectionString) |
Production, multi-replica, or when clients need to reconnect across pod restarts | Scales independently; serverless-friendly; handles reconnection and broadcast | Azure dependency and cost; additional latency hop |
In-process configuration, simply omit properties.connectionString:
kind: Reaction
apiVersion: v1
name: my-signalr-reaction
spec:
kind: SignalR
queries:
my-query:Azure SignalR Service configuration:
kind: Reaction
apiVersion: v1
name: my-signalr-reaction
spec:
kind: SignalR
properties:
connectionString: "Endpoint=https://<resource>.service.signalr.net;AccessKey=<key>;Version=1.0;"
queries:
my-query:For managed-identity-based auth to Azure SignalR Service, use:
properties:
connectionString: "Endpoint=https://<resource>.service.signalr.net;AuthType=azure"And configure the identity block per the Azure identity guidance.
Client connection: In both modes, clients connect via drasi tunnel reaction <name> <port> (development) or directly to the service endpoint (production). The reaction exposes a SignalR hub per subscribed query, and the client library (@microsoft/signalr for JS, Microsoft.AspNetCore.SignalR.Client for .NET) connects to /query/<query-id>.
SignalR reaction actor limitation (pre-1.0)
The SignalR reaction requires the query host to support a ReactionResource actor type that handles the configure method. When the Drasi API applies a SignalR reaction, the resource provider calls PUT /actors/ReactionResource/<name>/method/configure on the query host to set up the reaction. If the query host binary doesn't include SignalR actor support (e.g. main-branch builds where the actor plugin isn't compiled in), this call fails with:
Error configuring resource: GrpcError { ... "error invoke actor method: Put \"http://127.0.0.1:8080/actors/ReactionResource/dashboard/method/configure\": EOF" }This is a compile-time limitation, not a configuration error. The query-host binary either includes the SignalR reaction actor or it doesn't, there is no runtime flag or env var to enable it. If the query-host image was built without SignalR support, no configuration change can fix it.
Workarounds (none guaranteed):
- Use a different query-host image tag that includes SignalR support. In the
drasi-platformrepo, SignalR reaction support is compiled in when thesignalrfeature flag is enabled inquery-container/query-host/Cargo.toml. - Replace the SignalR reaction with an HTTP reaction connected to a standalone SignalR hub, or with the PostDaprPubSub reaction feeding a custom bridge service.
Verify: Run drasi apply -f signalr-reaction.yaml. If it succeeds (HTTP 200), the query host supports SignalR. If it returns HTTP 500 with the actor EOF error, the query host was built without SignalR support.
SignalR consumer reference implementation
The drasi-project/grafana-signalr repository provides a real-world SignalR consumer as a Grafana data source plugin. It validates:
- Operation types: Drasi SignalR reactions use single-letter operation codes,
i(insert/added),u(update),d(delete),x(control). These differ from the MCP reaction'sadded/updated/deletedstrings. Ensure your SignalR client handles all four. - Snapshot loading: The plugin calls a snapshot/query endpoint on initial connect to recover current state before streaming incremental changes. This is the SignalR equivalent of MCP's
resources/read. - Error handling: The plugin documents connection failure, missing query ID, and data processing error paths.
Cross-reference: github.com/drasi-project/grafana-signalr for the full implementation and operational experience.
SyncDaprStateStore reaction
Synchronises the current result set of a Continuous Query to a Dapr state store component. Each item in the result set is stored as a separate key-value entry, keyed by the query result's identity column.
Per-query configuration
kind: Reaction
apiVersion: v1
name: my-dapr-state-reaction
spec:
kind: SyncDaprStateStore
properties:
stateStoreName: my-statestore
queries:
my-query-id:Properties
| Property | Type | Required | Default | Description |
|---|---|---|---|---|
stateStoreName |
string | Yes | — | The metadata.name of the Dapr state store component. The component MUST be defined in the Drasi namespace (drasi-system by default) and point to the same state store used by your application. |
State store behaviour
- Added results → new key-value entry created.
- Updated results → existing entry overwritten with upsert semantics.
- Deleted results → existing entry deleted from the state store.
- The key is derived from the query result's identity, typically the first
RETURNcolumn. Verify the identity column is stable and unique before using this reaction.
Prerequisites
- Dapr must be installed in the cluster.
- Dapr Component YAML for the target state store (e.g.
state.azure.cosmosdb,state.redis) must be applied to the Drasi namespace. - Workload identity must be configured on the reaction pod if the state store uses Azure AD auth (see
## Reaction workload identity on AKS).
Validation
| Validation | Method |
|---|---|
| Reaction is healthy | drasi list reaction shows status Available |
| State store contains expected keys | Direct read from the Dapr state store API (localhost:<dapr-port>/v1.0/state/<stateStoreName>/<key>) returns the current query result for each identity |
| Deleted results are removed | Read the key for a deleted result — the store returns a null or 404 |
| Upsert idempotency | Repeated application of the same query result produces the same store entry without errors |
SyncVectorStore reaction
The SyncVectorStore reaction continuously synchronizes ContinuousQuery results to a vector database, generating embeddings for each result row and creating searchable vector collections. This is the primary reaction type for AI/RAG integration patterns.
Supported backends
| Backend | Type | Use case |
|---|---|---|
| Qdrant | Persistent vector DB | Production, high-scale semantic search |
| Azure AI Search | Cloud search service | Azure-native, hybrid (vector + full-text) search |
| InMemory | In-process memory | Development, testing, demos (data lost on restart) |
Embedding service configuration
| Setting | Required | Default | Notes |
|---|---|---|---|
embeddingServiceType |
Yes | — | Currently only AzureOpenAI and OpenAI supported |
embeddingEndpoint |
Yes | — | Full URL of the embeddings API endpoint |
embeddingApiKey |
Yes | — | API key; supports inline value or kind: Secret reference |
embeddingModel |
Yes | text-embedding-3-large |
Model or deployment name |
embeddingDimensions |
Yes | 3072 |
Vector dimension count (max 20000) |
Vector store configuration
| Setting | Required | Default | Notes |
|---|---|---|---|
vectorStoreType |
Yes | — | Qdrant, AzureAISearch, or InMemory |
connectionString |
Yes | — | Format varies by backend (see examples) |
distanceFunction |
No | CosineSimilarity |
Also: CosineDistance, EuclideanDistance, DotProductSimilarity, ManhattanDistance |
indexKind |
No | Hnsw |
Also: Flat, IvfFlat, DiskAnn |
isFilterable |
No | true |
Enables field-based filtering on vector search |
isFullTextSearchable |
No | false |
Full-text search on content fields (Azure AI Search only) |
Per-query configuration
Each query maps to one vector collection with a Handlebars template for document formatting:
queries:
my-query-id: |
{
"collectionName": "products",
"keyField": "product_id",
"documentTemplate": "Product: {{name}}\nDescription: {{description}}",
"titleTemplate": "{{name}}",
"vectorField": "content_vector",
"createCollection": true
}| Setting | Required | Notes |
|---|---|---|
collectionName |
Yes | Name of the collection in the vector store |
keyField |
Yes | Field from the query result used as the document key |
documentTemplate |
Yes | Handlebars template for the text to embed |
titleTemplate |
No | Optional title for search result display |
vectorField |
No | Default: embedding or backend-specific default |
createCollection |
No | true to auto-create on first sync |
Configuration example (Qdrant)
apiVersion: v1
kind: Reaction
name: product-catalog-qdrant
spec:
kind: SyncVectorStore
properties:
vectorStoreType: "Qdrant"
connectionString: "Endpoint=qdrant-service:6334"
embeddingServiceType: "AzureOpenAI"
embeddingEndpoint: "https://your-resource.openai.azure.com/"
embeddingApiKey:
kind: Secret
name: azure-openai-creds
key: api-key
embeddingModel: "text-embedding-3-large"
embeddingDimensions: 3072
distanceFunction: "CosineSimilarity"
indexKind: "Hnsw"
isFilterable: true
queries:
product-catalog: |
{
"collectionName": "products",
"keyField": "id",
"documentTemplate": "{{name}}: {{description}}",
"titleTemplate": "{{name}}",
"vectorField": "embedding",
"createCollection": true
}Configuration example (Azure AI Search)
spec:
kind: SyncVectorStore
properties:
vectorStoreType: "AzureAISearch"
connectionString: "Endpoint=https://your-search.search.windows.net;ApiKey=your-key"
embeddingServiceType: "AzureOpenAI"
embeddingEndpoint: "https://your-resource.openai.azure.com/"
embeddingApiKey:
kind: Secret
name: azure-openai-creds
key: api-key
embeddingModel: "text-embedding-3-large"
embeddingDimensions: 3072
isFullTextSearchable: true # Azure AI Search supports hybrid search
queries:
products: |
{
"collectionName": "products",
"keyField": "product_id",
"documentTemplate": "Product: {{name}}\n{{description}}",
"vectorField": "content_vector",
"createCollection": true
}Validation
| Check | Method |
|---|---|
| Reaction is healthy | drasi list reaction shows Ready |
| Vector store has documents | Query the vector store collection count (Qdrant /collections/{name}/points/count, Azure AI Search count parameter) |
| Document content matches query result | Read a document from the vector store and verify fields match the Handlebars template output |
| Embedding dimensions match config | Verify vector dimension matches embeddingDimensions setting |
| All three change types propagate | Insert, update, and delete a source row; assert added/updated/deleted in the vector store |
Reliability defaults
| Property | Baseline |
|---|---|
| Embedding retry | Bounded retries (3 attempts) on embedding service 5xx and 429 with exponential backoff |
| Vector store retry | Bounded retries (3 attempts) on write failures |
| Delivery guarantee | At-least-once per result change; duplicate upserts to the vector store MUST be idempotent (use keyField as the upsert key) |
| Backpressure | Reaction syncs asynchronously; vector store writes do not block the source change pipeline |
| Error handling | Failed embeddings are logged with query ID and result key; a failed vector write is retried then logged; poison documents are not automatically dropped |
See bundles/agent-integration/guide.md for the consumer-side architecture that reads from the vector store.
Output format patterns
Multiple Drasi reactions (EventGrid, EventBridge, StorageQueue, and custom reactions using the same pattern) support three output formats that control how query results are packaged into events.
Format comparison
| Format | Behaviour | Event count per change | Use when |
|---|---|---|---|
packed |
All query result rows are bundled into a single event on any change | 1 event per source change | Downstream wants a complete snapshot; fan-out is acceptable; latency is more important than per-row granularity |
unpacked |
Each added/updated/deleted row produces a separate event | N events per source change (one per affected row) | Downstream processes each row individually; per-row idempotency is needed; event-driven consumers that map one event to one action |
template |
A Handlebars template produces a custom-shaped event from the query result | 1 event per source change | Downstream requires a specific schema; field renaming, filtering, or enrichment is needed before delivery |
When to use which
Use packed when:
- The downstream is a cache or state store that replaces its entire dataset on each change.
- The downstream can efficiently diff the full payload against its current state.
- Event volume must be minimized (each source change = 1 event, not N).
Use unpacked when:
- Each result row maps to an independent downstream action (email per new order, alert per new sensor reading).
- The downstream needs per-row idempotency keys.
- A single source change affects a small, predictable number of rows.
Use template when:
- The downstream schema does not match the query result contract.
- You need to rename fields, combine fields, or add computed fields before delivery.
- You need to filter out fields from the result contract before they reach the downstream.
- The same query result needs different shapes for different reactions.
Template availability
| Reaction kind | packed |
unpacked |
template |
|---|---|---|---|
| EventGrid | Yes (default) | Yes | Yes |
| EventBridge | Yes (default) | Yes | Yes |
| StorageQueue | Yes (default) | Yes | Yes |
| Http | Per-query body template |
Per-query added/updated/deleted blocks |
Inherits query-level Handlebars |
| SyncVectorStore | Not applicable (always per-document) | Not applicable | Per-query documentTemplate |
Per-query template config on EventGrid and EventBridge
EventGrid and EventBridge reactions accept Handlebars templates per query with per-change-type blocks (verified against drasi-platform PR #343/#344, merged 2025-12/2026-01; shape confirmed from PR body 2026-08-24):
kind: Reaction
name: my-eventgrid-reaction
spec:
kind: EventGrid
# ... connection fields ...
format: template
queries:
my-query: |
added:
template: |
{ "eventType": "ItemAdded", "data": {{json after}} }
metadata: # becomes CloudEvent extension attributes
changeType: added
updated:
template: |
{ "before": {{json before}}, "after": {{json after}} }
deleted:
template: |
{ "id": "{{before.id}}" }Rules: each change-type block has template (Handlebars string) and optional metadata (key-value pairs applied as CloudEvent extension attributes); templates receive after (result data for added/updated) and before (previous result for updated/deleted); {{json <expr>}} emits JSON-safe output. Verify the exact field names against the pinned platform release docs before writing production YAML; this shape ships in the platform source, not yet in a tagged release newer than 0.10.0.
Per-reaction state store (direction of travel)
Platform PR #442 (merged 2026-08-15, not in any tagged release as of 2026-08-24) adds per-reaction Dapr state store support: reaction providers can declare state_store: true, the platform provisions and cleans up a per-reaction Dapr state store component, and injects a StateStoreName into reaction services. Do not write YAML against this surface until it lands in a release you pin; re-check the release notes when upgrading past 0.10.0.
Reliability contract
Every production reaction should document:
- Query subscriptions.
- Payload template and schema.
- Auth mechanism.
- Retry behavior.
- Timeout behavior.
- Idempotency key or duplicate handling.
- Dead-letter or manual replay process, where supported.
- Data classification and redaction rules.
- Downstream ownership and operational contact.
Restart durability: since the 2026-08-20 drasi-core change (PR #735, in the Server 0.2.2 line), the query-result sequence clock survives restarts via the configured state store; duplicate suppression no longer silently resets on pod restart. Details and re-verification guidance: bundles/recovery/guide.md "Restart semantics per runtime form".
Token and credential rotation
Use overlap rotation for webhook tokens and downstream credentials:
- Add new credential to downstream service while old credential remains valid.
- Update Reaction configuration to send the new credential.
- Validate successful delivery.
- Remove old credential.
- Record rotation evidence.
Do not perform a single cutover if failed delivery would be user-visible or hard to replay.
Azure identity guidance
For Azure Event Grid, SignalR, and other reactions using Microsoft Entra Workload ID, current Drasi docs document Workload ID flows. Verify current docs and ensure:
- AKS Workload Identity is enabled.
- Federated credential subject matches
system:serviceaccount:drasi-system:reaction.<reaction-name>(see## Reaction workload identity on AKSabove for the full apply-annotate-restart sequence). - The managed identity has the minimum role, for example EventGrid Data Contributor, SignalR App Server, or Azure Service Bus Data Sender where applicable.
- Role assignment propagation is accounted for in validation.
Logging and data minimization
Never log:
- Raw Authorization headers.
- Webhook tokens.
- Database passwords.
- Full payloads containing PII or sensitive operational data unless explicitly approved and redacted.
Prefer structured logs with correlation ids, query name, operation type, delivery status, latency, and redacted error details.
Validation
A reaction is only complete when all are true:
- Reaction resource is healthy.
- It subscribes to the expected query.
- A synthetic or real source change produces an expected reaction payload.
- The downstream system accepts and processes the payload.
- Failure behavior is tested or documented.
Per-reaction validation patterns
"Reaction delivered" must be defined per reaction type. Asserting only that Drasi sent the event is not sufficient - verify the downstream side effect.
| Reaction kind | Validation method | Sample assertion |
|---|---|---|
| Debug / Log | Inspect Drasi logs after a trigger | drasi logs reaction <id> --since 2m contains the expected change payload. |
| HTTP / Webhook | Capture the receiver call | Receiver returns 2xx; log line on the receiver shows the query ID, change kind, and a stable payload hash matching the query result. |
| gRPC | Receiver-side metric or log | gRPC server records a unary/stream call with the expected method and payload. |
| Server-Sent Events (SSE) | Open the SSE stream and trigger a change | curl -N http://<host>/api/v1/queries/<id>/stream emits a message event matching the query result within p95 ≤ 30 s. |
| Azure Event Grid | Event Grid subscription delivery | Event arrives on the subscription endpoint (e.g., Service Bus, Function) within the delivery SLA; eventType matches; data matches the query result contract. |
| Azure SignalR | Connected client receives message | A test SignalR client connected to the hub receives the expected message group/event within p95 ≤ 30 s. |
| MCP Reaction | MCP client receives notification | The MCP client gets notifications/resources/updated for drasi://query/<id> and resources/read returns the new result. |
| SyncVectorStore | Vector store has documents matching query | Query vector store collection count; assert document fields match Handlebars template output; assert all three change types propagate |
| AWS EventBridge | CloudEvent arrives on EventBus | EventBridge target (e.g., Step Function, Lambda) receives the expected CloudEvent with matching payload |
| Dapr StateStore | Dapr state store has expected key | Read state from Dapr state store component; key matches keyField and value matches query result |
| Dapr PubSub | Message arrives on Dapr pub/sub topic | Dapr subscriber receives the expected CloudEvent with matching topic and payload |
| Debezium | Kafka topic receives change event | Kafka consumer on the configured topic receives a Debezium-compatible payload event matching the query result |
| StorageQueue | Message appears in Azure Storage Queue | az storage message peek on the target queue shows a message matching the expected format |
| StoredProc | Target database shows side effect | Query the target stored procedure's effect table (e.g., audit log, summary row) and assert the expected state change |
Http Reaction manifest shape guardrail
When authoring Drasi for Kubernetes kind: Http reactions, this shape is required:
kind: Reaction
apiVersion: v1
name: my-http-reaction
spec:
kind: Http
properties:
baseUrl: "https://api.example.com"
token: "<optional-bearer-token>"
queries:
my-query-id: >
added:
url: "/api/path"
method: "POST"
body: |
{ "queryId": "my-query-id", "changeKind": "added", "result": {{json after}} }
headers:
Content-Type: "application/json"Failure pattern from field incidents:
- If
drasi apply reactionreturns400 ... missing field kind, check thatspec.kindexists and thatqueriesvalues are structured mappings (not empty/null entries).
Required evidence per release
For every reaction, record in the acceptance evidence (templates/drasi-acceptance-evidence.md):
- The reaction's idempotency strategy (e.g., dedup key, target-side upsert).
- The configured retry policy and timeout.
- A test trigger timestamp and the downstream observation timestamp.
- The verification that retried/duplicate deliveries are absorbed by the downstream system without double effects.
Common reaction failure modes
| Symptom | Likely cause | First action |
|---|---|---|
Reaction state shows running, no downstream effect |
Reaction not subscribed to the right query | Confirm reactions[].queries lists the query ID exactly. |
| Sporadic deliveries with same payload hash | Source emitting duplicate change notifications | Verify dedup on the reaction's downstream consumer. |
| Reaction errors only under load | Downstream rate limit or auth-token expiry race | Lower maxReplicas, rotate identity, confirm token lifetime. |
| Reaction stops emitting after secret rotation | Reaction container did not pick up the new secret | Confirm secret reload strategy (pod restart vs. env reload) for the runtime form in use. |
Per-reaction reliability defaults (baseline contract)
Treat the values below as a required contract the reaction author MUST document and validate before the reaction goes to production. The reaction author MUST establish concrete numbers per row; the table names what each row SHOULD specify, not what Drasi necessarily ships. If the upstream Drasi release publishes contradicting defaults, the published value wins - record both in the validation evidence.
| Reaction kind | Default retry policy expected | Backoff curve expected | Idempotency requirement | DLQ destination |
|---|---|---|---|---|
| HTTP webhook | Bounded retries (e.g., 5 attempts) on 5xx and 408/429; no retry on 4xx other than 408/429 | Exponential with full jitter, capped (e.g., base 1 s, max 30 s) | Receiver MUST dedupe on a stable key derived from query ID + change kind + payload hash | Operator-owned durable store (e.g., Service Bus dead-letter, S3 prefix); poison messages preserved with original headers |
| Event Grid | Honor Event Grid retry schedule; do not retry locally on top of it | Event Grid published curve (publisher MUST NOT add its own backoff) | Event ID + eventTime treated as unique by subscriber | Event Grid dead-letter container in storage |
| SignalR | Best-effort, no retry for transient client disconnects; reconnect handled by client | None at publisher; client SHOULD reconnect with exponential backoff | Client MUST tolerate redelivery of the last N messages on reconnect | Not applicable for fan-out; missed messages require client re-query against the result endpoint |
| MCP | No silent retry; surface upstream errors to the MCP client | Client-driven; server returns 503 with retry hint when query briefly inactive | resources/read MUST be safe to call repeatedly; result is the current snapshot |
None; the MCP client recovers by re-reading after notifications/resources/updated |
| SSE | No retry at server; client SHOULD reconnect using Last-Event-ID |
Client-driven exponential backoff | Stable event IDs so clients can resume without duplicates | None; gaps recovered by client re-read on reconnect |
| gRPC | Bounded retries on UNAVAILABLE, DEADLINE_EXCEEDED, RESOURCE_EXHAUSTED; no retry on INVALID_ARGUMENT, PERMISSION_DENIED |
Exponential with jitter, capped; honor server RetryInfo if returned |
Server MUST dedupe by request ID metadata | Operator-owned durable store; poison RPC payloads preserved |
| Debug | None (development only) | None | Not required | Not applicable |
| Log | None | None | Not required | Not applicable (logs themselves are the record) |
Record the actual configured values per row in templates/drasi-acceptance-evidence.md. If a row is left at "operator MUST establish" with no concrete value, the reaction is not production-ready.
Webhook reaction outbound integrity (HMAC, mTLS, allowlist)
Without integrity controls, a tampered Reaction (compromised container, hijacked egress, MITM on the path to the downstream) can exfiltrate row-level data from the query result. HTTP/webhook Reactions MUST enforce all of the following:
HMAC-SHA256 signature
- Header:
X-Drasi-Signature: sha256=<hex-digest>. - Computed over the raw request body using a shared secret.
- Secret rotated per the rotation runbook in
bundles/security/guide.md. Use overlap rotation (current + next valid simultaneously) so a rolling restart does not drop deliveries. - Include a key version header (e.g.,
X-Drasi-Signature-KeyId: 2) so the receiver can pick the correct secret during overlap.
Replay protection
- Header:
X-Drasi-Timestamp: <unix-seconds>. - Receiver rejects requests outside a 5-minute acceptance window (
abs(now - timestamp) > 300). - Timestamp MUST be inside the signed payload prefix (e.g., sign
"<timestamp>.<body>") so it cannot be rewritten in flight.
Receiver verification example (bash + openssl)
# $SECRET, $BODY_FILE, $TS, $SIG come from the incoming request
EXPECTED=$(printf '%s.' "$TS" | cat - "$BODY_FILE" \
| openssl dgst -sha256 -hmac "$SECRET" -binary | xxd -p -c 256)
# Constant-time compare
[ "$SIG" = "sha256=$EXPECTED" ] || { echo "bad signature"; exit 1; }
NOW=$(date +%s); [ $((NOW - TS)) -le 300 ] || { echo "stale"; exit 1; }mTLS for high-sensitivity downstreams
- Mutual TLS where the Reaction presents a client certificate issued from a private CA the downstream trusts.
- Certificate rotation tracked in the same evidence record as the HMAC secret.
- Required whenever payloads carry regulated data classes (PII, PHI, financial, secrets).
Outbound URL allowlist
- The Reaction pod's egress MUST be constrained to the explicit downstream FQDN(s).
- On AKS: enforce via Kubernetes
NetworkPolicywith an egresstoblock listing only the required FQDN/IP and port; combine with the cluster's egress firewall. - On Azure Container Apps: configure ACA egress rules (or place the ACA environment behind a controlled VNet egress) restricting outbound to the allowlisted endpoints.
- Point at
bundles/security/guide.mdegress-controls section for the canonical policy and review cadence. Wildcards are not acceptable - list FQDNs explicitly.
Validation evidence MUST record: HMAC key id in use, last rotation date, mTLS cert thumbprint and expiry, the exact allowlist set, and a synthetic negative test showing a tampered body or out-of-window timestamp is rejected.
Shadow / dry-run reaction rollout
Roll a new (or materially changed) reaction out in stages rather than pointing it at the production downstream on first deploy.
- Stage A - Debug shadow. Deploy the reaction subscribed to the production query but configured as a Debug or Log reaction (no external side effect). Verify it observes the same change stream the production reaction would.
- Stage B - Staging downstream. Re-point to a non-production downstream (test webhook receiver, dev Event Grid topic, dev SignalR hub, sandbox MCP client). Capture every emitted payload for a representative interval that covers normal traffic and at least one high-cardinality change burst.
- Stage C - Diff against expected. For each captured payload assert: query ID matches, change kind is one of
added|updated|deleted, payload hash matches the query result contract, idempotency key is present and stable, no PII outside the declared classification. - Stage D - Production promotion. Re-point to the production downstream. Keep Stage B running in parallel for one validation window so any drift is visible.
- Stage E - Decommission shadow. Remove the shadow reaction only after the production reaction has passed the freshness and delivery checks in
## Validation.
Validation evidence required
- Stage B capture file or log query showing N payloads observed.
- Stage C diff report showing zero schema or classification violations.
- Stage D timestamp of cutover and the first successful production downstream observation.
- Sign-off from the downstream owner that they expect and can absorb the traffic shape.
MCP-specific failure handling
The MCP Reaction sits on top of an active continuous query and a long-lived client transport. Three failure shapes MUST be handled explicitly.
Upstream query briefly inactive during resources/read
- The server SHOULD respond with HTTP 503 (or the JSON-RPC equivalent error) and include a retry hint (e.g.,
Retry-After: 2). - The error body MUST identify the query ID and the inactive reason (compiling, restarting, source disconnected) without leaking source credentials.
- Clients MUST treat 503 as transient and retry with bounded exponential backoff rather than re-subscribing.
Client disconnect during streaming
- On transport close the server MUST clean up the subscription state for that client within a bounded interval - recommend a 30-second budget - to avoid leaking subscription handles.
- Observable signal: a metric or log line on the MCP reaction showing
subscription_closed{reason="client_disconnect"}(or equivalent) within the budget. Validation evidence MUST cite the exact signal name and the measured cleanup latency. - A leaked subscription is recognizable as a monotonically growing subscription count on the reaction with no matching active clients.
Transport reconnect behaviour
- Clients SHOULD resume from the last observed
notifications/resources/updatedcursor by re-reading the resource and reconciling against locally cached state. - If the cursor was not persisted across the disconnect, the client MUST perform a full
resources/readand treat all results as potentially new - partial replay is not safe. - The server MUST NOT replay missed notifications on reconnect; the recovery path is always "client re-reads current state".