All skills

USE FOR: Drasi continuous-query solutions - real-time queries, change detection, reactive events, data-trigger pipelines on Drasi Server, Drasi for Kubernetes, or drasi-lib. Router: load bundle guides as needed. DO NOT USE for non-Drasi messaging (event-driven-messaging) or pure AKS/ACA hosting (aks-cluster-architecture, azure-container-apps).

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/drasi

This session only. Nothing lands on disk.

bundlessourcesguide.md

≈9.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Sources bundle

Use this bundle for Drasi Source design, provider schema selection, CDC setup, credentials, identity, lifecycle, and readiness validation.

Source authoring rules

  1. Verify the selected provider schema against current Drasi docs before writing YAML.
  2. Verify the source provider image exists at the platform version tag on GHCR BEFORE writing YAML. drasi list sourceprovider returns every provider definition bundled into the drasi init release, but the matching container image may not have been built for that version. Use drasi describe sourceprovider <ProviderName> to find the image name (spec.services.proxy.image / spec.services.reactivator.image), prepend ghcr.io/drasi-project/, append :<platform-version>, and verify it resolves on GHCR. See bundles/reactions/guide.md ## Reaction provider image verification gate for the PowerShell snippet (same process applies to source provider images). A missing image causes ImagePullBackOff at source creation time.
  3. Use apiVersion: v1 and kind: Source for Drasi for Kubernetes resources unless current docs say otherwise.
  4. Use names that are stable and externally safe. Queries reference the source name.
  5. Define the minimum table, resource, topic, or entity scope required.
  6. Do not create ContinuousQueries until the Source is available.
  7. Put secrets in the provider-supported secret mechanism. Do not inline production passwords.
  8. Prefer Microsoft Entra Workload ID only where the provider explicitly documents it.
  9. Document update behavior for the selected provider.
  10. For the Azure Event Hub Source, create a dedicated consumer group per Source. Do not use $Default, other consumers sharing the group will steal checkpoints and cause data loss. Each Drasi source maintains its own checkpoint position in the consumer group; if multiple readers share one group they advance each other's position and skip events.

Middleware pipeline for source changes

Use Drasi middleware when incoming source events need lightweight shaping before query evaluation. Prefer middleware over custom source/reaction code for common transformations that are supported by the current runtime.

Middleware is configured inside the ContinuousQuery definition (Drasi for Kubernetes: the ContinuousQuery resource manifest; Drasi Server: the config file; drasi-lib: the Rust API) as an ordered pipeline - each component receives the previous component's output. Full component reference with config shapes and examples (Unwind, JQ, Promote, ParseJson, Decoder, Relabel): https://drasi.io/reference/middleware/ (verified 2026-08-16).

Common middleware roles to verify against current docs before use:

Middleware Use when Notes
Unwind A source emits arrays and each element should become an independent change item Validate downstream result cardinality before production.
JQ JSON payloads need field extraction, reshaping, or filtering Keep expressions small and test with malformed payloads.
Promote Nested fields should become top-level properties for query matching Avoid promoting sensitive fields unless the query requires them.
ParseJson A string field contains JSON that queries need as structured data Treat parse failures as data-quality events, not silent drops.
Decoder Encoded payloads need conversion before query evaluation Verify codec and character set explicitly.
Relabel Source labels need normalization for query portability Record label mapping in the result contract.

A seventh component, map, is registered in the drasi-core middleware crate but has no page in the official reference (verified against middleware/src/map/mod.rs on main, 2026-08-24; the docs reference lists only the six above). It performs per-label reshape/projection: config keys are element labels mapping to insert / update / delete operation lists; each mapping supports op, a JSONPath selector (selected value exposed as $selected), label, id, elementType (node or relation { inNodeId, outNodeId }), a condition filter, and a properties map of property name to JSONPath expression. It is feature-gated in the middleware crate - confirm it is enabled in your runtime before using, and re-check whether official docs have documented it since.

Middleware output becomes query input. Apply the same data classification, redaction, and prompt-injection treatment to middleware output that you apply to raw source data.

Provider lifecycle matrix

Check current docs before relying on this matrix. It reflects the latest pass for this package.

Provider/resource Auth pattern to prefer Update behavior Notes
PostgreSQL Source MicrosoftEntraWorkloadID on AKS; otherwise secured credentials Delete and recreate, then recreate dependent queries Fields: host, port, user, database, ssl, tables. When using Workload ID, omit password; set user to the managed identity's display name (as registered in PostgreSQL Entra). The managed identity must be added as a PostgreSQL Flexible Server Entra admin before the source can connect.
Kubernetes Source Kubeconfig stored as a Kubernetes Secret Delete and recreate, then recreate dependent queries Requires credentials able to list/watch target resources. Avoid exec-based kubeconfigs in runtime containers.
SQL Server Source Microsoft Entra Workload ID when targeting Azure SQL and documented, otherwise secured credentials Re-apply same source name is documented as supported Ensure CDC and SQL permissions are configured.
Dataverse Source Microsoft Entra Workload ID is documented Verify current docs before updating Federated credential subject must match the per-source service account source.<source-name> in the Drasi namespace (NOT default — verified against Drasi for Kubernetes 0.10.0).
MySQL Source Secured database credentials Verify current docs before updating Confirm binlog/CDC prerequisites.
Azure Event Hub Source MicrosoftEntraWorkloadID on AKS; otherwise connection string via kind: Secret Re-apply same source name is documented as supported (verified 2026-08-16) Assign Azure Event Hubs Data Receiver RBAC role on the namespace to the managed identity. Fields: host (namespace FQDN), eventHubs (list), bootstrapWindow (integer minutes of backfill — verified 2026-08-16; not seconds), consumerGroup (default $Default) [VERIFY: consumerGroup is not in the current docs property table — confirm via drasi describe sourceprovider EventHub on the live cluster before relying on it]. Prefer a dedicated consumer group per Source. Event body becomes element data; system properties available for query predicates. Critical: Drasi uses Event Hub entity names as Cypher node labels — the hub name MUST match the label in MATCH (m:Label) (case-sensitive; escape dashed hub names with backticks). Workload ID SA: Drasi creates a per-source service account source.<source-name>, not default. The federated credential subject is system:serviceaccount:<drasi-namespace>:source.<source-name>. Annotate this SA with azure.workload.identity/client-id and azure.workload.identity/use=true AFTER drasi apply creates it, then restart source pods. The identity field is a top-level spec field (verified against Drasi 0.10.0 openapi ServiceIdentityDto), not inside spec.properties.

Readiness checks

For Drasi for Kubernetes:

drasi list source -n "$DRASI_NAMESPACE"
drasi describe source "$SOURCE_NAME" -n "$DRASI_NAMESPACE"
drasi wait source "$SOURCE_NAME" -n "$DRASI_NAMESPACE" -t 120

A Source must be available before dependent queries are created or trusted. If a query was created before the Source was available, recreate the query after the Source is ready.

Stability gate before query rollout (required)

Do not treat one successful drasi wait source as sufficient for production query rollout.

Run a short stability gate:

drasi list source -n "$DRASI_NAMESPACE"
# wait 5 minutes
drasi list source -n "$DRASI_NAMESPACE"
# wait 5 minutes
drasi list source -n "$DRASI_NAMESPACE"

Only proceed when the target source remains AVAILABLE=true across the full window and reactivator logs do not show reconnect/authentication loops.

If stability fails, stop query rollout and fix source prerequisites first.

PostgreSQL source checklist

Before applying a PostgreSQL Source:

  • Confirm database engine and hosting model.
  • Confirm logical replication or CDC prerequisites for that service.
  • Confirm the user has replication/read permissions required by the provider.
  • Use schema-qualified table names, for example public.Customer.
  • Confirm network path from Drasi runtime to the database.
  • Use TLS when required by the database service.
  • Store credentials securely and plan rotation.
  • Validate source availability before creating queries.
  • If a source fails with No primary key found for <table>, treat it as a provider/runtime compatibility defect until proven otherwise. Capture drasi describe source output and table DDL evidence, then temporarily remove the failing table from tables to restore pipeline health while escalating.

Mandatory SQL preflight for PostgreSQL Source

Before drasi apply on PostgreSQL sources, run SQL preflight against the target database and capture evidence:

-- 1) Target tables actually exist in the expected schema
SELECT table_schema, table_name
FROM information_schema.tables
WHERE table_schema = '<schema>'
ORDER BY table_name;

-- 2) Every table referenced by Drasi has a primary key
SELECT
  c.relname AS table_name,
  coalesce(bool_or(ct.contype='p'), false) AS has_pk,
  string_agg(CASE WHEN ct.contype='p' THEN a.attname END, ',' ORDER BY a.attnum)
    FILTER (WHERE ct.contype='p') AS pk_columns
FROM pg_class c
JOIN pg_namespace n ON n.oid = c.relnamespace
LEFT JOIN pg_constraint ct ON ct.conrelid = c.oid AND ct.contype = 'p'
LEFT JOIN pg_attribute a ON a.attrelid = c.oid AND a.attnum = ANY(ct.conkey)
WHERE n.nspname = '<schema>' AND c.relkind = 'r'
GROUP BY c.relname
ORDER BY c.relname;

If schema tables are missing, fix database initialization first (migrations/schema apply) before any Drasi source/query triage.

Azure PostgreSQL prerequisite gate for Drasi (required) - ordered checklist

Before enabling or re-enabling PostgreSQL Drasi queries on Azure Database for PostgreSQL Flexible Server, follow this exact order. Skipping or reordering steps produces hard-to-diagnose failures.

  1. Firewall/network. Confirm network path from AKS to PostgreSQL. For public access, verify firewall rules allow the AKS egress IP range. For private access, verify Private Endpoint or VNet integration.

  2. wal_level. Confirm wal_level is logical AND no restart is pending:

    az postgres flexible-server parameter show --resource-group <rg> --server-name <server> --name wal_level
    # Verify: value == "logical" AND isConfigPendingRestart == false

    If isConfigPendingRestart is true, restart the server before proceeding.

    az postgres flexible-server restart --resource-group <rg> --name <server>
  3. Replication role. Create or confirm a dedicated PostgreSQL role for Drasi CDC:

    CREATE ROLE drasi_replication WITH REPLICATION LOGIN PASSWORD '<alphanumeric-password-only>';
    GRANT CONNECT ON DATABASE <db> TO drasi_replication;
    GRANT USAGE ON SCHEMA <schema> TO drasi_replication;

    Password restriction. Use only alphanumeric characters in replication passwords. Special characters (!, @, #, $, %, ^, &, *) are transformed differently by azd env, Kubernetes Secret templates, and the PostgreSQL JDBC driver, causing silent auth failures.

  4. Table ownership. Transfer ownership of all CDC tables to the replication role. Without this, the Debezium engine cannot create the filtered publication and fails with must be owner of <table>:

    ALTER TABLE <schema>.<table> OWNER TO drasi_replication;
    -- Repeat for every table in the source's table list
  5. Role membership chain. If the replication role needs azure_pg_admin inherited privileges, grant the server admin role to the replication role:

    GRANT <server-admin-role> TO drasi_replication;

    Verify the chain:

    SELECT r.rolname, array_agg(m.rolname) AS member_of
    FROM pg_roles r
    LEFT JOIN pg_auth_members am ON r.oid = am.member
    LEFT JOIN pg_roles m ON am.roleid = m.oid
    WHERE r.rolname = 'drasi_replication'
    GROUP BY r.rolname;
  6. Cleanup stale publication and slot (if reconfiguring). If switching users or reconfiguring, drop the old publication and replication slot before applying the source:

    DROP PUBLICATION IF EXISTS <publication_name>;
    SELECT pg_drop_replication_slot('<slot_name>');

    Confirm the slot is gone: SELECT slot_name FROM pg_replication_slots WHERE slot_name = '<slot_name>';

  7. Apply Drasi source. Apply the PostgreSQL source manifest with password-auth credentials (not Workload ID - see the Workload Identity conflict note below).

  8. Wait and stabilise. drasi wait source, then run the 10-minute stability gate before adding queries.

  9. Apply queries. Only after the source has been stable for 10+ minutes with no auth or connectivity errors in reactivator logs.

Operational check commands:

az postgres flexible-server parameter show --resource-group <rg> --server-name <server> --name wal_level
az postgres flexible-server show --resource-group <rg> --name <server> --query "network" -o json
az postgres flexible-server execute --name <server> --admin-user <admin> --admin-password <pwd> --database-name <db> --querytext "SELECT slot_name, active FROM pg_replication_slots;"
kubectl logs -n <drasi-namespace> <postgres-reactivator-pod> -c reactivator --tail=200

If logs show wal_level ... is: 'replica', password authentication failed, must be owner of table, or repeated Connect timed out, treat source as unstable and do not onboard new queries yet.

Workload Identity conflict with password-auth sources (Azure PostgreSQL on AKS). When a source-specific service account (for example source.<source-name>) has the azure.workload.identity/client-id annotation AND the azure.workload.identity/use=true label, the Azure Workload Identity webhook injects AZURE_CLIENT_ID, AZURE_TENANT_ID, AZURE_FEDERATED_TOKEN_FILE, and AZURE_AUTHORITY_HOST environment variables into pods using that service account, including PostgreSQL source proxy and reactivator pods. When these env vars are present, the PostgreSQL JDBC driver may attempt Entra token authentication instead of using the password from the Kubernetes Secret, regardless of the password field in the source manifest. The result is FATAL: The access token has invalid format or password authentication failed in proxy/query-api logs, even when the source reactivator connects fine.

Fix: If the PostgreSQL source uses password-based auth from a Kubernetes Secret, remove the workload identity annotation and label from the source-specific service account before applying the source:

kubectl annotate serviceaccount source.<source-name> --namespace <drasi-namespace> azure.workload.identity/client-id-
kubectl label serviceaccount source.<source-name> --namespace <drasi-namespace> azure.workload.identity/use-

Then delete the source pods to force recreation without the env vars. The Event Hubs source can continue using Workload ID when the deployment script annotates the specific source service account (source.<source-name>) used by that source - verify via `kubectl get pod <eventhub-reactivator-pod> -n <ns> -o jsonpath='{.spec.serviceAccountName}'.

Do not hard-code wal_level or replication configuration blindly across all PostgreSQL hosting options. Azure Database for PostgreSQL, self-hosted PostgreSQL, and local test containers can require different setup steps.

Do not hard-code wal_level or replication configuration blindly across all PostgreSQL hosting options. Azure Database for PostgreSQL, self-hosted PostgreSQL, and local test containers can require different setup steps.

Workload Identity for Azure-backed sources (Drasi for Kubernetes on AKS)

Use this pattern when Drasi is deployed on AKS and sources connect to Azure-managed services (PostgreSQL Flexible Server, Event Hubs).

Prerequisites

  • AKS cluster has OIDC issuer enabled (oidcIssuerProfile.enabled: true) and workload identity enabled (securityProfile.workloadIdentity.enabled: true).
  • A user-assigned managed identity exists in Azure.
  • A federated identity credential is bound from the managed identity to the source-specific Drasi service account (source.<source-name>).

Important - Drasi for Kubernetes uses per-resource service accounts. Source pods use source.<source-name> and reaction pods use reaction.<reaction-name> (verified against Drasi for Kubernetes 0.10.0). The federated credential subject for sources is system:serviceaccount:<drasi-namespace>:source.<source-name>. Re-check with kubectl get pod <source-pod> -n <ns> -o jsonpath='{.spec.serviceAccountName}' before writing the federated credential.

Bicep setup (in identity-rbac.bicep or equivalent)

@description('OIDC issuer URL from the AKS cluster.')
param aksOidcIssuerUrl string

@description('Kubernetes namespace where Drasi is installed.')
param drasiNamespace string = 'drasi-system'

resource drasiSourcesIdentity 'Microsoft.ManagedIdentity/userAssignedIdentities@2023-01-31' = {
  name: 'id-${namePrefix}-drasi-sources-${suffix}'
  location: location
  tags: tags
}

resource drasiSourcesFederatedCredential 'Microsoft.ManagedIdentity/userAssignedIdentities/federatedIdentityCredentials@2023-01-31' = {
  parent: drasiSourcesIdentity
  name: 'drasi-system-source'
  properties: {
    issuer: aksOidcIssuerUrl
    subject: 'system:serviceaccount:${drasiNamespace}:source.<source-name>'
    audiences: ['api://AzureADTokenExchange']
  }
}

// Event Hubs Data Receiver — use identity.id in guid(), not principalId (not available at plan time)
resource drasiEventHubReceiverAssignment 'Microsoft.Authorization/roleAssignments@2022-04-01' = {
  name: guid(eventHubNamespace.id, drasiSourcesIdentity.id, eventHubsDataReceiverRoleId)
  scope: eventHubNamespace
  properties: {
    principalId: drasiSourcesIdentity.properties.principalId
    roleDefinitionId: eventHubsDataReceiverRoleId
    principalType: 'ServicePrincipal'
  }
}

output drasiSourcesIdentityClientId string = drasiSourcesIdentity.properties.clientId
output drasiSourcesIdentityPrincipalId string = drasiSourcesIdentity.properties.principalId
output drasiSourcesIdentityName string = drasiSourcesIdentity.name

Pass aksOidcIssuerUrl: drasiAks.outputs.oidcIssuerUrl where oidcIssuerUrl is an output of the AKS module: output oidcIssuerUrl string = cluster.properties.oidcIssuerProfile.issuerURL.

Post-provision script steps (before applying sources)

# 1. Add managed identity as PostgreSQL Flexible Server Entra admin
az postgres flexible-server microsoft-entra-admin create `
  --resource-group $resourceGroupName `
  --server-name $serverName `
  --display-name $identityName `
  --object-id $identityPrincipalId `
  --type ServicePrincipal

# 2. Annotate and label the Drasi source-specific service account
kubectl annotate serviceaccount source.<source-name> `
  --namespace drasi-system `
  "azure.workload.identity/client-id=$identityClientId" `
  --overwrite

kubectl label serviceaccount source.<source-name> `
  --namespace drasi-system `
  "azure.workload.identity/use=true" `
  --overwrite

If the source service account does not exist yet, run drasi apply first so the platform creates source.<source-name>, then run step 2, then restart source pods.

Source YAML shape

apiVersion: v1
kind: Source
name: my-postgres-source
spec:
  kind: PostgreSQL
  identity:
    kind: MicrosoftEntraWorkloadID
    clientId: <managed-identity-client-id>
  properties:
    host: <server>.postgres.database.azure.com
    port: 5432
    user: <managed-identity-display-name>   # as registered in PostgreSQL Entra admin
    database: <db-name>
    ssl: true
    tables:
      - schema.table_name
---
apiVersion: v1
kind: Source
name: my-eventhub-source
spec:
  kind: EventHub
  identity:
    kind: MicrosoftEntraWorkloadID
    clientId: <managed-identity-client-id>
  properties:
    host: <namespace>.servicebus.windows.net
    eventHubs:
      - <hub-name>
    bootstrapWindow: 0

Verifying token injection

After applying sources, confirm the workload identity webhook injected the token:

kubectl get pod <source-reactivator-pod> -n drasi-system \
  -o jsonpath='{.spec.containers[0].env[*].name}' | tr ' ' '\n' | grep AZURE
# Expected: AZURE_CLIENT_ID  AZURE_TENANT_ID  AZURE_FEDERATED_TOKEN_FILE  AZURE_AUTHORITY_HOST

If these env vars are absent, check that the source-specific SA has both the annotation and label and that the workload identity webhook is running in the cluster.

!CAUTION Workload Identity env vars can break password-auth sources. When AZURE_CLIENT_ID, AZURE_TENANT_ID, and AZURE_FEDERATED_TOKEN_FILE are present in the source pod environment, the PostgreSQL JDBC driver may prefer Entra token authentication over the password from the Kubernetes Secret, even when the source manifest specifies password from a kind: Secret reference. The proxy and query-api services fail with FATAL: The access token has invalid format. See the "Azure PostgreSQL prerequisite gate" note above for the fix - remove the annotation from the source-specific SA before applying password-auth sources.

azd output naming

When Bicep outputs managed identity properties, azd converts camelCase output names to UPPER_SNAKE_CASE environment variables automatically. For example, drasiSourcesIdentityClientId → DRASI_SOURCES_IDENTITY_CLIENT_ID. Use these as template variables in source YAML files and resolve them in provisioning scripts via azd env get-values or azd exec.

PostgreSQL replication slot operational runbook

Drasi PostgreSQL Source uses logical replication, which creates a replication slot in PostgreSQL that retains WAL until Drasi confirms it. Slots that go unconfirmed will bloat WAL and can take a database down. Treat replication slots as managed state, not throwaway plumbing.

Callout - PostgreSQL 19+ dynamic-effective wal_level side effect (cluster-wide). On PostgreSQL 19+, wal_level behaves as a floor: when the first logical replication slot is created on a server running wal_level = replica, Postgres automatically bumps the effective level to logical cluster-wide. This means a Drasi PostgreSQL Source that creates a slot on a PG19+ server has a CLUSTER-WIDE side effect - increased WAL volume affecting every database in the cluster, not just the Drasi-monitored one. Operators of PG19+ databases must be told this explicitly before a Drasi Source is provisioned; capacity-plan WAL disk and downstream archive throughput accordingly. Reference: https://thebuild.com/blog/2026/05/11/the-wallevel-you-set-is-not-the-wallevel-you-get/.

Outage class: WAL accumulation (slot retains WAL → disk fills → DB write outage). Slot leakage is not the outage - WAL accumulation is. A retained-but-unconsumed slot accumulates WAL on the upstream Postgres volume; once the volume fills, all writes to every database in that cluster fail, not just Drasi traffic. Tie the slot runbook explicitly to disk-space monitoring on the Postgres side, and alert on per-slot retained WAL as a leading indicator before disk-pressure becomes a write outage.

Sample alert SQL (run on a schedule against the upstream Postgres; warn at > 1 GiB, page at > 10 GiB per slot):

-- Per-slot retained WAL (leading indicator of WAL accumulation outage)
SELECT
  slot_name,
  plugin,
  active,
  pg_wal_lsn_diff(pg_current_wal_lsn(), confirmed_flush_lsn) AS retained_bytes,
  pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), confirmed_flush_lsn)) AS retained_pretty,
  CASE
    WHEN pg_wal_lsn_diff(pg_current_wal_lsn(), confirmed_flush_lsn) > 10 * 1024^3 THEN 'page'
    WHEN pg_wal_lsn_diff(pg_current_wal_lsn(), confirmed_flush_lsn) >  1 * 1024^3 THEN 'warn'
    ELSE 'ok'
  END AS severity
FROM pg_replication_slots
WHERE plugin = 'pgoutput';

Wire this query into the same alerting pipeline that watches Postgres disk-free percentage; either signal firing alone is insufficient - retained-WAL growth is the leading indicator, disk-free is the lagging one. Reference: POSETTE 2026 talk "Building event-driven systems with PostgreSQL logical replication and Drasi" - https://posetteconf.com/2026/talks/building-event-driven-systems-with-postgresql-logical-replication-and-drasi/.

Daily monitoring queries

-- Slots used by Drasi and their unconsumed WAL size
SELECT
  slot_name,
  plugin,
  active,
  restart_lsn,
  confirmed_flush_lsn,
  pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal
FROM pg_replication_slots
WHERE plugin = 'pgoutput';

-- Replication lag in bytes for active Drasi consumers
SELECT
  application_name,
  state,
  pg_size_pretty(pg_wal_lsn_diff(sent_lsn, flush_lsn)) AS flush_lag,
  pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn)) AS replay_lag
FROM pg_stat_replication;

Alert thresholds

Signal Warning Critical
retained_wal for a slot > 1 GB > 10 GB or > 25% of disk
Slot active = false for a known Drasi slot > 5 min > 30 min
flush_lag on an active replica > 1 min steady state > 5 min

Recovery actions (least-destructive first)

  1. Confirm the slot still maps to a running Drasi Source: drasi list source and match the slot name pattern documented in your Source config.
  2. If the Source is intentionally retired, drop the slot in PostgreSQL: SELECT pg_drop_replication_slot('<slot_name>');. Never drop a slot whose owner you cannot identify.
  3. If the Source pod has been pending for >30 min and the slot is bloating, capture diagnostics first (kubectl describe, recent logs, slot LSNs), then follow the recovery bundle's safe-recovery ladder. Do not delete the slot before exporting evidence.

Equivalent operational checks for other CDC-backed sources

  • SQL Server: query sys.dm_cdc_log_scan_sessions and sys.dm_cdc_errors weekly; verify CDC retention matches your Drasi catch-up window.
  • Dataverse: verify the change-tracking token has not expired (Dataverse expires tokens after extended inactivity).
  • Kubernetes Source: verify the watch cache has not been desynced via kubectl get events -n <ns> --field-selector reason=ResourceVersionTooOld.

Event Hub source checklist

Before applying an Azure Event Hub Source:

  • Confirm the Event Hubs namespace and event hub names exist in your Azure subscription.
  • Choose an auth strategy:
    • Microsoft Entra Workload ID (preferred on AKS): assign Azure Event Hubs Data Receiver RBAC role on the Event Hubs namespace to the managed identity. Set host to the namespace FQDN (e.g., your-namespace.servicebus.windows.net). Omit connectionString.
    • Connection string: use the namespace-level connection string from Azure Portal. Store as a Kubernetes Secret and reference via kind: Secret.
  • Configure consumerGroup (default: $Default). Create a dedicated consumer group per Drasi Source to avoid conflicts with other consumers. [VERIFY: consumerGroup is absent from the current docs property table for the EventHub source ; confirm it is still accepted via drasi describe sourceprovider EventHub on the live cluster before relying on it.]
  • Set bootstrapWindow (integer, default 0): number of minutes of historical events to replay when the query bootstraps (verified 2026-08-16 against the EventHub source docs). A non-zero value helps recover missed events after restarts but increases initial load. Recommended: set it to at least your maximum expected downtime in minutes.
  • List the event hub names in eventHubs. Each event hub in the namespace gets its own change stream. Each hub name becomes a Cypher node label; escape dashed names with backticks in queries (MATCH (e:`order-events`)).
  • Event Hub messages are treated as source changes in Drasi. The message body becomes the element data; system properties (enqueued time, offset, sequence number) are available for query predicates.
  • Confirm network access: the Event Hubs namespace must be reachable from the Drasi runtime (public endpoint with minimal TLS, Private Endpoint within the same VNet, or a configured firewall rule).
  • Start with bootstrapWindow: 0 during initial setup, then increase after confirming stable connectivity.

YAML example (Workload ID):

apiVersion: v1
kind: Source
name: order-events
spec:
  kind: EventHub
  properties:
    host: "orders-ns.servicebus.windows.net"
    consumerGroup: "drasi-orders"   # [VERIFY] see checklist note above
    eventHubs:
      - "order-placed"
      - "order-shipped"
    bootstrapWindow: 5   # MINUTES of backfill (300 would be 5 hours)
  identity:
    kind: MicrosoftEntraWorkloadID
    clientId: "00000000-0000-0000-0000-000000000000"

Validation:

Check Method
Source is available drasi wait source order-events -t 120
Events are flowing Send a test event to the Event Hub; verify the query result changes
Consumer group is isolated Check that no other consumer group with the same name exists on the same Event Hub
bootstrapWindow replay works Stop the Source, send events, restart; verify historical events are picked up within the window (minutes, not seconds)

Identity rules

Use this decision order:

  1. Provider documents Microsoft Entra Workload ID and the environment supports it: use workload identity.
  2. Provider documents Kubernetes Secret references: use Secret references and scope RBAC tightly.
  3. Provider only documents inline properties: keep examples for local/dev only, and add a production warning.
  4. Provider support is unclear: stop and verify docs before generating production code.

Dependency impact

Changing a Source can invalidate dependent queries. For providers that require delete/recreate:

  1. Disable or delete reactions that consume affected queries.
  2. Delete affected queries.
  3. Delete the Source.
  4. Apply the updated Source.
  5. Wait until the Source is available.
  6. Reapply queries.
  7. Reapply reactions.
  8. Run the validation bundle.

Exit criteria

Source work is complete only when:

  • The provider schema was checked against current docs.
  • CDC or watch prerequisites are enabled and evidenced.
  • Credentials or identity are least-privilege and stored correctly.
  • Network path and TLS expectations are verified.
  • The Source is available before dependent queries are trusted.
  • Provider-specific update behaviour and blast radius are documented.
  • Insert, update, and delete source changes are ready for validation through query and reaction consumers.

PostgreSQL source proxy client driver mismatch

The Drasi PostgreSQL source proxy is a Node.js app using knex for database connectivity. The resource provider passes the database connector type as the env var connector=PostgreSQL, but knex expects the client driver name in an env var called client (e.g. client=pg).

Symptom: The stepup-proxy pod crashes with:

Error: knex: Required configuration option 'client' is missing.

The pod logs show client: undefined in the dbConfig JSON.

Fix: Set the client env var on the proxy deployment:

kubectl set env deployment/stepup-proxy -n drasi-system client=pg

Root cause: The resource provider template uses connector as the config key name but the Node.js sourceProxy.js reads client. This is a naming inconsistency in the Drasi source-proxy image.

PostgreSQL source state store query compatibility

The Drasi source change-svc component uses the Dapr state query API to manage source offset state. The resource provider creates a Dapr component named rg-state with type state.redis, but the Dapr Redis state store only supports queries when Redis has the RedisJSON module (Redis Stack). Standard Redis returns:

state store rg-state query failed: redis-json server support is required for query capability

Fix (0.9.x platform): Replace the default state.redis component with state.mongodb pointing to the existing Drasi MongoDB:

kubectl apply -n drasi-system -f - <<YAML
apiVersion: dapr.io/v1alpha1
kind: Component
metadata:
  name: rg-state
spec:
  type: state.mongodb
  version: v1
  metadata:
    - name: host
      value: drasi-mongo:27017
    - name: databaseName
      value: Drasi
    - name: collectionName
      value: rg-state
YAML

Note: The reactivator's DaprOffsetBackingStore (Java) also uses rg-state for saving CDC offsets. Switching to MongoDB is transparent, the Dapr state API is backend-agnostic for basic CRUD operations.

Version scope: This workaround was verified against the 0.9.x platform, where a Drasi MongoDB (drasi-mongo) deployment exists. The 0.10.x platform is Redis-only (no drasi-mongo service ; see SKILL.md "Cross-minor architecture changes"), so this exact fix does not apply there. [VERIFY: On 0.10.x, check whether change-svc still uses a Dapr state query against rg-state before adapting this workaround ; inspect the running Dapr components in the Drasi namespace.]

apiVersion migration playbook

Drasi for Kubernetes is pre-1.0. The current resource apiVersion is v1; the platform may bump it (for example, v1 to a future v2) without a guaranteed backward-compatibility promise. When the platform release ships a new apiVersion, the operator must run this playbook before promoting the upgraded platform past dev:

  1. Detect drift. Inspect the active server's expected apiVersion against the manifests on disk:

    # [VERIFY: `-o yaml` output support is not in the current CLI docs; if the flag
    # is rejected at your CLI version, use `drasi describe <kind> <name>` per resource.]
    drasi list source -n "$DRASI_NAMESPACE" -o yaml | grep -E '^apiVersion:'
    drasi list query -n "$DRASI_NAMESPACE" -o yaml | grep -E '^apiVersion:'
    drasi list reaction -n "$DRASI_NAMESPACE" -o yaml | grep -E '^apiVersion:'

    Compare the observed apiVersion to the manifests committed in drasi/sources, drasi/queries, drasi/reactions.

  2. Export current manifests. Snapshot every Source, ContinuousQuery, and Reaction (drasi list <kind> -o yaml > snapshots/<kind>.yaml) before any conversion. Keep the prior-version manifests in version control until post-upgrade validation evidence is captured.

  3. Convert per release notes. Apply the per-release migration notes from the upstream changelog. Use the release-notes diff workflow described in bundles/currency-and-ai-context/guide.md to drive the field-by-field conversion. Do not invent migrations.

  4. Re-apply in dependency order. Apply converted manifests in forward dependency order: sources first (wait until READY), then queries (wait until ACTIVE), then reactions. Do not skip readiness checks.

  5. Validate end-to-end. Run the synthetic source change to query update to reaction effect test from the validation bundle on the new apiVersion before marking the migration complete.

  6. Retain prior-version manifests until validation evidence is captured. Only delete the prior-version snapshots after a green synthetic test, a delivery PR records the upgrade evidence (use templates/drasi-upgrade-evidence.md), and the rollback plan is no longer required.

Pre-1.0 apiVersion changes are NOT guaranteed to be backward-compatible. Treat every apiVersion bump as a breaking change until the release notes prove otherwise.

Source: SKILL.md on GitHub

1 warning8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Drasi skill package is a highly structured and security-conscious set of instructions for managing data change detection pipelines. It includes extensive documentation on threat modeling, workload identity setup on AKS, and specific guidance for preventing prompt injection when source data is fed into AI agents. All documented commands and scripts are legitimate operational tools for the Drasi platform, and no malicious patterns such as obfuscation, persistence, or data exfiltration were found.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: MEDIUM · 1 issue

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
metadata
{
  "last_verified": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/drasi