All skills

USE FOR: Drasi continuous-query solutions - real-time queries, change detection, reactive events, data-trigger pipelines on Drasi Server, Drasi for Kubernetes, or drasi-lib. Router: load bundle guides as needed. DO NOT USE for non-Drasi messaging (event-driven-messaging) or pure AKS/ACA hosting (aks-cluster-architecture, azure-container-apps).

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/drasi

This session only. Nothing lands on disk.

bundlesscaling-and-capacityguide.md

≈5.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Scaling and capacity bundle

Use this bundle for production sizing, scale-out design, throughput testing, backpressure handling, HPA/KEDA guidance, and multi-cluster considerations.

Scaling principle

Scale the bottleneck, not the whole solution blindly. In Drasi for Kubernetes, Sources, query containers, and Reactions can have different constraints and failure modes. In Drasi Server and drasi-lib, scaling often means scaling the hosting process, partitioning workloads, or changing the architecture.

Runtime-specific guidance

Drasi for Kubernetes

Assess these independently:

  • Sources: source system change rate, CDC/binlog/WAL throughput, network latency, permissions, connection limits, and source connector health.
  • Query containers: query complexity, number of active queries, join cardinality, result-set size, memory pressure, and CPU saturation.
  • Reactions: downstream throughput, retries, idempotency, dead-letter handling, rate limits, and backpressure.
  • Storage and state: query state growth, persistence settings, retention, and recovery behavior.

Use Kubernetes-native scaling only after confirming the Drasi component supports the intended replica pattern for the current version and provider.

Drasi Server

Assess:

  • CPU/memory saturation of the server process.
  • Source/reaction plugin contention.
  • Persistent state location and I/O.
  • Container Apps revision concurrency and scaling limits.
  • Whether the workload should remain on Drasi Server or move to Drasi for Kubernetes.

Do not assume multiple Drasi Server replicas can safely share the same state or source ownership unless the current server docs and architecture explicitly support that pattern.

drasi-lib

Assess:

  • Host application concurrency model and async runtime.
  • In-process state store behavior.
  • Source ingestion backpressure.
  • Reaction execution isolation.
  • Whether embedded processing should remain in-process or move to Drasi Server/Kubernetes.

Capacity questions

Before production, answer:

  • What is the expected source change rate at normal, peak, and burst traffic?
  • What is the largest expected query result set?
  • Which queries join across high-cardinality entities?
  • Which reaction is slowest or least reliable?
  • What happens if the downstream target is unavailable for 5, 30, or 120 minutes?
  • What metrics prove lag, backlog, processing latency, and reaction failure rate?
  • What is the rollback plan if scaling changes increase duplicate delivery or state drift?

Autoscaling guidance

  • Prefer measured scaling triggers over CPU-only assumptions.
  • Use HPA for Kubernetes resource-based scale decisions where supported.
  • Use KEDA for event-driven or custom/external metric scaling where it fits the workload and current Drasi component supports the scale target.
  • For reaction workloads, design idempotency before increasing concurrency.
  • Avoid scale-to-zero for components that must maintain source subscriptions, query state, or long-lived streams unless current docs explicitly support it.

Load-test pattern

  1. Start with one Source, one representative ContinuousQuery, and one Reaction.
  2. Seed realistic baseline data.
  3. Generate controlled changes at increasing rates.
  4. Measure source lag, query latency, result correctness, reaction latency, downstream errors, CPU, memory, and restarts.
  5. Add high-cardinality or join-heavy query cases.
  6. Add downstream throttling/failure tests.
  7. Document the tested ceiling and safe operating range.

Production gates

Do not call a Drasi solution production-ready until there is evidence for:

  • Sustained expected throughput.
  • Burst tolerance or backpressure behavior.
  • Query correctness under load.
  • Reaction idempotency and retry behavior.
  • Alert thresholds for lag, errors, restarts, and stale query results.
  • Resource requests/limits or Container Apps scaling settings.
  • Recovery behavior after restart, source reconnect, and downstream outage.

Warning signs

  • Queries rely on very broad MATCH patterns without selective predicates.
  • Reactions call slow APIs synchronously without retry/idempotency planning.
  • Database CDC permissions are correct but the source system cannot handle the extra load.
  • Metrics show healthy pods but stale query results.
  • Scaling increases duplicate downstream effects because reactions are not idempotent.
  • Load testing only validates resource creation, not source-change-to-reaction behavior.

Production scale exit criteria

Scaling work is complete only when:

  • The bottleneck is identified with evidence rather than assumed.
  • Load-test data covers expected and peak source-change rates.
  • Source, query, reaction, and downstream limits are considered together.
  • HPA/KEDA or manual scaling settings are documented with rollback values.
  • Backpressure, retry, throttling, and duplicate-delivery behaviour are tested or explicitly risk-accepted.
  • A lightweight synthetic canary or scheduled validation confirms event flow after scale changes.
  • Monitoring and alerts reflect the new capacity assumptions.

Autoscaling trigger matrix

Drasi components have different scaling shapes. Pick the trigger that matches the bottleneck, not the default CPU autoscaler.

Component Bottleneck Preferred trigger Notes
Source (CDC: PostgreSQL, SQL Server) WAL/log lag or producer throughput KEDA on connector lag (or replication slot lag) CPU-based HPA misses bursty CDC. Verify the source connector exposes a lag metric you can scrape before relying on KEDA.
Source (event-driven: Event Hub, webhook) Backlog depth KEDA on queue/topic lag Configure minReplicaCount=1 for sources that lose offsets at zero replicas.
Query container Evaluation CPU and memory pressure on hot queries HPA on CPU + memory; tune memory request based on result-set cardinality Set conservative maxReplicas; a runaway join can outscale a small cluster. Pair with a PodDisruptionBudget.
Reaction (HTTP, gRPC, SignalR) Downstream concurrency / fan-out HPA on CPU; KEDA on outbound queue depth if a buffer is involved Respect downstream rate limits - set maxReplicas low enough not to DoS the target.
Reaction (Event Grid, MCP, SSE) Mostly request-rate-bound HPA on CPU; KEDA on inbound request rate For SSE/MCP, scale-in must be graceful (drain open streams).

Verification before declaring autoscaling production-ready

  • Run a 30-minute load test that reaches at least 2× steady-state throughput; confirm scale-out occurs within the configured cooldown, then scale-in returns to baseline.
  • Confirm kubectl describe hpa and kubectl describe scaledobject show no metric-collection errors over the test window.
  • Capture before/after lag, p95 query freshness, and reaction delivery success during the load test as scaling evidence.

Re-check the Drasi platform release notes for the version you've pinned - supported KEDA scalers and exposed metrics can change between minor versions.

Per-component sizing template

Sizing by guesswork produces over-provisioned dev environments and under-provisioned production. Capture these inputs in a sizing.md per environment before requesting capacity. Verify all defaults against the connector and reaction reference docs at the pinned release.

Component Required inputs Starting allocation Scale signal
Source (PostgreSQL / SQL Server CDC) Events/sec at p50 and p99; row size in bytes; expected schema-change frequency 0.5 vCPU / 1 GiB; minReplicaCount=1 KEDA on slot/CDC lag bytes
Source (event-driven: Event Hub, webhook) Inbound requests/sec at p50 and p99; per-event payload size; max burst duration 0.5 vCPU / 0.5 GiB; minReplicaCount=1 KEDA on queue depth
Query container Active query count; p95 result-set cardinality per query; join fan-out; window/aggregation memory footprint 1 vCPU / 2 GiB per replica; one container per 5-10 small queries OR one container per heavy query HPA on memory + CPU; tune by measured p95 cardinality
Reaction (HTTP / gRPC) Outbound RPS at p50 and p99; downstream p95 latency; retry budget 0.25 vCPU / 0.5 GiB; concurrency = floor(target_RPS × downstream_p95_seconds × 1.5) HPA on CPU; KEDA on outbound queue depth if buffered
Reaction (SSE / MCP / SignalR) Concurrent connected clients; per-client message rate; idle timeout 0.5 vCPU / 1 GiB per 500 connected clients KEDA on active connection count

For every row, record the measured value next to the estimated value after the first month in production. The delta is the input to the next sizing review.

Storage budgeting (indexes, state stores, WAL retention)

Drasi's stateful surfaces consume disk you must budget - under-provisioning is a Sev-2 outage class.

Drasi state store (per component)

Confirm the backing store at the pinned release (RocksDB and Garnet are the two documented options in the platform window; check release notes for your version). Both grow with active query state.

  • Per-query state: roughly result_cardinality × per_row_bytes × 1.5 (1.5x accounts for index overhead and tombstones). Re-measure after the first week; if growth has not stabilised, the query likely has unbounded retention.
  • Retention: state stores do not auto-evict by default. If a query joins a slowly-growing dimension table, plan to either (a) re-bootstrap the query on a schedule (see operations bundle), or (b) pre-filter the dimension at the source.
  • Disk alert: 70% used → warn; 85% used → page. Capture both per-PV (Kubernetes) and per-container-volume (Drasi Server).

PostgreSQL replication-slot WAL retention

Slot-retained WAL is a budget item, not a free resource. Budget it explicitly:

  • Steady-state WAL: confirmed_flush_lsn should track within seconds of pg_current_wal_lsn(). If lag bytes climb past max_slot_wal_keep_size, Postgres will drop the slot (Drasi loses its position) rather than fill the disk.
  • Budget formula: max_slot_wal_keep_size ≥ peak_WAL_per_minute × max_acceptable_consumer_outage_minutes. Default suggestion: size for a 30-minute consumer outage at p99 WAL rate.
  • Disk headroom: the DB volume MUST have headroom = max_slot_wal_keep_size × (number_of_drasi_slots + 1) beyond steady-state needs. Without this headroom, a brief Drasi outage can trigger a source-DB write outage.

Event-driven sources (Event Hub, Kafka, etc.)

Retention is set on the source service, not on Drasi - but Drasi's tolerable outage budget is bounded by source-side retention. Record the source retention (hours/days) and the Drasi consumer-group's worst-case lag-tolerance against it in sizing.md.

Node pool placement (Drasi for Kubernetes)

Drasi components have different durability and scheduling requirements. Place them on the right node pool before production.

Component Node pool Spot / preemptible eligible? Rationale
Source containers (CDC, event-driven) User pool, on-demand No A source on a spot/preemptible node loses its replication-slot lease on eviction; reconnection may replay events or hit "slot in use" errors. The cost saving is rarely worth the operational tax.
Query containers User pool, on-demand for production No for production; yes for staging only Query state is in-memory at most pinned releases. Eviction triggers a re-bootstrap that may take minutes for high-cardinality queries; SLO violations are likely during the rebuild.
Reaction containers (HTTP, gRPC, Event Grid) User pool, on-demand Yes if reactions are idempotent and queue-bounded Stateless dispatchers tolerate eviction provided downstream idempotency holds. Confirm DLQ behaviour before scheduling on spot.
Reaction containers (SSE, MCP) User pool, on-demand No Connected-client streams are lost on eviction; clients must reconnect. Spot eviction during business hours is a visible UX regression.
Drasi operator / control plane System pool (or dedicated control-plane pool) No Control plane outage = no reconciliation. Keep on stable nodes.

nodeSelector / tolerations / affinity examples MUST be confirmed against the pinned platform release's chart values; defaults change between minor versions.

Cost-vs-resilience trade summary: Drasi's cost optimisation surfaces are reaction-tier spot eligibility and right-sizing query containers. Sources, query state, and streaming reactions are not safe spot candidates at the pinned release window.

Backpressure propagation (source → query → reaction)

Backpressure in Drasi is end-to-end. When a reaction's downstream slows down or fails, the pressure walks backward through the pipeline until it lands on the source system - frequently on the source database itself. Operators MUST set explicit bounds at each stage; an unbounded buffer at any one stage merely shifts the failure to the next.

The chain

  1. Reaction queue depth grows. The reaction is buffering events it cannot deliver.
    • Signal: per-reaction queue depth metric. Compare against the SLO baseline in bundles/observability/guide.md.
    • Bound: reaction queue max depth → DLQ when exceeded. Without this bound, the next stage absorbs the pressure.
  2. Query updates are throttled or buffered upstream of the reaction when the reaction queue is bounded.
    • Signal: query container memory %, CPU saturation, and update-fan-out latency.
    • Bound: query container memory limit and an alert on sustained memory pressure.
  3. Query container memory/CPU pressure rises if the query layer is the one buffering.
    • Signal: pod memory/CPU; restart count if OOM kills fire.
    • Bound: HPA on memory plus a hard resources.limits.memory so a single bad query cannot exhaust the node.
  4. Source consumption slows. For Postgres logical replication, the slot does NOT advance confirmed_flush_lsn until downstream confirms consumption, so WAL retention grows on the source DB.
    • Signal: pg_replication_slots lag bytes - query with SELECT slot_name, active, pg_wal_lsn_diff(pg_current_wal_lsn(), confirmed_flush_lsn) AS lag_bytes FROM pg_replication_slots;. Equivalent: SQL Server sys.dm_cdc_log_scan_sessions, Event Hub consumer-group lag.
    • Bound: configure the source DB with a maximum slot-retained WAL size (Postgres 13+: max_slot_wal_keep_size) so the DB itself protects against unbounded WAL growth.
  5. Source DB disk pressure. WAL retention is not free - at sustained backpressure, the source DB's volume fills.
    • Signal: source DB disk-free %, plus the pg_replication_slots lag bytes alert from step 4.
    • Bound: source-DB-side disk-space alert at a level that lets the operator act before write outage. An unhealthy Drasi reaction MUST NOT be allowed to take the source DB offline.

Operator checklist

  • Reaction queue: max depth set, DLQ destination configured, alert on depth.
  • Query container: memory limit set, alert on sustained memory pressure.
  • Source slot: max retained bytes set on the DB side, alert on slot lag.
  • Source DB volume: disk-free alert at an practical threshold.

The shape to remember: backpressure that has no DLQ at the reaction stage will eventually become a source-DB outage. Bound the cheapest stage to fail.

Pod disruption defaults

Drasi components are stateful in subtle ways - connector slot ownership, in-flight reaction queues, long-lived client streams. Voluntary disruptions (node drains, upgrades, autoscaler scale-in) MUST honour PDBs that match the component's failure mode.

Recommended PDBs

  • Source containers - minAvailable: 1 per source. A source connector typically holds the replication slot / capture-instance / consumer-group offset. Losing all replicas at once can drop the slot lease and force a re-establishment that may replay or skip events. Keep at least one replica available during voluntary disruption.
  • Query containers - maxUnavailable: 1 per query if the query container is stateful at the current release. If the query is stateless and re-derived from source replay, maxUnavailable: 1 still avoids a thundering-herd rebuild.
  • Reaction containers - maxUnavailable: 1 for SSE / MCP reactions to preserve client connections. For request/response reactions (HTTP, gRPC), maxUnavailable: 1 is also safe and avoids a delivery gap during rolling updates.

Example PDB shape:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: drasi-reaction-sse
  namespace: $DRASI_NAMESPACE
spec:
  maxUnavailable: 1
  selector:
    matchLabels:
      app.kubernetes.io/part-of: drasi
      drasi.io/reaction: <reaction-name>

Drain procedure for SSE / MCP scale-in

Scale-in for streaming reactions cannot be a hard pod kill - connected clients lose state.

  1. Mark the target pod as not-ready in the service selector (cordon equivalent) so new client connections route to remaining pods.
  2. Send a graceful-disconnect signal to active clients with a Retry-After (or protocol-specific) hint so they reconnect to the remaining pods.
  3. Poll active-connection count on the target pod. Wait for it to drop to zero or below an agreed threshold.
  4. kubectl drain the node or kubectl delete pod only after the count is drained.
  5. Verify with the synthetic test (bundles/validation/guide.md) that the source-to-reaction path is intact post-drain.

Concrete commands:

kubectl get pod -n "$DRASI_NAMESPACE" -l drasi.io/reaction=<reaction-name> -o wide
kubectl exec -n "$DRASI_NAMESPACE" <pod> -- curl -s localhost:<metrics-port>/metrics | grep active_connections
kubectl drain <node> --ignore-daemonsets --delete-emptydir-data --grace-period=120

Grace period must exceed the expected client reconnect window; tune from the reaction-specific docs at the pinned release.

Network partition behavior

Partitions are a fact of distributed operation. Drasi's behaviour at a partition depends on which leg fails. For each scenario, confirm at the current release whether the documented behaviour matches the table below.

(a) Partition between Drasi and the source DB

  • Source connector loses its connection to the DB.
  • On reconnect, it resumes from the last committed slot LSN / change-tracking position / consumer-group offset that the DB still holds.
  • No event loss is expected provided:
    • The replication slot was not dropped during the partition.
    • The source DB did not aggressively recycle WAL beyond confirmed_flush_lsn (Postgres protects this; verify equivalent on other engines).
  • Signals: source connector restart count, slot active=false while partitioned, slot lag bytes rising during the partition window.
  • Action: monitor; intervene only if the slot is reported lost. If the slot is lost, follow the destructive-recovery sign-off in bundles/recovery/guide.md before recreating.

(b) Partition between Drasi and the reaction endpoint

  • Reactions buffer or block per the graceful-degradation policy in bundles/recovery/guide.md ("Graceful degradation under downstream outage").
  • For buffer-and-retry reactions, queue depth grows; for block-source reactions (sync gRPC/SSE), source consumption stalls and WAL retention grows on the source DB.
  • Signals: reaction queue depth, reaction error rate, downstream health probe, plus the slot-lag chain from "Backpressure propagation" above.
  • Action: apply the outage-duration flowchart in the recovery bundle (5min / 30min / 2hr / sustained).

(c) Partition within the Drasi control plane

  • Pods may fail liveness, restart, and rejoin once the partition heals.
  • Verify that the platform's backing state store (etcd for vanilla Kubernetes, the managed-service equivalent on hosted Kubernetes) survives the partition. Drasi's CRDs and operator state depend on it.
  • Signals: control-plane pod restarts, operator reconciliation errors, CRD update latency.
  • Action: confirm at the current release whether the operator is leader-elected and whether a control-plane partition can cause split-brain CRD reconciliation. If the release notes do not document this, capture observed behaviour in your upgrade evidence template before declaring partition-tolerance for the deployment.

For every partition scenario, the operator MUST run the synthetic source-to-reaction test (bundles/validation/guide.md) after the partition heals before declaring recovery complete. Healthy pods do not prove healthy event flow.

Source: SKILL.md on GitHub

1 warning8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Drasi skill package is a highly structured and security-conscious set of instructions for managing data change detection pipelines. It includes extensive documentation on threat modeling, workload identity setup on AKS, and specific guidance for preventing prompt injection when source data is fed into AI agents. All documented commands and scripts are legitimate operational tools for the Drasi platform, and no malicious patterns such as obfuscation, persistence, or data exfiltration were found.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: MEDIUM · 1 issue

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
metadata
{
  "last_verified": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/drasi