Drasi Schema Evolution and Production Failure Patterns
Purpose
Reference for handling source schema changes in Drasi and common production failure patterns that arise from organizational (not just technical) issues.
Schema evolution
What happens when a Source schema changes
| Change type | Impact on continuous queries | Mitigation |
|---|---|---|
| Column added | Query unaffected if not referenced | No action needed |
| Column removed | Query fails if referencing removed column | Update query, redeploy |
| Column type changed | Query may produce errors or wrong results | Verify query compatibility, update |
| Table renamed | Source stops producing events for old name | Update Source config and query |
| New table added | Not visible to existing queries | Add to Source config, create new query |
| Table dropped | Source may error | Remove from Source config |
Best practices
- Version Source schemas. Track table/column definitions alongside Source YAML.
- Test queries against staging before production. Apply schema changes to a staging Source first.
- Use query defensive patterns. Avoid hardcoding column names that might change; use optional fields where possible.
- Monitor Source logs after schema changes. Look for deserialization errors or missing events.
- Coordinate schema changes with query owners. Schema changes are a cross-team concern.
PostgreSQL schema evolution
PostgreSQL Sources use logical replication. Schema changes (ADD COLUMN, ALTER COLUMN) are generally safe for the replication slot, but:
- Dropped columns may cause the Source to emit events with missing fields
- Type changes may cause deserialization failures in the query engine
- New tables require explicit addition to the Source's
tableslist
Production failure patterns
Query sprawl
Symptom: Dozens of unused queries consuming resources, nobody knows which are active.
Prevention:
- Tag queries with owner, department, and last-used date
- Review query inventory quarterly
- Delete unused queries (they still consume Source change-tracking resources)
- Use naming conventions:
{domain}-{purpose}-{owner}
Zombie queries
Symptom: Queries that are "running" but never produce results. Sources are connected but no changes flow.
Causes:
- Source CDC not configured (replication slot exists but no publication)
- Source pointing to wrong database/schema
- Query filters excluding all data
Detection: Query shows "Running" but result count stays at 0 for >24 hours.
Reaction storms
Symptom: A single data change triggers hundreds of reaction executions.
Causes:
- Query result flapping (result toggles between two states rapidly)
- Reaction retry loop (downstream 500 → retry → same 500 → retry)
- Bootstrap re-triggering after Source restart
Prevention:
- Implement idempotent reactions
- Add rate limiting on reaction endpoints
- Monitor reaction fire rate per query, alert on >10x baseline
- Use
skipControlSignals: trueon PostDaprPubSub to suppress bootstrap messages
Source lag accumulation
Symptom: Events arriving minutes/hours late, queries processing stale data.
Causes:
- Under-provisioned query containers
- Large batch of historical changes overwhelming the pipeline
- Source connection throttling
Detection: Compare Source event timestamp with query processing timestamp. Alert on >30s drift.
Capacity exhaustion
Symptom: Queries become slow, reactions timeout, Source events queue up.
Causes:
- Too many concurrent queries on one query container
- Complex queries with expensive joins
- Large result sets overwhelming reaction throughput
Prevention:
- Distribute queries across multiple query containers (see
scaling-and-capacitybundle) - Set resource limits per query container
- Monitor CPU/memory utilization, alert on >70% sustained