All skills
clickhouse avatar

/clickhouse-architecture-advisor

@5e162d6 official
by clickhouseclickhouse/agent-skills543 stars
39

MUST USE when designing ClickHouse architectures, selecting between ingestion or modeling patterns, or translating best practices into workload-specific system designs. Complements clickhouse-best-practices with decision frameworks and explicit provenance labels.

Use this Skill: https://skilld.dev/gh/clickhouse/agent-skills/clickhouse-architecture-advisor

This session only. Nothing lands on disk.

examplesobservability-high-throughput.md

≈664 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Example: Observability — High-throughput event ingestion

Scenario

  • Workload: observability / logs
  • Ingest rate: 300K events/sec
  • Producer shape: many agents, uneven bursts
  • Query pattern: time-range scans, grouped aggregations, service-level dashboards
  • Freshness target: under 5 seconds

Workload Summary

This is a high-ingest, append-friendly, time-series workload. The main architectural risks are:

  • excessive small parts
  • merge pressure
  • slow tail queries if rollups are not used for dashboards

Key Decisions

  1. Use a decoupled ingestion path
  2. Partition conservatively
  3. Preserve raw data while introducing focused rollups

Recommendations

1. Kafka engine + materialized view for ingestion

What
Use Kafka as the decoupling layer and load ClickHouse through Kafka engine tables and downstream MVs.

Why
The producer fleet is bursty and distributed. This pattern improves replayability and isolates producers from storage behavior.

How

  • Kafka topic per stream family
  • Kafka engine source table
  • MV into MergeTree raw table

Category
derived

Confidence
medium

Source

2. Monthly partitions on event time

What
Use PARTITION BY toYYYYMM(event_time) for the main raw table.

Why
This workload is time-bounded and retention-based, but daily partitioning would likely create unnecessary operational overhead at scale.

Category
derived

Confidence
medium

Source

3. Incremental MVs for hot service dashboards

What
Create rollup tables for repeated service-health queries.

Why
Dashboards and alerts should not repeatedly scan the raw log corpus.

Category
official

Confidence
high

Source

Example raw table

CREATE TABLE logs_raw
(
    event_time DateTime64(3),
    service LowCardinality(String),
    level LowCardinality(String),
    host String,
    message String,
    attrs JSON
)
ENGINE = MergeTree
PARTITION BY toYYYYMM(event_time)
ORDER BY (service, event_time, host);

Example rollup table

CREATE TABLE logs_rollup_1m
(
    bucket DateTime,
    service LowCardinality(String),
    level LowCardinality(String),
    count_state AggregateFunction(count)
)
ENGINE = AggregatingMergeTree
PARTITION BY toYYYYMM(bucket)
ORDER BY (service, level, bucket);

Source: SKILL.md on GitHub

No alerts17d4 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    The skill is a safe architectural advisor for ClickHouse workloads. It provides structured decision frameworks for ingestion, partitioning, and schema design based on official documentation. No malicious patterns, data exfiltration, or dangerous execution triggers were detected.

  • Socket17d

    No alerts

  • Snyk17d

    Risk: LOW · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 5e162d6. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 6 months ago
metadata
{
  "author": "ClickHouse Inc",
  "version": "0.1.0"
}
  • clickhouse
  • architecture
  • olap
  • ingestion
  • time-series
  • schema-design
  • partitioning
  • joins
  • telemetry

README badge

README badge for clickhouse/agent-skills/clickhouse-architecture-advisor

Guides ClickHouse architecture decisions for specific workloads—observability, analytics, IoT, financial services—by mapping workload shape to ingestion, partitioning, and join strategies with official documentation links. Classifies recommendations by provenance (official, derived, field) to separate documented behavior from heuristic field guidance.

Generated from the current SKILL.md.

Does this skill replace the clickhouse-best-practices skill?
No. This skill complements clickhouse-best-practices by adding workload-aware decision frameworks and provenance labels. Official documentation remains the source of truth for both.
What workload types does this skill cover?
Observability, security/SIEM, product analytics, IoT/telemetry, market data/financial services, and mixed OLAP with point-lookups. Each has scenario-specific rule files for ingestion, time-series retention, enrichment, and late-arriving events.
How does this skill distinguish between official, derived, and field guidance?
Official recommendations are directly from ClickHouse docs. Derived recommendations follow logically from documented behavior. Field recommendations are experience-based and include a disclaimer that they are heuristic and workload-dependent.
What should I do if a recommendation is uncertain?
The skill explicitly states when a recommendation is uncertain rather than presenting it as confident guidance.

Generated from the current SKILL.md. These answers refresh after source changes.