All skills
clickhouse avatar

/clickhouse-best-practices

@d284161 official
by clickhouseclickhouse/agent-skills543 stars
39

MUST USE when reviewing ClickHouse schemas, queries, or configurations. Contains 31 rules that MUST be checked before providing recommendations. Always read relevant rule files and cite specific rules in responses.

Use this Skill: https://skilld.dev/gh/clickhouse/agent-skills/clickhouse-best-practices

This session only. Nothing lands on disk.

rulesschema-partition-lifecycle.md

≈385 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Use Partitioning for Data Lifecycle Management

Impact: HIGH

Partitioning is primarily a data management technique, not a query optimization tool. It excels at:

  • Dropping data: Remove entire partitions as single metadata operations
  • TTL retention: Implement time-based retention policies efficiently
  • Tiered storage: Move old partitions to cold storage
  • Archiving: Move partitions between tables

Incorrect (no time alignment for lifecycle):

-- Cannot efficiently drop old data by time
CREATE TABLE events (...)
ENGINE = MergeTree()
PARTITION BY event_type  -- No time alignment
ORDER BY (timestamp);

-- Slow: must scan and delete row by row
DELETE FROM events WHERE timestamp < '2023-01-01';

Correct (time-based for lifecycle):

CREATE TABLE events (
    timestamp DateTime,
    event_type LowCardinality(String)
)
ENGINE = MergeTree()
PARTITION BY toStartOfMonth(timestamp)
ORDER BY (event_type, timestamp)
TTL timestamp + INTERVAL 1 YEAR DELETE;  -- Drops whole partitions

-- Fast: metadata-only operation
ALTER TABLE events DROP PARTITION '202301';

-- Archive to cold storage
ALTER TABLE events_archive ATTACH PARTITION '202301' FROM events;

Reference: Choosing a Partitioning Key

Source: SKILL.md on GitHub

No alerts17d5 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    This skill provides comprehensive best practices for ClickHouse database management, including schema design, query optimization, and ingestion strategies. It includes robust safety guardrails for AI agents, such as mandatory query limits, execution timeouts, and a structured schema discovery workflow. All external references and tools trace back to official ClickHouse vendor resources.

  • Socket17d

    No alerts

  • Snyk17d

    Risk: LOW · No issues

  • Runlayer7mo

    8/34 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at d284161. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 5 months ago
metadata
{
  "author": "ClickHouse Inc",
  "version": "0.4.0"
}

README badge

README badge for clickhouse/agent-skills/clickhouse-best-practices

Provides 31 ClickHouse-specific rules organized by priority across schema design, query optimization, data ingestion, and agent connectivity. Use this skill to validate schemas, review queries, and establish safe agent workflows with proper connection setup, schema discovery, and query safety procedures.

Generated from the current SKILL.md.

When should I use this skill?
Use this skill when reviewing ClickHouse schemas, queries, or data ingestion strategies. It contains 31 rules covering primary key design, data types, JOINs, partitioning, and insert batching that you must check before providing ClickHouse recommendations.
Does this skill help with AI agent connectivity to ClickHouse?
Yes. The skill includes rules for MCP and CLI connection setup, schema discovery workflows, and query safety (LIMIT, timeouts, progressive exploration) specific to AI agents querying ClickHouse.
What should I do if a rule doesn't exist for my question?
Fall back to the LLM's ClickHouse knowledge, search the official ClickHouse documentation, or use web search. Always cite your source in the response.
Are the rules mandatory or advisory?
The rules are mandatory checks before answering ClickHouse questions. They encode ClickHouse-specific behaviors (columnar storage, merge tree mechanics, sparse indexes) where general database intuition can be misleading.
Can I use this skill for INSERT performance tuning?
Yes. The skill covers batch sizing (10K-100K rows), async inserts for high-frequency small batches, mutation avoidance (ReplacingMergeTree instead of ALTER UPDATE), and OPTIMIZE TABLE risks.

Generated from the current SKILL.md. These answers refresh after source changes.