All skills
clickhouse avatar

/clickhouse-best-practices

@d284161 official
by clickhouseclickhouse/agent-skills543 stars
39

MUST USE when reviewing ClickHouse schemas, queries, or configurations. Contains 31 rules that MUST be checked before providing recommendations. Always read relevant rule files and cite specific rules in responses.

Use this Skill: https://skilld.dev/gh/clickhouse/agent-skills/clickhouse-best-practices

This session only. Nothing lands on disk.

rulesinsert-batch-size.md

≈382 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Batch Inserts Appropriately (10K-100K rows)

Impact: CRITICAL

Each INSERT creates a new data part. Single-row or small-batch inserts create thousands of tiny parts, overwhelming the merge process and causing cluster instability.

Incorrect (single-row or tiny batches):

# Single-row inserts - creates 10,000 parts!
for event in events:
    client.execute("INSERT INTO events VALUES", [event])

# Tiny batches - still too many parts
for batch in chunks(events, 100):  # 100 rows per INSERT
    client.execute("INSERT INTO events VALUES", batch)

Correct (proper batch size):

# Ideal batch size: 10,000-100,000 rows
BATCH_SIZE = 10_000
for batch in chunks(events, BATCH_SIZE):
    client.execute("INSERT INTO events VALUES", batch)

Recommended batch sizes:

Threshold Value
Minimum 1,000 rows
Ideal range 10,000-100,000 rows
Insert rate (sync) ~1 insert per second

Validation:

-- Monitor part count (>3000 per partition blocks inserts)
SELECT table, count() as parts, sum(rows) as total_rows
FROM system.parts
WHERE active AND database = 'default'
GROUP BY table
ORDER BY parts DESC;

Reference: Selecting an Insert Strategy

Source: SKILL.md on GitHub

No alerts17d5 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    This skill provides comprehensive best practices for ClickHouse database management, including schema design, query optimization, and ingestion strategies. It includes robust safety guardrails for AI agents, such as mandatory query limits, execution timeouts, and a structured schema discovery workflow. All external references and tools trace back to official ClickHouse vendor resources.

  • Socket17d

    No alerts

  • Snyk17d

    Risk: LOW · No issues

  • Runlayer7mo

    8/34 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at d284161. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 5 months ago
metadata
{
  "author": "ClickHouse Inc",
  "version": "0.4.0"
}

README badge

README badge for clickhouse/agent-skills/clickhouse-best-practices

Provides 31 ClickHouse-specific rules organized by priority across schema design, query optimization, data ingestion, and agent connectivity. Use this skill to validate schemas, review queries, and establish safe agent workflows with proper connection setup, schema discovery, and query safety procedures.

Generated from the current SKILL.md.

When should I use this skill?
Use this skill when reviewing ClickHouse schemas, queries, or data ingestion strategies. It contains 31 rules covering primary key design, data types, JOINs, partitioning, and insert batching that you must check before providing ClickHouse recommendations.
Does this skill help with AI agent connectivity to ClickHouse?
Yes. The skill includes rules for MCP and CLI connection setup, schema discovery workflows, and query safety (LIMIT, timeouts, progressive exploration) specific to AI agents querying ClickHouse.
What should I do if a rule doesn't exist for my question?
Fall back to the LLM's ClickHouse knowledge, search the official ClickHouse documentation, or use web search. Always cite your source in the response.
Are the rules mandatory or advisory?
The rules are mandatory checks before answering ClickHouse questions. They encode ClickHouse-specific behaviors (columnar storage, merge tree mechanics, sparse indexes) where general database intuition can be misleading.
Can I use this skill for INSERT performance tuning?
Yes. The skill covers batch sizing (10K-100K rows), async inserts for high-frequency small batches, mutation avoidance (ReplacingMergeTree instead of ALTER UPDATE), and OPTIMIZE TABLE risks.

Generated from the current SKILL.md. These answers refresh after source changes.