All skills
clickhouse avatar

/chdb-datastore

@46ef08c official
by clickhouseclickhouse/agent-skills543 stars
39

Use when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas. Provides chDB DataStore β€” same pandas API, ClickHouse engine underneath. Also handles reading from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake as DataFrames and joining across sources. TRIGGER when: user mentions DataFrame, parquet, csv, "fast pandas", "speed up pandas", or cross-source DataFrame joins; user imports `chdb.datastore` or `from datastore import DataStore`. SKIP this skill for raw SQL syntax (use chdb-sql instead), ClickHouse server administration, or non-Python DataStore API work.

Use this Skill: https://skilld.dev/gh/clickhouse/agent-skills/chdb-datastore

This session only. Nothing lands on disk.

SKILL.md

β‰ˆ174 tokens always: the name and description. β‰ˆ1.2k when used: this file. β‰ˆ7.1k more on demand in 5 files.

chdb DataStore β€” It's Just Faster Pandas

The Key Insight

# Change this:
import pandas as pd
# To this:
import chdb.datastore as pd
# Everything else stays the same.

DataStore is a lazy, ClickHouse-backed pandas replacement. Your existing pandas code works unchanged β€” but operations compile to optimized SQL and execute only when results are needed (e.g., print(), len(), iteration).

pip install chdb

Decision Tree: Pick the Right Approach

1. "I have a file/database and want to analyze it with pandas"
   β†’ DataStore.from_file() / from_mysql() / from_s3() etc.
   β†’ See references/connectors.md

2. "I need to join data from different sources"
   β†’ Create DataStores from each source, use .join()
   β†’ See examples/examples.md #3-5

3. "My pandas code is too slow"
   β†’ import chdb.datastore as pd β€” change one line, keep the rest

4. "I need raw SQL queries"
   β†’ Use the chdb-sql skill instead

Connect to Any Data Source β€” One Pattern

from datastore import DataStore

# Local file (auto-detects .parquet, .csv, .json, .arrow, .orc, .avro, .tsv, .xml)
ds = DataStore.from_file("sales.parquet")

# Database
ds = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")

# Cloud storage
ds = DataStore.from_s3("s3://bucket/data.parquet", nosign=True)

# URI shorthand β€” auto-detects source type
ds = DataStore.uri("mysql://root:pass@db:3306/shop/orders")

All 16+ sources and URI schemes β†’ connectors.md

After Connecting β€” Full Pandas API

result = ds[ds["age"] > 25]                                          # filter
result = ds[["name", "city"]]                                        # select columns
result = ds.sort_values("revenue", ascending=False)                  # sort
result = ds.groupby("dept")["salary"].mean()                         # groupby
result = ds.assign(margin=lambda x: x["profit"] / x["revenue"])     # computed column
ds["name"].str.upper()                                               # string accessor
ds["date"].dt.year                                                   # datetime accessor
result = ds1.join(ds2, on="id")                                      # join
result = ds.head(10)                                                 # preview
print(ds.to_sql())                                                   # see generated SQL

209 DataFrame methods supported. Full API β†’ api-reference.md

Cross-Source Join β€” The Killer Feature

from datastore import DataStore

customers = DataStore.from_mysql(host="db:3306", database="crm", table="customers", user="root", password="pass")
orders = DataStore.from_file("orders.parquet")

result = (orders
    .join(customers, left_on="customer_id", right_on="id")
    .groupby("country")
    .agg({"amount": "sum", "rating": "mean"})
    .sort_values("sum", ascending=False))
print(result)

More join examples β†’ examples.md

Writing Data

source = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")
target = DataStore("file", path="summary.parquet", format="Parquet")

target.insert_into("category", "total", "count").select_from(
    source.groupby("category").select("category", "sum(amount) AS total", "count() AS count")
).execute()

Troubleshooting

Problem Fix
ImportError: No module named 'chdb' pip install chdb
ImportError: cannot import 'DataStore' Use from datastore import DataStore or from chdb.datastore import DataStore
Database connection timeout Include port in host: host="db:3306" not host="db"
Join returns empty result Check key types match (both int or both string); use .to_sql() to inspect
Unexpected results Call ds.to_sql() to see the generated SQL and debug
Environment check Run python scripts/verify_install.py (from skill directory)

References

Note: This skill teaches how to use chdb DataStore. For raw SQL queries, use the chdb-sql skill. For contributing to chdb source code, see CLAUDE.md in the project root.

Source: SKILL.md on GitHub

No alerts17d4 checks Β· Risk SAFE
  • Gen Agent Trust Hub17d

    The skill is a legitimate tool provided by ClickHouse Inc to use the chdb DataStore API, which is an optimized, ClickHouse-backed replacement for the pandas library. It allows users to perform high-performance data analysis on various sources including local files (CSV, Parquet), cloud storage (S3, GCS), and databases (MySQL, PostgreSQL). The analysis found no evidence of malicious behavior, prompt injection, or unauthorized data exfiltration. All external resources and packages trace back to the official vendor infrastructure.

  • Socket17d

    No alerts

  • Snyk17d

    Risk: LOW Β· No issues

  • ZeroLeaks5mo

    Score: 93/100 Β· 2 sections analyzed

Signed by skilld at 46ef08c. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 4 months ago
compatibility
Requires Python 3.9+, macOS or Linux. pip install chdb.
Other metadata
metadata
{
  "author": "chdb-io",
  "version": "4.1",
  "homepage": "https://clickhouse.com/docs/chdb"
}
  • chdb
  • clickhouse
  • pandas
  • dataframe
  • parquet
  • csv
  • s3
  • mysql
  • postgresql
  • mongodb
  • lazy-evaluation
  • sql
  • data-analysis

README badge

README badge for clickhouse/agent-skills/chdb-datastore

Provides chdb DataStore, a ClickHouse-backed pandas replacement with the same API but lazy evaluation and SQL compilation underneath. Load tabular data from files, S3, MySQL, PostgreSQL, MongoDB, or other sources as DataFrames, then filter, group, aggregate, and join across sources using familiar pandas syntax β€” typically faster than pandas for large datasets.

Generated from the current SKILL.md.

Does DataStore work with my existing pandas code?
Yes. DataStore implements the pandas API β€” you can often replace `import pandas as pd` with `import chdb.datastore as pd` and keep the rest of your code unchanged. Operations are lazy and compile to SQL under the hood.
What data sources does DataStore support?
DataStore connects to 16+ sources including local files (parquet, csv, json, arrow, orc, avro, tsv, xml), MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, S3, Iceberg, and Delta Lake. Use `.from_file()`, `.from_mysql()`, `.from_s3()`, or the `.uri()` shorthand to auto-detect the source.
Can I join data across different sources?
Yes. Create separate DataStore instances for each source and use `.join()` to combine them. The skill includes examples of joining data from MySQL, parquet files, and S3 in a single query.
What Python versions does this require?
Python 3.9+, and only works on macOS or Linux. Install with `pip install chdb`.
Should I use this skill for raw SQL queries?
No. Use the chdb-sql skill instead. This skill is for the DataStore pandas-compatible API. If you need raw SQL syntax, switch to chdb-sql.

Generated from the current SKILL.md. These answers refresh after source changes.