Capability
- compatibility
- Requires Python 3.9+, macOS or Linux. pip install chdb.
Other metadata
- metadata
{
"author": "chdb-io",
"version": "4.1",
"homepage": "https://clickhouse.com/docs/chdb"
}
Topics
- chdb
- clickhouse
- pandas
- dataframe
- parquet
- csv
- s3
- mysql
- postgresql
- mongodb
- lazy-evaluation
- sql
- data-analysis
What it does
Provides chdb DataStore, a ClickHouse-backed pandas replacement with the same API but lazy evaluation and SQL compilation underneath. Load tabular data from files, S3, MySQL, PostgreSQL, MongoDB, or other sources as DataFrames, then filter, group, aggregate, and join across sources using familiar pandas syntax — typically faster than pandas for large datasets.
Generated from the current SKILL.md.
Frequently asked
Does DataStore work with my existing pandas code?
Yes. DataStore implements the pandas API — you can often replace `import pandas as pd` with `import chdb.datastore as pd` and keep the rest of your code unchanged. Operations are lazy and compile to SQL under the hood.
What data sources does DataStore support?
DataStore connects to 16+ sources including local files (parquet, csv, json, arrow, orc, avro, tsv, xml), MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, S3, Iceberg, and Delta Lake. Use `.from_file()`, `.from_mysql()`, `.from_s3()`, or the `.uri()` shorthand to auto-detect the source.
Can I join data across different sources?
Yes. Create separate DataStore instances for each source and use `.join()` to combine them. The skill includes examples of joining data from MySQL, parquet files, and S3 in a single query.
What Python versions does this require?
Python 3.9+, and only works on macOS or Linux. Install with `pip install chdb`.
Should I use this skill for raw SQL queries?
No. Use the chdb-sql skill instead. This skill is for the DataStore pandas-compatible API. If you need raw SQL syntax, switch to chdb-sql.
Generated from the current SKILL.md. These answers refresh after source changes.