Chdb Datastore

by ClickHouse2f6ec4b17a81Apache-2.0544 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 9 days ago

Use when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas. Provides chDB DataStore — same pandas API, ClickHouse engine underneath. Also handles reading from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake as DataFrames and joining across sources. TRIGGER when: user mentions DataFrame, parquet, csv, "fast pandas", "speed up pandas", or cross-source DataFrame joins; user imports `chdb.datastore` or `from datastore import DataStore`. SKIP this skill for raw SQL syntax (use chdb-sql instead), ClickHouse server administration, or non-Python DataStore API work.

Includes scriptsData & Analytics
AI-generated overview

Use chdb DataStore as a drop-in, ClickHouse-backed replacement for pandas to query, join and aggregate tabular data.

What it does
Explains how to swap pandas for chdb DataStore, a lazy ClickHouse-backed DataFrame library that keeps the pandas API while compiling operations to SQL. Covers connecting to files and databases, filtering, grouping, aggregating, joining across sources, and writing results out. Ships reference documents, examples, and a script that verifies the environment.
When to use it
Use when working with tabular data such as pandas DataFrames, parquet, CSV, Arrow or JSON and needing to filter, group, aggregate, join or speed up slow pandas code. Also relevant when reading from sources like S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg or Delta Lake as DataFrames.
Requirements
Python 3.9+ on macOS or Linux, with the chdb package installed via pip. Database and cloud connectors require the corresponding services and credentials, plus network access. Ships an executable verification script.

chdb DataStore — It's Just Faster Pandas

The Key Insight

python
# Change this:import pandas as pd# To this:import chdb.datastore as pd# Everything else stays the same.

DataStore is a lazy, ClickHouse-backed pandas replacement. Your existing pandas code works unchanged — but operations compile to optimized SQL and execute only when results are needed (e.g., print(), len(), iteration).

bash
pip install chdb

Decision Tree: Pick the Right Approach

1. "I have a file/database and want to analyze it with pandas"   → DataStore.from_file() / from_mysql() / from_s3() etc.   → See references/connectors.md
2. "I need to join data from different sources"   → Create DataStores from each source, use .join()   → See examples/examples.md #3-5
3. "My pandas code is too slow"   → import chdb.datastore as pd — change one line, keep the rest
4. "I need raw SQL queries"   → Use the chdb-sql skill instead

Connect to Any Data Source — One Pattern

python
from datastore import DataStore
# Local file (auto-detects .parquet, .csv, .json, .arrow, .orc, .avro, .tsv, .xml)ds = DataStore.from_file("sales.parquet")
# Databaseds = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")
# Cloud storageds = DataStore.from_s3("s3://bucket/data.parquet", nosign=True)
# URI shorthand — auto-detects source typeds = DataStore.uri("mysql://root:pass@db:3306/shop/orders")

All 16+ sources and URI schemes → connectors.md [blocked]

After Connecting — Full Pandas API

python
result = ds[ds["age"] > 25]                                          # filterresult = ds[["name", "city"]]                                        # select columnsresult = ds.sort_values("revenue", ascending=False)                  # sortresult = ds.groupby("dept")["salary"].mean()                         # groupbyresult = ds.assign(margin=lambda x: x["profit"] / x["revenue"])     # computed columnds["name"].str.upper()                                               # string accessords["date"].dt.year                                                   # datetime accessorresult = ds1.join(ds2, on="id")                                      # joinresult = ds.head(10)                                                 # previewprint(ds.to_sql())                                                   # see generated SQL

209 DataFrame methods supported. Full API → api-reference.md [blocked]

Cross-Source Join — The Killer Feature

python
from datastore import DataStore
customers = DataStore.from_mysql(host="db:3306", database="crm", table="customers", user="root", password="pass")orders = DataStore.from_file("orders.parquet")
result = (orders    .join(customers, left_on="customer_id", right_on="id")    .groupby("country")    .agg({"amount": "sum", "rating": "mean"})    .sort_values("sum", ascending=False))print(result)

More join examples → examples.md [blocked]

Writing Data

python
source = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")target = DataStore("file", path="summary.parquet", format="Parquet")
target.insert_into("category", "total", "count").select_from(    source.groupby("category").select("category", "sum(amount) AS total", "count() AS count")).execute()

Troubleshooting

ProblemFix
ImportError: No module named 'chdb'pip install chdb
ImportError: cannot import 'DataStore'Use from datastore import DataStore or from chdb.datastore import DataStore
Database connection timeoutInclude port in host: host="db:3306" not host="db"
Join returns empty resultCheck key types match (both int or both string); use .to_sql() to inspect
Unexpected resultsCall ds.to_sql() to see the generated SQL and debug
Environment checkRun python scripts/verify_install.py (from skill directory)

References

  • API Reference [blocked] — Full DataStore method signatures
  • Connectors [blocked] — All 16+ data source connection methods
  • Examples [blocked] — 10+ runnable examples with expected output
  • Verify Install [blocked] — Environment verification script
  • Official Docs

Note: This skill teaches how to use chdb DataStore. For raw SQL queries, use the chdb-sql skill. For contributing to chdb source code, see CLAUDE.md in the project root.

Source and attribution

Source:ClickHouse/agent-skillsinskills/chdb-datastoreat commit2f6ec4b

License: Apache-2.0

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from ClickHouse/agent-skills

Clickhouse Managed Postgres Rca

ClickHouse

Evidence-based root cause analysis workflow for performance issues on ClickHouse-managed Postgres instances.

DevOps & Cloud544updated 9 days ago

Clickhouse Js Node Troubleshooting

ClickHouse

Troubleshoots common errors and configuration issues with the ClickHouse Node.js client (@clickhouse/client).

Software Development544updated 9 days ago

Clickhouse Best Practices

ClickHouse

Reviews ClickHouse schemas, queries, and ingestion strategies against 31 documented best-practice rules.

Data & Analytics544updated 9 days ago

Clickhouse Architecture Advisor

ClickHouse

Guides ClickHouse architecture decisions for specific workloads, labeling each recommendation as official, derived, or field.

Data & Analytics544updated 9 days ago

Chdb Sql

ClickHouse

Use when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake) without setting up a server. Provides chDB — embedded ClickHouse SQL in Python with 1000+ functions, Session for stateful multi-step pipelines, parametrized queries, and cross-source joins via `s3()`, `mysql()`, `postgresql()`, `iceberg()`, `deltaLake()`, `remoteSecure()` table functions. TRIGGER when: user wants SQL on parquet/csv/files or across remote analytical sources; uses ClickHouse SQL features (window functions, windowFunnel, geoToH3, JSON path ops, Session, parametrized queries); imports `chdb` or calls `chdb.query()`. SKIP this skill for pandas-style DataFrame method-chaining (use chdb-datastore instead) or ClickHouse server administration.

Includes scripts
Awaiting classification544updated 9 days ago