Chdb Datastore

作者 ClickHouse2f6ec4b17a81Apache-2.0544 个星标收录于 2026年10月8日更新于 2026年10月8日仓库9天前更新

Use when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas. Provides chDB DataStore — same pandas API, ClickHouse engine underneath. Also handles reading from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake as DataFrames and joining across sources. TRIGGER when: user mentions DataFrame, parquet, csv, "fast pandas", "speed up pandas", or cross-source DataFrame joins; user imports `chdb.datastore` or `from datastore import DataStore`. SKIP this skill for raw SQL syntax (use chdb-sql instead), ClickHouse server administration, or non-Python DataStore API work.

包含脚本Data & Analytics
AI 生成的概览

使用 chdb DataStore 作为 pandas 的替代方案,基于 ClickHouse 查询、连接和聚合表格数据。

功能
说明如何用 chdb DataStore 替换 pandas:这是一个基于 ClickHouse 的惰性 DataFrame 库,保留 pandas API,同时把操作编译为 SQL。内容涵盖连接文件与数据库、筛选、分组、聚合、跨数据源连接以及写出结果。附带参考文档、示例和一个用于检查环境的脚本。
适用场景
适用于处理 pandas DataFrame、parquet、CSV、Arrow 或 JSON 等表格数据,并需要筛选、分组、聚合、连接或加速缓慢的 pandas 代码时。也适用于把 S3、MySQL、PostgreSQL、MongoDB、ClickHouse Cloud、Iceberg 或 Delta Lake 等来源读取为 DataFrame。
运行要求
需要 macOS 或 Linux 上的 Python 3.9+,并通过 pip 安装 chdb 包。数据库和云连接器需要相应服务与凭据,以及网络访问。附带一个可执行的验证脚本。

chdb DataStore — It's Just Faster Pandas

The Key Insight

python
# Change this:import pandas as pd# To this:import chdb.datastore as pd# Everything else stays the same.

DataStore is a lazy, ClickHouse-backed pandas replacement. Your existing pandas code works unchanged — but operations compile to optimized SQL and execute only when results are needed (e.g., print(), len(), iteration).

bash
pip install chdb

Decision Tree: Pick the Right Approach

1. "I have a file/database and want to analyze it with pandas"   → DataStore.from_file() / from_mysql() / from_s3() etc.   → See references/connectors.md
2. "I need to join data from different sources"   → Create DataStores from each source, use .join()   → See examples/examples.md #3-5
3. "My pandas code is too slow"   → import chdb.datastore as pd — change one line, keep the rest
4. "I need raw SQL queries"   → Use the chdb-sql skill instead

Connect to Any Data Source — One Pattern

python
from datastore import DataStore
# Local file (auto-detects .parquet, .csv, .json, .arrow, .orc, .avro, .tsv, .xml)ds = DataStore.from_file("sales.parquet")
# Databaseds = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")
# Cloud storageds = DataStore.from_s3("s3://bucket/data.parquet", nosign=True)
# URI shorthand — auto-detects source typeds = DataStore.uri("mysql://root:pass@db:3306/shop/orders")

All 16+ sources and URI schemes → connectors.md [blocked]

After Connecting — Full Pandas API

python
result = ds[ds["age"] > 25]                                          # filterresult = ds[["name", "city"]]                                        # select columnsresult = ds.sort_values("revenue", ascending=False)                  # sortresult = ds.groupby("dept")["salary"].mean()                         # groupbyresult = ds.assign(margin=lambda x: x["profit"] / x["revenue"])     # computed columnds["name"].str.upper()                                               # string accessords["date"].dt.year                                                   # datetime accessorresult = ds1.join(ds2, on="id")                                      # joinresult = ds.head(10)                                                 # previewprint(ds.to_sql())                                                   # see generated SQL

209 DataFrame methods supported. Full API → api-reference.md [blocked]

Cross-Source Join — The Killer Feature

python
from datastore import DataStore
customers = DataStore.from_mysql(host="db:3306", database="crm", table="customers", user="root", password="pass")orders = DataStore.from_file("orders.parquet")
result = (orders    .join(customers, left_on="customer_id", right_on="id")    .groupby("country")    .agg({"amount": "sum", "rating": "mean"})    .sort_values("sum", ascending=False))print(result)

More join examples → examples.md [blocked]

Writing Data

python
source = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")target = DataStore("file", path="summary.parquet", format="Parquet")
target.insert_into("category", "total", "count").select_from(    source.groupby("category").select("category", "sum(amount) AS total", "count() AS count")).execute()

Troubleshooting

ProblemFix
ImportError: No module named 'chdb'pip install chdb
ImportError: cannot import 'DataStore'Use from datastore import DataStore or from chdb.datastore import DataStore
Database connection timeoutInclude port in host: host="db:3306" not host="db"
Join returns empty resultCheck key types match (both int or both string); use .to_sql() to inspect
Unexpected resultsCall ds.to_sql() to see the generated SQL and debug
Environment checkRun python scripts/verify_install.py (from skill directory)

References

  • API Reference [blocked] — Full DataStore method signatures
  • Connectors [blocked] — All 16+ data source connection methods
  • Examples [blocked] — 10+ runnable examples with expected output
  • Verify Install [blocked] — Environment verification script
  • Official Docs

Note: This skill teaches how to use chdb DataStore. For raw SQL queries, use the chdb-sql skill. For contributing to chdb source code, see CLAUDE.md in the project root.

来源与署名

来源:ClickHouse/agent-skills位于skills/chdb-datastore提交2f6ec4b

许可证: Apache-2.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 ClickHouse/agent-skills 的技能

Clickhouse Managed Postgres Rca

ClickHouse

针对 ClickHouse 托管 Postgres 实例性能问题的基于证据的根因分析工作流。

DevOps & Cloud5449天前更新

Clickhouse Js Node Troubleshooting

ClickHouse

排查并解决 ClickHouse Node.js 客户端(@clickhouse/client)的常见错误与配置问题。

Software Development5449天前更新

Clickhouse Best Practices

ClickHouse

依据 31 条最佳实践规则审查 ClickHouse 的模式、查询与写入策略。

Data & Analytics5449天前更新

Clickhouse Architecture Advisor

ClickHouse

针对特定工作负载指导 ClickHouse 架构决策,并将每条建议标注为官方、推导或经验。

Data & Analytics5449天前更新

Chdb Sql

ClickHouse

Use when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake) without setting up a server. Provides chDB — embedded ClickHouse SQL in Python with 1000+ functions, Session for stateful multi-step pipelines, parametrized queries, and cross-source joins via `s3()`, `mysql()`, `postgresql()`, `iceberg()`, `deltaLake()`, `remoteSecure()` table functions. TRIGGER when: user wants SQL on parquet/csv/files or across remote analytical sources; uses ClickHouse SQL features (window functions, windowFunnel, geoToH3, JSON path ops, Session, parametrized queries); imports `chdb` or calls `chdb.query()`. SKIP this skill for pandas-style DataFrame method-chaining (use chdb-datastore instead) or ClickHouse server administration.

包含脚本
待分类5449天前更新