Chdb Datastore

作者 ClickHouse2f6ec4b17a81Apache-2.0544 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫9 天前更新

Use when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas. Provides chDB DataStore — same pandas API, ClickHouse engine underneath. Also handles reading from S3, MySQL, PostgreSQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake as DataFrames and joining across sources. TRIGGER when: user mentions DataFrame, parquet, csv, "fast pandas", "speed up pandas", or cross-source DataFrame joins; user imports `chdb.datastore` or `from datastore import DataStore`. SKIP this skill for raw SQL syntax (use chdb-sql instead), ClickHouse server administration, or non-Python DataStore API work.

包含腳本Data & Analytics
AI 產生的概覽

使用 chdb DataStore 作為 pandas 的替代方案,以 ClickHouse 為底層查詢、合併與彙總表格資料。

功能
說明如何以 chdb DataStore 取代 pandas:這是一個以 ClickHouse 為底層的惰性 DataFrame 函式庫,保留 pandas API,同時把操作編譯成 SQL。內容涵蓋連線檔案與資料庫、篩選、分組、彙總、跨資料來源合併,以及寫出結果。隨附參考文件、範例,以及一支檢查環境的指令碼。
適用情境
適用於處理 pandas DataFrame、parquet、CSV、Arrow 或 JSON 等表格資料,且需要篩選、分組、彙總、合併或加速緩慢的 pandas 程式碼時。也適用於把 S3、MySQL、PostgreSQL、MongoDB、ClickHouse Cloud、Iceberg 或 Delta Lake 等來源讀取為 DataFrame。
執行需求
需要 macOS 或 Linux 上的 Python 3.9+,並以 pip 安裝 chdb 套件。資料庫與雲端連接器需要對應的服務與憑證,以及網路存取。隨附一支可執行的驗證指令碼。

chdb DataStore — It's Just Faster Pandas

The Key Insight

python
# Change this:import pandas as pd# To this:import chdb.datastore as pd# Everything else stays the same.

DataStore is a lazy, ClickHouse-backed pandas replacement. Your existing pandas code works unchanged — but operations compile to optimized SQL and execute only when results are needed (e.g., print(), len(), iteration).

bash
pip install chdb

Decision Tree: Pick the Right Approach

1. "I have a file/database and want to analyze it with pandas"   → DataStore.from_file() / from_mysql() / from_s3() etc.   → See references/connectors.md
2. "I need to join data from different sources"   → Create DataStores from each source, use .join()   → See examples/examples.md #3-5
3. "My pandas code is too slow"   → import chdb.datastore as pd — change one line, keep the rest
4. "I need raw SQL queries"   → Use the chdb-sql skill instead

Connect to Any Data Source — One Pattern

python
from datastore import DataStore
# Local file (auto-detects .parquet, .csv, .json, .arrow, .orc, .avro, .tsv, .xml)ds = DataStore.from_file("sales.parquet")
# Databaseds = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")
# Cloud storageds = DataStore.from_s3("s3://bucket/data.parquet", nosign=True)
# URI shorthand — auto-detects source typeds = DataStore.uri("mysql://root:pass@db:3306/shop/orders")

All 16+ sources and URI schemes → connectors.md [blocked]

After Connecting — Full Pandas API

python
result = ds[ds["age"] > 25]                                          # filterresult = ds[["name", "city"]]                                        # select columnsresult = ds.sort_values("revenue", ascending=False)                  # sortresult = ds.groupby("dept")["salary"].mean()                         # groupbyresult = ds.assign(margin=lambda x: x["profit"] / x["revenue"])     # computed columnds["name"].str.upper()                                               # string accessords["date"].dt.year                                                   # datetime accessorresult = ds1.join(ds2, on="id")                                      # joinresult = ds.head(10)                                                 # previewprint(ds.to_sql())                                                   # see generated SQL

209 DataFrame methods supported. Full API → api-reference.md [blocked]

Cross-Source Join — The Killer Feature

python
from datastore import DataStore
customers = DataStore.from_mysql(host="db:3306", database="crm", table="customers", user="root", password="pass")orders = DataStore.from_file("orders.parquet")
result = (orders    .join(customers, left_on="customer_id", right_on="id")    .groupby("country")    .agg({"amount": "sum", "rating": "mean"})    .sort_values("sum", ascending=False))print(result)

More join examples → examples.md [blocked]

Writing Data

python
source = DataStore.from_mysql(host="db:3306", database="shop", table="orders", user="root", password="pass")target = DataStore("file", path="summary.parquet", format="Parquet")
target.insert_into("category", "total", "count").select_from(    source.groupby("category").select("category", "sum(amount) AS total", "count() AS count")).execute()

Troubleshooting

ProblemFix
ImportError: No module named 'chdb'pip install chdb
ImportError: cannot import 'DataStore'Use from datastore import DataStore or from chdb.datastore import DataStore
Database connection timeoutInclude port in host: host="db:3306" not host="db"
Join returns empty resultCheck key types match (both int or both string); use .to_sql() to inspect
Unexpected resultsCall ds.to_sql() to see the generated SQL and debug
Environment checkRun python scripts/verify_install.py (from skill directory)

References

  • API Reference [blocked] — Full DataStore method signatures
  • Connectors [blocked] — All 16+ data source connection methods
  • Examples [blocked] — 10+ runnable examples with expected output
  • Verify Install [blocked] — Environment verification script
  • Official Docs

Note: This skill teaches how to use chdb DataStore. For raw SQL queries, use the chdb-sql skill. For contributing to chdb source code, see CLAUDE.md in the project root.

來源與署名

來源:ClickHouse/agent-skills位於skills/chdb-datastore提交2f6ec4b

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 ClickHouse/agent-skills 的技能

Clickhouse Managed Postgres Rca

ClickHouse

針對 ClickHouse 代管 Postgres 執行個體效能問題、以證據為基礎的根因分析工作流程。

DevOps & Cloud5449 天前更新

Clickhouse Js Node Troubleshooting

ClickHouse

排查並解決 ClickHouse Node.js 客戶端(@clickhouse/client)的常見錯誤與設定問題。

Software Development5449 天前更新

Clickhouse Best Practices

ClickHouse

依據 31 條最佳實務規則審查 ClickHouse 的結構定義、查詢與寫入策略。

Data & Analytics5449 天前更新

Clickhouse Architecture Advisor

ClickHouse

針對特定工作負載指導 ClickHouse 架構決策,並將每項建議標註為官方、推導或經驗。

Data & Analytics5449 天前更新

Chdb Sql

ClickHouse

Use when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse Cloud, Iceberg, Delta Lake) without setting up a server. Provides chDB — embedded ClickHouse SQL in Python with 1000+ functions, Session for stateful multi-step pipelines, parametrized queries, and cross-source joins via `s3()`, `mysql()`, `postgresql()`, `iceberg()`, `deltaLake()`, `remoteSecure()` table functions. TRIGGER when: user wants SQL on parquet/csv/files or across remote analytical sources; uses ClickHouse SQL features (window functions, windowFunnel, geoToH3, JSON path ops, Session, parametrized queries); imports `chdb` or calls `chdb.query()`. SKIP this skill for pandas-style DataFrame method-chaining (use chdb-datastore instead) or ClickHouse server administration.

包含腳本
待分類5449 天前更新