Bigquery Bigframes

作者 gemini-cli-extensions2df10e25bbf7無授權條款215 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Generates Python code using BigQuery DataFrames (BigFrames), the pandas/scikit-learn-style API over BigQuery. Use when writing BigFrames code or doing pandas-style dataframe/ML work against BigQuery (e.g. in a notebook). Don't use for SQL-first workflows or the google-cloud-bigquery client library — use bigquery-basics.

AI 產生的概覽

指導使用 BigQuery DataFrames(BigFrames)撰寫 Python 程式碼,在 BigQuery 上進行 pandas 風格的資料處理與機器學習。

功能
提供使用 BigFrames(BigQuery 上 pandas/scikit-learn 風格的 Python API)的撰寫指南。內容涵蓋資料框實務,例如避免使用 to_pandas 和用 read_gbq 執行 SQL、優先使用存取器而非 UDF 和 lambda、檢查 schema 以及視覺化,還包括關於 bigframes.ml、預測輸出、自動縮放、超參數調校、ARIMA Plus、PCA 和模型持久化的模型開發說明。
適用情境
適用於撰寫 BigFrames 程式碼,或針對 BigQuery 進行 pandas 風格的資料框或機器學習工作時,例如在 notebook 中。不適用於以 SQL 為主的工作流程或 google-cloud-bigquery 用戶端程式庫。
執行需求
需要 BigFrames Python 程式庫和 BigQuery 存取權;不包含指令碼,只有說明。

BigFrames (BigQuery DataFrame) basics

BigFrames is a Python library that lets you take advantage of BigQuery data processing by using familiar Python APIs.

Generic Coding Guidelines

  • Avoid .to_pandas(): You MUST NOT use .to_pandas() to download the entire dataset into memory. There are some exceptions:
    • An error message explicitly requests you to use to_pandas()
    • You are going to visualize the data, and the visualization library does not accept BigFrames Dataframe/Series instances. In this case, reduce the amount of data you are going to download before calling .to_pandas()
  • Avoid read_gbq() for SQL: Do not write SQL queries and execute them with read_gbq(). Use BigFrames Dataframe/Series methods instead.
  • Use BigFrames ML package for Machine Learning Tasks: Do not use Scikit-learn or other ML libraries with BigFrames dataframes. Import your tools/classes from bigframes.ml.
  • Stay in the Cloud: Perform data cleaning, transformation, and analysis via BigFrames methods to leverage BigQuery's scale.
  • Accessors over UDFs/Lambdas:
    • Prefer built-in accessors (e.g., df.col.str.*, df.col.dt.*) over remote UDFs.
    • Do not use lambdas with Series.map() or DataFrame.apply().
  • Schema Verification: Do not assume schema of intermediate outputs. Check .dtypes after loading, and use display() with .head() or .peek().
  • Visualization: BigFrames Dataframe mostly works directly with Matplotlib, Seaborn, and other plotting libraries. If your attempt didn't work, try using the "plot" accessor. If that didn't work either, you MUST sample or aggregate your data to make it small enough before calling "to_pandas()".

Model Development

  • Unlike Scikit-learn: BigFrames' predict() method always returns a DataFrame containing both predictions and features (not just a series of predictions).
  • No random_state: Do not pass a random_state argument when instantiating BigFrames ML models.
  • Automatic Scaling: Do not use OneHotEncoder or StandardScaler unless explicitly requested (handled automatically).
  • Hyperparameter Tuning: You must write custom loops (BigFrames lacks GridSearchCV or RandomizedSearchCV).
  • ARIMA Plus (Forecasting):
    • Import from bigframes.ml.forecasting.
    • Sort data chronologically and split around a timepoint before training.
    • Prediction horizon must be less than or equal to training horizon.
  • PCA: BigFrames' PCA class lacks simple transform() method. Use predict() instead.
  • Model Persistence: To persist a model, use model.to_gbq(). To load a persisted model, use bpd.read_gbq_model().

來源與署名

來源:gemini-cli-extensions/data-agent-kit-starter-pack位於skills/bigquery-bigframes提交2df10e2

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 gemini-cli-extensions/data-agent-kit-starter-pack 的技能

Schema Mapping

gemini-cli-extensions

為 ETL、ELT 或資料整合任務規劃來源到目標的結構描述對應,產出文件化的對應宣言。

Data & Analytics215今天更新

Resolving Mcp Region Configs

gemini-cli-extensions

修復區域性 Google Cloud MCP 伺服器設定中未取代的區域佔位符,讓缺少的 MCP 工具得以註冊。

DevOps & Cloud215今天更新

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

待分類215今天更新

Ml Best Practices

gemini-cli-extensions

為機器學習筆記本提供逐步方案,涵蓋分群、預測、分類、迴歸與模型比較。

Data & Analytics215今天更新

Managing Python Dependencies

gemini-cli-extensions

指導代理偵測 Python 專案的相依性管理器並正確安裝套件,而不是使用全域 pip。

Software Development215今天更新

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

待分類215今天更新