Bigquery Bigframes

作者 gemini-cli-extensions2df10e25bbf7无许可证215 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Generates Python code using BigQuery DataFrames (BigFrames), the pandas/scikit-learn-style API over BigQuery. Use when writing BigFrames code or doing pandas-style dataframe/ML work against BigQuery (e.g. in a notebook). Don't use for SQL-first workflows or the google-cloud-bigquery client library — use bigquery-basics.

AI 生成的概览

指导使用 BigQuery DataFrames(BigFrames)编写 Python 代码,在 BigQuery 上进行 pandas 风格的数据处理与机器学习。

功能
提供使用 BigFrames(BigQuery 上 pandas/scikit-learn 风格的 Python API)的编码指南。内容涵盖数据框实践,例如避免使用 to_pandas 和用 read_gbq 执行 SQL、优先使用访问器而非 UDF 和 lambda、检查 schema 以及可视化,还包括关于 bigframes.ml、预测输出、自动缩放、超参数调优、ARIMA Plus、PCA 和模型持久化的模型开发说明。
适用场景
适用于编写 BigFrames 代码,或针对 BigQuery 进行 pandas 风格的数据框或机器学习工作时,例如在 notebook 中。不适用于以 SQL 为主的工作流或 google-cloud-bigquery 客户端库。
运行要求
需要 BigFrames Python 库和 BigQuery 访问权限;不包含脚本,只有说明。

BigFrames (BigQuery DataFrame) basics

BigFrames is a Python library that lets you take advantage of BigQuery data processing by using familiar Python APIs.

Generic Coding Guidelines

  • Avoid .to_pandas(): You MUST NOT use .to_pandas() to download the entire dataset into memory. There are some exceptions:
    • An error message explicitly requests you to use to_pandas()
    • You are going to visualize the data, and the visualization library does not accept BigFrames Dataframe/Series instances. In this case, reduce the amount of data you are going to download before calling .to_pandas()
  • Avoid read_gbq() for SQL: Do not write SQL queries and execute them with read_gbq(). Use BigFrames Dataframe/Series methods instead.
  • Use BigFrames ML package for Machine Learning Tasks: Do not use Scikit-learn or other ML libraries with BigFrames dataframes. Import your tools/classes from bigframes.ml.
  • Stay in the Cloud: Perform data cleaning, transformation, and analysis via BigFrames methods to leverage BigQuery's scale.
  • Accessors over UDFs/Lambdas:
    • Prefer built-in accessors (e.g., df.col.str.*, df.col.dt.*) over remote UDFs.
    • Do not use lambdas with Series.map() or DataFrame.apply().
  • Schema Verification: Do not assume schema of intermediate outputs. Check .dtypes after loading, and use display() with .head() or .peek().
  • Visualization: BigFrames Dataframe mostly works directly with Matplotlib, Seaborn, and other plotting libraries. If your attempt didn't work, try using the "plot" accessor. If that didn't work either, you MUST sample or aggregate your data to make it small enough before calling "to_pandas()".

Model Development

  • Unlike Scikit-learn: BigFrames' predict() method always returns a DataFrame containing both predictions and features (not just a series of predictions).
  • No random_state: Do not pass a random_state argument when instantiating BigFrames ML models.
  • Automatic Scaling: Do not use OneHotEncoder or StandardScaler unless explicitly requested (handled automatically).
  • Hyperparameter Tuning: You must write custom loops (BigFrames lacks GridSearchCV or RandomizedSearchCV).
  • ARIMA Plus (Forecasting):
    • Import from bigframes.ml.forecasting.
    • Sort data chronologically and split around a timepoint before training.
    • Prediction horizon must be less than or equal to training horizon.
  • PCA: BigFrames' PCA class lacks simple transform() method. Use predict() instead.
  • Model Persistence: To persist a model, use model.to_gbq(). To load a persisted model, use bpd.read_gbq_model().

来源与署名

来源:gemini-cli-extensions/data-agent-kit-starter-pack位于skills/bigquery-bigframes提交2df10e2

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 gemini-cli-extensions/data-agent-kit-starter-pack 的技能

Schema Mapping

gemini-cli-extensions

为 ETL、ELT 或数据集成任务规划源到目标的模式映射,产出文档化的映射宣言。

Data & Analytics215今天更新

Resolving Mcp Region Configs

gemini-cli-extensions

修复区域级 Google Cloud MCP 服务器配置中未替换的区域占位符,使缺失的 MCP 工具得以注册。

DevOps & Cloud215今天更新

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

待分类215今天更新

Ml Best Practices

gemini-cli-extensions

为机器学习笔记本提供分步方案,涵盖聚类、预测、分类、回归和模型比较。

Data & Analytics215今天更新

Managing Python Dependencies

gemini-cli-extensions

指导代理检测 Python 项目的依赖管理器并正确安装依赖,而不是使用全局 pip。

Software Development215今天更新

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

待分类215今天更新