Bigquery Bigframes

by gemini-cli-extensions2df10e25bbf7No license215 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Generates Python code using BigQuery DataFrames (BigFrames), the pandas/scikit-learn-style API over BigQuery. Use when writing BigFrames code or doing pandas-style dataframe/ML work against BigQuery (e.g. in a notebook). Don't use for SQL-first workflows or the google-cloud-bigquery client library — use bigquery-basics.

AI-generated overview

Guides writing Python code with BigQuery DataFrames (BigFrames) for pandas-style data work and ML on BigQuery.

What it does
Provides coding guidelines for using BigFrames, the pandas/scikit-learn-style Python API over BigQuery. It covers dataframe practices such as avoiding to_pandas and read_gbq for SQL, preferring accessors over UDFs and lambdas, checking schema, and visualization, plus model development notes on bigframes.ml, prediction output, scaling, hyperparameter tuning, ARIMA Plus, PCA, and model persistence.
When to use it
Use when writing BigFrames code or doing pandas-style dataframe or machine learning work against BigQuery, for example in a notebook. It is not intended for SQL-first workflows or the google-cloud-bigquery client library.
Requirements
Requires the BigFrames Python library and access to BigQuery; no scripts are included, only instructions.

BigFrames (BigQuery DataFrame) basics

BigFrames is a Python library that lets you take advantage of BigQuery data processing by using familiar Python APIs.

Generic Coding Guidelines

  • Avoid .to_pandas(): You MUST NOT use .to_pandas() to download the entire dataset into memory. There are some exceptions:
    • An error message explicitly requests you to use to_pandas()
    • You are going to visualize the data, and the visualization library does not accept BigFrames Dataframe/Series instances. In this case, reduce the amount of data you are going to download before calling .to_pandas()
  • Avoid read_gbq() for SQL: Do not write SQL queries and execute them with read_gbq(). Use BigFrames Dataframe/Series methods instead.
  • Use BigFrames ML package for Machine Learning Tasks: Do not use Scikit-learn or other ML libraries with BigFrames dataframes. Import your tools/classes from bigframes.ml.
  • Stay in the Cloud: Perform data cleaning, transformation, and analysis via BigFrames methods to leverage BigQuery's scale.
  • Accessors over UDFs/Lambdas:
    • Prefer built-in accessors (e.g., df.col.str.*, df.col.dt.*) over remote UDFs.
    • Do not use lambdas with Series.map() or DataFrame.apply().
  • Schema Verification: Do not assume schema of intermediate outputs. Check .dtypes after loading, and use display() with .head() or .peek().
  • Visualization: BigFrames Dataframe mostly works directly with Matplotlib, Seaborn, and other plotting libraries. If your attempt didn't work, try using the "plot" accessor. If that didn't work either, you MUST sample or aggregate your data to make it small enough before calling "to_pandas()".

Model Development

  • Unlike Scikit-learn: BigFrames' predict() method always returns a DataFrame containing both predictions and features (not just a series of predictions).
  • No random_state: Do not pass a random_state argument when instantiating BigFrames ML models.
  • Automatic Scaling: Do not use OneHotEncoder or StandardScaler unless explicitly requested (handled automatically).
  • Hyperparameter Tuning: You must write custom loops (BigFrames lacks GridSearchCV or RandomizedSearchCV).
  • ARIMA Plus (Forecasting):
    • Import from bigframes.ml.forecasting.
    • Sort data chronologically and split around a timepoint before training.
    • Prediction horizon must be less than or equal to training horizon.
  • PCA: BigFrames' PCA class lacks simple transform() method. Use predict() instead.
  • Model Persistence: To persist a model, use model.to_gbq(). To load a persisted model, use bpd.read_gbq_model().

Source and attribution

Source:gemini-cli-extensions/data-agent-kit-starter-packinskills/bigquery-bigframesat commit2df10e2

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from gemini-cli-extensions/data-agent-kit-starter-pack

Schema Mapping

gemini-cli-extensions

Plans source-to-target schema mappings for ETL, ELT, or data integration work, producing a documented Mapping Manifesto.

Data & Analytics215updated today

Resolving Mcp Region Configs

gemini-cli-extensions

Fixes unreplaced region placeholders in regional Google Cloud MCP server configs so missing MCP tools register.

DevOps & Cloud215updated today

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

Awaiting classification215updated today

Ml Best Practices

gemini-cli-extensions

Guides machine learning notebooks with step-by-step plans for clustering, forecasting, classification, regression and model comparison.

Data & Analytics215updated today

Managing Python Dependencies

gemini-cli-extensions

Guides agents to detect a Python project's dependency manager and install packages correctly instead of using global pip.

Software Development215updated today

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

Awaiting classification215updated today