Bigquery Bigframes

作者 google55b4e13eba6d無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Generates Python code using BigQuery DataFrames (BigFrames). Use by default for any Python data task involving BigQuery, including data processing, analysis, and machine learning. Don't use for SQL-first workflows or the google-cloud-bigquery client library — use bigquery-basics.

精選僅含說明Data & Analytics
AI 產生的概覽

指導使用 BigQuery DataFrames(BigFrames)撰寫 Python 程式碼,用於資料處理、分析與機器學習。

功能
此技能提供最佳實務說明,用於產生使用 BigQuery DataFrames(BigFrames)程式庫的 Python 程式碼。內容涵蓋 DataFrame API 的用法,例如部分排序模式、以 peek 預覽資料、避免在本機具體化資料、優先使用 DataFrame 方法而非原始 SQL,以及使用存取器取代 UDF 或 lambda。它也涵蓋使用 bigframes.bigquery.ml 套件與舊版 bigframes.ml 套件進行機器學習,包括迴歸、預測、PCA 與模型保存。另有兩個參考檔案說明線性迴歸與邏輯迴歸。
適用情境
適用於涉及 BigQuery 的 Python 資料任務,包括資料處理、分析與機器學習。它是這類任務的預設選擇,但不適用於以 SQL 為主的工作流程或 google-cloud-bigquery 用戶端程式庫。
執行需求
需要 BigQuery DataFrames(BigFrames)Python 程式庫以及對 BigQuery 的存取權。此技能不含指令碼,僅為指示文件,並附帶兩個可供模型讀取的參考檔案。

BigFrames (BigQuery DataFrame) basics

BigFrames is a Python library that lets you take advantage of BigQuery data processing by using familiar Python APIs.

Dataframe API best practices

  • Stay in the Cloud: Perform data cleaning, transformation, and analysis via BigFrames methods to leverage BigQuery's scale rather than downloading data.

  • Prefer partial ordering mode: Enable partial ordering mode right after importing BigFrames. This speeds up data processing significantly by relaxing row-sequence constraints.

    python
    import bigframes.pandas as bpdbpd.options.bigquery.ordering_mode = 'partial'
  • Use peek() for data preview: Use peek(n) to preview data instead of head(n). peek(n) randomly samples n rows and is significantly faster. head(n) returns rows in strict order and fails in partial ordering mode unless the DataFrame has been explicitly sorted.

  • Avoid materializing data locally: Methods like to_pandas() download all data to client memory, bypassing BigQuery’s distributed computation and risking Out of Memory (OOM) errors. Do not materialize data locally unless:

    • The dataset is small enough to fit safely in memory.
    • An error message explicitly requires local materialization.
  • Prefer Dataframe API over SQL queries: Do not write raw SQL queries via read_gbq() if a DataFrame/Series method achieves the same result, as it breaks the Pandas abstraction and prevents lazy query execution.

  • Accessors over UDFs/Lambdas:

    • Use built-in accessors (e.g., df.col.str.*, df.col.dt.*) instead of remote User Defined Functions (UDFs). UDFs require extra resources and time to deploy.
    • Do not use lambdas with Series.map() or DataFrame.apply(). These methods do not accept functions without udf or remote_function decorators.
    python
    # Avoid:df["upper"] = df["name"].map(lambda x: x.upper())
    # Prefer:df["upper"] = df["name"].str.upper()
  • Schema Verification: Do not assume the schema of intermediate outputs. Proactively verify schemas using .dtypes and inspect sample records using display() with .peek().

  • Visualization: Plot directly from the BigFrames DataFrame/Series when possible. BigFrames is compatible with Matplotlib and Seaborn. If direct plotting fails, use the .plot accessor. If the dataset is too large to plot, aggregate or sample the data before calling .to_pandas() to plot locally.

Machine Learning

  • Use bigframes.bigquery.ml package: Do not use Scikit-learn or other ML libraries with BigQuery DataFrames. Standard Scikit-learn models require bringing data into local client memory, whereas bigframes.bigquery.ml delegates training directly to BigQuery's scalable ML engine. Import functions from bigframes.bigquery.ml.

Reference Directory

  • Linear Regression [blocked]: Train a linear regression model to predict numerical values.
  • Logistic Regression [blocked]: Train a logistic regression model to predict boolean values.

BigFrames ML (Legacy)

The BigFrames ML package (bigframes.ml) is a legacy package that mimics the scikit-learn API but is no longer recommended for new projects. Only use this package if the user explicitly requests BigFrames ML.

  • Legacy Imports: When legacy BigFrames ML is requested, import tools and classes from bigframes.ml instead of bigframes.bigquery.ml.
  • DataFrame Return on Prediction: Unlike Scikit-learn, BigFrames' predict() method always returns a DataFrame containing both predictions and features, rather than a single series of predictions.
  • No random_state: Do not pass a random_state argument when instantiating BigFrames ML models, as this parameter is not supported in the BigFrames ML package.
  • Automatic Scaling: Do not use OneHotEncoder or StandardScaler unless explicitly requested, as scaling is handled automatically.
  • Hyperparameter Tuning: Write custom loops for hyperparameter tuning, as BigFrames lacks GridSearchCV or RandomizedSearchCV.
  • ARIMA Plus (Forecasting):
    • Import from bigframes.ml.forecasting.
    • Sort data chronologically and split around a timepoint before training.
    • Ensure the prediction horizon is less than or equal to the training horizon.
  • PCA: BigFrames' PCA class lacks a transform() method. Use predict() instead.
  • Model Persistence: To persist a model, use model.to_gbq(). To load a persisted model, use bpd.read_gbq_model().

來源與署名

來源:google/skills位於skills/cloud/bigquery-bigframes提交55b4e13

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 google/skills 的技能

Dpop Adoption

google

精選

Implement and debug OAuth 2.0 DPoP (RFC 9449) refresh token sender-constraining for WebCrypto, Node.js ES6, and browser runtimes integrating with Google's OAuth platform. Use when configuring non-extractable asymmetric key pairs (P-256), generating DPoP Proof JWTs for authorization code exchange and token refresh, or handling 400 use_dpop_nonce challenge retry loops at oauth2.googleapis.com/token. Don't use for unconstrained OAuth 2.0 flows (where refresh tokens are not bound to a client key pair), or for Google Cloud IAM / service account authentication.

待分類2026年10月8日

Finding Google Skills

google

精選

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

待分類2026年10月8日

Spanner Basics

google

精選

指導 Google Cloud Spanner 的執行個體與資料庫管理、結構定義設計、查詢與效能診斷。

Data & Analytics2026年10月8日

Secops Triage

google

精選

引導 SOC 分析師對 Google SecOps 安全警示進行分診,從調查到結案或升級。

Security2026年10月8日

Secops Investigate

google

精選

指導 SOC 分析師在 Google SecOps 中使用 UDM 查詢與時間軸進行深入的安全事件與實體調查。

Security2026年10月8日

Secops Hunt

google

精選

指導在 Google SecOps 中使用 UDM 查詢、IoC 回溯、普遍性與異常分析進行主動威脅狩獵。

Security2026年10月8日