Bigquery Optimization

作者 google55b4e13eba6d無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Provides workflows to optimize BigQuery environments (capacity planning, editions), storage assets (partitioning, clustering, storage lifecycles, billing models), and SQL queries. Use when optimizing cost, modeling Edition migrations, rightsizing reservations, evaluating logical vs. physical storage, designing table partitioning/clustering, generating table DDL, migrating unpartitioned tables, managing partition expiration, or optimizing individual SQL queries. Do not use for raw usage reporting (use bigquery-observability), query execution plan analysis, error troubleshooting, or diagnosing why a specific job was slow (use bigquery-troubleshooting).

AI 產生的概覽

指導 BigQuery 在容量版本、儲存配置與 SQL 查詢方面的成本與效能最佳化。

功能
此技能提供最佳化 BigQuery 環境的工作流程與參考指引。內容涵蓋容量規劃與 Editions 移轉建模、儲存計費模式、資料表分割與叢集 DDL、儲存生命週期與分割區到期,以及個別 SQL 查詢改寫。它產出建議、DDL 範本與主控台導覽指引,本身不執行變更。
適用情境
適用於評估 BigQuery 成本效益、建模從隨需計費到 Editions 的移轉、調整預留資源規模、選擇實體或邏輯儲存計費、設計分割或叢集、移轉未分割資料表,或改寫 SQL 查詢以減少槽位時間與資料讀取量的情境。
執行需求
需要安裝並設定 Google Cloud SDK,具備已驗證的 gcloud 工作階段或應用程式預設憑證、有效的帳單帳戶,以及 bigquery.admin、bigquery.resourceAdmin、bigquery.dataEditor 或 bigquery.jobUser 等 IAM 角色。需啟用 BigQuery 與 BigQuery Reservation API。此技能不附帶指令碼,僅為指示與參考文件。

BigQuery Optimization Workflow

Prerequisites & Environment Setup

Before executing optimization analyses, evaluating editions, or applying DDL modifications:

  1. Google Cloud SDK: Ensure the Google Cloud SDK is installed and configured.

  2. Project Selection: Set the active Google Cloud project:

    bash
    gcloud config set project {project_id}
  3. API Enablement: Ensure BigQuery and BigQuery Reservation APIs are enabled:

    bash
    gcloud services enable \    bigquery.googleapis.com bigqueryreservation.googleapis.com
  4. Authentication: Authenticate the environment:

    • CLI tools and bq commands: gcloud auth login
    • SDKs and automation: gcloud auth application-default login
    • Service accounts: Set GOOGLE_APPLICATION_CREDENTIALS="/path/to/key.json"
  5. Billing & IAM Roles:

    • Verify an active Google Cloud Billing account is attached to {project_id}.
    • Ensure appropriate IAM roles:
      • roles/bigquery.admin or roles/bigquery.resourceAdmin: Reservation and capacity commitment management.
      • roles/bigquery.dataEditor or roles/bigquery.admin: Modifying table schemas, partitioning, clustering, and storage billing models.
      • roles/bigquery.jobUser: Running evaluation queries.
  6. Companion Skills Installation: This skill is part of a 3-pillar operations suite (bigquery-observability, bigquery-optimization, bigquery-troubleshooting). If any companion skill is not yet installed in your environment, install the full suite:

    bash
    npx skills add google/skills --skill bigquery-observability --skill bigquery-optimization --skill bigquery-troubleshooting

    (If bigquery-observability is not installed, use the self-contained baseline formulas and query templates provided directly in the reference sections below).

Workflows

Determine the optimization focus of the user's request and follow the relevant workflow:

  • Telemetry & Observability Baseline: For direct raw usage telemetry, INFORMATION_SCHEMA queries, and baseline metric calculations, consult bigquery-observability (bigquery_observability). If the bigquery-observability companion skill is not available in the active environment, all optimization guidelines, DDL templates, and decision models across this skill and its reference guides are fully self-contained.
  • Capacity & Editions Modeling: Evaluate the cost-efficiency of migrating workloads from On-Demand to Editions, as well as rightsizing active Edition reservations, baseline commitments, and autoscaling caps.
    • Instructions: Read references/capacity_planning_editions.md to provide deep links to BigQuery's built-in recommendation UIs (e.g., Slot Estimator) and guide the user through UI navigation: 1. navigate to the Slot Estimator tab, 2. select 'On-Demand' as the source to analyze historical query volume, and 3. review the Cost-Optimized Recommendations and Slot Usage Chart.
  • Table & Storage Optimization: Optimize storage costs from a billing model, physical layout, and lifecycle perspective.
    • Billing Architecture: Read references/storage_billing_models.md for guidance on evaluating aggregate compression ratios (e.g. >2:1 threshold in US) to recommend Physical vs. Logical billing, noting that the break-even ratio depends on specific regional rates and custom enterprise contracts. When providing TABLE_STORAGE queries, always scope with WHERE table_schema = '{dataset_id}', use the regional dataset view, and warn that 0 rows indicates a region mismatch or lack of native tables rather than zero billable usage.
    • Partitioning & Clustering Strategy: Read references/table_partitioning_clustering.md to generate production DDL templates (CREATE TABLE, CTAS migrations for unpartitioned tables, and modifying clustering specifications), enforce pruning with require_partition_filter = true, and manage partition limits (up to 10,000 partitions/table).
    • Lifecycle Management: Read references/storage_lifecycle_management.md to pinpoint inactive data and define precise Time-to-Live (TTL) partition expirations, dataset expirations, and Time Travel window reductions.
  • SQL Optimization: Optimize individual SQL queries to reduce slot-time and the amount of data read.
    • Instructions: Follow the instructions in references/sql_optimization.md to provide recommendations to the user on how to rewrite their SQL query to reduce slot-time and the amount of data read.

Execution Guardrails

  • Terminology & Cost Framing: Never promise or guarantee "cost-reduction" or "reducing expenditure." Always frame recommendations using the terminology "optimizing your bill" or "improving cost-efficiency."
  • Explicit Scope Framing & Region Resolution: Always state the target project_id and region at the very top of your response so the user immediately knows the exact scope being evaluated. Follow this 3-tier resolution hierarchy:
    1. Explicit Region: Use the region specified in the user's prompt (e.g., europe-west1).
    2. Contextual Region: Resolve the region from the specific dataset or resource mentioned in the context.
    3. Unspecified Fallback: Default to us / region-us, explicitly state that us was assumed as the default, and instruct the user to substitute their region if their resources reside elsewhere. Region Formatting: In Cloud Console deep links, use the region identifier directly (e.g., region=us, region=europe-west1). In SQL queries against INFORMATION_SCHEMA, use the regional dataset qualifier (e.g., region-us, region-europe-west1).
  • Zero-Row Result Guard: If querying TABLE_STORAGE with WHERE table_schema = '{dataset_id}' returns 0 rows, do not proceed with an empty or zero-usage evaluation. Treat this as an indicator that the dataset may reside in a different region or have no native tables; stop and prompt the user to confirm the dataset's regional location.
  • Populate Concrete Parameters: When generating URLs and SQL queries, always substitute known project_id and region values directly into the code and links. Never leave literal {project_id} or {location} placeholders for the user to manually edit.
  • No Autonomous Purchasing or Financial Mutations: Never provide the user with executable scripts (e.g., gcloud or bq shell commands like bq update --storage_billing_model=...) designed to autonomously purchase annual commitments, alter edition tier bindings, or mutate storage billing models. Always guide the user to execute commitment purchases, reservation changes, and storage billing model updates manually via the Cloud Console UI.

來源與署名

來源:google/skills位於skills/cloud/bigquery-optimization提交55b4e13

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 google/skills 的技能

Google Cloud Solution N Tier Serverless Web App

google

精選

指導在 Google Cloud 上設計與實作安全的多層無伺服器網頁應用程式。

DevOps & Cloud2026年10月8日

Google Cloud Solution Hybrid Search Alloydb

google

精選

Discovers requirements and generates architectural, design, and deployment guidance for dynamic hybrid search systems by combining semantic search and keyword search. Optimized for AlloyDB hybrid search use cases in Google Cloud. Use when users need vector search combined with structured SQL filtering, faceted attributes, semantic reranking, in-database AI validation, or serverless hosting across transactional relational databases, analytical data warehouses, or managed database engines. DON'T use this skill for simple keyword-only search, or when a standalone non-relational vector database is required.

待分類2026年10月8日

Google Cloud Solution Architecture

google

精選

Interactively discovers requirements and designs holistic, multi-product system architectures, solution blueprints, and deployment recommendations for complex workloads on Google Cloud. Use when designing end-to-end cloud solutions, selecting and integrating Google Cloud services, generating architecture diagrams, or conducting requirements discovery for new cloud workloads or migrations. Don't use for single-product tasks (use product-specific skills), initial onboarding or authentication (use google-cloud-recipe-*), Well-Architected Framework reviews or audits (use google-cloud-waf-*), or workloads covered by specialized solution skills.

待分類2026年10月8日

Google Cloud Solution Agentic Ai Data Science Workflow

google

精選

Designs a tailored multi-product agentic data science architecture on Google Cloud that incorporates opinionated best practices. Use when architecting multi-product solutions for agent-based data analytics or ML workloads. Don't use for simple queries, non-agentic pipelines, general cloud reviews, or writing agent code.

待分類2026年10月8日

Google Cloud Solution Agentic Ai Borderless Data Lakehouse

google

精選

Discovers requirements and designs a borderless open data lakehouse using Lakehouse for Apache Iceberg and BigQuery data agents. Use when architecting multi-cloud storage infrastructure (Cloud Storage, AWS S3, Azure Blob), establishing ingestion and AI serving subsystems, configuring Cross-Cloud Interconnect, or deploying Gemini Enterprise Agent Platform and BigQuery data agents. Don't use for single-cloud data warehouses, or when the focus is on Knowledge Catalog metadata governance and Spark-driven IDE analytics workflows (use google-cloud-solution-agentic-analytics-spark-knowledge-catalog instead).

待分類2026年10月8日

Google Cloud Solution Agentic Ai Bidirectional Streaming

google

精選

Guides agents to interactively discover customer requirements for live, bidirectional multi-agent AI systems that process continuous streams of multimodal data for real-time technical guidance and safety monitoring. Generates a custom Google Cloud solution that uses opinionated best practices and architecture guidance. Use when users need agentic assistance to design and create a multi-product solution in the cloud for live bidirectional multimodal streaming workloads. Don't use for simple text-based chat applications or workloads without real-time streaming requirements.

待分類2026年10月8日