Bigquery Ai Ml

by gemini-cli-extensions2df10e25bbf7No license215 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Leverages BigQuery's built-in machine learning and GenAI capabilities for advanced data analytics. Use when you need to write SQL queries that perform time-series forecasting, detect outliers, find key drivers, or leverage generative AI capabilities in BigQuery.

Instructions onlyData & Analytics
AI-generated overview

Guides writing BigQuery SQL that uses built-in AI and ML functions for forecasting, anomaly detection, and GenAI.

What it does
This skill provides reference guidance for authoring BigQuery SQL that calls built-in AI and ML functions such as AI.FORECAST, AI.KEY_DRIVERS, AI.DETECT_ANOMALIES, and AI.GENERATE. It routes the agent to per-function reference documents covering classification, scoring, embeddings, semantic search, similarity, evaluation, and vector search. It also covers remote Vertex AI models and contribution analysis, which requires creating a model entity. The deliverable is written SQL query guidance rather than executed code.
When to use it
Use it when you need to write BigQuery SQL that performs time-series forecasting, outlier detection, key driver analysis, or generative AI operations. It fits analytics work where the AI or ML capability is invoked directly from SQL against BigQuery data.
Requirements
Requires BigQuery with access to its built-in AI and ML functions, and Vertex AI for remote models. No scripts ship with the skill; it is instructions and reference documents only.

BigQuery AI & ML

BigQuery integrates with Vertex AI to provide powerful machine learning and generative AI capabilities directly within SQL queries using built-in functions like AI.FORECAST, AI.KEY_DRIVERS, AI.DETECT_ANOMALIES, and AI.GENERATE.

[!IMPORTANT] You MUST read and follow the global constraints and mandatory function routing rules in ai_function_best_practices.md [blocked] before writing any BQML AI/ML SQL query.

Reference Directory

  • Best Practices: ai_function_best_practices.md [blocked]

  • Functions Reference:

    • AI.AGG: ai_agg.md [blocked] - Multi-row semantic aggregation and summarization.
    • AI.CLASSIFY: ai_classify.md [blocked] - Classify text.
    • AI.DETECT_ANOMALIES: ai_detect_anomalies.md [blocked] - Detect anomalies.
    • AI.EVALUATE: ai_evaluate.md [blocked] - Evaluate models.
    • AI.FORECAST: ai_forecast.md [blocked] - Time-series forecasting.
    • AI.GENERATE: ai_generate.md [blocked] - Generate text using LLMs.
    • AI.GENERATE_EMBEDDING: ai_generate_embedding.md [blocked] - Generate embeddings.
    • AI.GENERATE_TABLE: ai_generate_table.md [blocked] - Table-valued AI generation.
    • AI.IF: ai_if.md [blocked] - Evaluate semantic conditions.
    • AI.KEY_DRIVERS: ai_key_drivers.md [blocked] - Identifies key drivers, this is a TVF.
    • AI.SCORE: ai_score.md [blocked] - Score data.
    • AI.SEARCH: ai_search.md [blocked] - Semantic search.
    • AI.SIMILARITY: ai_similarity.md [blocked] - Semantic similarity.
    • Remote Models: remote_models.md [blocked] - Working with remote models (Vertex AI).
    • CONTRIBUTION_ANALYSIS: ml_contribution_analysis.md [blocked]
      • Finds contributing factors, key drivers of change. Requires creating a MODEL entity.
    • VECTOR_SEARCH: vector_search.md [blocked] - Vector search best practices.

Source and attribution

Source:gemini-cli-extensions/data-agent-kit-starter-packinskills/bigquery-ai-mlat commit2df10e2

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from gemini-cli-extensions/data-agent-kit-starter-pack

Schema Mapping

gemini-cli-extensions

Plans source-to-target schema mappings for ETL, ELT, or data integration work, producing a documented Mapping Manifesto.

Data & Analytics215updated today

Resolving Mcp Region Configs

gemini-cli-extensions

Fixes unreplaced region placeholders in regional Google Cloud MCP server configs so missing MCP tools register.

DevOps & Cloud215updated today

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

Awaiting classification215updated today

Ml Best Practices

gemini-cli-extensions

Guides machine learning notebooks with step-by-step plans for clustering, forecasting, classification, regression and model comparison.

Data & Analytics215updated today

Managing Python Dependencies

gemini-cli-extensions

Guides agents to detect a Python project's dependency manager and install packages correctly instead of using global pip.

Software Development215updated today

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

Awaiting classification215updated today