
Ml Best Practices
by gemini-cli-extensions2df10e25bbf7Apache-2.0215 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today
CRITICAL RULE: You MUST use this skill whenever the task involves any machine learning tasks or data analysis. Use this skill if the user's prompt or requirements mention any of the following: * Clustering * Classification * Regression * Time series forecasting * Statistical testing * Model comparison * ML * Data analysis SQL/BigQuery ML HANDOFF: If the user requires a SQL solution, use this skill to dictate the ANALYSIS STEPS (e.g., markdown analysis cells, visualization logic), but defer to `bigquery` for all SQL syntax.
Only the file list is public. File contents are available once the skill is installed in a workspace.
| Path | Size | Type |
|---|---|---|
| SKILL.md | 9.6 KB | text/markdown |
Source and attribution
Source:gemini-cli-extensions/data-agent-kit-starter-packinskills/ml-best-practicesat commit2df10e2
License: Apache-2.0
Content belongs to its original authors. SourceWeft indexes it from a public repository.
More from gemini-cli-extensions/data-agent-kit-starter-pack

Schema Mapping
gemini-cli-extensions
Plans source-to-target schema mappings for ETL, ELT, or data integration work, producing a documented Mapping Manifesto.

Resolving Mcp Region Configs
gemini-cli-extensions
Fixes unreplaced region placeholders in regional Google Cloud MCP server configs so missing MCP tools register.

Notebook Guidance
gemini-cli-extensions
This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

Managing Python Dependencies
gemini-cli-extensions
Guides agents to detect a Python project's dependency manager and install packages correctly instead of using global pip.

Google Cloud Storage Fuse
gemini-cli-extensions
Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

Google Cloud Auth Verification
gemini-cli-extensions
Mandatory Step 0 pre-flight execution order and authentication verification for Google Cloud Platform (GCP), Application Default Credentials (ADC), gcloud CLI, Spark, Dataproc, BigQuery, GCS, and notebook runtimes. Use whenever interacting with GCP resources, running Spark/PySpark pipelines, BigQuery queries, GCS paths (gs://), or creating/running notebooks.
More in Data & Analytics

Spanner Basics
Guides Google Cloud Spanner administration, schema design, querying and performance diagnosis.

Gke Cost Analysis
Answers natural-language questions about GKE cluster and workload costs using BigQuery billing exports and live cluster metrics.

Datalineage Bigquery Asset Impact Analysis
Guides an agent through downstream impact (blast radius) analysis for a BigQuery table or view using Data Lineage.

Bigquery Troubleshooting
Diagnoses failing, slow, or unexpectedly expensive BigQuery jobs through structured root-cause workflows.

Bigquery Optimization
Guides BigQuery cost and performance optimization across capacity editions, storage layout, and SQL queries.

Bigquery Bigframes
Guides writing Python code with BigQuery DataFrames (BigFrames) for data processing, analysis, and machine learning.