Ml Best Practices

作者 gemini-cli-extensions2df10e25bbf7Apache-2.0215 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

CRITICAL RULE: You MUST use this skill whenever the task involves any machine learning tasks or data analysis. Use this skill if the user's prompt or requirements mention any of the following: * Clustering * Classification * Regression * Time series forecasting * Statistical testing * Model comparison * ML * Data analysis SQL/BigQuery ML HANDOFF: If the user requires a SQL solution, use this skill to dictate the ANALYSIS STEPS (e.g., markdown analysis cells, visualization logic), but defer to `bigquery` for all SQL syntax.

AI 生成的概览

为机器学习笔记本提供分步方案,涵盖聚类、预测、分类、回归和模型比较。

功能
该技能为常见机器学习任务提供结构化的分析方案,包括聚类、时间序列预测、探索性数据分析与异常检测、分类、回归以及模型比较。每个方案列出有序步骤,例如理解数据结构、可视化、编码、划分数据集、训练、评估,并以最终 markdown 总结收尾。它还说明了特征化顺序和处理缺失值的基本实践。产出是带有 markdown 分析说明的笔记本式分析。
适用场景
当任务涉及机器学习或数据分析时使用,例如聚类、分类、回归、预测、统计检验或模型比较。它也适用于 SQL 或 BigQuery ML 请求,此时由它规定分析步骤,而 SQL 语法交由其他技能处理。
运行要求
仅为说明性内容,不包含脚本。需要能够运行分析代码并生成笔记本单元格的环境,并引用单独的 bigquery 技能来处理 SQL 语法。

ML Best Practices

I want to read a story about the data, not just run code. Ensure every code cell is followed by a markdown cell analyzing the results. End the notebook with a summary comprehensively answering the prompt.

If there is a good match between the user's request and a corresponding example plan, then adapt the example plan to fully answer the user's request:

Clustering:

Identify distinct groups based on their features.

  • Understand the schema and field descriptions.
  • Visualize features referenced in the prompt (e.g., with histograms, scatterplots).
  • Transform dates into timestamps.
  • Before applying encoders, check if the dataset already contains pre-encoded features and prefer existing numerical representations.
  • Prefer to keep data instead of dropping it when possible.
  • Transform ordinal data with an ordinal encoder.
  • Transform nominal data with a one hot encoder.
  • Standardize numerical features.
  • Perform clustering with a range of values, and collect the silhouette score.
  • Choose the optimal number of clusters based on the silhouette score.
  • Use dimensionality reduction (e.g., PCA) to project the data into two dimensions.
  • Scatterplot the samples in two dimensions with cluster labels as the hue.
  • Scatterplot the samples in two dimensions with a discrete feature as the hue.
  • Describe the clusters in text by feature distributions or typical feature values.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Time Series Forecasting:

Develop a predictive model to estimate future values based on historical trends. How might different modeling approaches impact the prediction accuracy?

  • Understand the schema and field descriptions.
  • Visualize the target feature over time at a reasonable granularity.
  • Always perform a chronological split on the data to create training, validation, and test sets.
  • Are there seasonal trends?
  • Test for stationarity.
  • Discuss possible modeling approaches. How might different modeling approaches impact the prediction accuracy?
  • Train two time series forecasting models to predict the target feature. Use previous seasonality and stationarity information as model hyperparameters.
  • Predict the target feature for the training and validation sets.
  • Optionally, hypertune models with the validation set.
  • Visualize the actual and predicted target feature vs time for each model on the training and validation sets.
  • Evaluate the validation performance with error metrics.
  • Select a model.
  • Retrain the selected model on the test and validation sets.
  • Predict the test values with the selected model.
  • Visualize the average target feature and the predicted test values.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Exploratory Data Analysis / Anomaly Detection:

Identify and describe any outliers, unusual patterns, or significant trends observed in the data. Provide visualizations to support your findings.

  • Understand the schema and field descriptions.
  • Visualize the target feature distribution in a way that shows outliers.
  • Identify and describe any outliers in the target feature.
  • Visualize relationships between the target feature and other features.
  • Identify and describe unusual patterns or significant trends.
  • Visualize patterns and trends.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Classification:

Given the data, can we classify by the target feature?

  • Understand the schema and field descriptions.
  • Identify rows that don't make sense. How many are there and what do they contain?
  • Identify rows without a target value. How many are there and what do they contain?
  • Drop rows that don't match the schema or don't have the target value (if it is reasonable to do so).
  • Split data into training, validation, and test sets.
  • Create features to represent when data are missing, if this is meaningful.
  • Handle missing data. Prefer to keep data instead of dropping it when possible.
  • Before applying encoders, check if the dataset already contains pre-encoded features and prefer existing numerical representations.
  • Transform ordinal data with an ordinal encoder.
  • Transform nominal data with a one hot encoder.
  • Standardize numerical features.
  • Train multiple models.
  • If there is evidence of overfitting, regularize and retrain the model.
  • If there is evidence of underfitting, consider adding or engineering features.
  • Evaluate the models.
  • Create confusion matrices.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Regression:

Predict the continuous valued target feature.

  • Understand the schema and field descriptions.
  • Identify rows that don't make sense. How many are there and what do they contain?
  • Identify rows without a target value. How many are there and what do they contain?
  • Develop an understanding of the data and determine how to handle missing values. This should make sense in the business context.
  • Identify any potential sources of group leakage. Aggregate where appropriate to prevent this.
  • Visualize target feature.
  • Split data into training, validation, and test sets.
  • Handle missing data. Prefer to keep data instead of dropping it when possible.
  • Before applying encoders, check if the dataset already contains pre-encoded features and prefer existing numerical representations.
  • Transform ordinal data with an ordinal encoder.
  • Transform nominal data with a one hot encoder. Restrict high cardinality categorical features to a tractable size.
  • Standardize numerical features.
  • Train multiple models.
  • Visualize the actual vs predicted values on training and validation data.
  • If there is evidence of overfitting, regularize and retrain the model.
  • If there is evidence of underfitting, consider adding or engineering features.
  • Evaluate the model error.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Comparing ML Models:

Evaluate and compare multiple models to determine which is most suitable for production based on predictive power, robustness, and viability.

  • Understand the schema and align metrics with business goals (e.g., cost of false positives vs. false negatives).
  • Establish baselines: define a naive baseline (majority class/mean) and a simple ML baseline (e.g., Logistic/Linear Regression).
  • Ensure rigorous validation: use identical, fixed data splits for all models and perform $k$-fold cross-validation.
  • If data is temporal, use chronological splits for validation.
  • Select and report metrics beyond accuracy (e.g., F1-Score, PR-AUC, MAE, RMSE) that reflect business impact.
  • Use bootstrapping to calculate 95% confidence intervals for key metrics to determine statistical significance.
  • Perform slice-based error analysis: evaluate model performance across key subpopulations and demographics to identify bias or specific failure modes.
  • Inspect and compare confusion matrices, residual plots, and calibration curves.
  • Evaluate operational trade-offs: consider inference latency, training time, compute cost, and model size.
  • Assess interpretability using tools like SHAP or LIME where transparency is required.
  • Conclusion: Recommend the optimal model for the specific use case, justifying the choice with both performance and production viability.

No match:

  • Understand the schema and field descriptions.
  • Identify rows that don't make sense. How many are there and what do they contain?
  • Identify rows without a target value. How many are there and what do they contain?
  • Drop rows that don't match the schema or don't have the target value (if it is reasonable to do so).
  • Create features to represent when data are missing, if this is meaningful.
  • Handle missing data. Prefer to keep data instead of dropping it when possible.
  • Before applying encoders, check if the dataset already contains pre-encoded features and prefer existing numerical representations.
  • Transform ordinal data with an ordinal encoder.
  • Transform nominal data with a one hot encoder.
  • Standardize numerical features.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Essential ML Practices

[!IMPORTANT] ALWAYS follow these ML practices

  • Strict Featurization Ordering: For supervised learning ALWAYS split the dataset into training and test data BEFORE fitting preprocessing pipelines (e.g. scaling, encoding). Fit the pipelines on the training data and test data independently.

  • Handling Missing or NULL Values: ALWAYS check for and handle missing and NULL values. First, analyze their frequency. Then, decide whether to keep them, drop them or impute them with a contextually appropriate value, and explain your reasoning.

来源与署名

来源:gemini-cli-extensions/data-agent-kit-starter-pack位于skills/ml-best-practices提交2df10e2

许可证: Apache-2.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 gemini-cli-extensions/data-agent-kit-starter-pack 的技能

Schema Mapping

gemini-cli-extensions

为 ETL、ELT 或数据集成任务规划源到目标的模式映射,产出文档化的映射宣言。

Data & Analytics215今天更新

Resolving Mcp Region Configs

gemini-cli-extensions

修复区域级 Google Cloud MCP 服务器配置中未替换的区域占位符,使缺失的 MCP 工具得以注册。

DevOps & Cloud215今天更新

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

待分类215今天更新

Managing Python Dependencies

gemini-cli-extensions

指导代理检测 Python 项目的依赖管理器并正确安装依赖,而不是使用全局 pip。

Software Development215今天更新

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

待分类215今天更新

Google Cloud Auth Verification

gemini-cli-extensions

Mandatory Step 0 pre-flight execution order and authentication verification for Google Cloud Platform (GCP), Application Default Credentials (ADC), gcloud CLI, Spark, Dataproc, BigQuery, GCS, and notebook runtimes. Use whenever interacting with GCP resources, running Spark/PySpark pipelines, BigQuery queries, GCS paths (gs://), or creating/running notebooks.

待分类215今天更新