Ml Best Practices

作者 gemini-cli-extensions2df10e25bbf7Apache-2.0215 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

CRITICAL RULE: You MUST use this skill whenever the task involves any machine learning tasks or data analysis. Use this skill if the user's prompt or requirements mention any of the following: * Clustering * Classification * Regression * Time series forecasting * Statistical testing * Model comparison * ML * Data analysis SQL/BigQuery ML HANDOFF: If the user requires a SQL solution, use this skill to dictate the ANALYSIS STEPS (e.g., markdown analysis cells, visualization logic), but defer to `bigquery` for all SQL syntax.

AI 產生的概覽

為機器學習筆記本提供逐步方案,涵蓋分群、預測、分類、迴歸與模型比較。

功能
此技能為常見機器學習任務提供結構化的分析方案,包括分群、時間序列預測、探索性資料分析與異常偵測、分類、迴歸以及模型比較。每個方案列出有序步驟,例如了解資料結構、視覺化、編碼、切分資料集、訓練、評估,並以最終 markdown 總結作結。它也說明了特徵化順序與處理缺失值的基本實務。產出是帶有 markdown 分析說明的筆記本式分析。
適用情境
當任務涉及機器學習或資料分析時使用,例如分群、分類、迴歸、預測、統計檢定或模型比較。它也適用於 SQL 或 BigQuery ML 請求,此時由它規定分析步驟,而 SQL 語法交由其他技能處理。
執行需求
僅為說明性內容,不包含指令碼。需要能執行分析程式碼並產生筆記本儲存格的環境,並引用獨立的 bigquery 技能來處理 SQL 語法。

ML Best Practices

I want to read a story about the data, not just run code. Ensure every code cell is followed by a markdown cell analyzing the results. End the notebook with a summary comprehensively answering the prompt.

If there is a good match between the user's request and a corresponding example plan, then adapt the example plan to fully answer the user's request:

Clustering:

Identify distinct groups based on their features.

  • Understand the schema and field descriptions.
  • Visualize features referenced in the prompt (e.g., with histograms, scatterplots).
  • Transform dates into timestamps.
  • Before applying encoders, check if the dataset already contains pre-encoded features and prefer existing numerical representations.
  • Prefer to keep data instead of dropping it when possible.
  • Transform ordinal data with an ordinal encoder.
  • Transform nominal data with a one hot encoder.
  • Standardize numerical features.
  • Perform clustering with a range of values, and collect the silhouette score.
  • Choose the optimal number of clusters based on the silhouette score.
  • Use dimensionality reduction (e.g., PCA) to project the data into two dimensions.
  • Scatterplot the samples in two dimensions with cluster labels as the hue.
  • Scatterplot the samples in two dimensions with a discrete feature as the hue.
  • Describe the clusters in text by feature distributions or typical feature values.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Time Series Forecasting:

Develop a predictive model to estimate future values based on historical trends. How might different modeling approaches impact the prediction accuracy?

  • Understand the schema and field descriptions.
  • Visualize the target feature over time at a reasonable granularity.
  • Always perform a chronological split on the data to create training, validation, and test sets.
  • Are there seasonal trends?
  • Test for stationarity.
  • Discuss possible modeling approaches. How might different modeling approaches impact the prediction accuracy?
  • Train two time series forecasting models to predict the target feature. Use previous seasonality and stationarity information as model hyperparameters.
  • Predict the target feature for the training and validation sets.
  • Optionally, hypertune models with the validation set.
  • Visualize the actual and predicted target feature vs time for each model on the training and validation sets.
  • Evaluate the validation performance with error metrics.
  • Select a model.
  • Retrain the selected model on the test and validation sets.
  • Predict the test values with the selected model.
  • Visualize the average target feature and the predicted test values.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Exploratory Data Analysis / Anomaly Detection:

Identify and describe any outliers, unusual patterns, or significant trends observed in the data. Provide visualizations to support your findings.

  • Understand the schema and field descriptions.
  • Visualize the target feature distribution in a way that shows outliers.
  • Identify and describe any outliers in the target feature.
  • Visualize relationships between the target feature and other features.
  • Identify and describe unusual patterns or significant trends.
  • Visualize patterns and trends.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Classification:

Given the data, can we classify by the target feature?

  • Understand the schema and field descriptions.
  • Identify rows that don't make sense. How many are there and what do they contain?
  • Identify rows without a target value. How many are there and what do they contain?
  • Drop rows that don't match the schema or don't have the target value (if it is reasonable to do so).
  • Split data into training, validation, and test sets.
  • Create features to represent when data are missing, if this is meaningful.
  • Handle missing data. Prefer to keep data instead of dropping it when possible.
  • Before applying encoders, check if the dataset already contains pre-encoded features and prefer existing numerical representations.
  • Transform ordinal data with an ordinal encoder.
  • Transform nominal data with a one hot encoder.
  • Standardize numerical features.
  • Train multiple models.
  • If there is evidence of overfitting, regularize and retrain the model.
  • If there is evidence of underfitting, consider adding or engineering features.
  • Evaluate the models.
  • Create confusion matrices.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Regression:

Predict the continuous valued target feature.

  • Understand the schema and field descriptions.
  • Identify rows that don't make sense. How many are there and what do they contain?
  • Identify rows without a target value. How many are there and what do they contain?
  • Develop an understanding of the data and determine how to handle missing values. This should make sense in the business context.
  • Identify any potential sources of group leakage. Aggregate where appropriate to prevent this.
  • Visualize target feature.
  • Split data into training, validation, and test sets.
  • Handle missing data. Prefer to keep data instead of dropping it when possible.
  • Before applying encoders, check if the dataset already contains pre-encoded features and prefer existing numerical representations.
  • Transform ordinal data with an ordinal encoder.
  • Transform nominal data with a one hot encoder. Restrict high cardinality categorical features to a tractable size.
  • Standardize numerical features.
  • Train multiple models.
  • Visualize the actual vs predicted values on training and validation data.
  • If there is evidence of overfitting, regularize and retrain the model.
  • If there is evidence of underfitting, consider adding or engineering features.
  • Evaluate the model error.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Comparing ML Models:

Evaluate and compare multiple models to determine which is most suitable for production based on predictive power, robustness, and viability.

  • Understand the schema and align metrics with business goals (e.g., cost of false positives vs. false negatives).
  • Establish baselines: define a naive baseline (majority class/mean) and a simple ML baseline (e.g., Logistic/Linear Regression).
  • Ensure rigorous validation: use identical, fixed data splits for all models and perform $k$-fold cross-validation.
  • If data is temporal, use chronological splits for validation.
  • Select and report metrics beyond accuracy (e.g., F1-Score, PR-AUC, MAE, RMSE) that reflect business impact.
  • Use bootstrapping to calculate 95% confidence intervals for key metrics to determine statistical significance.
  • Perform slice-based error analysis: evaluate model performance across key subpopulations and demographics to identify bias or specific failure modes.
  • Inspect and compare confusion matrices, residual plots, and calibration curves.
  • Evaluate operational trade-offs: consider inference latency, training time, compute cost, and model size.
  • Assess interpretability using tools like SHAP or LIME where transparency is required.
  • Conclusion: Recommend the optimal model for the specific use case, justifying the choice with both performance and production viability.

No match:

  • Understand the schema and field descriptions.
  • Identify rows that don't make sense. How many are there and what do they contain?
  • Identify rows without a target value. How many are there and what do they contain?
  • Drop rows that don't match the schema or don't have the target value (if it is reasonable to do so).
  • Create features to represent when data are missing, if this is meaningful.
  • Handle missing data. Prefer to keep data instead of dropping it when possible.
  • Before applying encoders, check if the dataset already contains pre-encoded features and prefer existing numerical representations.
  • Transform ordinal data with an ordinal encoder.
  • Transform nominal data with a one hot encoder.
  • Standardize numerical features.
  • Conclusion: comprehensively answer the prompt in a final markdown cell.

Essential ML Practices

[!IMPORTANT] ALWAYS follow these ML practices

  • Strict Featurization Ordering: For supervised learning ALWAYS split the dataset into training and test data BEFORE fitting preprocessing pipelines (e.g. scaling, encoding). Fit the pipelines on the training data and test data independently.

  • Handling Missing or NULL Values: ALWAYS check for and handle missing and NULL values. First, analyze their frequency. Then, decide whether to keep them, drop them or impute them with a contextually appropriate value, and explain your reasoning.

來源與署名

來源:gemini-cli-extensions/data-agent-kit-starter-pack位於skills/ml-best-practices提交2df10e2

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 gemini-cli-extensions/data-agent-kit-starter-pack 的技能

Schema Mapping

gemini-cli-extensions

為 ETL、ELT 或資料整合任務規劃來源到目標的結構描述對應,產出文件化的對應宣言。

Data & Analytics215今天更新

Resolving Mcp Region Configs

gemini-cli-extensions

修復區域性 Google Cloud MCP 伺服器設定中未取代的區域佔位符,讓缺少的 MCP 工具得以註冊。

DevOps & Cloud215今天更新

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

待分類215今天更新

Managing Python Dependencies

gemini-cli-extensions

指導代理偵測 Python 專案的相依性管理器並正確安裝套件,而不是使用全域 pip。

Software Development215今天更新

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

待分類215今天更新

Google Cloud Auth Verification

gemini-cli-extensions

Mandatory Step 0 pre-flight execution order and authentication verification for Google Cloud Platform (GCP), Application Default Credentials (ADC), gcloud CLI, Spark, Dataproc, BigQuery, GCS, and notebook runtimes. Use whenever interacting with GCP resources, running Spark/PySpark pipelines, BigQuery queries, GCS paths (gs://), or creating/running notebooks.

待分類215今天更新