Huggingface Trackio

作者 huggingfaceabc20ae526d8無授權條款11K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Track and visualize ML training experiments with Trackio. Use when logging metrics during training (Python API), firing alerts for training diagnostics, or retrieving/analyzing logged metrics (CLI). Supports real-time dashboard visualization, alerts with webhooks, HF Space syncing, and JSON output for automation.

僅含說明AI & Agents
AI 產生的概覽

指導使用 Trackio 實驗追蹤函式庫記錄、告警與擷取機器學習訓練指標。

功能
此技能說明如何使用 Trackio 追蹤並視覺化機器學習訓練實驗。內容涵蓋透過 Python API 記錄指標、為訓練診斷觸發結構化告警,以及透過 CLI 擷取指標與告警,並支援 JSON 輸出以便自動化。它也說明如何將結果同步到 Hugging Face Spaces 以產生即時儀表板。
適用情境
適用於在訓練指令碼中加入指標記錄、設定損失飆升或 NaN 梯度等診斷告警,或從命令列查詢已記錄的指標與告警。它也面向自主實驗迴圈,讓代理程式輪詢告警與指標以決定下一次執行。
執行需求
記錄與告警 API 需要 Trackio Python 套件,擷取需要 trackio 命令列工具。同步到 Hugging Face Space 需要 Hugging Face 帳號與 space 識別碼,Webhook 告警需要設定 Slack 或 Discord webhook。Space 同步與 Webhook 需要網路存取。此技能不附指令碼,僅包含說明與參考文件。

Trackio - Experiment Tracking for ML Training

Trackio is an experiment tracking library for logging and visualizing ML training metrics. It syncs to Hugging Face Spaces for real-time monitoring dashboards.

Three Interfaces

TaskInterfaceReference
Logging metrics during trainingPython APIreferences/logging_metrics.md [blocked]
Firing alerts for training diagnosticsPython APIreferences/alerts.md [blocked]
Retrieving metrics & alerts after/during trainingCLIreferences/retrieving_metrics.md [blocked]

When to Use Each

Python API → Logging

Use import trackio in your training scripts to log metrics:

  • Initialize tracking with trackio.init()
  • Log metrics with trackio.log() or use TRL's report_to="trackio"
  • Finalize with trackio.finish()

Key concept: For remote/cloud training, pass space_id — metrics sync to a Space dashboard so they persist after the instance terminates. Auto-created Spaces are public by default — pass private=True if the metrics should not be public.

→ See references/logging_metrics.md [blocked] for setup, TRL integration, and configuration options.

Python API → Alerts

Insert trackio.alert() calls in training code to flag important events — like inserting print statements for debugging, but structured and queryable:

  • trackio.alert(title="...", level=trackio.AlertLevel.WARN) — fire an alert
  • Three severity levels: INFO, WARN, ERROR
  • Alerts are printed to terminal, stored in the database, shown in the dashboard, and optionally sent to webhooks (Slack/Discord)

Key concept for LLM agents: Alerts are the primary mechanism for autonomous experiment iteration. An agent should insert alerts into training code for diagnostic conditions (loss spikes, NaN gradients, low accuracy, training stalls). Since alerts are printed to the terminal, an agent that is watching the training script's output will see them automatically. For background or detached runs, the agent can poll via CLI instead.

→ See references/alerts.md [blocked] for the full alerts API, webhook setup, and autonomous agent workflows.

CLI → Retrieving

Use the trackio command to query logged metrics and alerts:

  • trackio list projects/runs/metrics — discover what's available
  • trackio get project/run/metric — retrieve summaries and values
  • trackio list alerts --project <name> --json — retrieve alerts
  • trackio show — launch the dashboard
  • trackio sync — sync to HF Space

Key concept: Add --json for programmatic output suitable for automation and LLM agents.

→ See references/retrieving_metrics.md [blocked] for all commands, workflows, and JSON output formats.

Minimal Logging Setup

python
import trackio
# Spaces are PUBLIC by default (good for shareable dashboards);# pass private=True if the metrics should not be publictrackio.init(project="my-project", space_id="username/trackio", private=True)trackio.log({"loss": 0.1, "accuracy": 0.9})trackio.log({"loss": 0.09, "accuracy": 0.91})trackio.finish()

Minimal Retrieval

bash
trackio list projects --jsontrackio get metric --project my-project --run my-run --metric loss --json

Autonomous ML Experiment Workflow

When running experiments autonomously as an LLM agent, the recommended workflow is:

  1. Set up training with alerts — insert trackio.alert() calls for diagnostic conditions
  2. Launch training — run the script in the background
  3. Poll for alerts — use trackio list alerts --project <name> --json --since <timestamp> to check for new alerts
  4. Read metrics — use trackio get metric ... to inspect specific values
  5. Iterate — based on alerts and metrics, stop the run, adjust hyperparameters, and launch a new run
python
import trackio
trackio.init(project="my-project", config={"lr": 1e-4})
for step in range(num_steps):    loss = train_step()    trackio.log({"loss": loss, "step": step})
    if step > 100 and loss > 5.0:        trackio.alert(            title="Loss divergence",            text=f"Loss {loss:.4f} still high after {step} steps",            level=trackio.AlertLevel.ERROR,        )    if step > 0 and abs(loss) < 1e-8:        trackio.alert(            title="Vanishing loss",            text="Loss near zero — possible gradient collapse",            level=trackio.AlertLevel.WARN,        )
trackio.finish()

Then poll from a separate terminal/process:

bash
trackio list alerts --project my-project --json --since "2025-01-01T00:00:00"

來源與署名

來源:huggingface/skills位於skills/huggingface-trackio提交abc20ae

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架