All skills
huggingface avatar

/huggingface-trackio

@88e864e official
by Hugging Facehuggingface/skills11k stars
753

Track and visualize ML training experiments with Trackio. Use when logging metrics during training (Python API), firing alerts for training diagnostics, or retrieving/analyzing logged metrics (CLI). Supports real-time dashboard visualization, alerts with webhooks, HF Space syncing, and JSON output for automation.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-trackio

This session only. Nothing lands on disk.

SKILL.md

≈84 tokens always: the name and description. ≈1.2k when used: this file. ≈5.1k more on demand in 3 files.

Trackio - Experiment Tracking for ML Training

Trackio is an experiment tracking library for logging and visualizing ML training metrics. It syncs to Hugging Face Spaces for real-time monitoring dashboards.

Three Interfaces

Task Interface Reference
Logging metrics during training Python API references/logging_metrics.md
Firing alerts for training diagnostics Python API references/alerts.md
Retrieving metrics & alerts after/during training CLI references/retrieving_metrics.md

When to Use Each

Python API → Logging

Use import trackio in your training scripts to log metrics:

  • Initialize tracking with trackio.init()
  • Log metrics with trackio.log() or use TRL's report_to="trackio"
  • Finalize with trackio.finish()

Key concept: For remote/cloud training, pass space_id — metrics sync to a Space dashboard so they persist after the instance terminates. Auto-created Spaces are public by default — pass private=True if the metrics should not be public.

→ See references/logging_metrics.md for setup, TRL integration, and configuration options.

Python API → Alerts

Insert trackio.alert() calls in training code to flag important events — like inserting print statements for debugging, but structured and queryable:

  • trackio.alert(title="...", level=trackio.AlertLevel.WARN) — fire an alert
  • Three severity levels: INFO, WARN, ERROR
  • Alerts are printed to terminal, stored in the database, shown in the dashboard, and optionally sent to webhooks (Slack/Discord)

Key concept for LLM agents: Alerts are the primary mechanism for autonomous experiment iteration. An agent should insert alerts into training code for diagnostic conditions (loss spikes, NaN gradients, low accuracy, training stalls). Since alerts are printed to the terminal, an agent that is watching the training script's output will see them automatically. For background or detached runs, the agent can poll via CLI instead.

→ See references/alerts.md for the full alerts API, webhook setup, and autonomous agent workflows.

CLI → Retrieving

Use the trackio command to query logged metrics and alerts:

  • trackio list projects/runs/metrics — discover what's available
  • trackio get project/run/metric — retrieve summaries and values
  • trackio list alerts --project <name> --json — retrieve alerts
  • trackio show — launch the dashboard
  • trackio sync — sync to HF Space

Key concept: Add --json for programmatic output suitable for automation and LLM agents.

→ See references/retrieving_metrics.md for all commands, workflows, and JSON output formats.

Minimal Logging Setup

import trackio

# Spaces are PUBLIC by default (good for shareable dashboards);
# pass private=True if the metrics should not be public
trackio.init(project="my-project", space_id="username/trackio", private=True)
trackio.log({"loss": 0.1, "accuracy": 0.9})
trackio.log({"loss": 0.09, "accuracy": 0.91})
trackio.finish()

Minimal Retrieval

trackio list projects --json
trackio get metric --project my-project --run my-run --metric loss --json

Autonomous ML Experiment Workflow

When running experiments autonomously as an LLM agent, the recommended workflow is:

  1. Set up training with alerts — insert trackio.alert() calls for diagnostic conditions
  2. Launch training — run the script in the background
  3. Poll for alerts — use trackio list alerts --project <name> --json --since <timestamp> to check for new alerts
  4. Read metrics — use trackio get metric ... to inspect specific values
  5. Iterate — based on alerts and metrics, stop the run, adjust hyperparameters, and launch a new run
import trackio

trackio.init(project="my-project", config={"lr": 1e-4})

for step in range(num_steps):
    loss = train_step()
    trackio.log({"loss": loss, "step": step})

    if step > 100 and loss > 5.0:
        trackio.alert(
            title="Loss divergence",
            text=f"Loss {loss:.4f} still high after {step} steps",
            level=trackio.AlertLevel.ERROR,
        )
    if step > 0 and abs(loss) < 1e-8:
        trackio.alert(
            title="Vanishing loss",
            text="Loss near zero — possible gradient collapse",
            level=trackio.AlertLevel.WARN,
        )

trackio.finish()

Then poll from a separate terminal/process:

trackio list alerts --project my-project --json --since "2025-01-01T00:00:00"

Source: SKILL.md on GitHub

1 warning16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill documentation introduces Trackio, an experiment tracking library tailored for machine learning workflows. It outlines practices for logging metrics, monitoring alerts, and managing experimental configurations. A few architectural design details, such as default public visibility for new remote dashboards and flexible webhook parameter configurations, are worth reviewing to ensure they align with the user's data sensitivity preferences.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 88e864e. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 3 months ago
  • Python
  • CLI
  • trackio
  • experiment-tracking
  • ml
  • metrics
  • alerts
  • hugging-face
  • training
  • monitoring

README badge

README badge for huggingface/skills/huggingface-trackio

Logs and visualizes ML training metrics via Python API and CLI, with real-time dashboard syncing to Hugging Face Spaces. Includes alerts for training diagnostics, webhook notifications, and JSON output for programmatic queries during autonomous experiment iteration.

Generated from the current SKILL.md.

Does Trackio work with TRL (Transformers Reinforcement Learning)?
Yes. You can pass `report_to="trackio"` directly to TRL's training configuration, and Trackio will automatically log metrics without explicit `trackio.log()` calls.
Can I sync metrics to a Hugging Face Space for remote/cloud training?
Yes. Pass `space_id` to `trackio.init()` and metrics will sync to a Space dashboard in real-time, persisting even after the training instance terminates.
What alert severity levels does Trackio support?
Three levels: INFO, WARN, and ERROR. Alerts are printed to terminal, stored in the database, shown in the dashboard, and can be sent to webhooks (Slack/Discord).
Can I retrieve metrics and alerts programmatically for automation?
Yes. Use the `trackio` CLI with `--json` flags (e.g., `trackio list alerts --project <name> --json` or `trackio get metric ... --json`) to get structured output suitable for agents and scripts.
How do I poll for new alerts from a training script running in the background?
Use `trackio list alerts --project <name> --json --since <timestamp>` to retrieve only alerts after a given timestamp, enabling autonomous iteration workflows.

Generated from the current SKILL.md. These answers refresh after source changes.