All skills
huggingface avatar

/huggingface-trackio

@88e864e official
by Hugging Facehuggingface/skills11k stars
753

Track and visualize ML training experiments with Trackio. Use when logging metrics during training (Python API), firing alerts for training diagnostics, or retrieving/analyzing logged metrics (CLI). Supports real-time dashboard visualization, alerts with webhooks, HF Space syncing, and JSON output for automation.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-trackio

This session only. Nothing lands on disk.

referencesalerts.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Trackio Alerts

Alerts let you flag important training events directly from code. They are the primary mechanism for LLM agents to diagnose runs and iterate autonomously on ML experiments.

Alerts are printed to the terminal, stored in the database, displayed in the dashboard, and optionally sent to webhooks (Slack/Discord).

Core API

trackio.alert()

trackio.alert(
    title="Loss divergence",                    # Short title (required)
    text="Loss 5.2 still high after 200 steps", # Detailed description (optional)
    level=trackio.AlertLevel.WARN,               # INFO, WARN, or ERROR (default: WARN)
    webhook_url="https://hooks.slack.com/...",   # Per-alert webhook override (optional)
)

Alert Levels

Level Usage
trackio.AlertLevel.INFO Informational milestones (checkpoints saved, eval completed)
trackio.AlertLevel.WARN Potential issues (loss plateau, low accuracy, high gradient norm)
trackio.AlertLevel.ERROR Critical failures (NaN loss, divergence, OOM)

Webhook Support

Set a global webhook URL via trackio.init() or the TRACKIO_WEBHOOK_URL environment variable. Alerts are auto-formatted for Slack and Discord URLs.

trackio.init(
    project="my-project",
    webhook_url="https://hooks.slack.com/services/...",
    webhook_min_level=trackio.AlertLevel.WARN,  # Only send WARN+ to webhook
)

Per-alert override:

trackio.alert(
    title="Critical failure",
    level=trackio.AlertLevel.ERROR,
    webhook_url="https://hooks.slack.com/services/...",  # Overrides global URL
)

Environment variables:

  • TRACKIO_WEBHOOK_URL — global webhook URL
  • TRACKIO_WEBHOOK_MIN_LEVEL — minimum level for webhook delivery (info, warn, error)

Retrieving Alerts (CLI)

# List all alerts for a project
trackio list alerts --project my-project --json

# Filter by run or level
trackio list alerts --project my-project --run my-run --level error --json

# Poll for new alerts since a timestamp (efficient for agents)
trackio list alerts --project my-project --json --since "2025-06-01T12:00:00"

JSON Output Structure

{
  "project": "my-project",
  "run": null,
  "level": null,
  "since": "2025-06-01T12:00:00",
  "alerts": [
    {
      "run": "run-name",
      "title": "Loss divergence",
      "text": "Loss 5.2 still high after 200 steps",
      "level": "warn",
      "step": 200,
      "timestamp": "2025-06-01T12:05:30"
    }
  ]
}

Autonomous Agent Workflow

The recommended pattern for an LLM agent running ML experiments:

1. Insert Alerts Into Training Code

Add diagnostic trackio.alert() calls for conditions the agent should react to:

import trackio

trackio.init(project="hyperparam-sweep", config={"lr": lr, "batch_size": bs})

for step in range(num_steps):
    loss = train_step()
    trackio.log({"loss": loss, "step": step})

    if step > 200 and loss > 5.0:
        trackio.alert(
            title="Loss divergence",
            text=f"Loss {loss:.4f} still above 5.0 after {step} steps — learning rate may be too high",
            level=trackio.AlertLevel.ERROR,
        )

    if step > 500 and loss_delta < 0.001:
        trackio.alert(
            title="Training stall",
            text=f"Loss barely changed over last 100 steps (delta={loss_delta:.6f})",
            level=trackio.AlertLevel.WARN,
        )

    if math.isnan(loss):
        trackio.alert(
            title="NaN loss",
            text="Loss became NaN — training is broken",
            level=trackio.AlertLevel.ERROR,
        )
        break

trackio.finish()

2. Monitor Alerts

Alerts are automatically printed to the terminal when fired. If the agent is watching the training script's output (e.g. running in the foreground or tailing logs), it will see alerts immediately — no polling needed.

For background or detached runs, poll for alerts via CLI:

# Poll for alerts (run periodically)
trackio list alerts --project hyperparam-sweep --json --since "2025-06-01T00:00:00"

3. Inspect Metrics Around the Alert

When an alert fires, use trackio get snapshot to see all metrics at that point:

# Alert fired at step 200 — get all metrics in a ±5 step window
trackio get snapshot --project hyperparam-sweep --run run-1 --around 200 --window 5 --json

# Or inspect a single metric around the alert's timestamp
trackio get metric --project hyperparam-sweep --run run-1 --metric loss --around 200 --window 10 --json

4. React and Iterate

Based on alerts:

  • ERROR alerts → stop the run, adjust hyperparameters, relaunch
  • WARN alerts → inspect metrics with trackio get snapshot ..., decide whether to intervene
  • INFO alerts → note progress, continue monitoring

5. Compare Across Runs

# Check metrics from previous runs
trackio get run --project hyperparam-sweep --run run-1 --json
trackio get metric --project hyperparam-sweep --run run-1 --metric loss --json

# Launch new run with adjusted config
python train.py --lr 5e-5

Using Alerts with Transformers / TRL

When using report_to="trackio", you don't control the training loop directly. Use a TrainerCallback to fire alerts:

from transformers import TrainerCallback

class AlertCallback(TrainerCallback):
    def on_log(self, args, state, control, logs=None, **kwargs):
        if "trackio" not in args.report_to:
            return
        if logs and "loss" in logs:
            if logs["loss"] > 5.0 and state.global_step > 100:
                trackio.alert(
                    title="High loss",
                    text=f"Loss {logs['loss']:.4f} at step {state.global_step}",
                    level=trackio.AlertLevel.ERROR,
                )

trainer = SFTTrainer(
    model=model,
    args=SFTConfig(output_dir="./out", report_to="trackio"),
    callbacks=[AlertCallback()],
    ...
)

Source: SKILL.md on GitHub

1 warning16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill documentation introduces Trackio, an experiment tracking library tailored for machine learning workflows. It outlines practices for logging metrics, monitoring alerts, and managing experimental configurations. A few architectural design details, such as default public visibility for new remote dashboards and flexible webhook parameter configurations, are worth reviewing to ensure they align with the user's data sensitivity preferences.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 88e864e. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 3 months ago
  • Python
  • CLI
  • trackio
  • experiment-tracking
  • ml
  • metrics
  • alerts
  • hugging-face
  • training
  • monitoring

README badge

README badge for huggingface/skills/huggingface-trackio

Logs and visualizes ML training metrics via Python API and CLI, with real-time dashboard syncing to Hugging Face Spaces. Includes alerts for training diagnostics, webhook notifications, and JSON output for programmatic queries during autonomous experiment iteration.

Generated from the current SKILL.md.

Does Trackio work with TRL (Transformers Reinforcement Learning)?
Yes. You can pass `report_to="trackio"` directly to TRL's training configuration, and Trackio will automatically log metrics without explicit `trackio.log()` calls.
Can I sync metrics to a Hugging Face Space for remote/cloud training?
Yes. Pass `space_id` to `trackio.init()` and metrics will sync to a Space dashboard in real-time, persisting even after the training instance terminates.
What alert severity levels does Trackio support?
Three levels: INFO, WARN, and ERROR. Alerts are printed to terminal, stored in the database, shown in the dashboard, and can be sent to webhooks (Slack/Discord).
Can I retrieve metrics and alerts programmatically for automation?
Yes. Use the `trackio` CLI with `--json` flags (e.g., `trackio list alerts --project <name> --json` or `trackio get metric ... --json`) to get structured output suitable for agents and scripts.
How do I poll for new alerts from a training script running in the background?
Use `trackio list alerts --project <name> --json --since <timestamp>` to retrieve only alerts after a given timestamp, enabling autonomous iteration workflows.

Generated from the current SKILL.md. These answers refresh after source changes.