All skills
huggingface avatar

/huggingface-llm-trainer

@d0d3f43 official
by Hugging Facehuggingface/skills11k stars
753

Train or fine-tune language and vision models using TRL (Transformer Reinforcement Learning) or Unsloth with Hugging Face Jobs infrastructure. Covers SFT, DPO, GRPO and reward modeling training methods, plus GGUF conversion for local deployment. Includes guidance on the TRL Jobs package, UV scripts with PEP 723 format, dataset preparation and validation, hardware selection, cost estimation, Trackio monitoring, Hub authentication, model selection/leaderboards and model persistence. Use for tasks involving cloud GPU training, GGUF conversion, or when users mention training on Hugging Face Jobs without local GPU setup.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-llm-trainer

This session only. Nothing lands on disk.

referencestrackio_guide.md

≈1.6k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Trackio Integration for TRL Training

Trackio is an experiment tracking library that provides real-time metrics visualization for remote training on Hugging Face Jobs infrastructure.

⚠️ IMPORTANT: For Jobs training (remote cloud GPUs):

  • Training happens on ephemeral cloud runners (not your local machine)
  • Trackio syncs metrics to a Hugging Face Space for real-time monitoring
  • Without a Space, metrics are lost when the job completes
  • The Space dashboard persists your training metrics permanently
  • Auto-created Spaces are public by default — pass private=True to trackio.init() if the metrics should not be public

Setting Up Trackio for Jobs

Step 1: Add trackio dependency

# /// script
# dependencies = [
#     "trl>=0.12.0",
#     "trackio",  # Required!
# ]
# ///

Step 2: Create a Trackio Space (one-time setup)

Option A: Let Trackio auto-create (Recommended) Pass a space_id to trackio.init() and Trackio will automatically create the Space if it doesn't exist.

Option B: Create manually

  • Create Space via Hub UI at https://huggingface.co/new-space
  • Select Gradio SDK
  • OR use command: hf repos create my-trackio-dashboard --type space --space-sdk gradio

Step 3: Initialize Trackio with space_id

import trackio

trackio.init(
    project="my-training",
    space_id="username/trackio",  # CRITICAL for Jobs! Replace 'username' with your HF username
    private=True,                 # Spaces are PUBLIC by default; omit for a shareable dashboard
    config={
        "model": "Qwen/Qwen2.5-0.5B",
        "dataset": "trl-lib/Capybara",
        "learning_rate": 2e-5,
    }
)

Step 4: Configure TRL to use Trackio

SFTConfig(
    report_to="trackio",
    # ... other config
)

Step 5: Finish tracking

trainer.train()
trackio.finish()  # Ensures final metrics are synced

What Trackio Tracks

Trackio automatically logs:

  • ✅ Training loss
  • ✅ Learning rate
  • ✅ GPU utilization and memory (when an NVIDIA GPU is detected and nvidia-ml-py is installed — true for standard Jobs GPU flavors)
  • ✅ Training throughput
  • ✅ Custom metrics

How It Works with Jobs

  1. Training runs → Metrics stream to the running Space in small batches (sub-second)
  2. Space unreachable or still building → Metrics spill to an HF Bucket (auto-derived from space_id, or pinned with bucket_id=) every ~30 seconds
  3. Space dashboard → Stores its SQLite DB in the Bucket and ingests spilled metrics every ~15 seconds
  4. Job completes → trackio.finish() drains any pending metrics so everything is persisted

Default Configuration Pattern

Use sensible defaults for trackio configuration unless user requests otherwise.

Recommended Defaults

import trackio

trackio.init(
    project="qwen-capybara-sft",
    name="baseline-run",             # Descriptive name user will recognize
    space_id="username/trackio",     # Default space: {username}/trackio
    private=True,                    # Spaces are PUBLIC by default; omit for a shareable dashboard
    config={
        # Keep config minimal - hyperparameters and model/dataset info only
        "model": "Qwen/Qwen2.5-0.5B",
        "dataset": "trl-lib/Capybara",
        "learning_rate": 2e-5,
        "num_epochs": 3,
    }
)

Key principles:

  • Space ID: Use {username}/trackio with "trackio" as default space name
  • Run naming: Unless otherwise specified, name the run in a way the user will recognize
  • Config: Keep minimal - don't automatically capture job metadata unless requested
  • Grouping: Optional - only use if user requests organizing related experiments

Grouping Runs (Optional)

The group parameter helps organize related runs together in the dashboard sidebar. This is useful when user is running multiple experiments with different configurations but wants to compare them together:

# Example: Group runs by experiment type
trackio.init(project="my-project", run_name="baseline-run-1", group="baseline")
trackio.init(project="my-project", run_name="augmented-run-1", group="augmented")
trackio.init(project="my-project", run_name="tuned-run-1", group="tuned")

Runs with the same group name can be grouped together in the sidebar, making it easier to compare related experiments. You can group by any configuration parameter:

# Hyperparameter sweep - group by learning rate
trackio.init(project="hyperparam-sweep", run_name="lr-0.001-run", group="lr_0.001")
trackio.init(project="hyperparam-sweep", run_name="lr-0.01-run", group="lr_0.01")

Environment Variables for Jobs

You can configure trackio using environment variables instead of passing parameters to trackio.init(). This is useful for managing configuration across multiple jobs.

HF_TOKEN Required for creating Spaces and writing metrics to the HF Bucket (passed via secrets):

hf_jobs("uv", {
    "script": "...",
    "secrets": {
        "HF_TOKEN": "$HF_TOKEN"  # Enables Space creation and Hub push
    }
})

Example with Environment Variables

hf_jobs("uv", {
    "script": """
# Training script - trackio config from environment
import trackio
from datetime import datetime

# Auto-generate run name
timestamp = datetime.now().strftime("%Y-%m-%d_%H-%M")
run_name = f"sft_qwen25_{timestamp}"

# Project and space_id can come from environment variables
trackio.init(run_name=run_name, group="SFT")

# ... training code ...
trackio.finish()
""",
    "flavor": "a10g-large",
    "timeout": "2h",
    "secrets": {"HF_TOKEN": "$HF_TOKEN"}
})

When to use environment variables:

  • Managing multiple jobs with same configuration
  • Keeping training scripts portable across projects
  • Separating configuration from code

When to use direct parameters:

  • Single job with specific configuration
  • When clarity in code is preferred
  • When each job has different project/space

Viewing the Dashboard

After starting training:

  1. Navigate to the Space: https://huggingface.co/spaces/username/trackio
  2. The Gradio dashboard shows all tracked experiments
  3. Filter by project, compare runs, view charts with smoothing

Recommendation

  • Trackio: Best for real-time monitoring during long training runs
  • Weights & Biases: Best for team collaboration, requires account

Source: SKILL.md on GitHub

1 warning16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill includes some security considerations such as external downloads from trusted sources and command execution for model conversion tasks. While these warrant review, they are used within the skill's intended functionality for training and deploying machine learning models. See detailed analysis for context.

  • Socket16d

    1 alert: gptAnomaly

  • Snyk16d

    Risk: LOW · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at d0d3f43. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 6 months ago
  • trl
  • hugging-face
  • fine-tuning
  • llm
  • gpu
  • rlhf
  • lora
  • gguf
  • training

README badge

README badge for huggingface/skills/huggingface-llm-trainer

Trains language models using TRL (Supervised Fine-Tuning, Direct Preference Optimization, Group Relative Policy Optimization, or Reward Modeling) on Hugging Face Jobs infrastructure without local GPU setup. Includes GGUF conversion for local deployment, Trackio monitoring integration, and guidance on dataset preparation, hardware selection, and cost estimation. Targets developers who need cloud GPU training with automatic Hub persistence.

Generated from the current SKILL.md.

What training methods does this skill support?
SFT (Supervised Fine-Tuning), DPO (Direct Preference Optimization), GRPO (Group Relative Policy Optimization), and reward modeling. TRL documentation links are provided for each method.
Do I need a local GPU to use this skill?
No. Training runs on Hugging Face Jobs infrastructure with cloud GPUs, so no local GPU setup is required.
What plan do I need to use Hugging Face Jobs?
A Pro, Team, or Enterprise Hugging Face plan is required. Free accounts cannot access Jobs infrastructure.
Will my trained model be saved if training completes?
Only if you set `push_to_hub=True` in the training config and pass `HF_TOKEN` in the job secrets. The training environment is ephemeral; without Hub push, all results are lost.
Does this skill support Unsloth for faster training?
Yes. Unsloth is recommended when GPU memory is limited, speed matters, or you're training models larger than 13B. See the `references/unsloth.md` documentation for details.

Generated from the current SKILL.md. These answers refresh after source changes.