All skills
huggingface avatar

/huggingface-llm-trainer

@d0d3f43 official
by Hugging Facehuggingface/skills11k stars
753

Train or fine-tune language and vision models using TRL (Transformer Reinforcement Learning) or Unsloth with Hugging Face Jobs infrastructure. Covers SFT, DPO, GRPO and reward modeling training methods, plus GGUF conversion for local deployment. Includes guidance on the TRL Jobs package, UV scripts with PEP 723 format, dataset preparation and validation, hardware selection, cost estimation, Trackio monitoring, Hub authentication, model selection/leaderboards and model persistence. Use for tasks involving cloud GPU training, GGUF conversion, or when users mention training on Hugging Face Jobs without local GPU setup.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-llm-trainer

This session only. Nothing lands on disk.

referencestraining_methods.md

≈1.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

TRL Training Methods Overview

TRL (Transformer Reinforcement Learning) provides multiple training methods for fine-tuning and aligning language models. This reference provides a brief overview of each method.

Supervised Fine-Tuning (SFT)

What it is: Standard instruction tuning with supervised learning on demonstration data.

When to use:

  • Initial fine-tuning of base models on task-specific data
  • Teaching new capabilities or domains
  • Most common starting point for fine-tuning

Dataset format: Conversational format with "messages" field, OR text field, OR prompt/completion pairs

Example:

from trl import SFTTrainer, SFTConfig

trainer = SFTTrainer(
    model="Qwen/Qwen2.5-0.5B",
    train_dataset=dataset,
    args=SFTConfig(
        output_dir="my-model",
        push_to_hub=True,
        hub_model_id="username/my-model",
        eval_strategy="no",  # Disable eval for simple example
        # max_length=1024 is the default - only set if you need different length
    )
)
trainer.train()

Note: For production training with evaluation monitoring, see scripts/train_sft_example.py

Documentation: hf_doc_fetch("https://huggingface.co/docs/trl/sft_trainer")

Direct Preference Optimization (DPO)

What it is: Alignment method that trains directly on preference pairs (chosen vs rejected responses) without requiring a reward model.

When to use:

  • Aligning models to human preferences
  • Improving response quality after SFT
  • Have paired preference data (chosen/rejected responses)

Dataset format: Preference pairs with "chosen" and "rejected" fields

Example:

from trl import DPOTrainer, DPOConfig

trainer = DPOTrainer(
    model="Qwen/Qwen2.5-0.5B-Instruct",  # Use instruct model
    train_dataset=dataset,
    args=DPOConfig(
        output_dir="dpo-model",
        beta=0.1,  # KL penalty coefficient
        eval_strategy="no",  # Disable eval for simple example
        # max_length=1024 is the default - only set if you need different length
    )
)
trainer.train()

Note: For production training with evaluation monitoring, see scripts/train_dpo_example.py

Documentation: hf_doc_fetch("https://huggingface.co/docs/trl/dpo_trainer")

Group Relative Policy Optimization (GRPO)

What it is: Online RL method that optimizes relative to group performance, useful for tasks with verifiable rewards.

When to use:

  • Tasks with automatic reward signals (code execution, math verification)
  • Online learning scenarios
  • When DPO offline data is insufficient

Dataset format: Prompt-only format (model generates responses, reward computed online)

Example:

# Use TRL maintained script
hf_jobs("uv", {
    "script": "https://raw.githubusercontent.com/huggingface/trl/main/examples/scripts/grpo.py",
    "script_args": [
        "--model_name_or_path", "Qwen/Qwen2.5-0.5B-Instruct",
        "--dataset_name", "trl-lib/math_shepherd",
        "--output_dir", "grpo-model"
    ],
    "flavor": "a10g-large",
    "timeout": "4h",
    "secrets": {"HF_TOKEN": "$HF_TOKEN"}
})

Documentation: hf_doc_fetch("https://huggingface.co/docs/trl/grpo_trainer")

Reward Modeling

What it is: Train a reward model to score responses, used as a component in RLHF pipelines.

When to use:

  • Building RLHF pipeline
  • Need automatic quality scoring
  • Creating reward signals for PPO training

Dataset format: Preference pairs with "chosen" and "rejected" responses

Documentation: hf_doc_fetch("https://huggingface.co/docs/trl/reward_trainer")

Method Selection Guide

Method Complexity Data Required Use Case
SFT Low Demonstrations Initial fine-tuning
DPO Medium Paired preferences Post-SFT alignment
GRPO Medium Prompts + reward fn Online RL with automatic rewards
Reward Medium Paired preferences Building RLHF pipeline

Recommended Pipeline

For most use cases:

  1. Start with SFT - Fine-tune base model on task data
  2. Follow with DPO - Align to preferences using paired data
  3. Optional: GGUF conversion - Deploy for local inference

For advanced RL scenarios:

  1. Start with SFT - Fine-tune base model
  2. Train reward model - On preference data

Dataset Format Reference

For complete dataset format specifications, use:

hf_doc_fetch("https://huggingface.co/docs/trl/dataset_formats")

Or validate your dataset:

uv run https://huggingface.co/datasets/mcp-tools/skills/raw/main/dataset_inspector.py \
  --dataset your/dataset --split train

See Also

  • references/training_patterns.md - Common training patterns and examples
  • scripts/train_sft_example.py - Complete SFT template
  • scripts/train_dpo_example.py - Complete DPO template
  • Dataset Inspector - Dataset format validation tool

Source: SKILL.md on GitHub

1 warning16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill includes some security considerations such as external downloads from trusted sources and command execution for model conversion tasks. While these warrant review, they are used within the skill's intended functionality for training and deploying machine learning models. See detailed analysis for context.

  • Socket16d

    1 alert: gptAnomaly

  • Snyk16d

    Risk: LOW · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at d0d3f43. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 6 months ago
  • trl
  • hugging-face
  • fine-tuning
  • llm
  • gpu
  • rlhf
  • lora
  • gguf
  • training

README badge

README badge for huggingface/skills/huggingface-llm-trainer

Trains language models using TRL (Supervised Fine-Tuning, Direct Preference Optimization, Group Relative Policy Optimization, or Reward Modeling) on Hugging Face Jobs infrastructure without local GPU setup. Includes GGUF conversion for local deployment, Trackio monitoring integration, and guidance on dataset preparation, hardware selection, and cost estimation. Targets developers who need cloud GPU training with automatic Hub persistence.

Generated from the current SKILL.md.

What training methods does this skill support?
SFT (Supervised Fine-Tuning), DPO (Direct Preference Optimization), GRPO (Group Relative Policy Optimization), and reward modeling. TRL documentation links are provided for each method.
Do I need a local GPU to use this skill?
No. Training runs on Hugging Face Jobs infrastructure with cloud GPUs, so no local GPU setup is required.
What plan do I need to use Hugging Face Jobs?
A Pro, Team, or Enterprise Hugging Face plan is required. Free accounts cannot access Jobs infrastructure.
Will my trained model be saved if training completes?
Only if you set `push_to_hub=True` in the training config and pass `HF_TOKEN` in the job secrets. The training environment is ephemeral; without Hub push, all results are lost.
Does this skill support Unsloth for faster training?
Yes. Unsloth is recommended when GPU memory is limited, speed matters, or you're training models larger than 13B. See the `references/unsloth.md` documentation for details.

Generated from the current SKILL.md. These answers refresh after source changes.