All skills
huggingface avatar

/train-sentence-transformers

@0201949 official
by Hugging Facehuggingface/skills11k stars
753

Train or fine-tune sentence-transformers models across `SentenceTransformer` (bi-encoder, dense or static embedding model for retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal), `CrossEncoder` (reranker, pair scoring for two-stage retrieval / pair classification), `SparseEncoder` (SPLADE, sparse embedding model for learned-sparse retrieval), and `MultiVectorEncoder` (ColBERT / late-interaction, per-token embeddings scored with MaxSim). Covers loss selection, hard-negative mining, evaluators, distillation, LoRA, Matryoshka, and Hugging Face Hub publishing. Use for any sentence-transformers training task.

Use this Skill: https://skilld.dev/gh/huggingface/skills/train-sentence-transformers

This session only. Nothing lands on disk.

referenceshf_jobs_execution.md

≈1.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Hugging Face Jobs Execution

Run training on Hugging Face's managed GPUs without provisioning any local infrastructure. The same training script runs locally and on Jobs. This reference covers only the Jobs-specific concerns.

Prerequisites

  • Hugging Face account with a Pro, Team, or Enterprise plan. Jobs are paid.
  • HF_TOKEN with write permission. Log in once locally with hf auth login (the modern command from the hf CLI, replacing the deprecated huggingface-cli login).
  • Access to the hf_jobs() MCP tool, or the hf CLI (curl -LsSf https://hf.co/cli/install.sh | bash -s).

The three submission paths

1. Inline script via MCP (recommended in Claude Code)

Pass the full training script as script. Dependencies come from the PEP 723 header.

hf_jobs("uv", {
    "script": """
# /// script
# requires-python = ">=3.10"
# dependencies = ["sentence-transformers[train]>=5.0", "trackio"]
# ///

# <full training script content>
""",
    "flavor": "a10g-large",
    "timeout": "3h",
    "secrets": {"HF_TOKEN": "$HF_TOKEN"},
})

2. Script-from-URL via MCP

Upload the script to the Hub (as a model or dataset repo file) or a Gist, then reference by URL:

hf_jobs("uv", {
    "script": "https://huggingface.co/USERNAME/scripts/resolve/main/train_bi_encoder.py",
    "flavor": "a10g-large",
    "timeout": "3h",
    "secrets": {"HF_TOKEN": "$HF_TOKEN"},
})

Local file paths (./train.py, /path/to/train.py) do not work. Jobs run in isolated containers without access to your filesystem.

3. CLI

hf jobs uv run \
    --flavor a10g-large \
    --timeout 3h \
    --secrets HF_TOKEN \
    "https://huggingface.co/USERNAME/scripts/resolve/main/train.py"

Syntax gotchas:

  • Command order is hf jobs uv run, not hf jobs run uv.
  • Flags (--flavor, --timeout, --secrets) go before the script URL.
  • --secrets (plural), not --secret.

Required script modifications for Jobs

Add these to your TrainingArguments:

args = SentenceTransformerTrainingArguments(
    ...,
    push_to_hub=True,
    hub_model_id="your-username/my-model",
    hub_strategy="every_save",        # push each checkpoint; timeout-safe
    save_strategy="steps",
    save_steps=0.1,                   # 10 saves/pushes per epoch; scales with dataset size
)

Why each matters:

Argument Why
push_to_hub=True The Jobs container is destroyed after the job finishes. Without Hub push, all weights are lost.
hub_model_id Required to identify the destination repo.
hub_strategy="every_save" Default, but worth being deliberate about on Jobs: each checkpoint is pushed as it's written, so a timeout leaves all completed checkpoints on the Hub. "end" only pushes once trainer.train() returns, so a timeout loses everything.
save_strategy="steps" + save_steps=0.1 Checkpoints must actually be saved for hub_strategy="every_save" to push them. Fractional 0.1 = save every 10% of training, auto-scales with dataset size.

Secrets

Secrets are environment variables injected into the Jobs container. They never appear in logs and are not part of the script.

Secret Required when
HF_TOKEN Always, for Hub push. Also covers Trackio auth.
WANDB_API_KEY Using report_to="wandb".
MLFLOW_TRACKING_URI, MLFLOW_TRACKING_TOKEN Using MLflow with a remote server.

The $HF_TOKEN syntax in the job config references the value from your local environment at submission time. The literal string $HF_TOKEN is replaced with your token's value. Never hardcode tokens in the script itself.

Trackio (the default tracker in this skill) uses HF_TOKEN for auth, so no extra secrets are needed. Only switch to the W&B / MLflow rows above if you're using those trackers.

Timeout

Default is 30 minutes, which is too short for almost any real training. Set explicitly:

"timeout": "2h"       # 2 hours
"timeout": "90m"      # 90 minutes
"timeout": "1.5h"     # 90 minutes
"timeout": 7200       # seconds, as integer

Rule: estimated training time × 1.3. The extra buffer covers model loading, dataset caching, checkpoint saving, and Hub push.

On timeout, the container is killed immediately. Only data on the Hub (hub_strategy="every_save" saves you here) or in persistent volumes survives.

Dataset caching

Hugging Face datasets are cached at ~/.cache/huggingface/datasets by default. That's inside the container, which is destroyed after the job. Each Jobs run re-downloads the dataset.

For large datasets (>5 GB), this matters. Options:

  • Persistent /data volume (Jobs feature, check current documentation): set HF_DATASETS_CACHE=/data/datasets so caches persist across jobs.
  • Pre-cache locally, push to Hub: if the dataset is on Hub already, nothing to do. If it's local-only, dataset.push_to_hub(...) once so subsequent jobs load from Hub.

Monitoring a running job

hf jobs ps [--all]                        # running (or all) jobs
hf jobs inspect <job-id>                  # full config + status
hf jobs logs <job-id> [--follow|--tail N] # tail or stream
hf jobs cancel <job-id>
hf jobs hardware                          # list flavors + hourly rates

hf jobs logs <id> --follow under Bash run_in_background pairs nicely with a Monitor watching for the VERDICT: line emitted by your training script's verdict block.

MCP equivalents (signatures may vary by server version, so check the actual tool listing): hf_jobs("ps"), hf_jobs("logs", {"job_id": ...}), hf_jobs("cancel", {"job_id": ...}).

For recurring runs, hf jobs scheduled uv run "<cron>" <script> ... schedules. hf jobs scheduled ps/suspend/delete manages.

Common failures

"Model not found on Hub" after a successful-looking run

The run succeeded but push_to_hub was not enabled. The container is gone. The weights are gone.

Fix: always set push_to_hub=True + hub_model_id=... + secrets={"HF_TOKEN": "$HF_TOKEN"}.

Tracker not connecting

  • Trackio: HF_TOKEN missing or lacks write permission. Add "secrets": {"HF_TOKEN": "$HF_TOKEN"} and make sure the token has write access.
  • W&B: WANDB_API_KEY missing. Add "secrets": {"HF_TOKEN": "$HF_TOKEN", "WANDB_API_KEY": "$WANDB_API_KEY"}.

OOM on first step

Flavor too small. Move up one tier (see hardware_guide.md).

Training starts but eval hangs forever

eval_strategy="steps" with no eval_dataset. Always provide an eval dataset, or set eval_strategy="no".

Dataset download times out

Large dataset or slow cold-cache. Increase timeout or pre-cache to a persistent volume.

CachedMultipleNegativesRankingLoss + gradient_checkpointing=True crash

The cached losses are incompatible with gradient checkpointing. Disable gradient_checkpointing.

After submission, the MCP returns a job ID. Monitor with hf_jobs("logs", {"job_id": ...}) when you want an update. Don't poll in a tight loop. End-to-end submission templates live in scripts/train_sentence_transformer_example.py / scripts/train_cross_encoder_example.py / scripts/train_sparse_encoder_example.py. Wrap the script contents in the inline pattern from §1 above.

Source: SKILL.md on GitHub

1 warning16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides a comprehensive environment for training sentence-transformers models, including production-ready scripts and detailed documentation. It involves standard practices such as downloading packages from official registries and processing datasets from remote sources.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

Signed by skilld at 0201949. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated last month
  • sentence-transformers
  • training
  • fine-tuning
  • embeddings
  • retrieval
  • cross-encoder
  • sparse-encoder
  • ner
  • classification

README badge

README badge for huggingface/skills/train-sentence-transformers

Trains or fine-tunes sentence-transformers models across SentenceTransformer (bi-encoder for dense embeddings), CrossEncoder (reranker for pair scoring), and SparseEncoder (SPLADE for sparse vectors). Covers loss selection, hard-negative mining, evaluators, distillation, LoRA, and Hub publishing. Use this skill for any sentence-transformers training task.

Generated from the current SKILL.md.

Does this skill cover all three model types (SentenceTransformer, CrossEncoder, SparseEncoder)?
Yes. The skill routes you to type-specific references and production templates. Use section 1 to identify which model type matches your task, then load the corresponding references and example script.
Can I use this skill to fine-tune models with LoRA, distillation, or Matryoshka?
Yes. The skill includes variant scripts for `train_sentence_transformer_with_lora_example.py`, `train_sentence_transformer_distillation_example.py`, and `train_sentence_transformer_matryoshka_example.py`, plus distillation variants for CrossEncoder and SparseEncoder.
Do I need to write my own training script or can I copy from the templates?
Copy from the production templates (`scripts/train_<type>_example.py`). The skill explicitly states not to synthesize from the routing file alone; templates contain load-bearing scaffolding (autocast helpers, seed handling, version-compatible imports, required callbacks) that prior runs have missed when rolling their own.
What if my task involves hard-negative mining or training on multiple datasets?
The skill includes `scripts/mine_hard_negatives.py` for hard-negative mining and a `train_sentence_transformer_multi_dataset_example.py` variant. Check section 2 (Variant scripts) and `references/dataset_formats.md` for reshaping recipes.
Does this work with multimodal models or non-English languages?
For multimodal: install `sentence-transformers[train,image]` or add audio/video extras. For non-English: the skill references `references/base_model_selection.md` (non-English shortcuts) and `references/prompts_and_instructions.md` for prompt-tuned bases (E5, BGE, Qwen3-Embedding, etc.).

Generated from the current SKILL.md. These answers refresh after source changes.