All skills
huggingface avatar

/train-sentence-transformers

@0201949 official
by Hugging Facehuggingface/skills11k stars
753

Train or fine-tune sentence-transformers models across `SentenceTransformer` (bi-encoder, dense or static embedding model for retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal), `CrossEncoder` (reranker, pair scoring for two-stage retrieval / pair classification), `SparseEncoder` (SPLADE, sparse embedding model for learned-sparse retrieval), and `MultiVectorEncoder` (ColBERT / late-interaction, per-token embeddings scored with MaxSim). Covers loss selection, hard-negative mining, evaluators, distillation, LoRA, Matryoshka, and Hugging Face Hub publishing. Use for any sentence-transformers training task.

Use this Skill: https://skilld.dev/gh/huggingface/skills/train-sentence-transformers

This session only. Nothing lands on disk.

referencesevaluators_sparse_encoder.md

≈1.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Evaluators (Sparse Encoder)

All sparse-encoder evaluators live in sentence_transformers.sparse_encoder.evaluation. They mirror the bi-encoder versions with a Sparse prefix and default to dot product similarity (cosine on sparse vectors is less meaningful).

Choosing the right evaluator

Task Evaluator
Retrieval (nDCG, MRR, Recall), fast default SparseNanoBEIREvaluator
Retrieval on your own corpus / qrels SparseInformationRetrievalEvaluator
STS / continuous similarity SparseEmbeddingSimilarityEvaluator
Binary classification SparseBinaryClassificationEvaluator
Triplet accuracy SparseTripletEvaluator
Reranking (from retrieval candidates) SparseRerankingEvaluator
MSE vs. teacher (distillation) SparseMSEEvaluator
Translation (cross-lingual alignment) SparseTranslationEvaluator
Hybrid BM25 + sparse retrieval ReciprocalRankFusionEvaluator

Wrap multiple in SequentialEvaluator (from sentence_transformers.base.evaluation):

from sentence_transformers.base.evaluation import SequentialEvaluator
evaluator = SequentialEvaluator([sparse_nano_beir, my_custom_ir])

The default: SparseNanoBEIREvaluator

Small, fast subset of BEIR adapted for sparse retrieval. Typical runtime <1 minute on a mid-range GPU.

from sentence_transformers.sparse_encoder.evaluation import SparseNanoBEIREvaluator

evaluator = SparseNanoBEIREvaluator(
    dataset_names=["msmarco", "nfcorpus", "nq"],   # default: all 13 NanoBEIR datasets
    batch_size=32,
    show_progress_bar=False,
)

Output key for metric_for_best_model: eval_NanoBEIR_mean_dot_ndcg@10 (sparse defaults to dot product).

Sparsity tracking

Unlike the dense variant, the sparse evaluator also reports active dimension counts so you can monitor sparsity during training:

  • query_active_dims: non-zero entries per query vector
  • document_active_dims: non-zero entries per document vector

A healthy SPLADE checkpoint typically shows ~30-50 active dims for queries and ~150-250 for documents. If these drift toward the vocab size (~30k), the FLOPS regularization isn't doing its job. Raise query_regularizer_weight / document_regularizer_weight in SpladeLoss.

Retrieval on your own corpus

SparseInformationRetrievalEvaluator

Same shape as the dense version but operates on sparse vectors internally:

from sentence_transformers.sparse_encoder.evaluation import SparseInformationRetrievalEvaluator

evaluator = SparseInformationRetrievalEvaluator(
    queries={qid: text for qid, text in ...},
    corpus={doc_id: text for doc_id, text in ...},
    relevant_docs={qid: {doc_id, ...} for qid in ...},
    name="my-sparse-ir",
    ndcg_at_k=[10],
    mrr_at_k=[10],
    accuracy_at_k=[1, 5, 10],
    map_at_k=[100],
    batch_size=32,
)

Output keys: eval_{name}_dot_ndcg@10, eval_{name}_dot_mrr@10, etc. Also reports active-dims.

Heavy for large corpora. Use SparseNanoBEIREvaluator during training. Reserve full IR for post-training.

Hybrid retrieval

ReciprocalRankFusionEvaluator

Measures the performance of combining your sparse encoder with BM25 (or any other retriever) via reciprocal-rank fusion. Useful when shipping a hybrid system is the actual deployment target.

Other sparse evaluators

SparseEmbeddingSimilarityEvaluator

STS-style. Computes Pearson/Spearman between sparse vector similarities and gold labels. Uses dot product by default.

SparseBinaryClassificationEvaluator

For labeled pair classification with sparse embeddings.

SparseTripletEvaluator

For (anchor, positive, negative) triplets. Reports fraction where the positive is closer than the negative (by dot product).

SparseRerankingEvaluator

For custom re-ranking with sparse embeddings. Same semantics as the dense RerankingEvaluator.

SparseMSEEvaluator

For distillation setups. Compares sparse student embeddings against teacher outputs.

SparseTranslationEvaluator

For cross-lingual / make_multilingual-style alignment checking with sparse embeddings.

Writing metric_for_best_model

Pattern: f"eval_{evaluator.primary_metric}". Inspect after construction: print(evaluator.primary_metric). Common values:

  • eval_NanoBEIR_mean_dot_ndcg@10: SparseNanoBEIREvaluator default
  • eval_{name}_dot_ndcg@10: SparseInformationRetrievalEvaluator
  • eval_{name}_spearman_dot: SparseEmbeddingSimilarityEvaluator

Gotchas

  • Always run evaluator(model) once before training. This confirms the pipeline works (a fill-mask base scores ~0 on retrieval until trained).
  • Sparse evaluators default to dot product. Cosine on sparse vectors isn't meaningful.
  • Don't compare dense and sparse metrics directly: different scales (cosine ∈ [-1, 1] vs. dot ∈ [0, ∞)).
  • Always check query_active_dims / document_active_dims: thousands of active dims per doc means the FLOPS regularizer is mistuned, even if nDCG looks OK.

Source: SKILL.md on GitHub

1 warning16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides a comprehensive environment for training sentence-transformers models, including production-ready scripts and detailed documentation. It involves standard practices such as downloading packages from official registries and processing datasets from remote sources.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

Signed by skilld at 0201949. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated last month
  • sentence-transformers
  • training
  • fine-tuning
  • embeddings
  • retrieval
  • cross-encoder
  • sparse-encoder
  • ner
  • classification

README badge

README badge for huggingface/skills/train-sentence-transformers

Trains or fine-tunes sentence-transformers models across SentenceTransformer (bi-encoder for dense embeddings), CrossEncoder (reranker for pair scoring), and SparseEncoder (SPLADE for sparse vectors). Covers loss selection, hard-negative mining, evaluators, distillation, LoRA, and Hub publishing. Use this skill for any sentence-transformers training task.

Generated from the current SKILL.md.

Does this skill cover all three model types (SentenceTransformer, CrossEncoder, SparseEncoder)?
Yes. The skill routes you to type-specific references and production templates. Use section 1 to identify which model type matches your task, then load the corresponding references and example script.
Can I use this skill to fine-tune models with LoRA, distillation, or Matryoshka?
Yes. The skill includes variant scripts for `train_sentence_transformer_with_lora_example.py`, `train_sentence_transformer_distillation_example.py`, and `train_sentence_transformer_matryoshka_example.py`, plus distillation variants for CrossEncoder and SparseEncoder.
Do I need to write my own training script or can I copy from the templates?
Copy from the production templates (`scripts/train_<type>_example.py`). The skill explicitly states not to synthesize from the routing file alone; templates contain load-bearing scaffolding (autocast helpers, seed handling, version-compatible imports, required callbacks) that prior runs have missed when rolling their own.
What if my task involves hard-negative mining or training on multiple datasets?
The skill includes `scripts/mine_hard_negatives.py` for hard-negative mining and a `train_sentence_transformer_multi_dataset_example.py` variant. Check section 2 (Variant scripts) and `references/dataset_formats.md` for reshaping recipes.
Does this work with multimodal models or non-English languages?
For multimodal: install `sentence-transformers[train,image]` or add audio/video extras. For non-English: the skill references `references/base_model_selection.md` (non-English shortcuts) and `references/prompts_and_instructions.md` for prompt-tuned bases (E5, BGE, Qwen3-Embedding, etc.).

Generated from the current SKILL.md. These answers refresh after source changes.