All skills
huggingface avatar

/huggingface-vision-trainer

@386571e official
by Hugging Facehuggingface/skills11k stars
753

Trains and fine-tunes vision models for object detection (D-FINE, RT-DETR v2, DETR, YOLOS), image classification (timm models — MobileNetV3, MobileViT, ResNet, ViT/DINOv3 — plus any Transformers classifier), and SAM/SAM2 segmentation using Hugging Face Transformers on Hugging Face Jobs cloud GPUs. Covers COCO-format dataset preparation, Albumentations augmentation, mAP/mAR evaluation, accuracy metrics, SAM segmentation with bbox/point prompts, DiceCE loss, hardware selection, cost estimation, Trackio monitoring, and Hub persistence. Use when users mention training object detection, image classification, SAM, SAM2, segmentation, image matting, DETR, D-FINE, RT-DETR, ViT, timm, MobileNet, ResNet, bounding box models, or fine-tuning vision models on Hugging Face Jobs.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-vision-trainer

This session only. Nothing lands on disk.

referencestimm_trainer.md

≈887 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Using timm models with Hugging Face Trainer

Transformers has first-class support for timm models via the TimmWrapper classes. You can load any timm model and use it directly with the Trainer API for image classification. Here's how it works:

Loading a timm model

The TimmWrapperForImageClassification class (in transformers/src/transformers/models/timm_wrapper/modeling_timm_wrapper.py) wraps timm models so they're fully compatible with the Trainer API. You can load them via the Auto classes:

from transformers import AutoModelForImageClassification, AutoImageProcessor, Trainer, TrainingArguments

# Load a timm model for image classification
checkpoint = "timm/resnet50.a1_in1k"
image_processor = AutoImageProcessor.from_pretrained(checkpoint)
model = AutoModelForImageClassification.from_pretrained(
    checkpoint,
    num_labels=10,  # set to your number of classes
    ignore_mismatched_sizes=True,  # needed when changing num_labels from pretrained
)

Key details

  1. Image processor: The TimmWrapperImageProcessor automatically resolves the correct transforms from timm's config. It exposes both val_transforms and train_transforms (with augmentations), as noted in the code:
        # useful for training, see examples/pytorch/image-classification/run_image_classification.py
        self.train_transforms = timm.data.create_transform(**self.data_config, is_training=True)
  1. Loss computation is built-in: TimmWrapperForImageClassification.forward() accepts a labels argument and computes cross-entropy loss automatically, which is exactly what Trainer expects:
        loss = None
        if labels is not None:
            loss = self.loss_function(labels, logits, self.config)
  1. Returns ImageClassifierOutput: The output format is the standard transformers output, so Trainer handles it seamlessly.

Full training example

from transformers import AutoModelForImageClassification, AutoImageProcessor, Trainer, TrainingArguments
from datasets import load_dataset

# Load dataset
dataset = load_dataset("food101", split="train[:5000]")
dataset = dataset.train_test_split(test_size=0.2)

# Load timm model + processor
checkpoint = "timm/resnet50.a1_in1k"
image_processor = AutoImageProcessor.from_pretrained(checkpoint)
model = AutoModelForImageClassification.from_pretrained(
    checkpoint,
    num_labels=101,
    ignore_mismatched_sizes=True,
)

# Preprocessing
def transform(batch):
    batch["pixel_values"] = [image_processor(img)["pixel_values"][0] for img in batch["image"]]
    batch["labels"] = batch["label"]
    return batch

dataset["train"].set_transform(transform)
dataset["test"].set_transform(transform)

# Train
training_args = TrainingArguments(
    output_dir="./timm-finetuned",
    num_train_epochs=3,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    eval_strategy="epoch",
    save_strategy="epoch",
    logging_steps=50,
    remove_unused_columns=False,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=dataset["train"],
    eval_dataset=dataset["test"],
)

trainer.train()

Any timm checkpoint on the Hub (prefixed with timm/) works out of the box (ResNet, EfficientNet, ViT, ConvNeXt, etc). The wrapper handles all the translation between timm's interface and what Trainer expects.

Source: SKILL.md on GitHub

No alerts16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides a comprehensive suite for training and fine-tuning vision models on Hugging Face Jobs infrastructure. It includes production-ready scripts for object detection, image classification, and segmentation, along with tools for dataset validation and cost estimation. No security issues were detected, and the skill follows recommended practices for authentication and remote job management.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 386571e. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 6 months ago
  • huggingface
  • vision
  • object-detection
  • image-classification
  • segmentation
  • sam
  • detr
  • timm
  • transformers
  • training

README badge

README badge for huggingface/skills/huggingface-vision-trainer

Trains object detection, image classification, and SAM/SAM2 segmentation models on Hugging Face Jobs cloud GPUs without local setup. Supports D-FINE, RT-DETR v2, DETR, YOLOS, timm classifiers (MobileNetV3, ViT, ResNet), and segmentation with bbox or point prompts, with automatic dataset validation and Hub persistence.

Generated from the current SKILL.md.

Does this skill support SAM2 fine-tuning?
Yes. The skill covers both SAM and SAM2 segmentation training using bbox or point prompts. See scripts/sam_segmentation_training.py and references/finetune_sam2_trainer.md for SAM2-specific details.
What object detection models can I train?
D-FINE, RT-DETR v2, DETR, and YOLOS. The skill handles bbox format detection and conversion automatically, so you only need an objects column with bbox and category sub-fields.
Do I need a local GPU?
No. Training runs on Hugging Face Jobs cloud GPUs. You select hardware flavor (T4, L4, A10G, A100) at submission time, and results are automatically saved to the Hub.
What image classification models are supported?
Timm models (MobileNetV3, MobileViT, ResNet, ViT/DINOv3) and any Transformers classifier. Datasets must have an image column and a label column (integer or string).
What should I do before starting a training job?
Validate your dataset format first using the included dataset_inspector.py script to prevent format mismatches—the most common cause of training failures. Ensure your account has a paid plan and your token has write permissions.

Generated from the current SKILL.md. These answers refresh after source changes.