All skills
huggingface avatar

/huggingface-community-evals

@386571e official
by Hugging Facehuggingface/skills11k stars
753

Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware. Use for backend selection, local GPU evals, and choosing between vLLM / Transformers / accelerate. Not for HF Jobs orchestration, model-card PRs, .eval_results publication, or community-evals automation.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-community-evals

This session only. Nothing lands on disk.

examplesUSAGE_EXAMPLES.md

≈516 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Usage Examples

This document provides practical examples for running evaluations locally against Hugging Face Hub models.

What this skill covers

  • inspect-ai local runs
  • inspect-ai with vllm or Transformers backends
  • lighteval local runs with vllm or accelerate
  • smoke tests and backend fallback patterns

What this skill does NOT cover

  • model-index
  • .eval_results
  • community eval publication workflows
  • model-card PR creation
  • Hugging Face Jobs orchestration

If you want to run these same scripts remotely, use the hugging-face-jobs skill and pass one of the scripts in scripts/.

Setup

cd skills/hugging-face-evaluation
export HF_TOKEN=hf_xxx
uv --version

For local GPU runs:

nvidia-smi

inspect-ai examples

Quick smoke test

uv run scripts/inspect_eval_uv.py \
  --model meta-llama/Llama-3.2-1B \
  --task mmlu \
  --limit 10

Local GPU with vLLM

uv run scripts/inspect_vllm_uv.py \
  --model meta-llama/Llama-3.2-8B-Instruct \
  --task gsm8k \
  --limit 20

Transformers fallback

uv run scripts/inspect_vllm_uv.py \
  --model microsoft/phi-2 \
  --task mmlu \
  --backend hf \
  --trust-remote-code \
  --limit 20

lighteval examples

Single task

uv run scripts/lighteval_vllm_uv.py \
  --model meta-llama/Llama-3.2-3B-Instruct \
  --tasks "leaderboard|mmlu|5" \
  --max-samples 20

Multiple tasks

uv run scripts/lighteval_vllm_uv.py \
  --model meta-llama/Llama-3.2-3B-Instruct \
  --tasks "leaderboard|mmlu|5,leaderboard|gsm8k|5" \
  --max-samples 20 \
  --use-chat-template

accelerate fallback

uv run scripts/lighteval_vllm_uv.py \
  --model microsoft/phi-2 \
  --tasks "leaderboard|mmlu|5" \
  --backend accelerate \
  --trust-remote-code \
  --max-samples 20

Hand-off to Hugging Face Jobs

When local hardware is not enough, switch to the hugging-face-jobs skill and run one of these scripts remotely. Keep the script path and args; move the orchestration there.

Source: SKILL.md on GitHub

1 warning16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides a functional and secure environment for evaluating Hugging Face models locally using established frameworks like inspect-ai and lighteval. It adheres to standard security practices for handling authentication tokens and executing command-line tools. Users should be aware of standard considerations regarding the optional loading of custom model code from external repositories.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 386571e. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 6 months ago
  • hugging-face
  • inspect-ai
  • lighteval
  • evaluation
  • llm
  • vllm
  • transformers
  • gpu
  • benchmarking

README badge

README badge for huggingface/skills/huggingface-community-evals

Runs local evaluations of Hugging Face Hub models using inspect-ai or lighteval, with backends for vLLM, Transformers, and accelerate. Choose inspect-ai for explicit task control or lighteval for leaderboard-style benchmarks; start with smoke tests before scaling up on local GPU hardware.

Generated from the current SKILL.md.

Does this skill run evaluations remotely on Hugging Face Jobs?
No. This skill covers local GPU evaluation only. If you need remote execution on HF Jobs, use the hugging-face-jobs skill and pass it one of the local scripts from this skill.
What inference backends does this support?
vLLM (preferred for throughput), Hugging Face Transformers, and accelerate. The skill provides scripts to choose or fall back between them based on model support and hardware constraints.
Can I run evaluations without a local GPU?
Yes. Use inspect_eval_uv.py with Hugging Face Inference Providers for provider-backed evaluation without direct GPU control.
What evaluation frameworks does this cover?
inspect-ai and lighteval. Use inspect-ai for explicit task control; use lighteval when benchmarks map naturally to lighteval task strings, especially for leaderboard-style evaluations.
Does this handle publishing results to community evals?
No. This skill stops at local evaluation runs. For publishing results into the community evals workflow, hand off to the community-evals skill.

Generated from the current SKILL.md. These answers refresh after source changes.