All skills
huggingface avatar

/huggingface-zerogpu

@5e3b4d2 official
by Hugging Facehuggingface/skills11k stars
753

AI demos and GPU compute with Gradio Spaces and Hugging Face Spaces ZeroGPU. Use when writing or reviewing code that uses `@spaces.GPU`, configuring `python_version` or `requirements.txt` for a ZeroGPU Space, or handling ZeroGPU-specific code constraints — pickle-based process isolation, `gr.State` semantics across the worker boundary, no `torch.compile` (use AoTI instead), CUDA wheel-only builds (no `nvcc` at build or runtime), large vs xlarge sizing, and dynamic duration callables. Make sure to use this skill whenever the user mentions ZeroGPU, `@spaces.GPU`, or the `spaces` Python package, or hits ZeroGPU-specific code errors like `PicklingError` across the worker boundary, `illegal duration`, or `flash-attn` wheel-build failures — even when the user does not explicitly ask for ZeroGPU coding guidance. Trigger on `import spaces` or `@spaces.GPU` in code.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-zerogpu

This session only. Nothing lands on disk.

referencescuda-and-deps.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

CUDA Dependencies on ZeroGPU

Detailed guidance for installing CUDA-dependent packages on ZeroGPU. SKILL.md establishes the bottom line — wheels are the recommended path because the ZeroGPU build phase has no nvcc. This document covers wheel filename tag reading, the kernels-community fallback, and torch-family side-car drift.

When no wheel is available on PyPI

Common workarounds, in preference order:

  1. Pre-built wheel via direct URL. For flash-attn, the upstream project ships a fairly complete matrix at https://github.com/Dao-AILab/flash-attention/releases — check there first and pin the matching wheel URL.
  2. Build the wheel yourself and host it (e.g. on a public HF Hub repo) when no upstream wheel matches the Space environment.
  3. Use a kernels-community kernel (see below) — handles ABI matching for you, no version pinning needed.

Reading a CUDA wheel filename

A wheel filename like

flash_attn-2.8.0.post2+cu12torch2.8cxx11abiFALSE-cp312-cp312-linux_x86_64.whl

encodes four build-time choices:

Tag Meaning
cu12 CUDA major version
torch2.8 torch major.minor the wheel was compiled against
cxx11abiFALSE C++ stdlib ABI choice (TRUE or FALSE)
cp312-cp312 CPython version (3.12)

The wheel's compiled C-extension will ImportError on ABI/symbol mismatches if any of these drift at install time.

If you hand pip a wheel URL without pinning the surrounding environment, pip may resolve torch to a version different from the wheel's build target, and the Space will fail on first import. Therefore:

  • Pin torch==X.Y.Z in requirements.txt to match the wheel's torch2.X tag.
  • Set python_version: in the Space frontmatter to match the cp3XX tag.
  • Check the runtime's cxx11-ABI choice against the wheel; if unsure, try the opposite ABI wheel.

Prefer kernels-community when unsure

If you are not sure about the ZeroGPU runtime's torch / Python / ABI combination, prefer a kernels-community kernel (e.g. kernels-community/flash-attn2) instead of a raw wheel URL. The kernels runtime handles ABI matching on your behalf, so no version pinning is required in your Space.

torch-family side-car drift

torchvision, torchaudio, torchcodec, and similar side-car packages are built against a specific torch major.minor (and CUDA major). On ZeroGPU, the runtime's supported torch list lags behind PyPI, so projects often pin a non-latest torch — and a bare uv add <side-car> can silently resolve to a newer release that targets a different torch / CUDA, producing ABI/import failures even though uv lock succeeded without warnings.

Concretely observed (2026-04) with torch==2.9.1 pinned:

  • torchaudio resolves to 2.11.0, which targets torch 2.11 / CUDA 13. The 2.11.0 release dropped the Requires-Dist: torch==X.Y.Z line that every earlier release had, so uv sees no constraint and picks it.
  • torchcodec resolves to a release targeting torch 2.11. No torchcodec release on PyPI declares a torch dependency at all; the compatibility table lives only in the project README.
  • torchvision happens to resolve correctly because torchvision still declares Requires-Dist: torch==X.Y.Z. Which side-cars are affected changes over time — treat every torch-family package as suspect, not just these.

Verify at add/upgrade time

After any uv add <torch-side-car> or uv lock --upgrade, verify the resolved version targets the same torch major.minor as pinned. Two-step fallback because PyPI metadata is not always sufficient:

  1. Query PyPI for the resolved version's requires_dist:
    curl -s https://pypi.org/pypi/<pkg>/<version>/json \
      | python3 -c "import json,sys,re; rd=json.load(sys.stdin)['info'].get('requires_dist') or []; print('\n'.join(x for x in rd if re.match(r'^torch(?![a-z])', x)) or '(no torch constraint declared)')"
    If a torch==X.Y.Z line appears and matches the pinned torch, good. If it appears and does NOT match, the side-car is wrong — pin it down explicitly.
  2. If the query prints (no torch constraint declared), PyPI metadata is silent and cannot be trusted. Fall back to the project's own compatibility table (GitHub README / docs site) — torchcodec, for example, maintains one at https://github.com/pytorch/torchcodec. Pick the side-car version the table maps to the pinned torch major.minor, and pin it explicitly.

Preventive pin

Once the correct side-car version is known, pin it in pyproject.toml alongside torch so uv cannot drift on future uv lock --upgrade. The side-car version numbers for a given torch major.minor change each release; always re-verify, do not copy a mapping from an older project.

Source: SKILL.md on GitHub

No alerts16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides comprehensive guidance for developing on the Hugging Face ZeroGPU platform. It includes technical considerations such as the use of external dependency wheels and specific serialization patterns that are integral to the platform's architecture. See detailed analysis for context.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

Signed by skilld at 5e3b4d2. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 4 months ago
  • hugging-face
  • gradio
  • zerogpu
  • gpu
  • cuda
  • torch
  • spaces
  • ml-demos
  • inference

README badge

README badge for huggingface/skills/huggingface-zerogpu

Configures GPU compute for Gradio Spaces on Hugging Face ZeroGPU hardware using the `@spaces.GPU` decorator. Covers duration tuning, process isolation via pickle, CUDA availability semantics, and constraints like no `torch.compile` support and CUDA wheel-only builds.

Generated from the current SKILL.md.

Does this skill apply to Streamlit or Docker Spaces?
No. This skill covers only Gradio SDK Spaces running on ZeroGPU hardware. Streamlit apps now run as Docker Spaces, which cannot schedule onto ZeroGPU. For general Gradio coding, see the huggingface-gradio skill.
Can I use torch.compile with ZeroGPU?
No. torch.compile is not supported on ZeroGPU. Use PyTorch ahead-of-time compilation (AoTI) with torch 2.8+ instead.
Do I need to handle imports differently for local development vs ZeroGPU?
No. Import spaces unconditionally and add it to requirements.txt. The spaces package is already a no-op outside ZeroGPU — the decorator and monkey-patching are gated on the SPACES_ZERO_GPU environment variable.
Can I return CUDA tensors directly from a @spaces.GPU function?
No. Convert CUDA tensors to CPU before returning — call tensor.cpu() or tensor.cpu().numpy() — because unpickling a CUDA tensor in the main process will trigger torch.cuda._lazy_init(), which ZeroGPU blocks.
How do I choose between large and xlarge GPU size?
Use large (the default, half GPU) for most workloads. Reserve xlarge (full GPU, 2x quota cost) only when you genuinely need the extra memory or compute, as it also tends to queue longer.

Generated from the current SKILL.md. These answers refresh after source changes.