All skills
huggingface avatar

/huggingface-zerogpu

@5e3b4d2 official
by Hugging Facehuggingface/skills11k stars
753

AI demos and GPU compute with Gradio Spaces and Hugging Face Spaces ZeroGPU. Use when writing or reviewing code that uses `@spaces.GPU`, configuring `python_version` or `requirements.txt` for a ZeroGPU Space, or handling ZeroGPU-specific code constraints — pickle-based process isolation, `gr.State` semantics across the worker boundary, no `torch.compile` (use AoTI instead), CUDA wheel-only builds (no `nvcc` at build or runtime), large vs xlarge sizing, and dynamic duration callables. Make sure to use this skill whenever the user mentions ZeroGPU, `@spaces.GPU`, or the `spaces` Python package, or hits ZeroGPU-specific code errors like `PicklingError` across the worker boundary, `illegal duration`, or `flash-attn` wheel-build failures — even when the user does not explicitly ask for ZeroGPU coding guidance. Trigger on `import spaces` or `@spaces.GPU` in code.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-zerogpu

This session only. Nothing lands on disk.

referenceshow-zerogpu-works.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

How ZeroGPU works (mechanism)

Conceptual lifecycle of model weights, processes, and worker reuse on ZeroGPU. Useful when reasoning about cold-starts, why module-scope warmup does not carry over to requests, why returning CUDA tensors hangs the call, or why gr.State mutations do not persist across the worker boundary.

For numerical limits (concurrency slots per Space, queue priority by tier, etc.), see the ZeroGPU docs — those values change over time and are deliberately kept out of this skill.

Two processes, two lifetimes

A ZeroGPU Space runs as two separate processes:

  • Main web process — long-lived. Imports app.py, launches Gradio, stays up for the life of the Space. Holds no VRAM and, after the startup "pack" step, holds no model weights in RAM either.
  • GPU worker processes — short-lived. Forked per @spaces.GPU request (or reused if warm). Run the task and are eventually killed by the ZeroGPU scheduler when another Space needs the GPU slot. Your Space code never kills its own worker.

Module-scope .to("cuda") is captured to disk

When import spaces is active, model.to("cuda") at module scope is intercepted. The call is rewritten to to("cpu"), so the tensor data physically lives in main process RAM at this point. A "fake" CUDA-presenting tensor is registered alongside the original CPU tensor.

At a startup "pack" step, the backend writes those original CPU tensors to disk via direct I/O (O_DIRECT), then frees the corresponding RAM. After pack, the main process holds no model weights anywhere — the data lives only on disk.

This is why module-scope pipe(...) / model.generate(...) / model(...) calls do not run on a real GPU: there is no GPU attached to the main process, and after pack there are no weights to compute against either. Such calls either fail or silently fall back to CPU on the fake tensors.

Worker init: disk → pinned memory → VRAM

When a @spaces.GPU call lands, the scheduler routes it to a worker:

  1. Cold worker — forked from the main process. The patched torch is unpatched, real CUDA is initialized, and weights are read from the disk offload directory into pinned host memory and streamed onto VRAM through a double-buffered pipeline (essentially pin_memory().cuda(non_blocking=True) per batch). This is the "cold-start" cost.
  2. Warm worker (reused) — an alive worker bound to the same GPU slot is reused if the scheduler reports it idle. Init is skipped; weights stay on VRAM from the previous call. Subsequent requests within a burst hit this path.

A warm worker is eventually killed by the scheduler when another Space needs the GPU slot. The next call after that point pays the disk → VRAM cost again. Occasional cold-starts on a low-traffic Space are normal.

Why module-scope warmup does not help

A common instinct is to call pipe("warmup") at module scope to "prepare" the model. This does not work on ZeroGPU:

  • At module scope, no real GPU is attached. The fake CUDA tensors do not have data after pack, so pipe(...) either fails or silently runs on something other than a real GPU.
  • Even if you wrap the warmup in @spaces.GPU, the worker that ran the warmup will eventually be killed before the first real user request lands — leaving them with a cold worker anyway.

The right answer is to load eagerly at module scope (pipe = pipeline(..., device="cuda")) and accept that the first user request after a quiet period will be a cold worker. Cold-start is fast on ZeroGPU because of the pinned-memory disk pipeline; it is not free, but it is not "minutes of model download" either.

Why returning CUDA tensors hangs the call

The main process never has a CUDA context — it has no GPU attached and its torch never initialized CUDA. When a worker returns a CUDA tensor, unpickling it in the main process triggers torch.cuda._lazy_init(), which would attempt to initialize CUDA in the main process. ZeroGPU blocks this, and the call hangs.

The fix is purely client-side: convert to CPU before returning (.cpu(), .cpu().numpy(), etc.). See "Process Isolation and Pickle" in SKILL.md.

Why gr.State does not share by reference across the boundary

Worker processes are forked separately and exchange data with the main process via pickle. gr.State values cross this boundary on every yield, so mutations inside a @spaces.GPU generator are local to the worker until the mutated state is explicitly yielded back. The main process gets a fresh deserialized copy each time — id() differs, in-place mutations are invisible across the boundary.

See "Process Isolation and Pickle" in SKILL.md for the practical implications and references/concurrency.md for related parallel-handler concerns.

Source: SKILL.md on GitHub

No alerts16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides comprehensive guidance for developing on the Hugging Face ZeroGPU platform. It includes technical considerations such as the use of external dependency wheels and specific serialization patterns that are integral to the platform's architecture. See detailed analysis for context.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

Signed by skilld at 5e3b4d2. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 4 months ago
  • hugging-face
  • gradio
  • zerogpu
  • gpu
  • cuda
  • torch
  • spaces
  • ml-demos
  • inference

README badge

README badge for huggingface/skills/huggingface-zerogpu

Configures GPU compute for Gradio Spaces on Hugging Face ZeroGPU hardware using the `@spaces.GPU` decorator. Covers duration tuning, process isolation via pickle, CUDA availability semantics, and constraints like no `torch.compile` support and CUDA wheel-only builds.

Generated from the current SKILL.md.

Does this skill apply to Streamlit or Docker Spaces?
No. This skill covers only Gradio SDK Spaces running on ZeroGPU hardware. Streamlit apps now run as Docker Spaces, which cannot schedule onto ZeroGPU. For general Gradio coding, see the huggingface-gradio skill.
Can I use torch.compile with ZeroGPU?
No. torch.compile is not supported on ZeroGPU. Use PyTorch ahead-of-time compilation (AoTI) with torch 2.8+ instead.
Do I need to handle imports differently for local development vs ZeroGPU?
No. Import spaces unconditionally and add it to requirements.txt. The spaces package is already a no-op outside ZeroGPU — the decorator and monkey-patching are gated on the SPACES_ZERO_GPU environment variable.
Can I return CUDA tensors directly from a @spaces.GPU function?
No. Convert CUDA tensors to CPU before returning — call tensor.cpu() or tensor.cpu().numpy() — because unpickling a CUDA tensor in the main process will trigger torch.cuda._lazy_init(), which ZeroGPU blocks.
How do I choose between large and xlarge GPU size?
Use large (the default, half GPU) for most workloads. Reserve xlarge (full GPU, 2x quota cost) only when you genuinely need the extra memory or compute, as it also tends to queue longer.

Generated from the current SKILL.md. These answers refresh after source changes.