All skills
huggingface avatar

/huggingface-zerogpu

@5e3b4d2 official
by Hugging Facehuggingface/skills11k stars
753

AI demos and GPU compute with Gradio Spaces and Hugging Face Spaces ZeroGPU. Use when writing or reviewing code that uses `@spaces.GPU`, configuring `python_version` or `requirements.txt` for a ZeroGPU Space, or handling ZeroGPU-specific code constraints — pickle-based process isolation, `gr.State` semantics across the worker boundary, no `torch.compile` (use AoTI instead), CUDA wheel-only builds (no `nvcc` at build or runtime), large vs xlarge sizing, and dynamic duration callables. Make sure to use this skill whenever the user mentions ZeroGPU, `@spaces.GPU`, or the `spaces` Python package, or hits ZeroGPU-specific code errors like `PicklingError` across the worker boundary, `illegal duration`, or `flash-attn` wheel-build failures — even when the user does not explicitly ask for ZeroGPU coding guidance. Trigger on `import spaces` or `@spaces.GPU` in code.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-zerogpu

This session only. Nothing lands on disk.

referenceshow-quota-works.md

≈972 tokens on demand. Your agent reads this file only when SKILL.md points to it.

How ZeroGPU duration and quota are checked

Mechanism for duration validation and quota pre-checks. Useful when choosing duration values, debugging illegal duration vs quota exceeded errors, and understanding why the default 60s is pessimistic for short tasks.

For per-tier numerical thresholds (free vs Pro vs Team vs Enterprise quota minutes), the daily quota window length, runs-per-day limits, and pay-as-you-go pricing, see the ZeroGPU docs — those values change over time and are deliberately kept out of this skill.

What duration actually requests

Whatever value is passed to @spaces.GPU(duration=N) (or the default 60s when unspecified) becomes the requested duration the platform checks against. For xlarge, the request is doubled internally:

requested = N * 2 if size == "xlarge" else N

So @spaces.GPU(duration=60, size="xlarge") is internally a 120-second request — both for the tier-max check and the quota pre-check below.

Two distinct error modes

Two failure messages can come back from the scheduler before the call runs:

Error Trigger What helps
ZeroGPU illegal duration requested duration > visitor's tier per-call cap Lower duration. Sign in / upgrade tier. Waiting does not help.
ZeroGPU quota exceeded remaining quota < requested duration, OR runs-per-day cap reached Wait for the quota window to reset. For Pro / Team / Enterprise, pay-as-you-go credits cover the overflow.

The error wording for quota exceeded includes the explicit numbers, e.g.:

You have exceeded your Pro ZeroGPU quota
(60s requested vs. 30s left). Try again in 1:23:45.

The comparison is requested vs remaining — not actual run time vs remaining. A 10-second task left at the default 60s requests 60s of quota; once remaining < 60s the call fails even though the actual work would have fit.

Why the default 60s is pessimistic for short tasks

DEFAULT_SCHEDULE_DURATION in the spaces package is 60 seconds. So an undecorated @spaces.GPU (or @spaces.GPU() with no duration=) requests 60s of quota.

For a task that actually takes ~10 seconds:

  • The user's 60s quota gets reserved up front.
  • Once their remaining quota drops below 60s, your Space fails for them — even though they could have run many more 10s tasks if the request matched reality.
  • Your call also ranks lower in the queue than equivalent calls declaring smaller durations.

The fix is to declare the realistic duration explicitly:

@spaces.GPU(duration=15)
def fast_task(...):
    ...

For workloads where runtime depends on inputs, use a callable (per-request estimator):

def estimate_duration(prompt, steps):
    return int(steps * 3.5)

@spaces.GPU(duration=estimate_duration)
def variable_task(prompt, steps):
    ...

This preserves quota for light inputs and reserves more only when needed.

Quota window: 24h fixed from first use

The quota window's TTL is set when the first call of a fresh window lands and counts down unconditionally — it is not a sliding window, not a calendar-day reset, and not extended by subsequent use. A user who runs a call at 14:00 sees their next reset at 14:00 the following day, regardless of how heavily or lightly they use the Space in between.

For exact tier thresholds, runs-per-day caps, and pay-as-you-go billing rates, see the ZeroGPU docs.

Queue priority

The queue is node-level — requests from every Space scheduled on the same physical node compete for that node's GPU slots. Among queued requests, shorter declared duration ranks higher. So tight per-request duration estimates serve two goals at once: they preserve the user's quota and move the request up the queue.

Source: SKILL.md on GitHub

No alerts16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides comprehensive guidance for developing on the Hugging Face ZeroGPU platform. It includes technical considerations such as the use of external dependency wheels and specific serialization patterns that are integral to the platform's architecture. See detailed analysis for context.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

Signed by skilld at 5e3b4d2. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 4 months ago
  • hugging-face
  • gradio
  • zerogpu
  • gpu
  • cuda
  • torch
  • spaces
  • ml-demos
  • inference

README badge

README badge for huggingface/skills/huggingface-zerogpu

Configures GPU compute for Gradio Spaces on Hugging Face ZeroGPU hardware using the `@spaces.GPU` decorator. Covers duration tuning, process isolation via pickle, CUDA availability semantics, and constraints like no `torch.compile` support and CUDA wheel-only builds.

Generated from the current SKILL.md.

Does this skill apply to Streamlit or Docker Spaces?
No. This skill covers only Gradio SDK Spaces running on ZeroGPU hardware. Streamlit apps now run as Docker Spaces, which cannot schedule onto ZeroGPU. For general Gradio coding, see the huggingface-gradio skill.
Can I use torch.compile with ZeroGPU?
No. torch.compile is not supported on ZeroGPU. Use PyTorch ahead-of-time compilation (AoTI) with torch 2.8+ instead.
Do I need to handle imports differently for local development vs ZeroGPU?
No. Import spaces unconditionally and add it to requirements.txt. The spaces package is already a no-op outside ZeroGPU — the decorator and monkey-patching are gated on the SPACES_ZERO_GPU environment variable.
Can I return CUDA tensors directly from a @spaces.GPU function?
No. Convert CUDA tensors to CPU before returning — call tensor.cpu() or tensor.cpu().numpy() — because unpickling a CUDA tensor in the main process will trigger torch.cuda._lazy_init(), which ZeroGPU blocks.
How do I choose between large and xlarge GPU size?
Use large (the default, half GPU) for most workloads. Reserve xlarge (full GPU, 2x quota cost) only when you genuinely need the extra memory or compute, as it also tends to queue longer.

Generated from the current SKILL.md. These answers refresh after source changes.