All skills
huggingface avatar

/huggingface-zerogpu

@5e3b4d2 official
by Hugging Facehuggingface/skills11k stars
753

AI demos and GPU compute with Gradio Spaces and Hugging Face Spaces ZeroGPU. Use when writing or reviewing code that uses `@spaces.GPU`, configuring `python_version` or `requirements.txt` for a ZeroGPU Space, or handling ZeroGPU-specific code constraints — pickle-based process isolation, `gr.State` semantics across the worker boundary, no `torch.compile` (use AoTI instead), CUDA wheel-only builds (no `nvcc` at build or runtime), large vs xlarge sizing, and dynamic duration callables. Make sure to use this skill whenever the user mentions ZeroGPU, `@spaces.GPU`, or the `spaces` Python package, or hits ZeroGPU-specific code errors like `PicklingError` across the worker boundary, `illegal duration`, or `flash-attn` wheel-build failures — even when the user does not explicitly ask for ZeroGPU coding guidance. Trigger on `import spaces` or `@spaces.GPU` in code.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-zerogpu

This session only. Nothing lands on disk.

referencesconcurrency.md

≈641 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Concurrency Safety

Gradio handlers run in parallel by default on ZeroGPU. Code that works fine in single-user testing can silently corrupt or leak data in production. Always assume handlers execute concurrently.

No mutable global state

Per-request or per-user data must not live in module-level mutable variables. Concurrent requests will overwrite each other.

# BAD — concurrent requests overwrite each other
results = {}

def process(text):
    results["output"] = expensive_compute(text)  # race condition
    return results["output"]
# GOOD — pure function, no shared mutable state
def process(text):
    return expensive_compute(text)

For state that must persist within a single user session, use gr.State:

with gr.Blocks() as demo:
    history = gr.State(value=[])

    def add_message(msg, hist):
        hist.append(msg)
        return hist, hist

    btn.click(fn=add_message, inputs=[msg, history], outputs=[chatbot, history])

Note that on ZeroGPU, gr.State is pickled across the worker boundary on every yield — see "Process Isolation and Pickle" in SKILL.md for the implications.

No fixed file paths for outputs

Hardcoded output filenames cause concurrent requests to overwrite each other's files. This corrupts outputs and, worse, can leak one user's data to another.

# BAD — concurrent calls clobber the same file
def generate_image(prompt):
    image = pipe(prompt).images[0]
    image.save("output.png")
    return "output.png"
# GOOD — unique path per invocation
import tempfile

def generate_image(prompt):
    image = pipe(prompt).images[0]
    with tempfile.NamedTemporaryFile(suffix=".png", delete=False) as f:
        image.save(f.name)
        return f.name

The same applies to any intermediate files (audio, video, CSV exports). Always generate a unique path per invocation.

Read-only globals are safe

Model objects, tokenizers, and configs loaded once at startup and only read during requests are safe and encouraged. This is the standard ZeroGPU pattern: load at module scope, read inside @spaces.GPU handlers.

# SAFE — loaded once at module scope, read-only during requests
model = load_model().to("cuda")
tokenizer = load_tokenizer()

@spaces.GPU
def predict(text):
    tokens = tokenizer(text, return_tensors="pt").to("cuda")
    return model.generate(**tokens)

The "no mutable global state" rule targets writes from handlers, not reads. A handler that only reads from a global is concurrency-safe.

Source: SKILL.md on GitHub

No alerts16d3 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides comprehensive guidance for developing on the Hugging Face ZeroGPU platform. It includes technical considerations such as the use of external dependency wheels and specific serialization patterns that are integral to the platform's architecture. See detailed analysis for context.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

Signed by skilld at 5e3b4d2. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 4 months ago
  • hugging-face
  • gradio
  • zerogpu
  • gpu
  • cuda
  • torch
  • spaces
  • ml-demos
  • inference

README badge

README badge for huggingface/skills/huggingface-zerogpu

Configures GPU compute for Gradio Spaces on Hugging Face ZeroGPU hardware using the `@spaces.GPU` decorator. Covers duration tuning, process isolation via pickle, CUDA availability semantics, and constraints like no `torch.compile` support and CUDA wheel-only builds.

Generated from the current SKILL.md.

Does this skill apply to Streamlit or Docker Spaces?
No. This skill covers only Gradio SDK Spaces running on ZeroGPU hardware. Streamlit apps now run as Docker Spaces, which cannot schedule onto ZeroGPU. For general Gradio coding, see the huggingface-gradio skill.
Can I use torch.compile with ZeroGPU?
No. torch.compile is not supported on ZeroGPU. Use PyTorch ahead-of-time compilation (AoTI) with torch 2.8+ instead.
Do I need to handle imports differently for local development vs ZeroGPU?
No. Import spaces unconditionally and add it to requirements.txt. The spaces package is already a no-op outside ZeroGPU — the decorator and monkey-patching are gated on the SPACES_ZERO_GPU environment variable.
Can I return CUDA tensors directly from a @spaces.GPU function?
No. Convert CUDA tensors to CPU before returning — call tensor.cpu() or tensor.cpu().numpy() — because unpickling a CUDA tensor in the main process will trigger torch.cuda._lazy_init(), which ZeroGPU blocks.
How do I choose between large and xlarge GPU size?
Use large (the default, half GPU) for most workloads. Reserve xlarge (full GPU, 2x quota cost) only when you genuinely need the extra memory or compute, as it also tends to queue longer.

Generated from the current SKILL.md. These answers refresh after source changes.