All skills
huggingface avatar

/huggingface-spaces

@57d80c3 official
by Hugging Facehuggingface/skills11k stars
753

Build, deploy, and maintain applications on Hugging Face Spaces — Gradio / Docker / Static SDKs, ZeroGPU and dedicated hardware, model loading, debugging, buckets, inference providers, community grants. Use whenever the user asks to create or host an app on Hugging Face, port code onto ZeroGPU, fix a Space that won't build or run, or otherwise work with `hf spaces …`, `@spaces.GPU`, Space README frontmatter, or the `spaces` Python package.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-spaces

This session only. Nothing lands on disk.

referencesbuckets.md

≈1.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Persistent storage with Buckets

Spaces are stateless. All data is wiped on restart / rebuild. For state that must survive (user uploads, generations, dynamic feeds, logs, growing databases): mount an HF Bucket — S3-like object storage living at hf://buckets/<ns>/<bucket>.

Buckets are paid (per-TB storage). Check whoami.canPay and confirm with the user before creating one. Pricing + free tier: https://huggingface.co/storage.

Full docs: https://huggingface.co/docs/hub/storage-buckets.

Create + attach

hf buckets create <ns>/<bucket-name>                                # --private optional
hf spaces volumes set <ns>/<space> -v hf://buckets/<ns>/<bucket-name>:/data

After this, writes to /data/ in the Space are durable. Reads come from the bucket via the Xet storage backend.

To make the bucket files publicly addressable: leave the bucket public. Public bucket files are served at https://huggingface.co/buckets/<ns>/<bucket>/resolve/<path> (HTTP 302 redirect to a signed CDN URL). The Space writes once and the public URL works forever — no streaming proxy needed.

Write-durable, read-fast pattern

For a feed-style Space (e.g. a community jam where users save generations and browse a public timeline), don't re-scan disk on every request. Module-level disk scan → in-memory list → every write appends to both:

import os, json, uuid
from datetime import datetime, timezone

BUCKET_ID = "<ns>/<bucket-name>"
BUCKET_URL = f"https://huggingface.co/buckets/{BUCKET_ID}/resolve"
_feed = []

def _load_feed():
    root = "/data/songs"
    if not os.path.isdir(root):
        return
    for sid in os.listdir(root):
        meta = f"{root}/{sid}/meta.json"
        if os.path.isfile(meta):
            _feed.append(json.load(open(meta)))
    _feed.sort(key=lambda s: s["created_at"], reverse=True)

_load_feed()                                       # one scan at startup

@app.api(name="save", time_limit=60)
def save(audio_bytes: bytes, title: str):
    sid = uuid.uuid4().hex[:12]
    d = f"/data/songs/{sid}"; os.makedirs(d, exist_ok=True)
    open(f"{d}/audio.wav", "wb").write(audio_bytes)
    meta = {"id": sid, "title": title,
            "url": f"{BUCKET_URL}/songs/{sid}/audio.wav",
            "created_at": datetime.now(timezone.utc).isoformat()}
    json.dump(meta, open(f"{d}/meta.json", "w"))
    _feed.insert(0, meta)                          # cache stays current — no re-scan
    return meta

@app.api(name="feed", concurrency_limit=10)
def feed(): return _feed[:50]                      # zero disk I/O

Reference Space using this pattern: https://huggingface.co/spaces/victor/ace-step-jam

Anti-pattern: bucket as model-weights cache

Do NOT snapshot_download(..., local_dir="/data/weights") and load checkpoints from there. Bucket I/O is S3-paced; reading a 22 GB safetensors from /data during from_pretrained stalls past any @spaces.GPU duration cap.

For model weights, let HF Hub re-download to local container disk on each cold start. With HF_HUB_ENABLE_HF_TRANSFER=1 (set in the runtime by default) this is fast — typically much faster than streaming the same bytes through bucket I/O at request time.

Bucket I/O is fine for occasional metadata reads (the feed pattern above) or saving user information. It is not fine as the path your model loader streams gigabytes through every cold start.

Cache redirects

/home/user/.cache is read-only on ZeroGPU. Redirect transient caches at the top of app.py, before any library import that uses them:

import os
os.environ.setdefault("HF_HOME", "/data/.cache/huggingface")   # or /tmp on non-bucket Spaces
os.environ.setdefault("HF_MODULES_CACHE", "/tmp/hf_modules")
os.environ.setdefault("MPLCONFIGDIR", "/tmp/matplotlib")

Missing redirections fail silently or at first matplotlib / transformers / diffusers import.

Write access from the Space

The Space's HF_TOKEN secret needs write permission on the bucket. Set via Settings → Secrets in the Space UI, or hf spaces secrets set <id> HF_TOKEN=<token>.

Security note

Public bucket files are publicly accessible forever at their resolve URL. Don't write PII to a public bucket. If you need durable but private storage (e.g. per-user history requiring an HF login), keep the bucket private and gate reads through your Space's own auth.

Source: SKILL.md on GitHub

1 alert1mo3 checks · Risk SAFE
  • Gen Agent Trust Hub1mo

    This skill provides a comprehensive developer toolkit for building, deploying, and maintaining machine learning applications on Hugging Face Spaces. It includes detailed instructions for handling hardware configurations, specialized model dependencies (particularly for 3D generation), and persistent storage. No malicious patterns were detected, and all external resources are used within the context of standard machine learning development workflows on the Hugging Face platform.

  • Socket1mo

    No alerts

  • Snyk1mo

    Risk: CRITICAL · 3 issues

Signed by skilld at 57d80c3. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 2 months ago
  • hugging-face
  • spaces
  • gradio
  • zerogpu
  • docker
  • static
  • ml-deployment
  • model-hosting
  • inference

README badge

README badge for huggingface/skills/huggingface-spaces

Builds and deploys machine-learning applications on Hugging Face Spaces using Gradio, Docker, or Static SDKs. Covers ZeroGPU allocation, hardware selection, model loading, debugging, and the `hf spaces` CLI workflow. Use this skill when creating a Space from scratch, migrating code to ZeroGPU, or troubleshooting a Space that won't build or run.

Generated from the current SKILL.md.

Does this skill work with Docker and Static Spaces, or only Gradio?
It covers all three SDKs — Gradio, Docker, and Static. However, ZeroGPU is Gradio-only; Docker and Static Spaces use different hardware or no hardware at all.
What do I need to use ZeroGPU?
You must be on a PRO, Team, or Enterprise plan to create a ZeroGPU Space. Visitors consume their own daily quota (~5 min free / 40 min Pro / 60 min Enterprise) when they use the Space.
Can I use TensorFlow or ONNX as the main model on ZeroGPU?
No. ZeroGPU is PyTorch-first; non-PyTorch frameworks as the primary inference path require a dedicated paid GPU. Small non-torch tools (preprocessors, utilities) inside a PyTorch pipeline are fine on ZeroGPU.
What versions of Python and PyTorch does ZeroGPU support?
ZeroGPU officially supports Python 3.10.13 and 3.12.12. For PyTorch, it accepts 2.8.0, 2.9.1, 2.10.0, and 2.11.0; the runtime preinstalls the latest if you leave torch unpinned.
Should I test my Space locally before pushing to Hugging Face?
Minimal local checks (like `python3 -m py_compile app.py`) are fine, but the Space environment is the only one that matters. Build a release candidate locally, push it, then use the live URL as your test loop.

Generated from the current SKILL.md. These answers refresh after source changes.