All skills
huggingface avatar

/huggingface-lora-space-builder

@86cdeee official
by Hugging Facehuggingface/skills11k stars
753

Build and publish a Gradio demo on Hugging Face Spaces for a user-provided LoRA. Use when someone asks to create, generate, ship, or publish a Space, demo, Gradio app, or playground for a LoRA — including LoRAs for Qwen-Image, Qwen-Image-Edit, LTX-Video, Wan, FLUX, SDXL, or other diffusion base models. Also triggers when someone describes a LoRA they trained or hosts on the Hub and wants to share it. Covers picking the right base pipeline and `diffusers` inference recipe, designing a UI tailored to the LoRA's task and inputs (Union/multi-task control, edit, video, image, etc.), respecting model-card recommendations (trigger words, steps, guidance, LoRA scale, example inputs), and shipping to ZeroGPU hardware as a private Space by default.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-lora-space-builder

This session only. Nothing lands on disk.

referencesbase-modelskrea-2.md

≈1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Krea 2 reference

Krea 2 (K2) is a flow-matching text-to-image model: a 12B dense DiT (grouped-query attention) with a Qwen3-VL text encoder (multi-layer feature aggregation) and the Qwen-Image VAE (AutoencoderKLQwenImage). Pipeline class: Krea2Pipeline. It ships as two checkpoints designed to work together:

Repo Role Use it for
krea/Krea-2-Turbo 8-step distilled Inference / demos
krea/Krea-2-Raw base, non-distilled LoRA training — not for inference

For a LoRA Space, load Turbo. Krea 2 LoRAs are trained on RAW but run on Turbo (they "express strongly on Turbo"), and RAW is explicitly not meant for inference — expect poor quality if you load it directly. So a demo should almost always use krea/Krea-2-Turbo. The official LoRA cards confirm this ("To be used on krea/Krea-2-Turbo").

Needs a diffusers build that includes Krea2Pipeline (both repos are library_name: diffusers). If from diffusers import Krea2Pipeline fails, the installed diffusers predates the integration — update it.

Required dependencies

  • diffusers with Krea2Pipeline.
  • transformers recent enough for Qwen3-VL (the text encoder is a Qwen3VLModel, e.g. Qwen/Qwen3-VL-4B-Instruct).
  • torchvision — the Qwen3-VL processor pulls it in transitively; missing it is a startup ImportError. Always include it.
  • sentencepiece if you see tokenizer-related ImportErrors at startup.

Default load + LoRA (Turbo)

import torch
from diffusers import Krea2Pipeline

pipe = Krea2Pipeline.from_pretrained("krea/Krea-2-Turbo", torch_dtype=torch.bfloat16).to("cuda")
pipe.transformer.load_lora_adapter("user/my-krea2-lora", weight_name="my_lora.safetensors")
pipe.transformer.set_adapters("default", weights=1.0)

# include the LoRA's trigger word(s) from its card
image = pipe(
    "a deer grazing in a forest, <trigger words>",
    num_inference_steps=8, guidance_scale=0.0,
    generator=torch.Generator("cuda").manual_seed(0),
).images[0]

Krea 2 LoRAs load through the transformer's adapter API (pipe.transformer.load_lora_adapter(...) + pipe.transformer.set_adapters("default", weights=1.0)), per the official LoRA cards — not the pipeline-level pipe.load_lora_weights. Honor the LoRA's trigger word(s) and recommended weight (1.0 by default).

Inference recipe

  • Turbo (the demo default): num_inference_steps=8, guidance_scale=0.0 (guidance disabled), LoRA weight 1.0.
  • RAW: training only — don't ship a RAW inference demo; its quality is intentionally low (it's the malleable base you fine-tune on, then run on Turbo).

guidance_scale convention (gotcha)

Krea 2 enables guidance whenever guidance_scale > 0 and computes velocity as cond + guidance_scale * (cond − uncond) (≡ the usual CFG formulation with scale 1 + guidance_scale). So Turbo disables guidance with guidance_scale=0.0 — not 1.0.

Resolution

height/width must be divisible by 16 (vae_scale_factor * patch_size); the pipeline rounds up to a multiple of 16 (with a warning) otherwise. Default 1024×1024.

ZeroGPU duration

Standard T2I: place modules at module scope, pipe.to("cuda"), no torch.compile. Turbo 8-step 1024² is fast (≈ 20–40s) — set @spaces.GPU(duration=...) comfortably above that. The 12B DiT + Qwen3-VL encoder fits ZeroGPU; if you OOM on the default size, try @spaces.GPU(size="xlarge").

Things to watch

  • Load Turbo, not RAW, for demos. RAW is the training base and not meant for inference.
  • guidance_scale=0 disables guidance (Krea convention, unlike pipelines where 1.0 is "off"). Turbo = 8 steps, guidance_scale=0.0.
  • LoRAs use pipe.transformer.load_lora_adapter + set_adapters("default", …) (transformer-level), and have trigger words — read the card.
  • New pipeline → update diffusers if from diffusers import Krea2Pipeline fails.

Source: SKILL.md on GitHub

1 alert2mo3 checks · Risk SAFE
  • Gen Agent Trust Hub2mo

    This skill provides a comprehensive framework for building and publishing Gradio demos on Hugging Face Spaces. It follows established security best practices for managing authentication and secrets within the Hugging Face ecosystem and includes a human-in-the-loop review process before any code is published.

  • Socket2mo

    No alerts

  • Snyk2mo

    Risk: HIGH · 3 issues

Signed by skilld at 86cdeee. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 3 months ago
  • gradio
  • huggingface-spaces
  • lora
  • diffusers
  • text-to-image
  • image-to-image
  • text-to-video
  • image-to-video
  • video-to-video
  • qwen

README badge

README badge for huggingface/skills/huggingface-lora-space-builder

Builds and publishes a Gradio demo on Hugging Face Spaces for a user-provided LoRA model, handling inference with `diffusers` pipelines for base models like Qwen-Image, Qwen-Image-Edit, and LTX-Video. The skill reads the LoRA's model card to extract trigger words, recommended parameters, and task type, then designs a custom UI tailored to that LoRA's specific inputs and outputs before shipping the Space to ZeroGPU hardware.

Generated from the current SKILL.md.

What base models does this skill support?
First-class support for Qwen-Image, Qwen-Image-Edit, and LTX families. Other diffusion models can be attempted by analogy, but you should verify the pipeline class against the base model's own card before proceeding.
Do I need to authenticate to read a private LoRA repo?
The skill checks for a cached Hugging Face token first (from HF_TOKEN env var or CLI login). If one exists and can read the repo, no prompt is needed. If the repo is private and no valid token is cached, you'll be asked for one with write scope.
What happens if the LoRA model card has no base model or task info?
The skill will ask you directly for the base model, what the LoRA does, and recommended inference parameters (step count, guidance scale). You won't be asked piecemeal — all required info is batched into one message.
Is the published Space public or private?
Private by default. The Space is published to Hugging Face Spaces and the user can change visibility later.
Does this skill generate a local script or a published Space?
It generates and publishes a real Space on Hugging Face that runs in the browser, not a local script. The output is `app.py`, `requirements.txt`, and `README.md` bundled into a private Space.

Generated from the current SKILL.md. These answers refresh after source changes.