All skills
huggingface avatar

/huggingface-lora-space-builder

@86cdeee official
by Hugging Facehuggingface/skills11k stars
753

Build and publish a Gradio demo on Hugging Face Spaces for a user-provided LoRA. Use when someone asks to create, generate, ship, or publish a Space, demo, Gradio app, or playground for a LoRA — including LoRAs for Qwen-Image, Qwen-Image-Edit, LTX-Video, Wan, FLUX, SDXL, or other diffusion base models. Also triggers when someone describes a LoRA they trained or hosts on the Hub and wants to share it. Covers picking the right base pipeline and `diffusers` inference recipe, designing a UI tailored to the LoRA's task and inputs (Union/multi-task control, edit, video, image, etc.), respecting model-card recommendations (trigger words, steps, guidance, LoRA scale, example inputs), and shipping to ZeroGPU hardware as a private Space by default.

Use this Skill: https://skilld.dev/gh/huggingface/skills/huggingface-lora-space-builder

This session only. Nothing lands on disk.

referencesbase-modelsqwen-image.md

≈1.9k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Qwen-Image and Qwen-Image-Edit reference

The Qwen-Image family is fully supported in diffusers. Both base and edit variants accept LoRAs via the standard load_lora_weights interface.

Pipelines

Before using this table, verify against the base model's own card on the Hub. This table is best-effort and can lag a recent release. The diffusers snippet on the base model's Hub page is source of truth for which pipeline class to import. See SKILL.md Phase 2 for the procedure.

Base model Pipeline class Task
Qwen/Qwen-Image QwenImagePipeline Text-to-image
Qwen/Qwen-Image-Edit QwenImageEditPipeline Image editing (instruction-driven)
Qwen/Qwen-Image-Edit-2509 QwenImageEditPlusPipeline Image editing, multi-image input
Qwen/Qwen-Image-Edit-2511 QwenImageEditPlusPipeline Image editing, latest variant

The 2509 and 2511 variants use a different pipeline class than the original QwenImageEditPipeline — they take a list of input images and have different default parameters. Don't assume that variants in the same family share a pipeline class. Loading a 2511-trained LoRA onto QwenImageEditPipeline produces broken output; the failure is silent (no exception), so verifying against the base model card is the only way to catch it.

The 2511 variant integrates several popular community LoRAs into the base, which can mean a LoRA trained against earlier Qwen-Image-Edit may behave subtly differently when loaded against 2511; if the LoRA's model card specifies which Edit variant it was trained on, match it.

Required dependencies

Qwen-Image and Qwen-Image-Edit pipelines need extras beyond the standard diffusers/transformers/peft set, because the text encoder is Qwen2_5_VLForConditionalGeneration (Qwen 2.5-VL):

  • torchvision — required by Qwen2VLVideoProcessor, which the text encoder's processor pulls in transitively. Missing this is a startup-time ImportError ("Qwen2VLVideoProcessor requires the Torchvision library"). Always include in requirements.txt for any Qwen-Image Space.
  • sentencepiece — required by some Qwen tokenizer paths. Include if you see tokenizer-related ImportErrors at startup.

The 2511 variant in particular often requires the latest diffusers from git, since QwenImageEditPlusPipeline and 2511-specific fixes land before pip releases:

git+https://github.com/huggingface/diffusers

If from_pretrained("Qwen/Qwen-Image-Edit-2511", ...) fails with a class-not-found or attribute error, switch the requirement to git.

Default load (T2I)

import torch
from diffusers import QwenImagePipeline

pipe = QwenImagePipeline.from_pretrained(
    "Qwen/Qwen-Image",
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
pipe.load_lora_weights("user/my-qwen-lora")

pytorch_lora_weights.safetensors is the conventional filename. If the repo has a different name, pass weight_name="...".

For multiple adapters or when you want to control LoRA scale at inference time, use set_adapters:

pipe.load_lora_weights("user/my-qwen-lora", adapter_name="mylora")
pipe.set_adapters(["mylora"], adapter_weights=[0.9])

Default load (Image Edit)

For original Qwen-Image-Edit:

import torch
from diffusers import QwenImageEditPipeline

pipe = QwenImageEditPipeline.from_pretrained(
    "Qwen/Qwen-Image-Edit",
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
pipe.load_lora_weights("user/my-qwen-edit-lora")

For Qwen-Image-Edit-2509 and Qwen-Image-Edit-2511:

import torch
from diffusers import QwenImageEditPlusPipeline

pipe = QwenImageEditPlusPipeline.from_pretrained(
    "Qwen/Qwen-Image-Edit-2511",  # or 2509
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")
pipe.load_lora_weights("user/my-qwen-edit-lora")

QwenImageEditPlusPipeline accepts image=<PIL> or image=[<PIL>, <PIL>, ...] for multi-image edits. QwenImageEditPipeline accepts a single image. Default parameters differ slightly between the two — see "Inference defaults" below.

Inference defaults

For non-distilled Qwen-Image:

  • num_inference_steps: 50 by default; the LoRA's model card may recommend lower.
  • true_cfg_scale: typical 4.0.
  • width/height: multiples of 16, ideally 1024 or 1328 along the long axis.

For Qwen-Image-Edit (original):

  • num_inference_steps: 30–50 typical, often less for distilled variants.
  • true_cfg_scale: 4.0 typical.
  • Input image gets internally resized; passing a reasonable resolution (1024px on the long side) is fine.

For Qwen-Image-Edit-2509 / 2511 (QwenImageEditPlusPipeline):

  • num_inference_steps: 40 typical for 2511; 50 for 2509.
  • true_cfg_scale: 4.0.
  • guidance_scale: 1.0 (the new pipeline uses true_cfg_scale as the active CFG; standard guidance_scale is kept at 1.0).
  • Input is a list of one or more PIL images.

For Lightning / few-step LoRAs (e.g. lightx2v/Qwen-Image-Lightning-*):

  • num_inference_steps: 4 or 8 (read the LoRA's model card — they ship 4-step and 8-step variants).
  • true_cfg_scale: usually 1.0 (CFG disabled).
  • Often comes with a custom scheduler config — see the lightx2v README for the exact FlowMatchEulerDiscreteScheduler config to use.

Resolution buckets

Qwen-Image uses 16-pixel-aligned resolutions. When the user picks an aspect ratio, compute width and height as multiples of 16. A helper:

def round_to_bucket(w, h, multiple=16):
    return (w // multiple) * multiple, (h // multiple) * multiple

For image-edit pipelines, resize the input image to the nearest bucket while preserving aspect; don't crop.

ZeroGPU duration guidance

  • Standard 50-step Qwen-Image T2I at 1024×1024: 60–90 seconds.
  • 4-step Lightning Qwen-Image: 15–25 seconds.
  • Qwen-Image-Edit 30 steps: 60–90 seconds.

Set @spaces.GPU(duration=...) accordingly.

Common LoRA patterns on Qwen-Image

  • Style LoRAs (T2I). Standard load. Trigger word usually present. UI: prompt + aspect ratio.
  • Subject / character LoRAs (T2I). Standard load. Trigger word almost always present. UI: prompt with the trigger pre-prepended in code, possibly an example prompt highlighting the trigger.
  • Lighting / aesthetic LoRAs (T2I). Often paired with a recommended LoRA scale ≠ 1.0 — check the model card.
  • Edit LoRAs (image-to-image, on Qwen-Image-Edit). Specific instructions baked in. The LoRA might require a specific instruction phrasing — match the pattern from the model card. UI: input image + instruction textbox.
  • Lightning-distilled LoRAs. Lock step count and CFG to recommended values; hide the sliders.

Things to watch

  • VAE memory. For 1328×1328 outputs, consider pipe.enable_vae_tiling() and pipe.enable_vae_slicing() after loading. Keep them off for smaller resolutions to avoid quality loss.
  • Don't compile the transformer on ZeroGPU. torch.compile won't work; the speedup options on ZeroGPU are limited to reducing steps or using FP8-distilled variants.
  • Negative prompts work but defaults are often empty string. Don't expose a negative prompt in the UI unless the LoRA's behavior actually benefits from it.

Source: SKILL.md on GitHub

1 alert2mo3 checks · Risk SAFE
  • Gen Agent Trust Hub2mo

    This skill provides a comprehensive framework for building and publishing Gradio demos on Hugging Face Spaces. It follows established security best practices for managing authentication and secrets within the Hugging Face ecosystem and includes a human-in-the-loop review process before any code is published.

  • Socket2mo

    No alerts

  • Snyk2mo

    Risk: HIGH · 3 issues

Signed by skilld at 86cdeee. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 3 months ago
  • gradio
  • huggingface-spaces
  • lora
  • diffusers
  • text-to-image
  • image-to-image
  • text-to-video
  • image-to-video
  • video-to-video
  • qwen

README badge

README badge for huggingface/skills/huggingface-lora-space-builder

Builds and publishes a Gradio demo on Hugging Face Spaces for a user-provided LoRA model, handling inference with `diffusers` pipelines for base models like Qwen-Image, Qwen-Image-Edit, and LTX-Video. The skill reads the LoRA's model card to extract trigger words, recommended parameters, and task type, then designs a custom UI tailored to that LoRA's specific inputs and outputs before shipping the Space to ZeroGPU hardware.

Generated from the current SKILL.md.

What base models does this skill support?
First-class support for Qwen-Image, Qwen-Image-Edit, and LTX families. Other diffusion models can be attempted by analogy, but you should verify the pipeline class against the base model's own card before proceeding.
Do I need to authenticate to read a private LoRA repo?
The skill checks for a cached Hugging Face token first (from HF_TOKEN env var or CLI login). If one exists and can read the repo, no prompt is needed. If the repo is private and no valid token is cached, you'll be asked for one with write scope.
What happens if the LoRA model card has no base model or task info?
The skill will ask you directly for the base model, what the LoRA does, and recommended inference parameters (step count, guidance scale). You won't be asked piecemeal — all required info is batched into one message.
Is the published Space public or private?
Private by default. The Space is published to Hugging Face Spaces and the user can change visibility later.
Does this skill generate a local script or a published Space?
It generates and publishes a real Space on Hugging Face that runs in the browser, not a local script. The output is `app.py`, `requirements.txt`, and `README.md` bundled into a private Space.

Generated from the current SKILL.md. These answers refresh after source changes.