All skills
jimliu avatar

/baoyu-image-gen

@1567581 official
by Jim Liu 宝玉jimliu/baoyu-skills26k stars
2,896

AI image generation with OpenAI GPT Image 2.5, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream, Replicate and Agnes APIs. Supports text-to-image, reference images, aspect ratios, and batch generation from saved prompt files. Sequential by default; use batch parallel generation when the user already has multiple prompts or wants stable multi-image throughput. Use when user asks to generate, create, or draw images.

Use this Skill: https://skilld.dev/gh/jimliu/baoyu-skills/baoyu-image-gen

This session only. Nothing lands on disk.

referencesusage-examples.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Usage Examples

Extended CLI examples. SKILL.md shows the minimum set; read this file when the user asks about provider-specific invocation, batch generation, or less-common flags.

Core Patterns

# Basic text-to-image
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image cat.png

# With aspect ratio
${BUN_X} {baseDir}/scripts/main.ts --prompt "A landscape" --image out.png --ar 16:9

# High quality
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --quality 2k

# Prompt from files
${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png

# With reference images (any provider family that supports refs)
${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --ref source.png

Per-Provider

# OpenAI
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider openai --model gpt-image-2.5-flare

# Azure OpenAI (model = deployment name)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider azure --model gpt-image-2.5-flare

# OpenAI GPT Image 2 custom 4K size
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic landscape" --image out.png --provider openai --model gpt-image-2.5-sunburst --size 3840x2160

# Google with explicit model
${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --provider google --model gemini-3-pro-image --ref source.png

# OpenRouter (recommended default)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider openrouter

# OpenRouter with reference
${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --provider openrouter --model google/gemini-3.1-flash-image --ref source.png

# DashScope (default model)
${BUN_X} {baseDir}/scripts/main.ts --prompt "一只可爱的猫" --image out.png --provider dashscope

# DashScope Qwen-Image 2.0 Pro (custom size, Chinese text)
${BUN_X} {baseDir}/scripts/main.ts --prompt "为咖啡品牌设计一张 21:9 横幅海报,包含清晰中文标题" --image out.png --provider dashscope --model qwen-image-2.0-pro --size 2048x872

# DashScope legacy fixed-size
${BUN_X} {baseDir}/scripts/main.ts --prompt "一张电影感海报" --image out.png --provider dashscope --model qwen-image-max --size 1664x928

# DashScope Wan 2.7 Image Pro (4K text-to-image)
${BUN_X} {baseDir}/scripts/main.ts --prompt "一间有着精致窗户的花店" --image out.png --provider dashscope --model wan2.7-image-pro --size 4096x4096

# DashScope Wan 2.7 Image with reference image (multi-image fusion)
${BUN_X} {baseDir}/scripts/main.ts --prompt "把图2的涂鸦喷绘在图1的汽车上" --image out.png --provider dashscope --model wan2.7-image-pro --ref car.webp paint.webp

# Z.AI GLM-image
${BUN_X} {baseDir}/scripts/main.ts --prompt "一张带清晰中文标题的科技海报" --image out.png --provider zai

# Z.AI with custom size
${BUN_X} {baseDir}/scripts/main.ts --prompt "A science illustration with labels" --image out.png --provider zai --model glm-image --size 1472x1088

# MiniMax
${BUN_X} {baseDir}/scripts/main.ts --prompt "A fashion editorial portrait" --image out.jpg --provider minimax

# MiniMax with subject reference (character/portrait consistency)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A girl by the library window" --image out.jpg --provider minimax --model image-01 --ref portrait.png --ar 16:9

# Replicate (default: google/nano-banana-2)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider replicate

# Replicate Seedream 4.5
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic portrait" --image out.png --provider replicate --model bytedance/seedream-4.5 --ar 3:2

# Replicate Wan 2.7 Image Pro
${BUN_X} {baseDir}/scripts/main.ts --prompt "A concept frame" --image out.png --provider replicate --model wan-video/wan-2.7-image-pro --size 2048x1152

# Codex CLI (uses Codex / ChatGPT subscription — no OPENAI_API_KEY; requires `codex` on PATH and `codex login`)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic portrait" --image out.png --provider codex-cli --ar 16:9

# Codex CLI with reference images (style/composition guidance)
${BUN_X} {baseDir}/scripts/main.ts --prompt "Match this color palette" --image out.png --provider codex-cli --ref source.png --ar 1:1

# Agnes (default model)
${BUN_X} {baseDir}/scripts/main.ts --prompt "A detailed infographic" --image out.png --provider agnes

# Agnes with aspect ratio and URL output
${BUN_X} {baseDir}/scripts/main.ts --prompt "A cinematic scene" --image out.txt --provider agnes --ar 16:9 --response-format url

# Agnes with reference image
${BUN_X} {baseDir}/scripts/main.ts --prompt "Apply this style" --image out.png --provider agnes --ref source.png

Notes on codex-cli:

  • Never auto-selected — pin via --provider codex-cli or default_provider: codex-cli in EXTEND.md.
  • Only n=1 supported (Codex image_gen returns one image per call); --size, --imageSize, --quality, and --imageApiDialect are ignored or rejected.
  • Typically 5–10× slower than direct OpenAI / Google API calls (except on cache hits). Tune via BAOYU_CODEX_IMAGEGEN_TIMEOUT_MS, BAOYU_CODEX_IMAGEGEN_RETRIES, and BAOYU_CODEX_IMAGEGEN_CACHE_DIR.

Batch Mode

# Batch from saved prompt files
${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json

# Batch with explicit worker count
${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json --jobs 4 --json

Batch File Format

{
  "jobs": 4,
  "tasks": [
    {
      "id": "hero",
      "promptFiles": ["prompts/hero.md"],
      "image": "out/hero.png",
      "provider": "replicate",
      "model": "google/nano-banana-2",
      "ar": "16:9",
      "quality": "2k"
    },
    {
      "id": "diagram",
      "promptFiles": ["prompts/diagram.md"],
      "image": "out/diagram.png",
      "ref": ["references/original.png"]
    }
  ]
}

Paths in promptFiles, image, and ref are resolved relative to the batch file's directory. jobs is optional (overridden by CLI --jobs). A top-level array without the jobs wrapper is also accepted.

Source: SKILL.md on GitHub

1 alert16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill facilitates official API-based image generation across various known providers (OpenAI, Google, Azure, OpenRouter, DashScope, Z.AI, MiniMax, Replicate, Agnes). Analysis confirmed no malicious operations, prompt injections, or dynamic code execution flaws.

  • Socket16d

    2 alerts: gptAnomaly, gptSecurity

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    2/9 files flagged

  • ZeroLeaks5mo

    1 finding · Score: 86/100

Signed by skilld at 1567581. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 3 weeks ago
version
2.2.0
Other metadata
metadata
{
  "openclaw": {
    "homepage": "https://github.com/JimLiu/baoyu-skills#baoyu-image-gen",
    "requires": {
      "anyBins": [
        "bun",
        "npx"
      ]
    }
  }
}
  • API
  • image-generation
  • openai
  • google
  • azure
  • replicate
  • text-to-image
  • batch-generation
  • multimodal

README badge

README badge for jimliu/baoyu-skills/baoyu-image-gen

Generates images via OpenAI, Google, Azure OpenAI, OpenRouter, DashScope, Replicate, and 5+ other APIs. Supports text-to-image, reference images, aspect ratios, and batch generation with configurable worker concurrency. Routes through a provider-agnostic CLI that requires Bun or Node.js and API credentials per provider.

Generated from the current SKILL.md.

Which image generation APIs does this skill support?
OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream, Replicate, Codex CLI, and Agnes.
Does this skill support reference images?
Yes, but support varies by provider. Google multimodal, OpenAI GPT Image edits, Azure OpenAI edits, OpenRouter multimodal, Replicate, MiniMax, and Seedream 4.0+ support references. Jimeng, Seedream 3.0, and most DashScope models do not.
Can I generate multiple images at once?
Yes. Use batch mode with `--batchfile` and `--jobs` for parallel generation, or the `--n` option for single-call multi-image requests (Replicate requires `--n 1`).
Do I need an OpenAI API key to use this skill?
Only if you use the `openai` provider or have it as your default. Other providers require their own API keys. If using Codex CLI without an OpenAI key, use `--provider codex-cli` instead.
Does this skill require bun or npm?
Yes, the skill requires either `bun` or `npx` to run the TypeScript scripts.

Generated from the current SKILL.md. These answers refresh after source changes.