Skills Plugins MCP Prompt Model 博客 我的中心
Development #image #api #ai

baoyu-image-gen

AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream, Replicate and Agnes APIs. Supports text-to-image, reference images, aspect ratios, and batch generation from saved prompt files. Sequential by default; use batch parallel generation when the user already has multiple prompts or wants stable multi-image throughput. Use when user asks to generate, create, or draw images.

DeepseekModel Curated skill Quality Excellent · 90 v1.0.0

Get

https://deepseekmodel.com/api/download.php?id=jimliu-baoyu-skills-skills-baoyu-image-gen-skill-md&format=skill
Download .skill Standard format with system_prompt and model_config, ready for any agent framework
The actual content of the system_prompt field in the .skill file.
name baoyu-image-gen description AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream, Replicate and Agnes APIs. Supports text-to-image, reference images, aspect ratios, and batch generation from saved prompt files. Sequential by default; use batch parallel generation when the user already has multiple prompts or wants stable multi-image throughput. Use when user asks to generate, create, or draw images. version 2.1.0 metadata {"openclaw":{"homepage":"https://github.com/JimLiu/baoyu-skills#baoyu-image-gen","requires":{"anyBins":"[Truncated]"}}} Image Generation (AI SDK) Official API-based image generation. Supports OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope (阿里通义万象), Z.AI GLM-Image, MiniMax, Jimeng (即梦), Seedream (豆包), Replicate and Agnes. User Input Tools When this skill prompts the user, follow this tool-selection rule (priority order): Prefer built-in user-input tools exposed by the current agent runtime — e.g., AskUserQuestion , request_user_input , clarify , ask_user , or any equivalent. Fallback : if no such tool exists, emit a numbered plain-text message and ask the user to reply with the chosen number/answer for each question. Batching : if the tool supports multiple questions per call, combine all applicable questions into a single call; if only single-question, ask them one at a time in priority order. Concrete AskUserQuestion references below are examples — substitute the local equivalent in other runtimes. Script Directory {baseDir} = this SKILL.md's directory. All scripts/... paths below are relative to {baseDir} . Main script: {baseDir}/scripts/main.ts . Batch payload helper: {baseDir}/scripts/build-batch.ts . Resolve ${BUN_X} : prefer bun ; else npx -y bun ; else suggest brew install oven-sh/bun/bun . Step 0: Load Preferences ⛔ BLOCKING This step MUST complete before any image generation — generation is blocked until EXTEND.md exists. Check these paths in order; first hit wins: Path Scope .baoyu-skills/baoyu-image-gen/EXTEND.md Project ${XDG_CONFIG_HOME:-$HOME/.config}/baoyu-skills/baoyu-image-gen/EXTEND.md XDG $HOME/.baoyu-skills/baoyu-image-gen/EXTEND.md User home Found → load, parse, apply. If default_model.[provider] is null → ask model only. Not found → run first-time setup ( references/config/first-time-setup.md ) using AskUserQuestion to collect provider + model + quality + save location. Save EXTEND.md, then continue. Do not generate images before this completes. Legacy compatibility: if .baoyu-skills/baoyu-imagine/EXTEND.md exists and the new path doesn't, the runtime renames it to baoyu-image-gen . If both exist, the runtime leaves them alone and uses the new path. EXTEND.md keys : default provider, default quality, default aspect ratio, default image size, OpenAI image API dialect, default models, batch worker cap, provider-specific batch limits. Schema: references/config/preferences-schema.md . Usage Minimum working examples — see references/usage-examples.md for the full set including per-provider invocations and batch mode. Identity-preserving reference prompts When the user wants a real person/character/object preserved from reference images, do not replace the reference with a long generic description. Prefer short, hard identity-preservation language: "Use the person/object in the reference image(s) as the same identity. Do not redesign it or create a similar-looking new subject." "Only change scene, clothing, pose, lighting, rendering style, and composition. Keep the face/proportions/hair/key accessories/overall identity from the references." If using multiple references, state that they are the same subject and should jointly define identity. Pitfall: long descriptions like "young East Asian woman, oval face, clear eyes..." can cause the model to synthesize a new person matching the description instead of preserving the referenced person. # Basic ${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image cat.png # With aspect ratio and high quality ${BUN_X} {baseDir}/scripts/main.ts --prompt "A landscape" --image out.png --ar 16:9 --quality 2k # Prompt from files ${BUN_X} {baseDir}/scripts/main.ts --promptfiles system.md content.md --image out.png # With reference image ${BUN_X} {baseDir}/scripts/main.ts --prompt "Make blue" --image out.png --ref source.png # Specific provider ${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider dashscope --model qwen-image-2.0-pro # OpenAI GPT Image 2 ${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider openai --model gpt-image-2 # Codex CLI (uses logged-in Codex subscription — no OPENAI_API_KEY required; requires `codex` on PATH) ${BUN_X} {baseDir}/scripts/main.ts --prompt "A cat" --image out.png --provider codex-cli --ar 16:9 # Batch mode ${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json -- jobs 4 # Build a batch file from outline.md + prompts/ (e.g. baoyu-article-illustrator output) ${BUN_X} {baseDir}/scripts/build-batch.ts --outline outline.md --prompts prompts --output batch.json --images-dir attachments ${BUN_X} {baseDir}/scripts/main.ts --batchfile batch.json -- jobs 4 Reference-Image Identity Preservation When the user wants a person/object preserved from reference images: Prefer a small curated set of existing source references (usually 2–4) over many images; large multi-megabyte refs can destabilize streaming providers. Make the prompt say the references are the same subject and the output must use that identity. Avoid long generic facial-feature descriptions that can cause the model to synthesize a new similar-looking person. Do not use newly generated outputs as references unless the user explicitly asks; generated refs compound drift. If results become too polished or influencer-like, reduce stylized refs and add explicit anti-beautification constraints (no face slimming, eye enlargement, heavy makeup, commercial travel shoot, over-smoothing). If the subject should look younger/older, preserve the face and express age through clothing, posture, scene, and styling; do not ask the model to change facial identity. Options Option Description --prompt <text> , -p Prompt text --promptfiles <files...> Read prompt from files (concatenated) --image <path> Output image path (required in single-image mode) --batchfile <path> JSON batch file for multi-image generation --jobs <count> Worker count for batch mode (default: auto, max from config, built-in default 10) --provider google|openai|azure|openrouter|dashscope|zai|minimax|jimeng|seedream|replicate|codex-cli|agnes Force provider (default: auto-detect; codex-cli is never auto-selected — must be pinned via CLI or EXTEND.md) --model <id> , -m Model ID — see provider references for defaults and allowed values --ar <ratio> Aspect ratio ( 16:9 , 1:1 , 4:3 , …) --size <WxH> Explicit size (e.g., 1024x1024 ; for gpt-image-2 , width/height must be multiples of 16, max edge 3840px, ratio no wider than 3:1) --quality normal|2k Quality preset (default: 2k ) --imageSize 1K|2K|4K Image size for Google/OpenRouter (default: from quality) --imageApiDialect openai-native|ratio-metadata OpenAI-compatible endpoint dialect — use ratio-metadata for gateways that expect aspect-ratio size plus metadata.resolution --ref <files...> Reference images. Supported by Google multimodal, OpenAI GPT Image edits, Azure OpenAI edits (PNG/JPG only), OpenRouter multimodal models, Replicate supported families, MiniMax subject-reference, Seedream 5.0/4.5/4.0, DashScope wan2.7-image-pro / wan2.7-image . Not supported by Jimeng, Seedream 3.0, SeedEdit 3.0, or any DashScope model outside the wan2.7-image* family --n <count> Number of images. Replicate requires --n 1 (single-output save semantics) --json JSON output Environment Variables Variable Description OPENAI_API_KEY OpenAI API key AZURE_OPENAI_API_KEY Azure OpenAI API key OPENROUTER_API_KEY OpenRouter API key GOOGLE_API_KEY Google API key DASHSCOPE_API_KEY DashScope API key ZAI_API_KEY (alias BIGMODEL_API_KEY ) Z.AI API key MINIMAX_API_KEY MiniMax API key REPLICATE_API_TOKEN Replicate API token JIMENG_ACCESS_KEY_ID , JIMENG_SECRET_ACCESS_KEY Jimeng (即梦) Volcengine credentials ARK_API_KEY Seedream (豆包) Volcengine ARK API key <PROVIDER>_IMAGE_MODEL Per-provider model override ( OPENAI_IMAGE_MODEL , GOOGLE_IMAGE_MODEL , DASHSCOPE_IMAGE_MODEL , ZAI_IMAGE_MODEL / BIGMODEL_IMAGE_MODEL , MINIMAX_IMAGE_MODEL , OPENROUTER_IMAGE_MODEL , REPLICATE_IMAGE_MODEL , JIMENG_IMAGE_MODEL , SEEDREAM_IMAGE_MODEL , AGNES_IMAGE_MODEL ) AZURE_OPENAI_DEPLOYMENT (alias AZURE_OPENAI_IMAGE_MODEL ) Azure default deployment <PROVIDER>_BASE_URL Per-provider endpoint override AZURE_API_VERSION Azure image API version (default 2025-04-01-preview ) JIMENG_REGION Jimeng region (default cn-north-1 ) OPENAI_IMAGE_API_DIALECT openai-native | ratio-metadata OPENROUTER_HTTP_REFERER , OPENROUTER_TITLE Optional OpenRouter attribution BAOYU_IMAGE_GEN_MAX_WORKERS Override batch worker cap BAOYU_IMAGE_GEN_<PROVIDER>_CONCURRENCY Per-provider concurrency (e.g., BAOYU_IMAGE_GEN_REPLICATE_CONCURRENCY ; for codex-cli use BAOYU_IMAGE_GEN_CODEX_CLI_CONCURRENCY ) BAOYU_IMAGE_GEN_<PROVIDER>_START_INTERVAL_MS Per-provider start-gap BAOYU_CODEX_IMAGEGEN_BIN Override the codex-imagegen wrapper path for the codex-cli provider (default: bundled scripts/codex-imagegen/main.ts ; accepts .ts or legacy .sh /binary) BAOYU_CODEX_IMAGEGEN_CACHE_DIR Enable idempotency cache for the codex-cli provider (off by default) BAOYU_CODEX_IMAGEGEN_TIMEOUT_MS Per-attempt codex exec timeout for the codex-cli provider (default: 300000 ms) BAOYU_CODEX_IMAGEGEN_RETRIES Wrapper-side retry attempts on retryable errors for the codex-cli provider (default: 2) BAOYU_CODEX_IMAGEGEN_LOG_FILE Append JSONL diagnostic log for the codex-cli provider Load priority : CLI args > EXTEND.md > env vars > <cwd>/.baoyu-skills/.env > ~/.baoyu-skills/.env Codex/ChatGPT OAuth is not an OpenAI API key --provider openai --model gpt-image-2 uses the standard OpenAI Images API ( /v1/images/generations or /v1/images/edits ) and requires OPENAI_API_KEY . A Codex or ChatGPT desktop login is a different entitlement and is not a drop-in replacement for OPENAI_API_KEY ; do not paste a Codex OAuth token into OPENAI_API_KEY or only set OPENAI_BASE_URL to a Codex backend. If the user wants to use their Codex subscription / GPT Image 2 entitlement without an OpenAI API key, route through a Codex-native backend instead of this skill's openai provider: In Codex runtime: use the native imagegen skill/tool. In non-Codex runtimes with codex CLI installed and logged in: use baoyu-image-gen --provider codex-cli (preferred — it gives you the same retry / cache / batch flow as every other provider). The provider spawns the bundled scripts/codex-imagegen/main.ts ; the same code lives upstream at packages/baoyu-codex-imagegen/src/main.ts for standalone callers. In Hermes runtimes with a native image_generate tool: use that tool as a fallback, and state whether reference images were passed directly or reconstructed from extracted traits. Do not modify the existing openai provider to silently consume Codex OAuth. The first-class Codex-CLI path is the dedicated codex-cli provider, which has its own auth (Codex login), route ( codex exec ), request shape, and tests. See references/codex-oauth-vs-openai-api-key.md . Model Resolution Priority (highest → lowest) applies to every provider: CLI flag --model <id> EXTEND.md default_model.[provider] Env var <PROVIDER>_IMAGE_MODEL Built-in default For OpenAI, the built-in default is gpt-image-2 . gpt-image-1.5 , gpt-image-1 , and GPT Image snapshots remain selectable with --model or OPENAI_IMAGE_MODEL . For Azure, --model / default_model.azure is the Azure deployment name. AZURE_OPENAI_DEPLOYMENT is the preferred env var; AZURE_OPENAI_IMAGE_MODEL is kept as a backward-compatible alias. If your Azure deployment is named after the underlying model, use gpt-image-2 ; otherwise use the exact custom deployment name. EXTEND.md overrides env vars: if EXTEND.md sets default_model.google: "gemini-3-pro-image" and the env var sets GOOGLE_IMAGE_MODEL=gemini-3.1-flash-image , EXTEND.md wins. Display model info before each generation : Using [provider] / [model] Switch model: --model <id> | EXTEND.md default_model.[provider] | env <PROVIDER>_IMAGE_MODEL OpenAI-Compatible Gateway Dialects provider=openai means the auth and routing entrypoint is OpenAI-compatible. It does not guarantee the upstream image API uses OpenAI native semantics. When a gateway expects a different wire format, set default_image_api_dialect in EXTEND.md, OPENAI_IMAGE_API_DIALECT , or --imageApiDialect : openai-native : pixel size ( 1536x1024 ) and native OpenAI quality fields ratio-metadata : aspect-ratio size ( 16:9 ) plus metadata.resolution ( 1K|2K|4K ) and metadata.orientation Use openai-native for the OpenAI native API or strict clones; try ratio-metadata for compatibility gateways in front of Gemini or similar models. Current limitation: ratio-metadata applies only to text-to-image; reference-image edits still need openai-native or a provider with first-class edit support. Provider-Specific Guides Each provider has its own quirks (model families, size rules, ref support, limits). Read these when the user picks that provider or asks for non-default behavior: Provider Reference DashScope (Qwen-Image families, custom sizes) references/providers/dashscope.md Z.AI (GLM-Image, cogview-4) references/providers/zai.md MiniMax (image-01, subject-reference) references/providers/minimax.md
Keywords that activate this skill. Click one to copy it.

This skill does not provide trigger words.

The downloaded .skill package contains the following fields.
Field Description
formatFormat tag (skill/v1)
skill_idUnique skill ID
nameSkill name
versionVersion
descriptionDescription
categoryCategories (array)
trigger_wordsTrigger words
tagsTags
sourceSource
source_urlSource URL (this page)
exported_atExported at (set per download)
system_promptSystem prompt body
model_configModel config: provider / model / temperature / max_tokens / top_p
examplesExamples
install_guideImport guide for Coze / Dify / Claude / custom frameworks
The same skill can be exported in different platform formats.
.skill Standard format with system_prompt and model_config, ready for any agent framework Download
.skillpro Enhanced format with scripts, tools, dependencies and hooks Download
.json Plain JSON export with system_prompt and model parameters only Download
Coze Markdown with frontmatter, for Coze platform import Download
Dify Dify DSL, import directly after creating an app Download

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。