Skills Plugins MCP Prompt Model 博客 我的中心
内容创作 #image #api #ai

gpt-image-2

Generate and edit images using OpenAI's GPT Image 2 API. Interactive skill that guides users through image creation with style presets, cost-aware draft/final workflow, thinking mode, carousels, and photo editing. This skill should be used when the user requests image generation via OpenAI/GPT Image 2, wants to create social media carousels, edit photos into artistic styles, or needs images with readable text (infographics, diagrams, posters).

DeepseekModel 官方收录技能 质量 优秀 · 78 v1.0.0

获取

https://deepseekmodel.com/api/download.php?id=glebis-claude-skills-gpt-image-2-skill-md&format=skill
下载 .skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用
.skill 文件中 system_prompt 字段的实际内容。
name gpt-image-2 description Generate and edit images using OpenAI's GPT Image 2 API. Interactive skill that guides users through image creation with style presets, cost-aware draft/final workflow, thinking mode, carousels, and photo editing. This skill should be used when the user requests image generation via OpenAI/GPT Image 2, wants to create social media carousels, edit photos into artistic styles, or needs images with readable text (infographics, diagrams, posters). GPT Image 2 — Interactive Image Generation Generate and edit images via OpenAI's GPT Image 2 API with an interactive, guided workflow. Interactive Flow When the user invokes this skill, guide them through these steps using AskUserQuestion. Do not skip steps — the interactive flow is the core experience. Step 1: What are we making? Ask the user what they want to create. Offer these options: Single image — one image from a text prompt Photo edit — transform an existing photo into a style Carousel — 5-10 cohesive slides for LinkedIn/Instagram Variants — multiple versions of the same concept Quick generate — skip questions, just run the prompt If the user already provided a clear prompt (e.g. "generate an editorial image of a rocket"), skip to Step 3. Step 2: Style selection Show the user available presets grouped by category. Read presets.yaml and present them: Visual styles (no text in image): editorial, blueprint, ink, risograph, wireframe, constellation, brutalist, grain Text-heavy (leverages GPT Image 2 text rendering): infographic, slide, diagram, poster, menu, manga Community favorites: trading-card, pixar, app-mockup, isometric, action-figure, cinematic, panorama Reference-anchored: vhs — 1980s late-night infomercial title card: scanline-striped gradient italic caps on pure black. It auto-attaches a bundled reference image ( references/vhs-infomercial.png ), so the look stays consistent batch-to-batch. Pass the ad copy as the subject; for multi-line copy separate lines with / (e.g. --preset vhs "THEY TRUSTED YOU / NOW / PROVE IT" ). Custom — user describes their own style Ask: "Which style? Or describe your own." Step 3: Platform & sizing Ask where this will be used: YouTube thumbnail (1280×720) Instagram square (1080×1080) Slides/presentation (1920×1080) Blog hero (1200×630) X/Twitter (1600×900) Story (1080×1920) Custom size No resize (use API default) Aspect-ratio caveat: --platform does NOT change the generation size — it generates at the configured size (default 1024×1024) and resizes/stretches afterwards, which distorts non-square targets (e.g. --platform story stretches a square to 1080×1920, cropping the composition's edges). For portrait or landscape compositions, pass the API-native size directly: --size 1024x1536 (portrait) or --size 1536x1024 (landscape). Preflight false positives: the background-conflict heuristic trips on color words applied to non-background elements (e.g. "off-white text" in a dark-background prompt reads as a second background). If the flagged conflict is spurious, re-run with --force , or rephrase ("pale gray text"). Step 3.5: Preflight prompt check (automatic) Before any generation spend, the script now composes the final prompt first (preset + subject + style), then checks it for internal contradictions — most often a preset that hard-codes something the subject overrides (e.g. the editorial preset forces "on pure black background" while your subject asks for a warm off-white ground). The check prefers a fast Haiku call via the llm CLI; if Haiku is unavailable (no llm , no Anthropic credit) it falls back to the configured llm default model, then to a built-in static heuristic. The resolved prompt and the verdict are printed. If a conflict is found, generation is aborted before spending — fix the prompt or preset and re-run, or override with --force (generate anyway) or --no-preflight (skip the check). This is what prevents the "generated on the wrong background, now regenerate" waste. When composing prompts that set a background/palette, don't combine a background-fixing preset ( editorial , blueprint , etc.) with a different requested background — either drop the preset and specify the full style yourself, or accept the preset's background. Step 4: Draft first, then final Always generate a draft first unless the user says "skip draft" or uses --draft false . Generate with --draft (quality=low, ~$0.006/image) Show the image to the user using the Read tool Ask: "Like this direction? I can: (a) generate final quality, (b) adjust the prompt, (c) try a different style, (d) regenerate with a new seed" If approved, generate final with --quality high (~$0.21/image) Use --seed from the draft to maintain composition when upgrading to final This draft→final flow saves ~97% on iteration costs. Step 5: Show result and offer next actions After generation, always: Show the image using the Read tool Open it with open <path> for full-resolution preview Report the cost Offer: "Want to (a) generate variants, (b) edit this further, (c) use as reference for more images, (d) done?" Carousel Workflow When the user wants a carousel (5-10 slides): 1. Story arc Ask: "What's the story? Give me the key message and I'll draft a 10-slide arc." Then propose a slide-by-slide plan like: Slide 1: [Cover] — hook headline + hero image Slide 2: [Problem] — bold statement Slide 3: [Context] — illustration + explanation ... Slide 10: [CTA] — call to action with URL Ask the user to approve or modify the plan. 2. Style consistency Use the same preset + seed range across all slides. For carousels: Pick one visual style for all slides Use --seed to lock composition patterns Include pagination dots in prompts (e.g., "10 small dots at bottom, third dot highlighted orange") Maintain consistent color palette and typography 3. Draft batch Generate all slides as drafts first ($0.006 × 10 = $0.06 total). Show them all to the user as a contact sheet or one by one. Ask which ones to regenerate or adjust. 4. Final batch Only generate finals for approved slides. Offer to generate all at once with -y flag. Photo Edit Workflow When the user wants to transform a photo: Ask for the source image (file path or clipboard) For clipboard: save with osascript to a temp file Show available styles and ask which to try Generate a draft edit first Show result, ask if they want adjustments Generate final when approved Use --edit <path> for the API call. Cost Awareness Always communicate costs before generating: Quality Per image 10-slide carousel --draft (low) $0.006 $0.06 medium $0.05 $0.50 high (default) $0.21 $2.10 high + thinking $0.25-0.42 $2.50-4.20 Thinking mode adds 20-100% cost. Only suggest it for text-heavy or complex compositions. The script auto-confirms when cost < $0.50. Above that, it prompts the user. Prompt Engineering Tips When helping users write prompts, apply these patterns: Structure : Scene → Subject → Detail → Lighting → Constraint Front-load the subject : put the main thing first For text in images : quote exact text with single quotes: 'with the headline "Hello World"' Character consistency : maintain a 5-tuple: age + appearance + hairstyle + distinctive features + clothing Style tags at end : append tags like editorial-magazine , studio-product to converge batches Use --seed for iteration : lock composition, vary only the prompt details CLI Reference # Basic generation scripts/gpt_image_2.py "prompt" output.png # With preset and platform scripts/gpt_image_2.py --preset editorial --platform square "subject" out.png # Draft mode (~$0.006/image) scripts/gpt_image_2.py --draft "prompt" out.png # With thinking for complex layouts scripts/gpt_image_2.py --thinking medium --preset diagram "OAuth flow" out.png # Seed for reproducibility scripts/gpt_image_2.py --seed 42 "prompt" out.png # Edit existing photo scripts/gpt_image_2.py --edit photo.png "transform into constellation style" out.png # Reference-anchored preset (auto-attaches its bundled reference image) scripts/gpt_image_2.py --preset vhs --platform youtube "THEY TRUSTED YOU / NOW / PROVE IT" ad.png # Variants with contact sheet scripts/gpt_image_2.py --n 4 --preset ink "mountain" out.png # Cost estimate scripts/gpt_image_2.py --estimate --n 10 --quality high "batch test" # Skip confirmation scripts/gpt_image_2.py -y --n 10 "batch" out.png # Dry run (show prompt without API call) scripts/gpt_image_2.py --dry-run --preset editorial "test" out.png # Preflight runs automatically before spend; override if needed scripts/gpt_image_2.py --force "prompt with a known conflict" out.png # generate anyway scripts/gpt_image_2.py --no-preflight "prompt" out.png # skip the check Files scripts/gpt_image_2.py — main CLI (Python, requires PyYAML) presets.yaml — style presets (visual + text-heavy + community + reference-anchored). A preset may declare a reference: path (relative to the skill dir); it auto-attaches as a style anchor unless the user passes their own --reference . See the vhs preset. platforms.yaml — 8 platform sizing presets references/api_reference.md — full API documentation references/vhs-infomercial.png — bundled style anchor for the vhs preset ~/.config/gpt-image-2/config.yaml — user defaults ~/.config/gpt-image-2/history.jsonl — generation log ~/.config/gpt-image-2/last.json — last run (for again )
Agent 识别该技能的关键词,点击任意一个即可复制。

该技能未提供触发词。

下载的 .skill 包内含以下字段。
字段 说明
format格式标识(skill/v1)
skill_id技能唯一 ID
name技能名称
version版本号
description技能描述
category所属分类(数组)
trigger_words触发词列表
tags标签列表
source来源标识
source_url来源链接(本页地址)
exported_at导出时间(每次下载生成)
system_prompt系统提示词正文
model_config模型参数:provider / model / temperature / max_tokens / top_p
examples示例
install_guide各平台导入说明(Coze / Dify / Claude / 自定义框架)
同一份技能可按不同平台格式导出。
.skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用 下载
.skillpro 增强格式,额外含脚本 / 工具 / 依赖 / 钩子占位 下载
.json 纯 JSON 导出,只含 system_prompt 与模型参数 下载
Coze 带 frontmatter 的 Markdown,Coze 平台导入用 下载
Dify Dify DSL,创建应用后直接导入 下载

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。