学习教育
#agent
academic-figure-generation
Generates publication-quality academic figures (framework diagrams, pipeline illustrations, system architectures, method overviews) from a paper's method text and a target caption, using a local PaperBanana multi-agent pipeline (Retriever → Planner → Stylist → Visualizer → Critic).
DeepseekModel
官方收录技能
质量 优秀 · 78
v1.0.0
获取
https://deepseekmodel.com/api/download.php?id=jxtse-scientific-research-skills-skills-academic-figure-generation-skill-md&format=skill
下载 .skill
标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用
.skill 文件中 system_prompt 字段的实际内容。
name academic-figure-generation description Generates publication-quality academic figures (framework diagrams, pipeline illustrations, system architectures, method overviews) from a paper's method text and a target caption, using a local PaperBanana multi-agent pipeline (Retriever → Planner → Stylist → Visualizer → Critic). Academic Figure Generation Thin CLI wrapper around PaperBanana (a.k.a. PaperVizAgent), a multi-agent figure-generation pipeline for academic papers. The skill provides exactly one script: scripts/generate.py . It feeds your method text + caption into PaperBanana and writes N candidate PNGs. Model selection and API keys come from PaperBanana's own configs/model_config.yaml — the wrapper does not override them. One-time setup Clone PaperBanana somewhere convenient: git clone https://github.com/dwzhu-pku/PaperBanana.git ~/PaperBanana cd ~/PaperBanana uv venv && uv pip install -r requirements.txt Configure configs/model_config.yaml — set the image model and the matching API key. Two common setups: defaults: image_model_name: "gemini-3-pro-image-preview" # or "openai/gpt-5.4-image-2" model_name: "gemini-3.1-pro-preview" # text model for Planner/Stylist/Critic api_keys: google_api_key: "..." # required for Gemini models openrouter_api_key: "" # required for openai/gpt-5.4-image-2 Use Gemini if you have a Google AI key; use GPT-Image-2 via OpenRouter if you have an OpenRouter key. Pick one — there's nothing else to wire up. Workflow Step 1: Gather inputs You need: Method text : the relevant section of the paper describing the approach ( ./method.md or ./method.tex ). Figure caption : the target caption, e.g. "Figure 1: Overview of our framework" . If the user only gives a vague request, ask: What aspect of the method should the figure focus on? Style? (block diagram, flowchart, pipeline, architecture, comparison) Venue / column width? (ACL ≤ 7.5", NeurIPS single-column 5.5") Step 2: Generate ~/PaperBanana/.venv/bin/python scripts/generate.py \ --paperbanana-root ~/PaperBanana \ --method-file ./method.md \ --caption "Figure 1: Overview of our framework" \ --out-dir ./figures/v1 \ --candidates 3 \ --aspect-ratio 16:9 Flag Default Notes --paperbanana-root (required) Path to your PaperBanana checkout --method-file (required) Method section as a text/markdown file --caption (required) Target figure caption --out-dir (required) Where PNGs land --candidates 3 Independent diagram candidates --max-concurrent 2 Cap concurrent runs (be gentle on quota) --exp-mode demo_full Full pipeline (Planner+Stylist+Visualizer+Critic). Use demo_planner_critic to skip Stylist, or vanilla for single-shot. --aspect-ratio 16:9 One of 21:9 , 16:9 , 3:2 , 1:1 --max-critic-rounds 2 Critique → revise loops (early-exits if critic says "No changes needed") Step 3: Present & iterate Show all candidates to the user. Common refinements: color scheme, layout, label text, font size. Re-run with a tweaked caption or more candidates. Step 4: Export PNGs are written as candidate_0.png , candidate_1.png , … in --out-dir . For camera-ready PDFs: magick candidate_0.png candidate_0.pdf . Style guidelines Color : consistent, colorblind-friendly palette Fonts : match the paper's body font (Times for ACL/EMNLP, Helvetica/Arial for many ML venues) Labels : concise; no full sentences inside the diagram Arrows : solid for data flow, dashed for optional / feedback loops Whitespace : don't overcrowd — reviewers skim figures in seconds Common figure types Type When to use Key elements Pipeline / Flowchart Sequential processing Boxes + arrows, L→R or T→B Architecture System overview Nested boxes, clear module boundaries Comparison Before/after, baseline vs proposed Side-by-side panels Ablation Component contributions Bar charts, highlighted rows Framework High-level conceptual overview Abstract shapes, minimal detail Troubleshooting 429 RESOURCE_EXHAUSTED on Gemini : monthly Google AI Studio spending cap hit. Raise it at https://ai.studio/spend or switch image_model_name to openai/gpt-5.4-image-2 and set OPENROUTER_API_KEY . OpenRouter Client not initialized : OPENROUTER_API_KEY not in env and openrouter_api_key not in yaml. No PNGs in output dir : check out_dir/results.json for the raw per-candidate response and any error messages. Long latency (>5 min) : most wall time is the image model. Lower --candidates or use --exp-mode vanilla for faster iteration. Links PaperBanana repo: https://github.com/dwzhu-pku/PaperBanana PaperVizAgent (Google Research version of the same project): https://github.com/google-research/papervizagent
Agent 识别该技能的关键词,点击任意一个即可复制。
该技能未提供触发词。
下载的 .skill 包内含以下字段。
| 字段 | 说明 |
|---|---|
| format | 格式标识(skill/v1) |
| skill_id | 技能唯一 ID |
| name | 技能名称 |
| version | 版本号 |
| description | 技能描述 |
| category | 所属分类(数组) |
| trigger_words | 触发词列表 |
| tags | 标签列表 |
| source | 来源标识 |
| source_url | 来源链接(本页地址) |
| exported_at | 导出时间(每次下载生成) |
| system_prompt | 系统提示词正文 |
| model_config | 模型参数:provider / model / temperature / max_tokens / top_p |
| examples | 示例 |
| install_guide | 各平台导入说明(Coze / Dify / Claude / 自定义框架) |