vox-director
Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.
Alisa0808
@Alisa0808
⬇ 0
★ 1,900
Python
安装
dsh plugin add github:Alisa0808/vox-director
需要可复现安装时,可在仓库后追加 #commit 固定提交。
Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.
该插件未提供要点说明,请参考仓库 README。
aiai-videoclaude-codeclaude-skillcollage-video
- 安装并启动 DeepSeek Harness:
npx @deepseek-ai/dsh web - 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
- 用 dsh plugins list 确认已安装,必要时重启 Harness 生效
插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。
| 代码仓库 | github.com/Alisa0808/vox-director |
| 许可证 | MIT |
| 主要语言 | Python |
| 下载量 | 0 |
| GitHub 星标 | 1,900 |
| 最近推送 | 2026-08-11 |
| 收录日期 | 2026-09-19 |
| 分类 | 趣味 |
事实信息来自公开插件目录快照(2026-09-19),介绍文案由本站再加工。
以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。
English · 简体中文
🎬 Vox Director
Turn one topic into a finished Vox-style paper-collage explainer / ad video — script, collage keyframes, motion, voice-over, music and captions, all automated.
An agent skill that runs end to end on the Atlas Cloud API + local ffmpeg, usable by any coding agent (Claude Code, Codex, etc.). You give it a one-line topic; it gives you an mp4.
[图片: License: MIT]
[图片: Powered by Atlas Cloud]
[图片: Agent Skill]
https://github.com/user-attachments/assets/ed08d230-7bcb-4b48-a17d-23c079208f9f
▶ "The evolution of Chinese civilization" · 30s
[图片: How football conquered the world]
[图片: Mexican street food]
[图片: A brief history of money]
[图片: A brief history of Silicon Valley]
Football history · 60s
Mexican street food · 60s
A brief history of money · 60s
Silicon Valley history · 60s
▶ more films — click any thumbnail to play
What it is
The look is the modern editorial paper-collage popularized by Vox explainers: hand-cut paper cut-outs, torn edges, tape, halftone dots, newspaper clippings, bold flat color per beat, big cut-out headlines — brought to life with motion, a narrator, music and captions.
How it works
One topic flows through one script per stage, all driven by a single beats.json per project:
topic
│
├─ 1. beat map pick a narrative arc → write beats.json ◀── GATE 1: you approve the beat map
├─ 2. style bake-off render the same beat in 3–4 themes ◀── GATE 2: you pick the look by eye
├─ 3. keyframes one collage poster per beat (nano-banana-2)
├─ 4. motion animate each poster (gemini-omni-flash i2v)
├─ 5. voice + music one narrator (xai/tts) + BGM (minimax/music)
├─ 6. assemble ffmpeg: concat, duck music under VO, burn captions + watermark
└─ final.mp4
That flow is B-roll — a topic in, everything generated. Two more input modalities reuse the same engine:
A-roll — you already have a talking-head video. It is ASR-segmented into beats and re-styled into the collage look, keeping the real face, lip-sync and gestures frame-for-frame (gemini-omni-flash/video-edit, auto-retrying on seedance-2.0/reference-to-video).
C-roll — you have one still photo (a selfie, a product shot). The subject is cut out as a photographic sticker — never redrawn — and each beat's poster is generated around it (nano-banana-2/edit). The narration can be cloned into the subject's own voice.
Two ideas make or break the result, and the skill is built around both:
The look is born in the image step. Each beat is a finished collage poster. All the collage DNA (torn paper, cut-outs, halftone, headline text) lives in that image — if the poster isn't a rich collage, nothing downstream saves it.
The motion is added after. By default an AI video model animates the whole poster (the "living poster" path). For dramatic piece-by-piece assembly, an optional local keyframe engine cuts the poster into parts and drives them frame-by-frame (no content filters, pixel-exact — great for real people).
Two human decision gates keep you in control (approve the beat map; pick the style); everything else is automated.
Models (verified on Atlas Cloud)
Job
Model
Keyframe / collage poster
google/nano-banana-2/text-to-image
Animate (non-real content)
google/gemini-omni-flash/image-to-video
Animate (real people / brands)
kwaivgi/kling-video-o3-pro/image-to-video
Re-style a talking-head (A-roll)
google/gemini-omni-flash/video-edit
Anchor a photo in the collage (C-roll)
google/nano-banana-2/edit
Narration
xai/tts-v1
Narration in a real person's voice
bytedance/seed-audio-1.0 (voice cloning)
Music
minimax/music-2.6
Cut out an element (advanced path)
youchuan/v8.1/remove-background
Model IDs drift — the skill fetches the live list from GET https://api.atlascloud.ai/api/v1/models before running.
Install
This is an agent skill — it works with any coding agent that can read a workflow and run scripts (Claude Code, Codex, …). Claude Code auto-discovers it as a skill; other agents read AGENTS.md → SKILL.md.
Option A — from this repo:
git clone https://github.com/Alisa0808/vox-director.git ~/.claude/skills/vox-director
Option B — from the packaged skill: download vox-director.skill and install it via your Claude skills UI.
Then set your Atlas Cloud API key (get one at atlascloud.ai/console/api-keys):
export ATLASCLOUD_API_KEY="sk-..."
Quick start
Just ask your coding agent, with the skill installed:
"Make me a Vox-style collage video introducing Mexican street food — English, 16:9, 15 seconds."
The agent will draft a beat map for your approval, run a style bake-off for you to pick from, then generate keyframes → motion → voice → music and assemble out/<project>/final.mp4.
Requirements
A coding agent — Claude Code, Codex, or similar
Atlas Cloud API key
ffmpeg + ffprobe (brew install ffmpeg)
Python 3 with Pillow (pip install pillow) — for caption/watermark overlays
What's in the box
SKILL.md the skill (English) — the workflow the agent follows
SKILL.zh.md the same skill in Chinese
AGENTS.md entry point for non-Claude agents (Codex, …)
references/ the creative engine
prompt-guide.md the LOOK layer — prompt structures, vocab & 9 theme presets
beat-layer.md 14 narrative arcs + hook/pacing + shot patterns
voices.md xai/tts voice roster — pick a voice_id per language/tone
models-and-gotchas.md every API / ffmpeg gotcha, already solved
local-engine.md the advanced element-level motion engine
scripts/ one script per pipeline stage
examples/ ready-to-run beats.json examples
assets/ the showcase film
Credits
Built by @alisaqqt — follow for more agent-skill experiments.
Inspired by the collage-ad workflows of Stav Zilber, rom1trs and Higgsfield, and by Vox's explainer visual language.
Built end to end on Atlas Cloud — one prompt, one film.
License
MIT © 2026 Alisa Qian
数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。