Skills Plugins MCP Prompt Model 博客 我的中心
趣味 #ai#ai-video#claude-code#claude-skill#collage-video

vox-director

Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.

Alisa0808 @Alisa0808 ⬇ 0 ★ 1,900 Python

安装

dsh plugin add github:Alisa0808/vox-director
下载安装清单

需要可复现安装时,可在仓库后追加 #commit 固定提交。

Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.

该插件未提供要点说明,请参考仓库 README。

aiai-videoclaude-codeclaude-skillcollage-video
  1. 安装并启动 DeepSeek Harness:npx @deepseek-ai/dsh web
  2. 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
  3. 用 dsh plugins list 确认已安装,必要时重启 Harness 生效

插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。

代码仓库github.com/Alisa0808/vox-director
许可证MIT
主要语言Python
下载量0
GitHub 星标1,900
最近推送2026-08-11
收录日期2026-09-19
分类趣味

事实信息来自公开插件目录快照(2026-09-19),介绍文案由本站再加工。

以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。

English · 简体中文

🎬 Vox Director

Turn one topic into a finished Vox-style paper-collage explainer / ad video — script, collage keyframes, motion, voice-over, music and captions, all automated.

An agent skill that runs end to end on the Atlas Cloud API + local ffmpeg, usable by any coding agent (Claude Code, Codex, etc.). You give it a one-line topic; it gives you an mp4.

[图片: License: MIT]
 [图片: Powered by Atlas Cloud]
 [图片: Agent Skill]

https://github.com/user-attachments/assets/ed08d230-7bcb-4b48-a17d-23c079208f9f

▶ "The evolution of Chinese civilization" · 30s

    [图片: How football conquered the world]

    [图片: Mexican street food]

    [图片: A brief history of money]

    [图片: A brief history of Silicon Valley]

    Football history · 60s
    Mexican street food · 60s
    A brief history of money · 60s
    Silicon Valley history · 60s

▶ more films — click any thumbnail to play

What it is

The look is the modern editorial paper-collage popularized by Vox explainers: hand-cut paper cut-outs, torn edges, tape, halftone dots, newspaper clippings, bold flat color per beat, big cut-out headlines — brought to life with motion, a narrator, music and captions.

How it works

One topic flows through one script per stage, all driven by a single beats.json per project:

topic
  │
  ├─ 1. beat map        pick a narrative arc → write beats.json      ◀── GATE 1: you approve the beat map
  ├─ 2. style bake-off  render the same beat in 3–4 themes           ◀── GATE 2: you pick the look by eye
  ├─ 3. keyframes       one collage poster per beat  (nano-banana-2)
  ├─ 4. motion          animate each poster          (gemini-omni-flash i2v)
  ├─ 5. voice + music   one narrator (xai/tts) + BGM (minimax/music)
  ├─ 6. assemble        ffmpeg: concat, duck music under VO, burn captions + watermark
  └─ final.mp4

That flow is B-roll — a topic in, everything generated. Two more input modalities reuse the same engine:

A-roll — you already have a talking-head video. It is ASR-segmented into beats and re-styled into the collage look, keeping the real face, lip-sync and gestures frame-for-frame (gemini-omni-flash/video-edit, auto-retrying on seedance-2.0/reference-to-video).

C-roll — you have one still photo (a selfie, a product shot). The subject is cut out as a photographic sticker — never redrawn — and each beat's poster is generated around it (nano-banana-2/edit). The narration can be cloned into the subject's own voice.

Two ideas make or break the result, and the skill is built around both:

The look is born in the image step. Each beat is a finished collage poster. All the collage DNA (torn paper, cut-outs, halftone, headline text) lives in that image — if the poster isn't a rich collage, nothing downstream saves it.

The motion is added after. By default an AI video model animates the whole poster (the "living poster" path). For dramatic piece-by-piece assembly, an optional local keyframe engine cuts the poster into parts and drives them frame-by-frame (no content filters, pixel-exact — great for real people).

Two human decision gates keep you in control (approve the beat map; pick the style); everything else is automated.

Models (verified on Atlas Cloud)

Job
Model

Keyframe / collage poster
google/nano-banana-2/text-to-image

Animate (non-real content)
google/gemini-omni-flash/image-to-video

Animate (real people / brands)
kwaivgi/kling-video-o3-pro/image-to-video

Re-style a talking-head (A-roll)
google/gemini-omni-flash/video-edit

Anchor a photo in the collage (C-roll)
google/nano-banana-2/edit

Narration
xai/tts-v1

Narration in a real person's voice
bytedance/seed-audio-1.0 (voice cloning)

Music
minimax/music-2.6

Cut out an element (advanced path)
youchuan/v8.1/remove-background

Model IDs drift — the skill fetches the live list from GET https://api.atlascloud.ai/api/v1/models before running.

Install

This is an agent skill — it works with any coding agent that can read a workflow and run scripts (Claude Code, Codex, …). Claude Code auto-discovers it as a skill; other agents read AGENTS.md → SKILL.md.

Option A — from this repo:

git clone https://github.com/Alisa0808/vox-director.git ~/.claude/skills/vox-director

Option B — from the packaged skill: download vox-director.skill and install it via your Claude skills UI.

Then set your Atlas Cloud API key (get one at atlascloud.ai/console/api-keys):

export ATLASCLOUD_API_KEY="sk-..."

Quick start

Just ask your coding agent, with the skill installed:

"Make me a Vox-style collage video introducing Mexican street food — English, 16:9, 15 seconds."

The agent will draft a beat map for your approval, run a style bake-off for you to pick from, then generate keyframes → motion → voice → music and assemble out/<project>/final.mp4.

Requirements

A coding agent — Claude Code, Codex, or similar

Atlas Cloud API key

ffmpeg + ffprobe (brew install ffmpeg)

Python 3 with Pillow (pip install pillow) — for caption/watermark overlays

What's in the box

SKILL.md              the skill (English) — the workflow the agent follows
SKILL.zh.md           the same skill in Chinese
AGENTS.md             entry point for non-Claude agents (Codex, …)
references/           the creative engine
  prompt-guide.md       the LOOK layer — prompt structures, vocab & 9 theme presets
  beat-layer.md         14 narrative arcs + hook/pacing + shot patterns
  voices.md             xai/tts voice roster — pick a voice_id per language/tone
  models-and-gotchas.md every API / ffmpeg gotcha, already solved
  local-engine.md       the advanced element-level motion engine
scripts/              one script per pipeline stage
examples/             ready-to-run beats.json examples
assets/               the showcase film

Credits

Built by @alisaqqt — follow for more agent-skill experiments.

Inspired by the collage-ad workflows of Stav Zilber, rom1trs and Higgsfield, and by Vox's explainer visual language.

Built end to end on Atlas Cloud — one prompt, one film.

License

MIT © 2026 Alisa Qian

数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。