Agent skill

Vox Explainer

by CK42BB in CK42BB/vox-explainer-skill

End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt.

MITAuto-check passedMedia & Creative

Install Vox Explainer

skills CLI
$ npx skills add CK42BB/vox-explainer-skill --skill vox-explainer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install CK42BB/vox-explainer-skill vox-explainer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vox-explainer
GitHub stars
110
Token cost
~2.6k tokens
SKILL.md length
1,179 words
Files
6 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt.

  • Works in 6 steps: Script → Voiceover (xAI TTS v1) → Keyframes (Seedream 5.0 Pro) → …
  • The user asks for an explainer video
  • SKILL.md covers Prerequisites, Project structure, Stage 1 — Script and Stage 2 — Voiceover (xAI TTS v1), plus 6 more sections
  • Calls ffprobe; needs ATLASCLOUD_API_KEY

What it does

Vox Explainer is an agent skill from CK42BB/vox-explainer-skill. End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt. One topic in, finished film out — script, keyframes (Seedream 5.0 Pro), animation (Gemini Omni Flash), voiceover (xAI TTS), music (MiniMax Music 2.6), all assembled locally with ffmpeg via the Atlas Cloud API. Use this skill whenever the user asks for an explainer video, a Vox-style video, a documentary short, an educational video essay, a "one prompt to video" pipeline, or wants to turn a topic/article/report into a narrated…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `references/atlas-cloud-api.md` and `references/ffmpeg-assembly.md`).

It sits in Media & Creative, covering Video production, Text to speech and voice and Educational content. It works with MiniMax and FFmpeg. The repository describes itself as: A Claude Code skill that turns one topic prompt into a finished Vox-style explainer video. Script, keyframes, animation, voiceover, music, and local assembly — end to end. The licence is MIT.

When your agent uses it

  • The user asks for an explainer video
  • A Vox-style video
  • A documentary short
  • An educational video essay

Example prompts

  • “one prompt to video”
  • “t say”
  • “/vox-explainer”

Requirements

  • Python 3
  • A credential in ATLASCLOUD_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Script
  2. Voiceover (xAI TTS v1)
  3. Keyframes (Seedream 5.0 Pro)
  4. Animation (Gemini Omni Flash)
  5. Music (MiniMax Music 2.6)
  6. Assembly (ffmpeg, local)

What it can do on your machine

Read from SKILL.md and the folder at commit 7003225. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ffprobe

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • atlascloud.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ATLASCLOUD_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vox Explainer loads about 2.6k tokens when it runs, and up to ~6.7k if it reads all its reference files. Until then it costs about 180 tokens; SKILL.md has 1,179 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~180
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from CK42BB/vox-explainer-skill at commit 7003225, republished under its MIT licence (© CK42BB). 1,179 words, ~2,634 tokens.

Download SKILL.mdSave it as .claude/skills/vox-explainer/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
vox-explainer
description
End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt. One topic in, finished film out — script, keyframes (Seedream 5.0 Pro), animation (Gemini Omni Flash), voiceover (xAI TTS), music (MiniMax Music 2.6), all assembled locally with ffmpeg via the Atlas Cloud API. Use this skill whenever the user asks for an explainer video, a Vox-style video, a documentary short, an educational video essay, a "one prompt to video" pipeline, or wants to turn a topic/article/report into a narrated video with subtitles and music. Also trigger when the user mentions Atlas Cloud video pipelines, keyframe-to-animation workflows, or automated video essays — even if they don't say "Vox."

Vox-Style Explainer Video Pipeline

Produce a complete, narrated, subtitled, scored explainer video from a single topic. The pipeline runs in six stages, each producing artifacts the next stage consumes. Voiceover is generated BEFORE animation because VO duration drives all timing decisions.

TOPIC → 1.Script → 2.Voiceover → 3.Keyframes → 4.Animation → 5.Music → 6.Assembly
         (Claude)   (xAI TTS)    (Seedream)    (Omni Flash)  (MiniMax)  (ffmpeg)

Prerequisites

  • ATLASCLOUD_API_KEY environment variable set (get one at atlascloud.ai — free credits on signup)
  • ffmpeg and ffprobe installed (with libass for subtitle burning)
  • Python 3.10+ with requests
  • ~$3–8 in API credit per 60-second film (see cost table in references/atlas-cloud-api.md)

Before making any API calls, verify current endpoint paths and model IDs against https://www.atlascloud.ai/docs — media model APIs evolve quickly. The patterns in references/atlas-cloud-api.md were verified July 2026.

Project structure

Create this layout for every film and keep it — every stage reads from and writes to it:

project/
├── brief.md            # topic, angle, target length, audience
├── script.json         # beats with narration, shot descriptions, captions
├── audio/
│   ├── vo/             # beat_01.mp3 ... beat_NN.mp3
│   ├── vo_durations.json
│   └── music.mp3
├── frames/             # keyframe_01.png ... style_anchor.png
├── clips/              # clip_01.mp4 ... (silent animated clips)
├── subs/               # captions.ass
└── final/              # film.mp4

Stage 1 — Script

You write this yourself. No API call. This is the stage that most determines quality — a great script with average visuals beats the reverse.

Read references/vox-style-guide.md (Writing section) before drafting. Core rules:

  1. Structure as beats. A beat = one narration paragraph (1–3 sentences, 8–14 seconds spoken) + one visual. A 60-second film is 6–8 beats; 3 minutes is 16–22 beats.
  2. Vox narrative arc: cold-open hook (a surprising concrete fact or question) → context ("to understand X, you have to go back to...") → escalating explanation with one clear throughline → complication or twist → resolution that reframes the hook.
  3. Voice: second person welcome ("you've probably seen..."), present tense for historical narrative, concrete numbers over abstractions, short declarative sentences. No throat-clearing, no "in this video we will."
  4. One idea per beat. If a sentence introduces a second concept, split the beat.

Write script.json:

json
{
  "title": "Tang Golden Age",
  "topic": "How Tang Dynasty China built the world's largest city",
  "style_seed": "mixed-media paper collage, ink-wash Chinese mountains, vermillion red and aged cream palette",
  "accent_color": "#C0392B",
  "beats": [
    {
      "id": 1,
      "narration": "In the seventh century, Tang China built Chang'an — the largest, richest city on Earth. The whole world came to its gates.",
      "visual": "Emperor Taizong enthroned at center as paper cutout, court figures flanking, pagodas and ink-wash mountains behind, red seal stamp upper right reading 盛唐",
      "caption_text": null,
      "motion": "slow push-in on emperor, clouds drift left, subtle parallax between cutout layers"
    }
  ]
}

caption_text is for on-screen kinetic typography moments (a key stat or term) — most beats leave it null; narration subtitles are handled at assembly. motion becomes the animation prompt in Stage 4.

Stage 2 — Voiceover (xAI TTS v1)

Generate one audio file per beat, then measure durations. These durations are the master clock for the whole film.

Call the Atlas Cloud audio endpoint per beat (model: xAI TTS v1 — see references/atlas-cloud-api.md for the exact request shape and voice selection guidance). Pick ONE voice for the whole film. Vox register: measured, warm, slightly wry — avoid "movie trailer" voices.

After generation, probe every file and write audio/vo_durations.json:

bash
ffprobe -v error -show_entries format=duration -of csv=p=0 audio/vo/beat_01.mp3
json
{"beats": [{"id": 1, "file": "audio/vo/beat_01.mp3", "duration": 9.83}], "total": 61.2}

Sanity-check total length against the brief. If a beat runs long, tighten the narration and regenerate that beat only.

Stage 3 — Keyframes (Seedream 5.0 Pro)

One keyframe per beat, all in a consistent visual style. Consistency is the hard problem; solve it with the style anchor pattern:

  1. Generate the anchor. Use bytedance/seedream-v5.0-pro/text-to-image for beat 1 (usually the title/hook frame). Build the prompt from the style block in references/vox-style-guide.md (Visual Grammar section) + the beat's visual + the film's style_seed and accent_color. Request 16:9, 2K.
  2. Review the anchor before proceeding. Regenerate until the style is right — every other frame inherits it.
  3. Generate remaining frames with the edit model. Use bytedance/seedream-v5.0-pro/edit with the anchor as a reference image (up to 10 refs supported; anchor + optionally the previous frame). Prompt: "Keep the exact art style, palette, paper-collage treatment, and border framing of the reference. New scene: {beat.visual}"

This locks palette, texture, and framing across the film the way a human art director would.

Every prompt should end with the style suffix (see style guide), which encodes the Vox look: paper cutouts with white borders, halftone dots, tape strips, bold flat geometric accents, generous margins, single accent color.

Frames with caption_text set: instruct Seedream to render the text in a bold condensed sans, since Seedream 5.0 Pro's typography rendering is strong. Keep it under 5 words per frame.

Stage 4 — Animation (Gemini Omni Flash)

Animate each keyframe into a clip via image-to-video. Target clip duration = that beat's VO duration + 0.5s of breathing room (round up to the model's supported increments).

Motion prompts for the collage aesthetic should be SUBTLE — this is the most common failure mode. The Vox look is "motion graphics," not "footage." Good motion vocabulary:

  • slow push-in / pull-out (2D camera, not 3D dolly)
  • parallax drift between cutout layers
  • clouds/smoke drifting laterally
  • paper elements sliding in from frame edge and settling
  • halftone dots or texture shimmering gently
  • a single element animating (flag waving, water rippling) while everything else holds

Explicitly forbid in every prompt: "no camera shake, no 3D rotation, no morphing of faces or text, no style drift, elements remain flat paper cutouts."

Poll each generation task to completion and download to clips/clip_NN.mp4. Generations fail or drift sometimes — review each clip; regenerate any where text warps or the style breaks. Budget for ~15% regeneration.

Show full SKILL.md (424 more words)Show less

Stage 5 — Music (MiniMax Music 2.6)

One instrumental bed for the whole film. Request is_instrumental: true — vocals fight the narration.

Prompt formula: {mood} {genre-adjacent texture}, {tempo}, {instrumentation}, {arc}. Example for a history piece: "contemplative cinematic underscore, felt piano and soft strings with light percussion pulse, 90bpm, builds gradually from sparse to full, documentary style". Match instrumentation to subject (guzheng/dizi textures for the Tang piece; analog synth pulse for a tech topic).

Music 2.6 generates fixed-length songs; if the track is shorter than the film, loop it at assembly with a crossfade; if longer, trim with fade-out. Never let the music arc fight the film arc — a mid-track drop landing on a quiet beat is jarring, so audition the result against the film's shape.

Stage 6 — Assembly (ffmpeg, local)

Read references/ffmpeg-assembly.md for full recipes. The sequence:

  1. Conform each clip to its beat's VO duration (trim or hold last frame), normalize to 1920x1080 30fps.
  2. Concat clips in order.
  3. Build the VO track by concatenating beat audio with 0.5s gaps matching the video timing.
  4. Generate subs/captions.ass from the script — Vox-style burned-in subtitles: bold sans, white with heavy black outline, bottom-centered, max 2 lines, split narration at clause boundaries and time each chunk within its beat's window.
  5. Mix audio: VO at full level, music ducked -12 to -15dB under narration (sidechain compression or a simple volume automation — recipes in the reference).
  6. Burn subtitles + attribution ("Made with ..." small lower-right if desired) and encode: H.264, CRF 18, AAC 192k, -movflags +faststart.

Final QC pass before delivering: watch for (a) subtitle/VO sync drift, (b) music overpowering narration, (c) a clip whose motion loops visibly, (d) style drift between frames. Fix at the stage where the problem originated, not with assembly hacks.

Orchestration notes

  • Run stages strictly in order the first time; after that, any stage can be re-run in isolation because all state lives in the project files.
  • API generation is slow (video clips: 1–4 min each). Submit all animation tasks concurrently, then poll — don't serialize.
  • Keep every prompt you send in a prompts.log file — reproducibility matters and users will want to tweak-and-regenerate single beats.
  • Cost scales linearly with beats. Quote the user a rough cost before Stage 3 (the first paid-heavy stage): beats × ($0.10 image + ~$1.20 per 10s clip) + ~$0.50 audio.

References

  • references/vox-style-guide.md — read before Stage 1 and Stage 3. The visual grammar, writing voice, and prompt templates.
  • references/atlas-cloud-api.md — read before Stage 2. Endpoints, model IDs, async task pattern, polling code, costs.
  • references/ffmpeg-assembly.md — read before Stage 6. Conform, concat, subtitle, ducking, and encode recipes.

© CK42BB, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in the repository root of CK42BB/vox-explainer-skill.

  • SKILL.md
  • LICENSE
  • README.md
  • references/atlas-cloud-api.md
  • references/ffmpeg-assembly.md
  • references/vox-style-guide.md

Open the folder on GitHubat commit 7003225

Compare with similar skills

Vox Explainer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vox Explainer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vox Explainer this skillCK42BB/vox-explainer-skill110—~2.6kAutomated safety check: PassMIT
Videohub Story Editorcacity/VideoHub167—~2.3kAutomated safety check: NotesMIT
Explain Videolimin112/min-skill359—~2.4kAutomated safety check: PassNone
Qiaomu Cutjoeseesun/qiaomu-cut-skill369—~6.8kAutomated safety check: NotesMIT
Video Productionspeechlab0210/video-production-skill105—~4.1kAutomated safety check: NotesMIT
ShowtimeFavioVazquez/showtime158—~3kAutomated safety check: PassMIT

Similar skills

  • Videohub Story Editor

    cacity/VideoHub

    把长视频或已有字幕转成有完整叙事的几分钟短片。先基于原文字幕和画面证据理解、选段与重排,再对最终时间轴重新翻译和可选润色;既可输出保留原声的双语字幕版,也可把原声降到 30% 并用 MiniMax 或豆包 TTS 生成影视解说、短剧混剪、播客串讲或知识解读版。已有项目可进入本地五轨时间线继续调整切点、旁白、原声窗口、字幕、音量和转场,并按修订版本渲染。用于“把长视频讲成短故事”“按字幕自动剪辑”…

    167 GitHub stars~2.3k tokensUpdated 5 days ago
    Media & CreativeAuto-check: notes
  • Explain Video

    limin112/min-skill

    Build a narrated explainer video from a concept — discussion → structure → HTML slide deck → narration script → TTS voice → subtitles → background music → Playwright screen recording → ffmpeg…

    359 GitHub stars~2.4k tokensUpdated 14 days ago
    Media & CreativeAuto-check passed
  • Qiaomu Cut

    joeseesun/qiaomu-cut-skill

    把一句话需求转成可复现、可验收视频工程的乔木智能剪辑导演。Use when the user asks to create, plan, edit, remix, explain, narrate, subtitle, animate, composite, or render a video—including one-line requests such as “制作一个科普视频:介绍…

    369 GitHub stars~6.8k tokensUpdated 9 days ago
    Media & CreativeAuto-check: notes
  • Video Production

    speechlab0210/video-production-skill

    AI educational video production pipeline. An agent skill from speechlab0210/video-production-skill.

    105 GitHub stars~4.1k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Showtime

    FavioVazquez/showtime

    A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…

    158 GitHub stars~3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Content To Video

    architectds/modeldock

    Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video.

    117 GitHub stars~2.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed

Works with

Questions about Vox Explainer

What does Vox Explainer do?

End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt. Vox Explainer is an agent skill from CK42BB/vox-explainer-skill. End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt.

When should I use Vox Explainer?

Vox Explainer fits situations like: the user asks for an explainer video; A Vox-style video; A documentary short; an educational video essay.

How do I install Vox Explainer in Claude Code?

Run `npx skills add CK42BB/vox-explainer-skill --skill vox-explainer -a claude-code`. Or copy the skill folder (the CK42BB/vox-explainer-skill repository) into .claude/skills/vox-explainer in your project. Claude Code loads it when a task matches its description.

How do I install Vox Explainer in Codex?

Run `npx skills add CK42BB/vox-explainer-skill --skill vox-explainer -a codex`. Or copy the skill folder (the CK42BB/vox-explainer-skill repository) into .agents/skills/vox-explainer in your project. Codex loads it when a task matches its description.

Can I use Vox Explainer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add CK42BB/vox-explainer-skill --skill vox-explainer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vox-explainer, .gemini/skills/vox-explainer, .github/skills/vox-explainer and .opencode/skills/vox-explainer in your project.

What does Vox Explainer need to run?

Going by SKILL.md and its folder, Vox Explainer needs the command-line tools its instructions call (ffprobe) and credentials named ATLASCLOUD_API_KEY. Our summary lists: Python 3; A credential in ATLASCLOUD_API_KEY.

Does Vox Explainer access the network?

SKILL.md names 1 domain. As links in the text: atlascloud.ai. This is read from the text; nothing was executed.

Is Vox Explainer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vox Explainer use?

Vox Explainer is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vox Explainer use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Vox Explainer?

Skills that share tags, products or a category with Vox Explainer: Videohub Story Editor (cacity/VideoHub, 167 stars), Explain Video (limin112/min-skill, 359 stars), Qiaomu Cut (joeseesun/qiaomu-cut-skill, 369 stars) and Video Production (speechlab0210/video-production-skill, 105 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vox Explainer?

CK42BB (a GitHub user) maintains it in CK42BB/vox-explainer-skill, which has 110 GitHub stars. The repository was last updated on July 11, 2026.

Source: CK42BB/vox-explainer-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.