Agent skill

Content To Video

by architectds in architectds/modeldock

Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video.

Apache-2.0Auto-check passedMedia & Creative

Install Content To Video

skills CLI
$ npx skills add architectds/modeldock --skill content-to-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install architectds/modeldock content-to-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/architectds/modeldock.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/content-to-video .claude/skills/content-to-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
content-to-video
GitHub stars
117
Token cost
~2.4k tokens
SKILL.md length
1,081 words
Files
18 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video.

  • Works in 6 steps: CLASSIFY - Read… → PLAN - Write the one-line message,… → SCRIPT + NARRATION - Narration-first,… → …
  • The user asks to make a video
  • SKILL.md covers Pipeline, Classification quick matrix, Choose a rendering backend and Bundled machinery (read these…, plus 4 more sections
  • Runs JavaScript and Python scripts from its folder; calls node

What it does

Content To Video is an agent skill from architectds/modeldock. Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video. Classifies the content type and user intent, picks the right pipeline, format preset (16:9 / 9:16 / 1:1), pacing, visual strategy, and TTS voice, then runs production end-to-end (message + storyboard with a per-shot tech-stack decision, narration with measured durations, HTML/three.js or HyperFrames scenes, image-gen…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/beat-sync.md` and `references/classification.md`).

It sits in Media & Creative, covering Text to speech and voice, Motion graphics and Transcription. It works with HeyGen, Three.js and FFmpeg. The repository describes itself as: ModelDock makes the AI tools you already own work together. The licence is Apache-2.0.

When your agent uses it

  • The user asks to make a video
  • Turn this into a video
  • Produce a promo/ explainer/tutorial/story/social short/data video from any content
  • Wants automatically-produced higher-quality video without specifying the full production plan

Example prompts

  • “make a video”
  • “turn this into a video”
  • “produce a promo/ explainer/tutorial/story/social short/data video”
  • “/content-to-video”

Requirements

  • Python 3
  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. CLASSIFY - Read references/classification.md. Run
  2. PLAN - Write the one-line message, audience, distribution, CTA.
  3. SCRIPT + NARRATION - Narration-first, one line per shot, into a
  4. SCENES + ASSETS - Build scenes per the pipeline's visual strategy and
  5. SOUND - After picture locks: pick BGM per sound-design.md (audition
  6. RENDER + ASSEMBLE + QA - Preview 3 frames per scene, render clips,

What it can do on your machine

Read from SKILL.md and the folder at commit e34a709. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (JavaScript and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Content To Video loads about 2.4k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 256 tokens; SKILL.md has 1,081 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~256
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from architectds/modeldock at commit e34a709, republished under its Apache-2.0 licence (© architectds). 1,081 words, ~2,432 tokens.

Download SKILL.mdSave it as .claude/skills/content-to-video/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
content-to-video
description
Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video. Classifies the content type and user intent, picks the right pipeline, format preset (16:9 / 9:16 / 1:1), pacing, visual strategy, and TTS voice, then runs production end-to-end (message + storyboard with a per-shot tech-stack decision, narration with measured durations, HTML/three.js or HyperFrames scenes, image-gen atmosphere and transparent sprites, asset strategy, ffmpeg assembly with subtitles, and single-frame vision QA with quality gates). Use when the user asks to "make a video", "turn this into a video", or "produce a promo/ explainer/tutorial/story/social short/data video" from any content, or wants automatically-produced higher-quality video without specifying the full production plan. Supersedes the former promo-film skill; HyperFrames (HTML-to-MP4) is an optional second rendering backend.

Content to Video

Produce a finished, high-quality MP4 video from arbitrary source content. This skill is the router and orchestrator: classify -> plan -> produce -> QA. It bundles the full promo-film machinery (three.js/HTML scene contract, render commands, ffmpeg assembly, TTS, i18n) and adds per-content-type presets, adaptations (screen capture, data-viz, vertical format), an optional HyperFrames rendering backend, and stricter quality gates.

Pipeline

  1. CLASSIFY - Read references/classification.md. Run node scripts/classify.mjs <source> for a fast first pass. Confirm content_type, format, duration, pacing, and visual strategy (intent beats signals; follow confidence rules in classification.md).
  2. PLAN - Write the one-line message, audience, distribution, CTA. Design the storyboard with image-gen (planning reference only - panels are never final frames). Structure the film per the chosen pipeline in references/pipelines.md. For EVERY storyboard panel, decide and record the shot's tech stack - primary technique + layers - using references/tech-stack.md ("shot 3: three.js core + GSAP title + sprite embers"). A panel with no tech stack is a Q1 failure. Pass gate Q1.
  3. SCRIPT + NARRATION - Narration-first, one line per shot, into a narration JSON. Synthesize with TTS, measure real durations, write them back. PAD > FADE so speech is never cut (see references/pipeline.md).
  4. SCENES + ASSETS - Build scenes per the pipeline's visual strategy and the chosen rendering backend (below). Follow the identity rules in references/methodology.md: labeled real UI + 3D forms + text; image-gen only for atmosphere and planning panels.
  5. SOUND - After picture locks: pick BGM per sound-design.md (audition inside the cut, bed ~0.34), pin SFX with a declarative relative-frame table (genre vocabulary: whoosh/impact/riser/sparkle/transition; no game-pack timbres). If the user chose a strong-beat BGM, cut the timeline in beatF(n) per beat-sync.md and verify <= 3f error after render.
  6. RENDER + ASSEMBLE + QA - Preview 3 frames per scene, render clips, assemble with ffmpeg (Ken Burns, crossfades, audio mix with loudnorm, ASS subs) or render a HyperFrames composition, deliver BGM + no-BGM versions, then run Q2 per-shot vision checks and Q3 watch-through. Repeat per language from one parameterized builder.

Classification quick matrix

Source looks likeContent typePipelineFormat
README / product / launchpromopromo16:9
article / "what is X"explainerexplainer16:9
how-to / steps / codetutorialtutorial16:9
novel / script / dialoguestorystory16:9
.pptx / .key / .pdf deckslidesslides-to-video16:9
report / metrics / chartsdatadata-story16:9 or 1:1
listicle / hacks / hashtagssocialsocial-vertical9:16
audio / transcriptpodcastpodcast-video16:9

Full decision tree, presets, and overrides: references/classification.md. Per-pipeline structures and visual strategies: references/pipelines.md.

Choose a rendering backend

Two scene-authoring backends; pick per pipeline and content:

  • Bundled machinery (default): three.js/HTML scenes with a deterministic frame(t) contract + build_film.py assembly. Best for cinematic 3D, real-UI heroes, particles. Promo/explainer/story/tutorial/slides default here.
  • HyperFrames: HTML + GSAP timelines rendered by the hyperframes CLI. Best for kinetic motion graphics, vertical social, data charts, and caption/overlay-heavy pieces. See references/hyperframes.md (verified commands, composition contract, --variables/--batch for language variants). The official HyperFrames skills are installed locally (hyperframes, hyperframes-core, hyperframes-animation, hyperframes-cli, ...) - delegate composition authoring to them via references/hyperframes.md.

Both share the same upstream (classification, narration JSON with measured durations, Q1-Q3 gates) - only scene authoring and the render command change.

Bundled machinery (read these first)

  • references/methodology.md - the full method: decisions, failure modes, asset sourcing, checklists. Read it before starting any film.
  • references/pipeline.md - scene contract, render/build commands, timing math (FADE/PAD/TAIL), ffmpeg filtergraph, TTS, i18n.
  • references/tech-stack.md - per-shot tech-stack catalog and decision rule; read it at storyboard time, panel by panel.
  • references/sprites.md - image-gen sprite playbook (chroma-key, slicing, atlases, manifests, frame-sequence puppets); read before any sprite layer.
  • references/sound-design.md - BGM + SFX + mix: genre-based vocabulary, declarative relative-frame SFX table, volume math, two-version delivery. Read at sound stage; picture locks before sound.
  • references/beat-sync.md - when the user picks a strong-beat BGM: analyze the grid first (librosa), write the timeline in beatF(n), verify cut errors <= 3f on the rendered cut.

Scene contract: every scene is a standalone HTML page at 1920x1080 exposing window.__modeldock = { frame(t), duration, frameAsync?(t) } and __ready. Render environment: headless MS Edge via Playwright (SwiftShader flags), local static server with HTTP Range (206). Scripts (adapt config, do not edit the engine):

  • scripts/classify.mjs - classify source content -> production recommendation.
  • scripts/build_film.py - ffmpeg assembly (Ken Burns, xfade, audio mix, ASS subs, --lang).
  • scripts/render-clip.mjs - render an animated scene to an mp4 clip at 25 fps.
  • scripts/preview-scenes.mjs - sample 0.3/0.55/0.8 frames per scene, capture page/console errors, write 3-up strips.
  • scripts/qa-frames.mjs - extract per-shot and mid-fade frames for QA.
  • scripts/static-server-range.mjs - minimal static server with Range support.
Show full SKILL.md (377 more words)Show less

Non-negotiable rules

  • Narration drives the timeline; measure real TTS durations, never guess.
  • PAD > FADE (e.g. 1.5s vs 1.0s) so speech finishes before the next shot.
  • Never let image-gen carry product identity: no "device" heroes. Identity = labeled real UI + clear 3D forms + text sprites. image-gen is for atmosphere backgrounds, concept stills, validated hub layouts, and transparent-background decorative sprites (smoke, embers, glow orbs, sparkles, dust, generic non-brand icons). Sprites are powerful in animation, but they never carry identity.
  • Plan-stage image-gen designs the OUTLINE and STORYBOARD (one panel per planned shot, plus 2-3 style anchors when shots share a look). Panels are planning references, never final frames.
  • Composite image layers with heavy blur + radial feather mask; never hard crop edges (they read as panels/windows).
  • Vision-QA one image per call, never contact strips (cross-frame misreads).
  • Real UI footage: object-fit contain so controls are never cropped.
  • Keep the builder parameterized by language (--lang).
  • Windows: never pass non-ASCII text through a PowerShell pipe; read UTF-8 files from disk or write with Node fs utf8 + unicode escapes.
  • A static server used for video scenes must support HTTP Range (206).
  • Add a data: favicon link to every scene so console-error checks stay clean.

Quality gates (the "better" in better video)

Three gates; do not proceed past a failing gate. QA on single frames only.

  • Q1 - Plan: one-line message is a sentence, audience + distribution stated, CTA defined, every claim has a visible proof point, storyboard reviewed, voice sample signed off.
  • Q2 - Production (per shot, single frames): labeled identity, no modem/router reading, text readable and in safe margins, no hard edges / double exposure, clean transition midpoints, footage shows all key controls.
  • Q3 - Final (watch with audio): no cut speech, audio clean, subtitles fit and match, correct aspect/fps/duration, h264 yuv420p +faststart.

Full checklists and failure handling: references/quality.md.

Language variants

One builder, per-language narration JSON and audio, shared language-neutral scenes, translated text stills and subs. With HyperFrames, per-language text can live in --variables/--batch instead of duplicated compositions. Never pass non-ASCII text through a PowerShell pipe.

QA loop

  1. Q1 before production; fix the flow while a panel costs one generation.
  2. Preview every scene: zero page/console errors (HyperFrames: check + --strict render).
  3. Q2 per shot on single frames; check transition midpoints.
  4. Q3 watch-through with audio; only then deliver.

© architectds, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts, references) in skills/content-to-video of architectds/modeldock.

  • SKILL.md
  • agents/openai.yaml
  • references/beat-sync.md
  • references/classification.md
  • references/hyperframes.md
  • references/methodology.md
  • references/pipeline.md
  • references/pipelines.md
  • references/quality.md
  • references/sound-design.md
  • references/sprites.md
  • references/tech-stack.md
  • scripts/build_film.py
  • scripts/classify.mjs
  • scripts/preview-scenes.mjs
  • scripts/qa-frames.mjs
  • scripts/render-clip.mjs
  • scripts/static-server-range.mjs

Open the folder on GitHubat commit e34a709

Compare with similar skills

Content To Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Content To Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Content To Video this skillarchitectds/modeldock117—~2.4kAutomated safety check: PassApache-2.0
Qiaomu Cutjoeseesun/qiaomu-cut-skill372—~6.8kAutomated safety check: NotesMIT
Painted Animationtuzhechen2005/opus-video-skills1371 repos~1.9kAutomated safety check: PassCustom licence
Hyperframes CLIOpenMinis/MinisSkills444—~3.5kAutomated safety check: PassMIT
Article Explainer Videowwwzhouhui/skills_collection283—~2.2kAutomated safety check: PassNone
Video TalkcraftVincentwei1021/video-talkcraft1.4k—~8.7kAutomated safety check: NotesCustom licence

Similar skills

  • Qiaomu Cut

    joeseesun/qiaomu-cut-skill

    把一句话需求转成可复现、可验收视频工程的乔木智能剪辑导演。Use when the user asks to create, plan, edit, remix, explain, narrate, subtitle, animate, composite, or render a video—including one-line requests such as “制作一个科普视频:介绍…

    372 GitHub stars~6.8k tokensUpdated 11 days ago
    Media & CreativeAuto-check: notes
  • Painted Animation

    tuzhechen2005/opus-video-skills

    Make hand-painted watercolour-and-ink cartoon videos (MP4) with code — p5.js + p5.brush rendered frame by frame in headless Chrome, encoded with ffmpeg — starring Clawd or any character.

    137 GitHub starsUsed in 1 repo~1.9k tokens
    Media & CreativeAuto-check passed
  • Hyperframes CLI

    OpenMinis/MinisSkills

    HyperFrames CLI and Minis rendering. An agent skill from OpenMinis/MinisSkills.

    444 GitHub stars~3.5k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Article Explainer Video

    wwwzhouhui/skills_collection

    把一篇技术长文/论文解读自动做成章节式解说视频(1080p, 5-8 分钟)。双主题:warm(奶油底+珊瑚红+cozy-handdrawn 透明插图,亲和感)和 midnight(深蓝黑底+琥珀金+宋体标题+executive-tech 插图,AI 科技感),storyboard 一个 theme 字段切换。每章三种 layout 混排:illustration(左文右图+Ken…

    283 GitHub stars~2.2k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Video Talkcraft

    Vincentwei1021/video-talkcraft

    终极口播视频 skill:中文口播稿 + 成品配音 → CPU 字级时间戳 → SHOTBOOK 层矩阵分镜 → Remotion 电影感成片(横屏默认/竖屏)。当用户要"做口播视频"、"解说/科普视频"、"把文案变成视频"、"给配音配画面动效"时使用。默认使用成品配音,可选 Fish Audio 从稿子合成配音与时间戳;数字人生成技术不在本 skill…

    1.4k GitHub stars~8.7k tokensUpdated 8 days ago
    Media & CreativeAuto-check: notes
  • Showtime

    FavioVazquez/showtime

    A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…

    206 GitHub stars~3k tokensUpdated yesterday
    Media & CreativeAuto-check passed

Questions about Content To Video

What does Content To Video do?

Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video. Content To Video is an agent skill from architectds/modeldock. Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video.

When should I use Content To Video?

Content To Video fits situations like: the user asks to make a video; turn this into a video; produce a promo/ explainer/tutorial/story/social short/data video from any content; wants automatically-produced higher-quality video without specifying the full production plan.

How do I install Content To Video in Claude Code?

Run `npx skills add architectds/modeldock --skill content-to-video -a claude-code`. Or copy the skill folder (skills/content-to-video in architectds/modeldock) into .claude/skills/content-to-video in your project. Claude Code loads it when a task matches its description.

How do I install Content To Video in Codex?

Run `npx skills add architectds/modeldock --skill content-to-video -a codex`. Or copy the skill folder (skills/content-to-video in architectds/modeldock) into .agents/skills/content-to-video in your project. Codex loads it when a task matches its description.

Can I use Content To Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add architectds/modeldock --skill content-to-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/content-to-video, .gemini/skills/content-to-video, .github/skills/content-to-video and .opencode/skills/content-to-video in your project.

What does Content To Video need to run?

Going by SKILL.md and its folder, Content To Video needs JavaScript and Python for the scripts in its folder and the command-line tools its instructions call (node). Our summary lists: Python 3; Node.js.

Does Content To Video access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Content To Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Content To Video use?

Content To Video is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Content To Video use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.

What are the alternatives to Content To Video?

Skills that share tags, products or a category with Content To Video: Qiaomu Cut (joeseesun/qiaomu-cut-skill, 372 stars), Painted Animation (tuzhechen2005/opus-video-skills, 137 stars), Hyperframes CLI (OpenMinis/MinisSkills, 444 stars) and Article Explainer Video (wwwzhouhui/skills_collection, 283 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Content To Video?

architectds (a GitHub user) maintains it in architectds/modeldock, which has 117 GitHub stars. The repository was last updated on October 6, 2026.

Source: architectds/modeldock on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.