Playwright Browser Demo Recording
digitalsamba/claude-code-video-toolkit
Records browser interactions as video with Playwright, covering viewport sizing, cursor highlighting, and converting output for Remotion.
A skill your agent uses when the user wants to create an explainer, documentary, knowledge-sharing, news-broadcast, product-introduction, or data-report video from a topic.
$ npx skills add calcuforge/explainer-video-maker --skill explainer-video-maker -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install calcuforge/explainer-video-maker explainer-video-maker --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/calcuforge/explainer-video-maker.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/explainer-video-maker .claude/skills/explainer-video-maker && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "explainer-video-maker" agent skill from https://github.com/calcuforge/explainer-video-maker/tree/main/skills/explainer-video-maker into .claude/skills/explainer-video-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explainer-video-maker", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/calcuforge/explainer-video-maker/tree/main/skills/explainer-video-makerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add calcuforge/explainer-video-maker --skill explainer-video-maker -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install calcuforge/explainer-video-maker explainer-video-maker --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calcuforge/explainer-video-maker.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/explainer-video-maker .agents/skills/explainer-video-maker && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "explainer-video-maker" agent skill from https://github.com/calcuforge/explainer-video-maker/tree/main/skills/explainer-video-maker into .agents/skills/explainer-video-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explainer-video-maker", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calcuforge/explainer-video-maker --skill explainer-video-maker -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install calcuforge/explainer-video-maker explainer-video-maker --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calcuforge/explainer-video-maker.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/explainer-video-maker .cursor/skills/explainer-video-maker && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "explainer-video-maker" agent skill from https://github.com/calcuforge/explainer-video-maker/tree/main/skills/explainer-video-maker into .cursor/skills/explainer-video-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explainer-video-maker", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/calcuforge/explainer-video-maker.git --path skills/explainer-video-maker--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add calcuforge/explainer-video-maker --skill explainer-video-maker -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install calcuforge/explainer-video-maker explainer-video-maker --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calcuforge/explainer-video-maker.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/explainer-video-maker .gemini/skills/explainer-video-maker && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "explainer-video-maker" agent skill from https://github.com/calcuforge/explainer-video-maker/tree/main/skills/explainer-video-maker into .gemini/skills/explainer-video-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explainer-video-maker", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install calcuforge/explainer-video-maker explainer-video-makerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add calcuforge/explainer-video-maker --skill explainer-video-maker -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/calcuforge/explainer-video-maker.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/explainer-video-maker .github/skills/explainer-video-maker && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "explainer-video-maker" agent skill from https://github.com/calcuforge/explainer-video-maker/tree/main/skills/explainer-video-maker into .github/skills/explainer-video-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explainer-video-maker", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calcuforge/explainer-video-maker --skill explainer-video-maker -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install calcuforge/explainer-video-maker explainer-video-maker --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calcuforge/explainer-video-maker.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/explainer-video-maker .opencode/skills/explainer-video-maker && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "explainer-video-maker" agent skill from https://github.com/calcuforge/explainer-video-maker/tree/main/skills/explainer-video-maker into .opencode/skills/explainer-video-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explainer-video-maker", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
explainer-video-makerA skill your agent uses when the user wants to create an explainer, documentary, knowledge-sharing, news-broadcast, product-introduction, or data-report video from a topic.
Explainer Video Maker is an agent skill from calcuforge/explainer-video-maker. Use when the user wants to create an explainer, documentary, knowledge-sharing, news-broadcast, product-introduction, or data-report video from a topic. Trigger keywords: "make a documentary", "make a video about", "make an explainer", "make a news video", "make a product intro", "make a knowledge video", "help me make a ... video". Also trigger for Chinese equivalents like "帮我制作一个...纪录片/视频", "做一个...解说视频". Supports auto topic selection when user only names a category. Produces video via research → struct design →…
Its SKILL.md is about 6.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 84 other files, including scripts and reference files (for example `references/expression_intent_mapping.md`, `references/natural-narration.md` and `references/search-providers.md`).
It sits in Media & Creative, covering Video production. It works with Remotion, ComfyUI and Playwright. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 3900f32. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
gitpippython3npmplaywrightbashpythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Explainer Video Maker loads about 6.5k tokens when it runs, and up to ~45k if it reads all its reference files. Until then it costs about 163 tokens; SKILL.md has 2,473 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from calcuforge/explainer-video-maker at commit 3900f32, republished under its Apache-2.0 licence (© calcuforge). 2,473 words, ~6,550 tokens.
.claude/skills/explainer-video-maker/SKILL.md (or your agent's skills folder). This skill also uses 81 other files; get the full folder from GitHub.Automated pipeline for narration-driven explainer videos from any topic. Supports documentaries, knowledge sharing, news, data reports, product introductions, and any format suitable for narrated explanation.
Audio drives visuals: each section carries exactly one narration; its audio
duration sets the narration's total frame count, which is split across the
section's 1-N scenes by each scene's percentage.
Resolve SKILL_DIR to the directory containing this SKILL.md.
SKILL_DIR="${SKILL_DIR:-${CLAUDE_SKILL_DIR}}"
python3 "${SKILL_DIR}/scripts/tool/check_prereqs.py"External dependencies:
${SKILL_DIR}/requirements.txt
(pip install -r "${SKILL_DIR}/requirements.txt"). The Playwright browser
download (playwright install chromium) is needed ONLY in standalone
environments — when PLAYWRIGHT_CDP_URL is set (e.g. the hermes-desktop
container), Playwright attaches to the shared desktop Chromium over CDP and
no download is requiredffmpeg, ffprobe on PATHnpxcomfyui-scheduler,
remotion-video-template) — installed under the workspace's dep/ directorycomfyui-scheduler (the ComfyUI CLI) and remotion-video-template (the
Remotion render backend) are separate repositories. Install them under the
workspace's dep/ directory — i.e. dep/ inside the workspace root (the
directory that contains projects/). Configs store these as workspace-relative
paths (e.g. dep/remotion-video-template); the scripts resolve a relative path
against the workspace at runtime. init_project.py pre-fills dependence_paths
with these paths, so no manual path editing is normally needed.
Repository URLs:
comfyui-scheduler — https://github.com/calcuforge/comfyui-scheduler.gitremotion-video-template — https://github.com/calcuforge/remotion-video-template.git# Run from the workspace root (the directory that contains projects/)
mkdir -p dep
git clone https://github.com/calcuforge/comfyui-scheduler.git dep/comfyui-scheduler
git clone https://github.com/calcuforge/remotion-video-template.git dep/remotion-video-template
pip install -e dep/comfyui-scheduler # comfyui-scheduler CLI on PATH
( cd dep/remotion-video-template && npm install ) # node_modulesIf the dependencies already exist somewhere else, set
dependence_paths.remotion_template / dependence_paths.comfyui_scheduler in
project_config.yaml to their paths (a ~-path, absolute, or relative to the
workspace) instead of the dep/ default.
All projects MUST live under the workspace projects/ directory.
Project directory naming: scripts/tool/init_project.py creates the project
directory from the scripts/project_config_tpl.yaml template, named via its
--project-dir-name argument. The project name is categorized across THREE
dimensions (joined with _):
air_crash
(空难事件), history (历史), tech (科技), hardware (硬件), ... Derive it
from the user's topic.documentary (纪录片),
knowledge_sharing (知识分享), news_broadcast (新闻播报), product_intro,
data_report, tutorial, ...resolution (1080p/4k), orientation
(horizontal/vertical), theme, language, target audience, ...If the name already exists, a numeric suffix is appended automatically:
air_crash_documentary_1080p_horizontal,
air_crash_documentary_1080p_horizontal2, ...
Example: user asks "制作一个1080P横屏的空难纪录片" → content=air_crash,
structure=documentary, params=1080p+horizontal →
projects/air_crash_documentary_1080p_horizontal/. Check whether a project whose
all three dimensions match already exists; if not, create a new one.
After creation, edit project_config.yaml to set project.name (the same
categorized name, e.g. air_crash_documentary_1080p_horizontal),
project.video_style, project.target_audience, and other request-dependent
fields. project.name MUST be the categorized name — never the specific video
title.
Video directory — never reuse. Every video-making request creates a NEW
video{N}/ directory (video1, video2, ...). Each time the user asks to make
a video, create the next available video{N}/; never reuse or overwrite an
existing one. (Resuming an interrupted pipeline continues the same in-progress
video{N}/ — that is recovery, not reuse.)
Project directory reuse (the table below is about the project dir, not the video dir):
| Situation | Action |
|---|---|
| First video request | Run Step 1: create projects/{content}_{structure}_{params}/ |
| Same content + structure + params | Reuse existing project, then create a new video{N}/ inside it |
| Any dimension differs | Run Step 1: create a new project directory |
To determine project reuse: compare ALL categorization dimensions — content
category, narrative structure, and every explicitly-specified fine parameter
(resolution, orientation, theme, language, target audience). Reuse only if
every dimension matches. If the request maps to a different content category or
structure, or adds/removes a fine parameter (e.g. 1080p vs 4K, horizontal vs
vertical), create a new project — even when one dimension (e.g. documentary) is
shared. Each video request still creates a fresh video{N}/ inside.
projects/
├── project1/
│ ├── project_config.yaml # Project global preferences
│ ├── voice_file.wav # TTS reference voice (shared by all videos)
│ ├── bgm.mp3 # Background music (Step 11, shared by all videos)
│ ├── ad_video/ # Ad short videos for Step 14 (pre-created, empty; drop ad files here)
│ ├── video1/
│ │ ├── result.mp4 # Final rendered video (Step 13)
│ │ ├── origin_result.mp4 # Step 14 — pre-insertion render, preserved when ads are inserted
│ │ ├── video_config.yaml # Topic (Step 2) + content summaries (Step 5)
│ │ ├── video_struct.yaml # Video structure (stories → sections (one narration) → scenes (1-N, percentage))
│ │ ├── video_tasks.yaml # AIGC task list
│ │ ├── remotion_sections.yaml # Remotion render config
│ │ ├── tmp/ # General temporary files (cache, discovery results, etc.)
│ │ ├── search_results/
│ │ │ ├── result1.md # Research result 1
│ │ │ └── result2.md # Research result 2
│ │ └── stories/
│ │ ├── story1/
│ │ │ ├── script.md # Chapter narration script (Step 5)
│ │ │ └── narration1/
│ │ │ ├── speech.wav # Narration audio
│ │ │ └── scenes/
│ │ │ ├── video_prompt_{scene_id}.yaml # Structured video prompt per scene (Step 8a)
│ │ │ ├── origin_scene1.png # AIGC raw output
│ │ │ ├── scene1.png # Upscaled asset
│ │ │ ├── origin_scene2.mp4
│ │ │ └── scene2.mp4
│ │ └── story2/
│ └── video2/
└── project2/Detailed structure reference: templates/demo_projects/
The agent makes all decisions autonomously across all 13 steps. No user interaction is required until the final video is ready. Infer sensible defaults from the user's request (language, style, audience, duration).
The agent pauses for user confirmation at key points:
| When | What to ask / report |
|---|---|
| Before Step 1 (project_config generation) | Ask user to confirm: video_style, target_audience, language, orientation, resolution, duration, tts backend |
| Before Step 2 (topic selection) | Present the chosen topic (auto) or confirm the user's topic; ask user to approve before proceeding |
| After each step completes | Report which artifacts were generated (file paths), then wait for user confirmation before starting the next step |
In manual mode, never proceed to the next step until the user explicitly confirms (e.g., "ok", "continue", "next", "确认", "继续").
人机协作的远程桌面栈(VNC/虚拟桌面/Chromium): 人机协作仅发生在登录和
验证码场景(访问被登录墙/验证码墙拦住时,如搜索/页面抓取被要求登录、安全验证、
滑块验证)。其余环节全部由程序自动完成,无需也不应启动远程桌面栈。流程:
脚本/agent 遇到登录墙或验证码墙 → 报告被拦的 URL → python scripts/tool/ensure_remote_desktop.py start --url <被拦页面>(幂等,启动缺失服务,
经 CDP :9222 在共享 Chromium 打开该页)→ 提醒用户:优先 hermes channel 推送
(ensure_remote_desktop.py notify "<提醒内容>",即 hermes send --to <平台> 推到
home channel;自动识别已配置平台——channel directory(hermes send --list)→
~/.hermes/config.yaml platforms 段(优先有 home channel 的),可用
--to/HERMES_NOTIFY_TO 指定;退出码非 0 表示推送不可用/失败,报错会附修复提示
如 hermes config set <平台>_HOME_CHANNEL <id>)→ 推送失败才退回对话提醒 → 用户经 VNC 完成登录/验证:
提醒中的 VNC 地址一律运行脚本获取——取 ensure_remote_desktop.py start(或
status)JSON 输出的 vnc_viewer_url 字段(脚本优先输出环境变量
VNC_VIEWER_URL 的外网地址,未设置时回退容器本地 noVNC 地址),agent 不要自行
拼接地址 → 确认完成后脚本/agent 继续(搜索脚本经 CDP 使用共享 Chromium 的持久
profile——用户刚完成的登录/验证会被直接复用)→ 全部结束后 ... stop 关闭。
脚本同理:只有登录/验证码才需要协作,且在使用前后自行 start/stop、提醒同样推送
优先。机制:栈由 hermes-desktop 容器(镜像仓库 hermes-hitl-environment)内的
supervisord 按需管理(pulseaudio/xvfb/openbox/chromium/x11vnc/novnc,等价于
bash /scripts/launch-desktop.sh),start/stop 驱动 supervisorctl。
project.creation_mode.project.creation_mode from project_config.yaml to
determine behavior. Do NOT rely on conversational memory — always re-read the
field from the file.Detailed step-by-step instructions:
- Steps 1–4 (Setup): references/workflow-setup.md
- Steps 5–7 (Content): references/workflow-content.md
- Steps 8–14 (Production): references/workflow-production.md
| # | Step | Key Script | Output |
|---|---|---|---|
| 1 | Project initialization | scripts/tool/init_project.py, scripts/verify/verify_project_config.py, scripts/tool/check_environment.py | project_config.yaml |
| 2 | Define topic | — (agent research) | video_config.yaml |
| 3 | Topic research | scripts/search_provider/search.py, scripts/search_provider/search_rss.py | search_results/*.md |
| 4 | Design chapter list | scripts/verify/verify_stories.py | video_struct.yaml (stories only) |
| 5 | Write chapter scripts | scripts/verify/verify_story_scripts.py | stories/{story_id}/script.md, video_config.yaml (summary + chapter_summaries) |
| 6 | Design scene list | scripts/tool/generate_scene_list.py, scripts/verify/verify_video_struct.py | video_struct.yaml (full structure) |
| 7 | TTS + frame calculation | scripts/tool/run_tts.py, scripts/verify/verify_audio.py | speech.wav per scene |
| 8 | Search stock media | scripts/search_provider/search_stock_media.py, scripts/verify/verify_stock_assets.py | scenes/origin_* stock assets |
| 9 | Design AIGC prompts + plan tasks | scripts/tool/build_video_prompt.py, scripts/verify/verify_video_tasks.py | video_prompt_{scene_id}.yaml per scene, video_tasks.yaml |
| 10 | Execute AIGC tasks | scripts/tool/run_aigc.py, scripts/tool/run_upscale.py, scripts/verify/verify_aigc_assets.py | scenes/ assets |
| 11 | Generate background music | scripts/tool/run_bgm.py | projects/{name}/bgm.mp3 |
| 12 | Generate remotion config | scripts/tool/generate_remotion_sections.py, scripts/verify/verify_remotion_sections.py, scripts/verify/verify_remotion_data.py | remotion_sections.yaml |
| 13 | Render video | scripts/tool/render.py | result.mp4 |
| 14 | Insert ad videos | scripts/tool/insert_ad_videos.py, then re-run Step 13 render | remotion_sections.yaml with ad sections; result.mp4 re-rendered with ads (original kept as origin_result.mp4) |
Mandatory validation gates:
verify_project_config.py must exit 0check_environment.py must exit 0 (ComfyUI + TTS nodes reachable — the ONLY way node availability is determined)verify_stories.py must exit 0 (re-do step if not)verify_story_scripts.py must exit 0verify_video_struct.py must exit 0verify_audio.py must exit 0verify_stock_assets.py must exit 0verify_video_tasks.py must exit 0verify_aigc_assets.py must exit 0run_bgm.py must exit 0 (skipped when bgm.enabled: false; fill bgm.prompt in project_config.yaml first — it is empty at init)verify_remotion_sections.py must exit 0, then verify_remotion_data.py must exit 0insert_ad_videos.py must exit 0 (skipped when ad_video.enabled: false or no ad videos found; for insert_position: middle fill ad_video.insert_after_story first), then RE-RUN the Step 13 render so result.mp4 includes the ads| Rule | Requirement |
|---|---|
| Projects under workspace | All project directories MUST be under projects/ in the workspace. Never create outside. |
| Project name = categorized | project.name MUST be the categorized project name derived from content + narrative structure + fine params (e.g. air_crash_documentary_1080p_horizontal) — never the specific video title. |
| New video dir per request | Every video-making request creates a NEW video{N}/ directory. Never reuse or overwrite an existing video{N}/ — always pick the next available N. |
| project_config.yaml 已有值不改 | At project creation (Step 1), NEVER modify a field in project_config.yaml that already has a value (template default or script-generated) — only fill empty/placeholder fields. Change an existing value ONLY when the user explicitly asks. |
| Audio-master clock | Each narration's audio duration (plus tts.pause_seconds silence, default 0.5s) determines that narration's total frames: narration.total_frame = ceil((audio_duration + pause_seconds) × fps). The narration's scenes split it by percentage (largest-remainder, Σ scene frames == narration.total_frame). Never hand-estimate. |
| One narration = one section, split into scenes | Each section has exactly one narration and 1-N scenes. A narration should usually drive MULTIPLE scenes — split long (>30-40 chars) or multi-idea narrations into 2+ scenes and set each scene's integer percentage (Σ = 100). A single scene is the exception for short, single-idea narrations. |
| Scene percentage split | Each section's scene percentage values MUST be integers summing to exactly 100. Enforced by verify_video_struct.py. |
| Script = merged narrations | A chapter's script.md MUST equal all its narration contents concatenated in section order. Splitting a narration into scenes must not add/drop/reword text. Enforced by verify_video_struct.py. |
| Data fields ≠ narration | In data/text components, label/title/suffix/headers hold SHORT labels only — never a narration sentence (no sentence punctuation). The full sentence stays in narration.content. Don't make a StatCounter/DataBar just because narration contains a number. Enforced by verify_remotion_data. |
| Per-style scene mix | Footage-led styles are visual-majority: documentary ≥ 75%, knowledge_sharing ≥ 60% of ALL scenes are visual (AssetImage/AssetVideo/KenBurnsImage/MediaSection) — default each narration to a visual, ask "能换成画面吗?" before a text component. Data-led styles are data/text-majority: news_broadcast / data_report ≤ 50% visual — facts live in structured components (StatCounter/MetricsRow/DataBar/DataTable/Timeline/QuoteBlock), visual accents use STOCK footage, and AIGC NEVER depicts real people/products/UIs (news credibility). Enforced by verify_video_struct.py (Step 6); per-style strategy in references/expression_intent_mapping.md + references/special-rules.md. |
| Locale-aware search | Detect network locale by REACHABILITY (Baidu reachable + Google blocked ⇒ China), not just system locale. In a domestic China network, NEVER use Google/Wikipedia (unreachable) — use Baidu/Bing/Baike only. search.py auto-drops google/wikipedia in China networks. |
| Playwright for web | All website access uses Playwright Chromium (headless), except where curl is explicitly specified (RSS feeds). |
| 人机协作远程桌面(仅登录/验证码) | Human-machine collaboration happens ONLY for login & CAPTCHA walls (search/page access blocked by a login form, security check, or slider captcha). Everything else is fully automated — never start the desktop stack for previews/demos. Flow: script/agent hits a login/CAPTCHA wall → report the blocked URL → python scripts/tool/ensure_remote_desktop.py start --url <page> (idempotent; checks supervisord services pulseaudio/xvfb/openbox/chromium/x11vnc/novnc in the hermes-desktop container, starts only missing ones, opens the page in the shared Chromium via CDP) → REMIND the user: PREFER hermes-channel push (ensure_remote_desktop.py notify "<text>", i.e. hermes send --to <platform> to the home channel — auto-detects configured platforms: channel directory (hermes send --list) → ~/.hermes/config.yaml platforms section (home-channel ones first); --to/HERMES_NOTIFY_TO overrides; non-zero exit = push unavailable/failed, error carries a fix hint such as hermes config set <PLATFORM>_HOME_CHANNEL <id>) → ONLY fall back to a conversation reminder when the push fails → user completes login/CAPTCHA via VNC — the address in ANY reminder MUST be the vnc_viewer_url field from ensure_remote_desktop.py start/status output (the script prefers the VNC_VIEWER_URL env var and falls back to the container-local noVNC address; never compose the address yourself) → after confirmation, continue → ALWAYS ... stop when done. Scripts follow the same rule: collaboration for login/CAPTCHA only, start-before / stop-after, push-first reminders, script-sourced VNC address. |
| Anti-slop narration | Narration text MUST follow references/natural-narration.md. No AI-sounding filler, no rhetorical hooks, no rule-of-three abuse. |
| Narration length | No hard character cap on narration content. Write substantive sentences and vary their length; split a narration into multiple scenes for visual reasons, not for length. |
| Verify before proceed | Each step's verify script must pass before moving to the next step. |
| Node availability = script only | Determine ComfyUI/TTS node availability EXCLUSIVELY via scripts/tool/check_environment.py (Step 1 gate): run it and read its JSON data.guidance on failure. NEVER probe nodes yourself — no curl/nc/Test-NetConnection/manual TCP/socket checks, no guessing from config values. If a backend step (7/9/10/11) later fails with a connectivity error, re-run check_environment.py to diagnose, then act on its guidance. |
| Long tasks: background + 3h timeout | All long-running scripts (TTS, AIGC, upscale, BGM, render) MUST be launched in the background (Bash run_in_background: true) so the conversation is never blocked waiting; continue with other ready work instead, and proceed when the completion notification arrives (then check exit status + artifacts). Set generous timeouts — at least 3 hours (10800s): run_aigc.py --total-timeout 10800, render.py --timeout 10800, run_tts.py --timeout 10800, run_bgm.py --timeout 10800. Never use default 1-2h timeouts for these steps. |
| Absolute paths | All script path arguments (--config, --video-struct, --output, etc.) MUST be absolute paths. Scripts reject relative paths with an error. |
| Output confined to project | ALL agent-produced files (search results, scripts, audio, AIGC assets, remotion configs, rendered video) MUST be written under the project directory's pre-defined resource dirs or its tmp/ directory. Scripts that produce output files MUST expose a --output (or equivalent) parameter so output paths are explicit. NEVER write to system temp dirs (/tmp, %TEMP%, TMPDIR), the workspace root, or any path outside the project. |
| Faststart progressive playback | When presenting the finished video as a player in the chat, the mp4 MUST be faststart (moov atom at the front) and embedded for PROGRESSIVE playback — e.g. <video controls preload="metadata" src="...">. Do NOT load the whole file at once (preload="auto"). The render pipeline already emits faststart mp4s. |
| AIGC cross-scene consistency | For subjects that appear across multiple AIGC scenes (recurring characters, specific objects, branded items, consistent environments), the common.subject.description and common.style fields in all their video_prompt_{scene_id}.yaml files MUST use the SAME appearance description (same wording, same visual attributes). This prevents ComfyUI from generating visually inconsistent outputs for the same subject across scenes. If a character/object appears in N scenes, write the description once, then reuse it verbatim in all N prompt files. |
| Stock media for generic visuals | For scenes showing generic, non-specific visuals (atmosphere, mood, environment — NOT specific people/events/products), prefer asset_generation_method: stock over AIGC — but only when the corresponding flag is enabled: stock_media.search_image (default true) for image scenes, stock_media.search_video (default false) for video scenes. If a flag is false, use AIGC for that type. Also requires stock_media.sources to be non-empty. Configure sources in project_config.yaml (each entry: provider + api_key). See expression_intent_mapping.md for when stock is appropriate. |
| Ad insertion (Step 14) | When ad_video.enabled (default true) and ad videos exist under {project_root}/ad_video or ad_video.directories, insert ALL found videos at the configured ad_video.insert_position (beginning | middle | end, default middle): the script SPLICES them into remotion_sections.yaml as sections (AssetVideo scene + section-level audio carrying the ad's own sound; subtitles after the insertion point shift accordingly) — NO ffmpeg post-processing. Then RE-RUN the Step 13 render so result.mp4 includes the ads; the pre-ad render is preserved as origin_result.mp4. For middle, the AGENT must pick the chapter boundary and fill ad_video.insert_after_story (a story id) before running the script. Always deliver result.mp4. |
Load on demand — do NOT load all at once:
| File | Load when |
|---|---|
| references/workflow-setup.md | Steps 1–4 — project init, topic, research, chapters |
| references/topic-selection.md | Step 2 — only when auto topic selection is needed (user only names a category: web-search candidates, de-dupe against existing project topics, then apply the full selection strategy) |
| references/workflow-content.md | Steps 5–7 — scripts, scene design, TTS |
| references/workflow-production.md | Steps 8–14 — stock media, AIGC, bgm, remotion config, render, ad insertion |
| references/natural-narration.md | Step 5 — writing chapter narration scripts |
| references/search-providers.md | Step 3 — topic research |
| references/expression_intent_mapping.md | Step 6 — choosing scene types and components |
| references/stock_image_mapping.md | Step 6 — only if stock_media.search_image: true |
| references/stock_video_mapping.md | Step 6 — only if stock_media.search_video: true (default false) |
| references/special-rules.md | Step 6 — style-specific scene constraints (e.g. documentary opens on video) |
| templates/demo_projects/ | Any step — reference for config file structure |
© calcuforge, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 81 other files (scripts, references) in skills/explainer-video-maker of calcuforge/explainer-video-maker.
Open the folder on GitHubat commit 3900f32
Explainer Video Maker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Explainer Video Maker this skillcalcuforge/explainer-video-maker | 100 | — | ~6.5k | Automated safety check: Pass | Apache-2.0 | |
| Playwright Browser Demo Recordingdigitalsamba/claude-code-video-toolkit | 2.2k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Whiteboard Videotrustfuture/simon-skills | 353 | — | ~1.8k | Automated safety check: Notes | MIT | |
| Immersive Short Videoxrkseek/XRK-AGT | 140 | — | ~1k | Automated safety check: Pass | MIT | |
| Eva Feature Screenshotvvedantb/eva | 101 | — | ~3.2k | Automated safety check: Notes | MIT | |
| Cs Digital Human Product Video PipelineChenShuo2004/cs-skills | 193 | — | ~1.3k | Automated safety check: Pass | MIT |
digitalsamba/claude-code-video-toolkit
Records browser interactions as video with Playwright, covering viewport sizing, cursor highlighting, and converting output for Remotion.
trustfuture/simon-skills
手绘白板风"边画边讲"讲解视频出片 skill(Excalidraw 风格逐笔动画 + 火山引擎配音 + 烧录字幕 + 品牌水印与片尾卡 + 横竖两张封面 + 各平台发布文案)。当用户说"做一期白板视频 / 边画边讲 / 手绘讲解视频 / 用 excalidraw 做视频 / whiteboard video",或要在本仓库里新建一期、改场景、换贴纸或真实…
xrkseek/XRK-AGT
Produce immersive vertical short videos (口播/科普/产品讲解) without AI-slop aesthetics.
vvedantb/eva
Grab a single clean HD screenshot of a new eva feature from the real running app (Playwright at deviceScaleFactor 2, 1280 layout captured crisp at 2560×1440, dev overlays hidden) and write a tweet…
ChenShuo2004/cs-skills
当用户要把产品事实、口播、数字人、产品界面和 CTA 制作成可验收的数字人产品介绍视频时使用:统一预检 ChatCut、FFmpeg、ComfyUI、Fish/TTS 与 Remotion,按 plan、sample、batch 三种模式编排,先完成可审批样片再批量或出成片。用于有明确产品包的横版/竖版产品视频流水线;不要用于通用自动剪辑、电商短视频复刻、纯视频选题策划或未经确认的批量生成与发布。
ChenShuo2004/cs-skills
把真实网页做成 30–60 秒产品演示宣传片:Playwright 采集页面长截图,Remotion 用虚拟浏览器窗口做推进、平移、滚动和点击,渲染可交付 mp4,并可附口播稿。触发词包括 $cs-web-promo-film、网页宣传片、产品演示片、把这个网页做成视频、产品功能演示。不用于网页幻灯片、静态海报、真人出镜、电商对标复刻(改用 $cs-auto-videl)、ChatCut…
Works with
Categories
A skill your agent uses when the user wants to create an explainer, documentary, knowledge-sharing, news-broadcast, product-introduction, or data-report video from a topic. Explainer Video Maker is an agent skill from calcuforge/explainer-video-maker. Use when the user wants to create an explainer, documentary, knowledge-sharing, news-broadcast, product-introduction, or data-report video from a topic.
Explainer Video Maker fits situations like: the user wants to create an explainer; knowledge-sharing; product-introduction; data-report video from a topic.
Run `npx skills add calcuforge/explainer-video-maker --skill explainer-video-maker -a claude-code`. Or copy the skill folder (skills/explainer-video-maker in calcuforge/explainer-video-maker) into .claude/skills/explainer-video-maker in your project. Claude Code loads it when a task matches its description.
Run `npx skills add calcuforge/explainer-video-maker --skill explainer-video-maker -a codex`. Or copy the skill folder (skills/explainer-video-maker in calcuforge/explainer-video-maker) into .agents/skills/explainer-video-maker in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calcuforge/explainer-video-maker --skill explainer-video-maker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/explainer-video-maker, .gemini/skills/explainer-video-maker, .github/skills/explainer-video-maker and .opencode/skills/explainer-video-maker in your project.
Going by SKILL.md and its folder, Explainer Video Maker needs Python for the scripts in its folder and the command-line tools its instructions call (git, pip, python3, npm, playwright and bash). Our summary lists: Python 3; Node.js.
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Explainer Video Maker is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.5k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 38k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Explainer Video Maker: Playwright Browser Demo Recording (digitalsamba/claude-code-video-toolkit, 2.2k stars), Whiteboard Video (trustfuture/simon-skills, 353 stars), Immersive Short Video (xrkseek/XRK-AGT, 140 stars) and Eva Feature Screenshot (vvedantb/eva, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
calcuforge (a GitHub organization) maintains it in calcuforge/explainer-video-maker, which has 100 GitHub stars. The repository was last updated on September 16, 2026.
Source: calcuforge/explainer-video-maker on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.