Anything2explainer
Vincentwei1021/anything2explainer
给一个主题,产出一条黑底 MG 风格(幕底可选星点或点阵波)、有配音字幕章节进度条的科普讲解视频(中文或英文;Remotion 代码动画;时长由用户定,常用 3–5 分钟)。内含可编译模板、图元库、配音/分镜/渲染工具、风格与动效规范、多 agent 分工协议与 QC 判据,以及一条完整样片(《RAG 与知识库》)作为质量标尺。Turn any topic into a narrated…
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
$ npx skills add ericrisco/rsc-harness --skill ai-media -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness ai-media --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-media .claude/skills/ai-media && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ai-media" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ai-media into .claude/skills/ai-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-media", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/ai-mediaType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill ai-media -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness ai-media --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ai-media .agents/skills/ai-media && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ai-media" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ai-media into .agents/skills/ai-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-media", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill ai-media -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness ai-media --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ai-media .cursor/skills/ai-media && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ai-media" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ai-media into .cursor/skills/ai-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-media", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/ai-media--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill ai-media -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness ai-media --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ai-media .gemini/skills/ai-media && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ai-media" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ai-media into .gemini/skills/ai-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-media", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness ai-mediaInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill ai-media -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ai-media .github/skills/ai-media && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ai-media" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ai-media into .github/skills/ai-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-media", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill ai-media -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness ai-media --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ai-media .opencode/skills/ai-media && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ai-media" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/ai-media into .opencode/skills/ai-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-media", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ai-mediaA skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
AI Media is an agent skill from ericrisco/rsc-harness. Use when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with ffmpeg (mux, duck, loudnorm, concat). NOT still-image generation/editing (that is replicate-images); NOT code-rendered React compositing (that is remotion-video).
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/ffmpeg-assembly.md`).
It sits in Media & Creative, covering Video production, Text to speech and voice and AI video generation. It works with FFmpeg, Remotion and React. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
11 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
ffmpegFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
ELEVENLABS_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AI Media loads about 3.3k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 88 tokens; SKILL.md has 1,476 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,476 words, ~3,290 tokens.
.claude/skills/ai-media/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.You are the cross-modal director. You decide which generative-media model to call per modality, in what order, with what params, then assemble the pieces with ffmpeg into one finished file. You do not own a single provider's API surface and you do not prompt still images — you orchestrate and glue.
Map the goal to modalities and an ordered step list, and lock that plan before you generate a single asset — media generation is slow and metered, so a re-roll of a 10 s Veo clip or a 90 s music track costs real money and minutes. Fixing the scene list, aspect ratio, target loudness and model per modality first is cheaper than discovering at mux time that your clips are 9:16 and your VO is the wrong sample rate. The "delegate to" column is where the actual call mechanics live — you pick the model and params, those skills run the call.
| Goal | Needs | Ordered steps | Delegate calls to |
|---|---|---|---|
| Narrated explainer | stills + img→video + VO + music | script → per-scene stills → clip per scene → VO → music → conform → concat → mix+duck → loudnorm → MP4 | replicate-images, fal/replicate |
| Product teaser (1 hero) | 1 still + img→video + music | still → clip → music → mix → loudnorm → MP4 | replicate-images, fal/replicate |
| Faceless short | stills + img→video + VO + music + captions | (explainer pipeline) + burn captions | replicate-images; ../video-shorts/SKILL.md for the script |
| Just a voiceover | VO only | script → TTS → loudnorm | — |
| Just a clip from a still | img→video only | still (input) → clip | fal/replicate |
| Code-rendered explainer | none of the above | render from React/TS | stop — route to remotion-video |
If the video is rendered from data/code (charts, timelines, JSON-driven scenes), this is not your job → ../remotion-video/SKILL.md. You handle model-generated + ffmpeg-glued.
ElevenLabs Python SDK. The call is convert(text, voice_id, model_id, output_format); auth via ELEVENLABS_API_KEY.
from elevenlabs.client import ElevenLabs
client = ElevenLabs() # reads ELEVENLABS_API_KEY
audio = client.text_to_speech.convert(
text="Your narration script here.",
voice_id="JBFqnCBsd6RMkjVDRZzb",
model_id="eleven_multilingual_v2", # final-quality VO
output_format="mp3_44100_128", # codec_samplerate_bitrate
)
with open("vo.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)Pick the model tier by what the job needs:
| Model | When | Latency | Cost lever |
|---|---|---|---|
eleven_v3 | most expressive, hero final VO — verify availability first (see caveat) | higher | most credits/char |
eleven_multilingual_v2 | high-quality multilingual VO (default for finals) | medium | medium |
eleven_flash_v2_5 | real-time / batch / scale / drafts | ~75 ms | cheapest |
Do not hardcode eleven_v3 blind. It shipped to the API in alpha (Aug 2025) and the docs model list now carries it, but the text_to_speech.convert API reference still documents the default as eleven_multilingual_v2 and does not enumerate eleven_v3 as a guaranteed value. Before you build a final pass on it, confirm it returns from GET /v1/models for your key (or just call once and check) — otherwise default to eleven_multilingual_v2, which is the safe, always-available hero tier.
output_format is codec_samplerate_bitrate — e.g. mp3_44100_128, mp3_22050_32. Match the VO sample rate to your assembly target, do not master the VO loud and hope.
Bad → Good:
mp3_22050_32, then mux onto a 48 kHz video — ffmpeg silently resamples, you get artifacts and a level mismatch.mp3_44100_128), and set final loudness with loudnorm in assembly, not by cranking the TTS.TTS is billed per character/token (~0.5–1 credit/char on the Flash/Turbo lines). ElevenLabs cut TTS API pricing up to 55% on 2026-05-07 (e.g. Flash on Creator $0.11→$0.05 / 1k tokens) — that figure is TTS-specific, not the Music cut. Pricing staling fast: these are point-in-time numbers from elevenlabs.io/pricing/api as of 2026-06-02 — re-check the page before quoting a budget. Shorter scripts and Flash on drafts are the cost levers.
The still is an input, not your output. Generate or edit the source image in ../replicate-images/SKILL.md, then animate it here. Reality check: every serious 2026 model does 1080p or native 4K — resolution is no longer the differentiating axis. The hard limit is per-generation duration (~5–15 s, model-dependent). Long pieces are one clip per scene, then concat — never one long take.
Durations below are from each vendor's own pages (as of 2026-06; see references/models-and-params.md for the citations) — they move with releases, so verify on the catalog before a final run:
| Model | Duration | Aspect / max res | Control surface | Native audio | Open-source |
|---|---|---|---|---|---|
| Google Veo 3.1 | 8 s / generation | 16:9 / 9:16, up to 4K | high | yes — synced 48 kHz dialogue/SFX | no |
| Kling 3.0 | up to 15 s | flexible, 4K | strong identity/temporal, lip-sync | no | no |
| Runway Gen-4.5 | 2–10 s | flexible | best — motion brushes, camera control, reference image | no | no |
| MiniMax Hailuo 02 | 6 s or 10 s (1080p caps at 6 s) | up to 1080p | medium | no | no |
| Wan 2.6 | up to 15 s | up to 1080p | first/last-frame control, A/V sync | no (sync) | yes (Apache) |
Choose by the binding constraint: need synced dialogue → Veo 3.1; need precise camera/motion control → Runway Gen-4.5; need identity consistency across scenes or the longest single take → Kling 3.0 / Wan 2.6; need open-source/self-host → Wan 2.6; cost-sensitive 1080p → Hailuo 02. Endpoint ids and per-call mechanics live in ../fal/SKILL.md / ../replicate/SKILL.md (both rails carry these models). See references/models-and-params.md for endpoint ids and current limits.
Costs are per-minute and plan-dependent — treat them as approximate and verify on the vendor pricing page (figures as of 2026-06; sources in references/models-and-params.md):
| Model | Cost (approx, verify) | Licensing story | Control |
|---|---|---|---|
| ElevenLabs Music v2 | per-minute, ~$0.15–0.50/min depending on plan (Music API pricing cut up to 50% at v2 launch — separate from the 55% TTS cut) | cleanest — vendor states trained only on licensed data, cleared for commercial use (Believe collaboration named at launch) | genre-switch mid-track |
| Suno v5 | plan-based | usage rights on paid plans post Nov-2025 label settlements (rights, not ownership) | vendor blind-test benchmark ELO ~1293 |
| Udio | $30/mo Pro plan (commercial rights); no official public API — third-party gateways only | UMG-licensed platform announced for 2026 | — |
Confirm commercial rights before you ship. Licensing differs per model and per plan; "I generated it" is not "I may sell the ad with it." For a clean commercial story with an official API, ElevenLabs Music v2 is the safe default — Udio has no first-party API, so do not plan a programmatic pipeline around it. The rest is the same fal/replicate call mechanics.
Four operations. Each is a copy-paste recipe; full filter graphs, caption burning and pitfalls are in references/ffmpeg-assembly.md.
(a) Mux VO onto video — map both streams, copy video, take the shorter duration:
ffmpeg -i scene.mp4 -i vo.mp3 \
-map 0:v -map 1:a -c:v copy -shortest out.mp4(b) Duck music under the VO — sidechaincompress keys the music off the voice so it drops when narration plays (pro DAW ducking, no manual keyframes):
ffmpeg -i vo.mp3 -i music.mp3 -filter_complex \
"[1:a][0:a]sidechaincompress=threshold=0.03:ratio=8:attack=20:release=300[duck]; \
[0:a][duck]amix=inputs=2:duration=longest[aout]" \
-map "[aout]" -c:a aac mix.m4aCheaper static fallback when sidechain is overkill — fix the music low under a full VO:
ffmpeg -i vo.mp3 -i music.mp3 -filter_complex \
"[1:a]volume=0.3[m];[0:a][m]amix=inputs=2:duration=longest[aout]" \
-map "[aout]" mix.m4a(c) Loudnorm to a target LUFS (two-pass) — measure, then apply. Target -14 LUFS for social/streaming, -16 for podcast-style VO. Normalize per track before mixing.
# pass 1: measure (read the JSON it prints)
ffmpeg -i mix.m4a -af loudnorm=I=-14:TP=-1.5:LRA=11:print_format=json -f null -
# pass 2: apply with the measured values
ffmpeg -i mix.m4a -af \
loudnorm=I=-14:TP=-1.5:LRA=11:measured_I=-20.1:measured_TP=-4.2:measured_LRA=6.0:measured_thresh=-30.8:offset=0.5:linear=true \
master.m4a(d) Concat scenes — conform first. Same-codec/res/fps clips → fast demuxer with -c copy. Mismatched clips → re-encode and scale first, then concat. Never -c copy-concat mismatched clips — you get desync or a corrupt stream.
# all clips identical codec/res/fps:
printf "file '%s'\n" scene1.mp4 scene2.mp4 scene3.mp4 > list.txt
ffmpeg -f concat -safe 0 -i list.txt -c copy joined.mp4
# mismatched: conform each, then concat filter
ffmpeg -i s1.mp4 -i s2.mp4 -filter_complex \
"[0:v]scale=1920:1080,fps=30,setsar=1[v0];[1:v]scale=1920:1080,fps=30,setsar=1[v1]; \
[v0][1:a?][v1][1:a?]concat=n=2:v=1:a=0[v]" -map "[v]" joined.mp4Ordered command list — generate once, assemble deterministically:
../replicate-images/SKILL.md (one prompt per scene).fal/replicate.convert(...) at the master sample rate.body.mp4.mix.m4a.master.m4a.master.m4a onto body.mp4 with -shortest → final.mp4.Emit this as a runnable script. scripts/verify.sh lints it (loudnorm present, conform-before-concat, final MP4 target).
../fal/SKILL.md / ../replicate/SKILL.md for per-call cost; treat budget as a constraint you set before generating.| Anti-pattern | Why it bites | Do instead |
|---|---|---|
| Generating assets before locking the pipeline | aspect/sample-rate/duration mismatches surface at mux time, forcing paid re-rolls | lock scene list, aspect, LUFS, models first |
| One long video-gen call for the whole piece | models cap at ~5–15 s; you fight the limit and waste rolls | one clip per scene, then concat |
Mixing tracks without per-track loudnorm | VO buried or blasting over music; inconsistent levels | two-pass loudnorm each track before mix |
-c copy-concat of mismatched clips | desync, corrupt stream, wrong frame timing | conform res/fps/SAR, then concat |
| Music low set by ear / static only when VO needs space | narration gets masked under the bed | sidechaincompress keyed off the VO |
| Shipping generated music without checking rights | "generated" ≠ "licensed to sell" — legal exposure | confirm commercial rights per model/plan |
| Prompting/editing the still inside this skill | duplicates replicate-images' job, worse prompts | delegate the still, consume it here |
| Mastering loudness by cranking the TTS | clipping, no true-peak control | set level with loudnorm, not the generator |
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/ai-media of ericrisco/rsc-harness.
Open the folder on GitHubat commit e3d5b33
AI Media next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AI Media this skillericrisco/rsc-harness | 174 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Anything2explainerVincentwei1021/anything2explainer | 2.4k | — | ~2.7k | Automated safety check: Pass | Custom licence | |
| Super Video MakerBomx/super-video-maker-skill | 309 | — | ~11k | Automated safety check: Notes | None | |
| Beatdesign WorkspaceBeatAPI/BeatDesign | 132 | 1 repos | ~632 | Automated safety check: Pass | Apache-2.0 | |
| Remotion Video Factorywwwzhouhui/skills_collection | 283 | — | ~990 | Automated safety check: Pass | None | |
| Remotion Best Practiceslyonjs/shortvid.io | 147 | 33 repos | ~1k | Automated safety check: Pass | MIT |
Vincentwei1021/anything2explainer
给一个主题,产出一条黑底 MG 风格(幕底可选星点或点阵波)、有配音字幕章节进度条的科普讲解视频(中文或英文;Remotion 代码动画;时长由用户定,常用 3–5 分钟)。内含可编译模板、图元库、配音/分镜/渲染工具、风格与动效规范、多 agent 分工协议与 QC 判据,以及一条完整样片(《RAG 与知识库》)作为质量标尺。Turn any topic into a narrated…
Bomx/super-video-maker-skill
End-to-end AI video production skill for agentic frameworks.
BeatAPI/BeatDesign
操作本地 BeatDesign 项目中的 Canvas、Assets、生成任务、视频时间线和字幕;在用户要求创作、整理、检查或剪辑媒体时使用。
wwwzhouhui/skills_collection
视觉从代码生长的技术讲解视频工厂:给一个主题,产出一条"图解动画 + AI 配音"的 MP4——结构图/矩阵/连线拓扑/图表/数字滚动全部用 Remotion + React/SVG 代码精确绘制(不用 AI 视频模型,杜绝公式乱码与连线漂移,可参数化、改数据自动重排),旁白由 TTS 逐段生成 + ffprobe 实测时长(默认 edge-tts 神经音色;--engine clone…
lyonjs/shortvid.io
Best practices for Remotion - Video creation in React. An agent skill from lyonjs/shortvid.io.
heygen-com/hyperframes
Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
ericrisco/rsc-harness
A skill your agent uses when building, refactoring, or debugging Angular (v20/21+): standalone components, signals, zoneless change detection, @if/@for/@defer control flow, inject() DI…
Categories
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…. AI Media is an agent skill from ericrisco/rsc-harness. Use when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with ffmpeg (mux, duck, loudnorm, concat).
AI Media fits situations like: A creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover; image-to-video clips; score — then glue them with ffmpeg (mux.
Run `npx skills add ericrisco/rsc-harness --skill ai-media -a claude-code`. Or copy the skill folder (skills/ai-media in ericrisco/rsc-harness) into .claude/skills/ai-media in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill ai-media -a codex`. Or copy the skill folder (skills/ai-media in ericrisco/rsc-harness) into .agents/skills/ai-media in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill ai-media -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-media, .gemini/skills/ai-media, .github/skills/ai-media and .opencode/skills/ai-media in your project.
Going by SKILL.md and its folder, AI Media needs a shell for the scripts in its folder, the command-line tools its instructions call (ffmpeg) and credentials named ELEVENLABS_API_KEY. Our summary lists: Python 3; A Bash shell; A credential in ELEVENLABS_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
AI Media is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with AI Media: Anything2explainer (Vincentwei1021/anything2explainer, 2.4k stars), Super Video Maker (Bomx/super-video-maker-skill, 309 stars), Beatdesign Workspace (BeatAPI/BeatDesign, 132 stars) and Remotion Video Factory (wwwzhouhui/skills_collection, 283 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 174 GitHub stars. The repository holds 233 skills in this directory. The repository was last updated on October 7, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.