Video Understand
calesthio/OpenMontage
Understand video content locally using ffmpeg frame extraction and Whisper transcription.
Call nodetool.media or nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a…
$ npx skills add nodetool-ai/nodetool --skill api-media -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install nodetool-ai/nodetool api-media --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/system-skills/api-media .claude/skills/api-media && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "api-media" agent skill from https://github.com/nodetool-ai/nodetool/tree/main/packages/system-skills/api-media into .claude/skills/api-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-media", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/nodetool-ai/nodetool/tree/main/packages/system-skills/api-mediaType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add nodetool-ai/nodetool --skill api-media -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install nodetool-ai/nodetool api-media --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .agents/skills && cp -r skills-src/packages/system-skills/api-media .agents/skills/api-media && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "api-media" agent skill from https://github.com/nodetool-ai/nodetool/tree/main/packages/system-skills/api-media into .agents/skills/api-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-media", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add nodetool-ai/nodetool --skill api-media -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install nodetool-ai/nodetool api-media --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/packages/system-skills/api-media .cursor/skills/api-media && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "api-media" agent skill from https://github.com/nodetool-ai/nodetool/tree/main/packages/system-skills/api-media into .cursor/skills/api-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-media", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/nodetool-ai/nodetool.git --path packages/system-skills/api-media--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add nodetool-ai/nodetool --skill api-media -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install nodetool-ai/nodetool api-media --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/packages/system-skills/api-media .gemini/skills/api-media && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "api-media" agent skill from https://github.com/nodetool-ai/nodetool/tree/main/packages/system-skills/api-media into .gemini/skills/api-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-media", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install nodetool-ai/nodetool api-mediaInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add nodetool-ai/nodetool --skill api-media -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .github/skills && cp -r skills-src/packages/system-skills/api-media .github/skills/api-media && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "api-media" agent skill from https://github.com/nodetool-ai/nodetool/tree/main/packages/system-skills/api-media into .github/skills/api-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-media", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add nodetool-ai/nodetool --skill api-media -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install nodetool-ai/nodetool api-media --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/packages/system-skills/api-media .opencode/skills/api-media && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "api-media" agent skill from https://github.com/nodetool-ai/nodetool/tree/main/packages/system-skills/api-media into .opencode/skills/api-media/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "api-media", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
api-mediaCall nodetool.media or nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a…
API Media is an agent skill from nodetool-ai/nodetool. Call nodetool.media or nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a video, run ffmpeg or ffprobe, download a video, and read the status and cost of a generation. Load before the first call into these namespaces.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative, covering Transcription, Video production and Image editing. It works with FFmpeg. The repository describes itself as: Agent-first Creative Workspace. The licence is AGPL-3.0.
Read from SKILL.md and the folder at commit 515bd28. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are javascript).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
API Media loads about 1.9k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 824 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from nodetool-ai/nodetool at commit 515bd28, republished under its AGPL-3.0 licence (© nodetool-ai). 824 words, ~1,946 tokens.
.claude/skills/api-media/SKILL.md (or your agent's skills folder).nodetool.media makes one result with one picked model and saves it as an
asset. No workflow is needed. nodetool.generations is the record of every
media generation: status, cost, and the assets it made. The model comes from
nodetool.models.pick (see api-models). How to word a prompt for one model
line is the *-prompting skill that find_model names in prompting_skill.
Each call takes model as a pick/find result, {provider, model_id}, or
a "provider/model_id" string.
| Call | Options | Capability to pick |
|---|---|---|
generateImage(prompt, model, opts) | width, height, quality, negative_prompt, output_file, background | text_to_image |
editImage(inputFile, prompt, model, opts) | reference_files, strength, target_width, target_height, negative_prompt, output_file, background | image_to_image |
generateVideo(prompt, model, opts) | duration_seconds, aspect_ratio, resolution, num_frames, negative_prompt, output_file, background | text_to_video |
animateImage(inputFile, model, opts) | prompt, duration_seconds, aspect_ratio, resolution, num_frames, output_file, background | image_to_video |
videoFromReferences(referenceFiles, model, opts) | prompt, use_reference_video_audio, duration_seconds, aspect_ratio, resolution, output_file, background | reference_to_video |
speak(text, model, opts) | voice, speed, output_file, background | text_to_speech |
generateMusic(prompt, model, opts) | lyrics, duration_seconds, output_file, background | text_to_music |
transcribe(inputFile, model, opts) | language, prompt | automatic_speech_recognition |
embed(text, model, opts) | dimensions. text is a string or an array. | generate_embedding |
A generation answers with asset_id, asset_uri (asset://…),
generation_id, mime_type and bytes, plus path when you passed
output_file. An input file (inputFile, reference_files) is an
asset:// URI or a workspace path.
editImage is the call for "make it like this". Pass the attached image or
an approved earlier result as inputFile, and further images to match in
reference_files.videoFromReferences is for a shot that more than one reference defines: a
character plus a garment, a product plus a location. The prompt then says
the action and the camera, not the subjects.duration_seconds loosely. Measure the result with
ffprobe before you cut to it.createCanvas and
fillText, or with renderText from
@nodetool-ai/sandbox-flow/lib.image.draw.background: true returns at once with
{generation_id, status: "running", background: true}. Collect the result
with nodetool.generations.wait(receipt). A run can have 16 open at most.
Start several, then wait for all of them in the same action:
const model = await nodetool.models.pick("text_to_video");
const receipts = await Promise.all(prompts.map((p) =>
nodetool.media.generateVideo(p, model, { background: true })));
const done = await Promise.all(receipts.map((r) =>
nodetool.generations.wait(r, { timeout_seconds: 900 })));image.*, audio.* and video.* (guest globals, no nodetool. prefix) take
a generation result, its asset_uri or an asset id, and answer run-local
handles. video.addAudio(videoHandle, audioHandle) combines media, for
example. Save a finished handle with nodetool.media.toImage(handle),
toAudio(handle) or toVideo(handle) before the action ends. A handle is dead
in the next action. Do not pull bytes into the guest.
Hold each result in a local variable and feed it straight into the next call.
Record the uris a later action or turn will need with nodetool.memory.save,
and never re-run generation for something already saved.
The judge is a chat model that reads images (pick("generate_message") on a
vision model), not the model that made the picture.
| Call | Answers |
|---|---|
critique(image, brief, visionModel, {taste_profile}) | A pass/revise verdict and concrete defects with locations and fixes. Feed the fixes into the next prompt. When it names no defect, make fresh variations. |
compare(images, brief, visionModel, {taste_profile}) | The winner of 2–8 candidates by a pairwise knockout, each match judged twice with the order swapped, plus every verdict |
scoreAdherence(image, brief, visionModel, {questions}) | Yes/no answers to up to 12 checks made from the brief, and the fraction that passed |
understandVideo(video, prompt, videoModel, {max_tokens}) | {text, truncated}. Gemini reads the whole clip with audio. Other models get still frames with no audio. A truncated answer is half an answer: raise max_tokens or use a model that does not reason. |
| Call | Does |
|---|---|
ffmpeg(args, {inputs, output_file, timeout_seconds}) | Runs ffmpeg in the workspace with no shell. args is the argv after the binary. Paths are workspace-relative, and URLs are refused. inputs stages assets first as {"a.mp4": "asset://…"} (8 files at most). output_file is saved as an asset. The default timeout is 180 s, the maximum 600 s. |
ffprobe(path, {inputs, timeout_seconds}) | Reads the format and streams. path may be an asset:// URI. The answer has a summary with duration_seconds, width, height and has_audio as real numbers and booleans. |
downloadVideo(url, outputFile, {format, timeout_seconds}) | Downloads a public video with yt-dlp, up to 2 GiB. |
// Concatenate two assets in one call.
await nodetool.media.ffmpeg(
["-i", "a.mp4", "-i", "b.mp4", "-filter_complex",
"[0:v][0:a][1:v][1:a]concat=n=2:v=1:a=1[v][a]", "-map", "[v]", "-map", "[a]",
"out.mp4"],
{ inputs: { "a.mp4": uriA, "b.mp4": uriB }, output_file: "out.mp4" }
);Every generation result carries generation_id. Read the cost with get
instead of guessing it.
| Call | Answers |
|---|---|
list({status, provider, capability, thread_id, job_id, since, limit, start_key}) | Generations, newest first, with status, cost, errors and assets |
get(idOrResult) | One generation in full: request parameters, cost and how it was priced, the provider request id, the assets |
wait(idOrResult, {timeout_seconds}) | Waits for a background generation. The default is 300 s. A timeout answers the current record; call again to keep waiting. |
cancel(idOrResult) | Stops a running generation. It answers cancelled: false when the generation already settled. |
reconcile(idOrResult) | Asks the provider what it billed and replaces the estimate |
fromProvider(provider, {model, status, since, until, limit, cursor}) | The provider's own history, from any machine. fal_ai answers. |
getFromProvider(provider, requestId, {model}) | One provider-side generation, with the output urls it still hosts. This recovers a lost asset. |
Statuses are pending, running, recovering, completed, failed,
cancelled, needs_attention and interrupted. completed means the
output is durable. A failed, cancelled or interrupted generation can still be
billed.
© nodetool-ai, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in packages/system-skills/api-media of nodetool-ai/nodetool.
Open the folder on GitHubat commit 515bd28
API Media next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| API Media this skillnodetool-ai/nodetool | 556 | — | ~1.9k | Automated safety check: Pass | AGPL-3.0 | |
| Video Understandcalesthio/OpenMontage | 65k | — | ~841 | Automated safety check: Pass | AGPL-3.0 | |
| Karaoke CaptionsAI-Builder-Club/skills | 1.3k | 1 repos | ~850 | Automated safety check: Pass | None | |
| Media Processingynulihao/AgentSkillOS | 617 | — | ~817 | Automated safety check: Notes | MIT | |
| Ffmpegrendi-api/ffmpeg-cheatsheet | 1.7k | — | ~1.2k | Automated safety check: Pass | None | |
| AutoshortsUpload-Post/skill-autoshorts | 151 | — | ~5.3k | Automated safety check: Notes | MIT |
calesthio/OpenMontage
Understand video content locally using ffmpeg frame extraction and Whisper transcription.
AI-Builder-Club/skills
Generate TikTok/Shorts-style karaoke captions using MLX Whisper, ASS subtitles, and FFmpeg libass.
ynulihao/AgentSkillOS
Process multimedia files with FFmpeg (video/audio encoding, conversion, streaming, filtering, hardware acceleration), ImageMagick (image manipulation, format conversion, batch processing, effects…
rendi-api/ffmpeg-cheatsheet
A skill your agent uses when the user asks for FFmpeg or FFprobe commands, video/audio conversion, trimming, resizing, padding, overlays, subtitles, thumbnails, GIFs, storyboards, slideshows…
Upload-Post/skill-autoshorts
Daily pipeline that picks one long video from a folder, transcribes it with Whisper, uses Gemini 3 Flash multimodal to find every viral short-form moment, cuts each candidate with FFmpeg, adds a…
Orkas-AI/Orkas-VideoStudio
Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…
nodetool-ai/nodetool
Cut a NodeTool timeline to music and shape its pacing — detect the beat grid, place cuts on phrases, pick a cut type, build speed ramps with time remap, and give the piece an arc.
nodetool-ai/nodetool
Add and animate a consistent text layer on an existing NodeTool timeline.
nodetool-ai/nodetool
Choose and animate colour on a NodeTool timeline, including shape and text gradients, colour grades, 3D LUTs, and dither.
nodetool-ai/nodetool
Write a shootable, precisely timed commercial beat sheet and store it as a NodeTool storyboard, with a consistent entity roster behind every shot.
nodetool-ai/nodetool
Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead…
nodetool-ai/nodetool
Stage the frame on a NodeTool timeline — grids, focal placement, safe areas per aspect ratio, depth layers and parallax, camera moves, and where elements enter and leave.
Works with
Categories
Call nodetool.media or nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a…. API Media is an agent skill from nodetool-ai/nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a video, run ffmpeg or ffprobe, download a video, and read the status and cost of a generation.
API Media fits situations like: tasks that involve Transcription; tasks that involve Video production; tasks that involve Image editing.
Run `npx skills add nodetool-ai/nodetool --skill api-media -a claude-code`. Or copy the skill folder (packages/system-skills/api-media in nodetool-ai/nodetool) into .claude/skills/api-media in your project. Claude Code loads it when a task matches its description.
Run `npx skills add nodetool-ai/nodetool --skill api-media -a codex`. Or copy the skill folder (packages/system-skills/api-media in nodetool-ai/nodetool) into .agents/skills/api-media in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nodetool-ai/nodetool --skill api-media -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/api-media, .gemini/skills/api-media, .github/skills/api-media and .opencode/skills/api-media in your project.
SKILL.md names no scripts, command-line tools or credentials: API Media is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
API Media is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with API Media: Video Understand (calesthio/OpenMontage, 65k stars), Karaoke Captions (AI-Builder-Club/skills, 1.3k stars), Media Processing (ynulihao/AgentSkillOS, 617 stars) and Ffmpeg (rendi-api/ffmpeg-cheatsheet, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
nodetool-ai (a GitHub organization) maintains it in nodetool-ai/nodetool, which has 556 GitHub stars. The repository holds 127 skills in this directory. The repository was last updated on October 8, 2026.
Source: nodetool-ai/nodetool on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.