Agent skill

API Media

by nodetool-ai in nodetool-ai/nodetool

Call nodetool.media or nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a…

AGPL-3.0Auto-check passedMedia & Creative

Install API Media

skills CLI
$ npx skills add nodetool-ai/nodetool --skill api-media -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nodetool-ai/nodetool api-media --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/system-skills/api-media .claude/skills/api-media && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
api-media
GitHub stars
556
Token cost
~1.9k tokens
SKILL.md length
824 words
Files
1
Skills in repo
127
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Call nodetool.media or nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a…

  • Tasks that involve Transcription
  • SKILL.md covers Generation calls, Handles and saving, Judging and Host binaries, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Video production

What it does

API Media is an agent skill from nodetool-ai/nodetool. Call nodetool.media or nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a video, run ffmpeg or ffprobe, download a video, and read the status and cost of a generation. Load before the first call into these namespaces.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription, Video production and Image editing. It works with FFmpeg. The repository describes itself as: Agent-first Creative Workspace. The licence is AGPL-3.0.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Video production
  • Tasks that involve Image editing

Example prompts

  • “/api-media”

What it can do on your machine

Read from SKILL.md and the folder at commit 515bd28. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are javascript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

API Media loads about 1.9k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 824 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nodetool-ai/nodetool at commit 515bd28, republished under its AGPL-3.0 licence (© nodetool-ai). 824 words, ~1,946 tokens.

Download SKILL.mdSave it as .claude/skills/api-media/SKILL.md (or your agent's skills folder).
name
api-media
description
Call nodetool.media or nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a video, run ffmpeg or ffprobe, download a video, and read the status and cost of a generation. Load before the first call into these namespaces.

nodetool.media and nodetool.generations

nodetool.media makes one result with one picked model and saves it as an asset. No workflow is needed. nodetool.generations is the record of every media generation: status, cost, and the assets it made. The model comes from nodetool.models.pick (see api-models). How to word a prompt for one model line is the *-prompting skill that find_model names in prompting_skill.

Generation calls

Each call takes model as a pick/find result, {provider, model_id}, or a "provider/model_id" string.

CallOptionsCapability to pick
generateImage(prompt, model, opts)width, height, quality, negative_prompt, output_file, backgroundtext_to_image
editImage(inputFile, prompt, model, opts)reference_files, strength, target_width, target_height, negative_prompt, output_file, backgroundimage_to_image
generateVideo(prompt, model, opts)duration_seconds, aspect_ratio, resolution, num_frames, negative_prompt, output_file, backgroundtext_to_video
animateImage(inputFile, model, opts)prompt, duration_seconds, aspect_ratio, resolution, num_frames, output_file, backgroundimage_to_video
videoFromReferences(referenceFiles, model, opts)prompt, use_reference_video_audio, duration_seconds, aspect_ratio, resolution, output_file, backgroundreference_to_video
speak(text, model, opts)voice, speed, output_file, backgroundtext_to_speech
generateMusic(prompt, model, opts)lyrics, duration_seconds, output_file, backgroundtext_to_music
transcribe(inputFile, model, opts)language, promptautomatic_speech_recognition
embed(text, model, opts)dimensions. text is a string or an array.generate_embedding

A generation answers with asset_id, asset_uri (asset://…), generation_id, mime_type and bytes, plus path when you passed output_file. An input file (inputFile, reference_files) is an asset:// URI or a workspace path.

  • editImage is the call for "make it like this". Pass the attached image or an approved earlier result as inputFile, and further images to match in reference_files.
  • videoFromReferences is for a shot that more than one reference defines: a character plus a garment, a product plus a location. The prompt then says the action and the camera, not the subjects.
  • Models honour duration_seconds loosely. Measure the result with ffprobe before you cut to it.
  • To put words on a picture, draw them yourself with createCanvas and fillText, or with renderText from @nodetool-ai/sandbox-flow/lib.image.draw.
Background generations

background: true returns at once with {generation_id, status: "running", background: true}. Collect the result with nodetool.generations.wait(receipt). A run can have 16 open at most. Start several, then wait for all of them in the same action:

js
const model = await nodetool.models.pick("text_to_video");
const receipts = await Promise.all(prompts.map((p) =>
  nodetool.media.generateVideo(p, model, { background: true })));
const done = await Promise.all(receipts.map((r) =>
  nodetool.generations.wait(r, { timeout_seconds: 900 })));

Handles and saving

image.*, audio.* and video.* (guest globals, no nodetool. prefix) take a generation result, its asset_uri or an asset id, and answer run-local handles. video.addAudio(videoHandle, audioHandle) combines media, for example. Save a finished handle with nodetool.media.toImage(handle), toAudio(handle) or toVideo(handle) before the action ends. A handle is dead in the next action. Do not pull bytes into the guest.

Hold each result in a local variable and feed it straight into the next call. Record the uris a later action or turn will need with nodetool.memory.save, and never re-run generation for something already saved.

Judging

The judge is a chat model that reads images (pick("generate_message") on a vision model), not the model that made the picture.

CallAnswers
critique(image, brief, visionModel, {taste_profile})A pass/revise verdict and concrete defects with locations and fixes. Feed the fixes into the next prompt. When it names no defect, make fresh variations.
compare(images, brief, visionModel, {taste_profile})The winner of 2–8 candidates by a pairwise knockout, each match judged twice with the order swapped, plus every verdict
scoreAdherence(image, brief, visionModel, {questions})Yes/no answers to up to 12 checks made from the brief, and the fraction that passed
understandVideo(video, prompt, videoModel, {max_tokens}){text, truncated}. Gemini reads the whole clip with audio. Other models get still frames with no audio. A truncated answer is half an answer: raise max_tokens or use a model that does not reason.
Show full SKILL.md (263 more words)Show less

Host binaries

CallDoes
ffmpeg(args, {inputs, output_file, timeout_seconds})Runs ffmpeg in the workspace with no shell. args is the argv after the binary. Paths are workspace-relative, and URLs are refused. inputs stages assets first as {"a.mp4": "asset://…"} (8 files at most). output_file is saved as an asset. The default timeout is 180 s, the maximum 600 s.
ffprobe(path, {inputs, timeout_seconds})Reads the format and streams. path may be an asset:// URI. The answer has a summary with duration_seconds, width, height and has_audio as real numbers and booleans.
downloadVideo(url, outputFile, {format, timeout_seconds})Downloads a public video with yt-dlp, up to 2 GiB.
js
// Concatenate two assets in one call.
await nodetool.media.ffmpeg(
  ["-i", "a.mp4", "-i", "b.mp4", "-filter_complex",
   "[0:v][0:a][1:v][1:a]concat=n=2:v=1:a=1[v][a]", "-map", "[v]", "-map", "[a]",
   "out.mp4"],
  { inputs: { "a.mp4": uriA, "b.mp4": uriB }, output_file: "out.mp4" }
);

nodetool.generations

Every generation result carries generation_id. Read the cost with get instead of guessing it.

CallAnswers
list({status, provider, capability, thread_id, job_id, since, limit, start_key})Generations, newest first, with status, cost, errors and assets
get(idOrResult)One generation in full: request parameters, cost and how it was priced, the provider request id, the assets
wait(idOrResult, {timeout_seconds})Waits for a background generation. The default is 300 s. A timeout answers the current record; call again to keep waiting.
cancel(idOrResult)Stops a running generation. It answers cancelled: false when the generation already settled.
reconcile(idOrResult)Asks the provider what it billed and replaces the estimate
fromProvider(provider, {model, status, since, until, limit, cursor})The provider's own history, from any machine. fal_ai answers.
getFromProvider(provider, requestId, {model})One provider-side generation, with the output urls it still hosts. This recovers a lost asset.

Statuses are pending, running, recovering, completed, failed, cancelled, needs_attention and interrupted. completed means the output is durable. A failed, cancelled or interrupted generation can still be billed.

© nodetool-ai, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/system-skills/api-media of nodetool-ai/nodetool.

Open the folder on GitHubat commit 515bd28

Compare with similar skills

API Media next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

API Media compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
API Media this skillnodetool-ai/nodetool556—~1.9kAutomated safety check: PassAGPL-3.0
Video Understandcalesthio/OpenMontage65k—~841Automated safety check: PassAGPL-3.0
Karaoke CaptionsAI-Builder-Club/skills1.3k1 repos~850Automated safety check: PassNone
Media Processingynulihao/AgentSkillOS617—~817Automated safety check: NotesMIT
Ffmpegrendi-api/ffmpeg-cheatsheet1.7k—~1.2kAutomated safety check: PassNone
AutoshortsUpload-Post/skill-autoshorts151—~5.3kAutomated safety check: NotesMIT

Similar skills

  • Video Understand

    calesthio/OpenMontage

    Understand video content locally using ffmpeg frame extraction and Whisper transcription.

    65k GitHub stars~841 tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Karaoke Captions

    AI-Builder-Club/skills

    Generate TikTok/Shorts-style karaoke captions using MLX Whisper, ASS subtitles, and FFmpeg libass.

    1.3k GitHub starsUsed in 1 repo~850 tokens
    Media & CreativeAuto-check passed
  • Media Processing

    ynulihao/AgentSkillOS

    Process multimedia files with FFmpeg (video/audio encoding, conversion, streaming, filtering, hardware acceleration), ImageMagick (image manipulation, format conversion, batch processing, effects…

    617 GitHub stars~817 tokensUpdated 7 mo ago
    Media & CreativeAuto-check: notes
  • Ffmpeg

    rendi-api/ffmpeg-cheatsheet

    A skill your agent uses when the user asks for FFmpeg or FFprobe commands, video/audio conversion, trimming, resizing, padding, overlays, subtitles, thumbnails, GIFs, storyboards, slideshows…

    1.7k GitHub stars~1.2k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Autoshorts

    Upload-Post/skill-autoshorts

    Daily pipeline that picks one long video from a folder, transcribes it with Whisper, uses Gemini 3 Flash multimodal to find every viral short-form moment, cuts each candidate with FFmpeg, adds a…

    151 GitHub stars~5.3k tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Stage Edit

    Orkas-AI/Orkas-VideoStudio

    Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

    498 GitHub stars~2.4k tokensUpdated 16 days ago
    Media & CreativeAuto-check passed

More from nodetool-ai/nodetool

All 127 skills in this repo
  • Beat Sync Editing

    nodetool-ai/nodetool

    Cut a NodeTool timeline to music and shape its pacing — detect the beat grid, place cuts on phrases, pick a cut type, build speed ramps with time remap, and give the piece an arc.

    556 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Caption Titles

    nodetool-ai/nodetool

    Add and animate a consistent text layer on an existing NodeTool timeline.

    556 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Color Motion

    nodetool-ai/nodetool

    Choose and animate colour on a NodeTool timeline, including shape and text gradients, colour grades, 3D LUTs, and dither.

    556 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Commercial Beat Sheet

    nodetool-ai/nodetool

    Write a shootable, precisely timed commercial beat sheet and store it as a NodeTool storyboard, with a consistent entity roster behind every shot.

    556 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Elevenlabs Audio Prompting

    nodetool-ai/nodetool

    Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead…

    556 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Frame Composition

    nodetool-ai/nodetool

    Stage the frame on a NodeTool timeline — grids, focal placement, safe areas per aspect ratio, depth layers and parallax, camera moves, and where elements enter and leave.

    556 GitHub stars~3.5k tokensUpdated today
    Auto-check passed

Works with

Questions about API Media

What does API Media do?

Call nodetool.media or nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a…. API Media is an agent skill from nodetool-ai/nodetool.generations from a code action: generate or edit images, video, speech and music with a picked model, transcribe and embed, judge images with a vision model, read a video, run ffmpeg or ffprobe, download a video, and read the status and cost of a generation.

When should I use API Media?

API Media fits situations like: tasks that involve Transcription; tasks that involve Video production; tasks that involve Image editing.

How do I install API Media in Claude Code?

Run `npx skills add nodetool-ai/nodetool --skill api-media -a claude-code`. Or copy the skill folder (packages/system-skills/api-media in nodetool-ai/nodetool) into .claude/skills/api-media in your project. Claude Code loads it when a task matches its description.

How do I install API Media in Codex?

Run `npx skills add nodetool-ai/nodetool --skill api-media -a codex`. Or copy the skill folder (packages/system-skills/api-media in nodetool-ai/nodetool) into .agents/skills/api-media in your project. Codex loads it when a task matches its description.

Can I use API Media in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nodetool-ai/nodetool --skill api-media -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/api-media, .gemini/skills/api-media, .github/skills/api-media and .opencode/skills/api-media in your project.

What does API Media need to run?

SKILL.md names no scripts, command-line tools or credentials: API Media is instructions for the agent only.

Does API Media access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is API Media safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does API Media use?

API Media is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does API Media use?

About 1.9k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to API Media?

Skills that share tags, products or a category with API Media: Video Understand (calesthio/OpenMontage, 65k stars), Karaoke Captions (AI-Builder-Club/skills, 1.3k stars), Media Processing (ynulihao/AgentSkillOS, 617 stars) and Ffmpeg (rendi-api/ffmpeg-cheatsheet, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains API Media?

nodetool-ai (a GitHub organization) maintains it in nodetool-ai/nodetool, which has 556 GitHub stars. The repository holds 127 skills in this directory. The repository was last updated on October 8, 2026.

Source: nodetool-ai/nodetool on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.