Video Understand
calesthio/OpenMontage
Understand video content locally using ffmpeg frame extraction and Whisper transcription.
Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.
$ npx skills add calesthio/generative-media-skills --skill ffmpeg-media-finishing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install calesthio/generative-media-skills ffmpeg-media-finishing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/production/runtime-assembly/ffmpeg-media-finishing .claude/skills/ffmpeg-media-finishing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ffmpeg-media-finishing" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/ffmpeg-media-finishing into .claude/skills/ffmpeg-media-finishing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ffmpeg-media-finishing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/ffmpeg-media-finishingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add calesthio/generative-media-skills --skill ffmpeg-media-finishing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install calesthio/generative-media-skills ffmpeg-media-finishing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/production/runtime-assembly/ffmpeg-media-finishing .agents/skills/ffmpeg-media-finishing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ffmpeg-media-finishing" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/ffmpeg-media-finishing into .agents/skills/ffmpeg-media-finishing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ffmpeg-media-finishing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill ffmpeg-media-finishing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install calesthio/generative-media-skills ffmpeg-media-finishing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/production/runtime-assembly/ffmpeg-media-finishing .cursor/skills/ffmpeg-media-finishing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ffmpeg-media-finishing" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/ffmpeg-media-finishing into .cursor/skills/ffmpeg-media-finishing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ffmpeg-media-finishing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/calesthio/generative-media-skills.git --path skills/production/runtime-assembly/ffmpeg-media-finishing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add calesthio/generative-media-skills --skill ffmpeg-media-finishing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install calesthio/generative-media-skills ffmpeg-media-finishing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/production/runtime-assembly/ffmpeg-media-finishing .gemini/skills/ffmpeg-media-finishing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ffmpeg-media-finishing" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/ffmpeg-media-finishing into .gemini/skills/ffmpeg-media-finishing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ffmpeg-media-finishing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install calesthio/generative-media-skills ffmpeg-media-finishingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add calesthio/generative-media-skills --skill ffmpeg-media-finishing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/production/runtime-assembly/ffmpeg-media-finishing .github/skills/ffmpeg-media-finishing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ffmpeg-media-finishing" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/ffmpeg-media-finishing into .github/skills/ffmpeg-media-finishing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ffmpeg-media-finishing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill ffmpeg-media-finishing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install calesthio/generative-media-skills ffmpeg-media-finishing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/production/runtime-assembly/ffmpeg-media-finishing .opencode/skills/ffmpeg-media-finishing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ffmpeg-media-finishing" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/ffmpeg-media-finishing into .opencode/skills/ffmpeg-media-finishing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ffmpeg-media-finishing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ffmpeg-media-finishingProvider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.
Ffmpeg Media Finishing is an agent skill from calesthio/generative-media-skills. Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables. Use when finalizing images, image sequences, video, audio, captions, overlays, social/export variants, checksums, manifests, delivery specs, and QA with ffmpeg/ffprobe, including transcode versus stream-copy decisions, scaling, frame-rate handling, color metadata caution, loudness normalization, subtitle burn-in or sidecars, concat/trim, GIFs, batch variants, and reproducible command reporting.
Its SKILL.md is about 8.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `EVAL.md`, `scripts/media_probe.py` and `tests/test_media_probe.py`).
It sits in Media & Creative, covering Video production and Transcription. It works with FFmpeg. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
ffmpegffprobepythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
ffmpeg.orgw3.orgitu.inttech.ebu.chsupport.google.comads.tiktok.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ffmpeg Media Finishing loads about 8.2k tokens when it runs. Until then it costs about 133 tokens; SKILL.md has 2,740 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 2,740 words, ~8,181 tokens.
.claude/skills/ffmpeg-media-finishing/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Use this skill at the finishing stage: after creative generation or editing has produced source media and before the user receives deliverables. Your job is to turn fragile, inconsistent media outputs into reproducible, spec-aware files with an audit trail.
This skill is provider-independent. It assumes local ffmpeg and ffprobe are available, but not that every codec, filter, hardware encoder, or subtitle renderer is compiled into the installed build. FFmpeg build capabilities, encoder names, platform upload rules, and social specs are volatile; re-check them at production time with the commands below and with the current delivery spec.
Do not guess the state of media. Probe first, decide, run deterministic commands, then probe outputs.
Every finishing handoff should include:
ffprobe;This skill includes a deterministic Python 3.11+ standard-library helper at scripts/media_probe.py. Use it when you need a repeatable technical audit or a quick delivery-spec gate before or after finishing. It is assistance, not a replacement for perceptual review, color-managed approval, audio listening, caption/accessibility review, rights review, or platform/client signoff.
Documented fact, verified 2026-07-11: ffprobe is designed to gather multimedia stream information and can print JSON with -show_format, -show_streams, and -print_format json; if an input cannot be opened or recognized it returns a positive exit code. The helper invokes ffprobe with a subprocess argv list, never through a shell, and never runs ffmpeg or mutates media.
Invocation:
python skills/production/runtime-assembly/ffmpeg-media-finishing/scripts/media_probe.py input.mp4 --pretty
python skills/production/runtime-assembly/ffmpeg-media-finishing/scripts/media_probe.py input.mp4 --expect expected.json --prettyOptions:
media: required media path to inspect; the file is opened only by ffprobe.--expect expected.json: optional expectation spec. Exit code 0 means the audit ran and every expectation passed; exit code 1 means the audit ran but one or more expectations failed.--ffprobe PATH_OR_NAME: override the ffprobe executable. Default is ffprobe on PATH.--timeout SECONDS: ffprobe timeout. Default is 30.--pretty: emit two-space formatted JSON instead of compact JSON.Exit codes:
0 audit completed and expectations passed, or no expectations were provided
1 audit completed but at least one expectation failed
2 usage or input-file error
3 ffprobe executable unavailable
4 ffprobe returned a non-zero exit code
5 ffprobe timed out
6 ffprobe output was not valid or usable JSON
7 expectation JSON could not be read, parsed, or validatedNormalized audit schema, version 1.0:
{
"schema_version": "1.0",
"status": "pass",
"tool": "media_probe.py",
"ffprobe": {
"executable": "ffprobe",
"argv": ["ffprobe", "-v", "error", "-show_format", "-show_streams", "-print_format", "json", "input.mp4"]
},
"media": {
"path": "input.mp4",
"format_name": "mov,mp4,m4a,3gp,3g2,mj2",
"format_long_name": "QuickTime / MOV",
"duration_seconds": 1.0,
"size_bytes": 12345,
"bit_rate": 98765
},
"streams": {
"counts": {"video": 1, "audio": 1, "subtitle": 0, "other": 0},
"video": {
"index": 0,
"codec_name": "h264",
"profile": "High",
"width": 1920,
"height": 1080,
"pix_fmt": "yuv420p",
"r_frame_rate": {"raw": "30/1", "value": 30.0},
"avg_frame_rate": {"raw": "30/1", "value": 30.0},
"duration_seconds": 1.0
},
"audio": {
"index": 1,
"codec_name": "aac",
"sample_rate": 48000,
"channels": 2,
"channel_layout": "stereo",
"sample_fmt": "fltp",
"duration_seconds": 1.0
}
},
"expectations": null
}If a video or audio stream is absent, its normalized value is null and the corresponding count is 0. Unknown numeric values are emitted as null rather than invented.
Expectation schema:
{
"checks": [
{"field": "media.duration_seconds", "op": "tolerance", "value": {"value": 10.0, "tolerance": 0.05}},
{"field": "streams.video.codec_name", "op": "equals", "value": "h264"},
{"field": "streams.video.width", "op": "equals", "value": 1920},
{"field": "streams.video.height", "op": "equals", "value": 1080},
{"field": "streams.video.avg_frame_rate.value", "op": "tolerance", "value": {"value": 30.0, "tolerance": 0.01}},
{"field": "streams.video.pix_fmt", "op": "equals", "value": "yuv420p"},
{"field": "streams.audio.codec_name", "op": "equals", "value": "aac"},
{"field": "streams.audio.sample_rate", "op": "equals", "value": 48000},
{"field": "streams.audio.channels", "op": "equals", "value": 2}
]
}Supported op values:
equals: exact JSON equality.in: actual value must be one of the values in the expected array.range: numeric actual value must be within {"min": number, "max": number}; either boundary may be omitted.tolerance: numeric actual value must be within value +/- tolerance.Malformed expectation numeric values, non-finite numbers, and negative tolerance values are usage errors in the expectation file. The helper reports them as exit code 7 with structured error JSON instead of treating them as failed media checks.
Fields are dotted paths into the normalized report. Supported expectation targets intentionally cover common finishing gates: duration tolerance, video codec, width, height, frame rate, pixel format, audio codec, sample rate, and channel count. The schema can also address any stable field shown above, but do not treat this helper as proof of perceptual quality, color correctness, loudness compliance, or rights status.
Record the FFmpeg build and encoder/filter availability because behavior depends on the installed binary, compile flags, and external libraries.
ffmpeg -hide_banner -version
ffmpeg -hide_banner -encoders
ffmpeg -hide_banner -muxers
ffmpeg -hide_banner -filters
ffmpeg -hide_banner -h encoder=libx264
ffmpeg -hide_banner -h filter=loudnormRun ffprobe on every source and save machine-readable JSON when the job has more than one input, a client spec, or a risk of mismatch:
ffprobe -v error \
-show_format -show_streams -show_chapters \
-print_format json input.mp4 > input.ffprobe.jsonFor a compact human audit:
ffprobe -v error \
-select_streams v:0 \
-show_entries stream=codec_name,profile,width,height,pix_fmt,r_frame_rate,avg_frame_rate,time_base,field_order,color_range,color_space,color_transfer,color_primaries,sample_aspect_ratio,display_aspect_ratio,duration,nb_frames \
-of default=noprint_wrappers=1 input.mp4
ffprobe -v error \
-select_streams a:0 \
-show_entries stream=codec_name,sample_rate,channels,channel_layout,sample_fmt,bits_per_sample,duration,bit_rate \
-of default=noprint_wrappers=1 input.mp4Audit at least:
If ffprobe reports missing or inconsistent metadata, treat it as a warning, not as permission to invent values. Ask the user or infer cautiously from the creative pipeline only when necessary, and document the inference.
Use stream copy when the encoded streams are already acceptable and only container-level operations are needed:
ffmpeg -y -i input.mov -map 0 -c copy -movflags +faststart output.mp4Stream copy is appropriate for remuxing, trimming on keyframe boundaries, adding compatible metadata, or extracting tracks. It is not appropriate when changing resolution, codec, pixel format, frame rate by resampling, loudness, captions burn-in, overlays, color conversion, or any filtergraph output. FFmpeg documents -codec copy as copying streams without re-encoding; filters require decoded frames and therefore a re-encode.
Transcode when the deliverable requires a different codec/container, resolution, frame rate, pixel format, audio layout, loudness, burned captions, overlays, or compatibility profile.
When quality matters, prefer a constant-quality setting over arbitrary bitrates unless a platform or client spec requires a bitrate cap. For H.264 with libx264, a common mezzanine-to-delivery pattern is:
ffmpeg -y -i input.mov \
-map 0:v:0 -map 0:a? \
-c:v libx264 -preset slow -crf 18 -pix_fmt yuv420p \
-c:a aac -b:a 192k -ar 48000 \
-movflags +faststart output.mp4Production heuristic: use lower CRF for higher quality/larger files and higher CRF for smaller files. Verify visually; CRF values are not delivery guarantees. If a client gives a target bitrate, max bitrate, GOP length, profile, chroma, or level, follow the spec instead of this heuristic.
Distinguish container from codec:
Common delivery choices:
yuv420p.Before using a codec, check the local build:
ffmpeg -hide_banner -encoders | findstr /i "264 265 av1 prores aac opus"On non-Windows shells, use grep -Ei instead of findstr.
Hardware encoders such as h264_nvenc, hevc_videotoolbox, h264_qsv, or h264_amf are machine- and driver-dependent. They can be fast but may differ in rate-control behavior and quality. Do not assume availability; test a short sample.
Pick the geometry behavior intentionally:
Fit inside 1920x1080 with black padding:
ffmpeg -y -i input.mp4 \
-vf "scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1" \
-c:v libx264 -crf 18 -preset slow -pix_fmt yuv420p \
-c:a copy output_1080p_letterbox.mp4Fill 1080x1920 vertical by cropping:
ffmpeg -y -i input.mp4 \
-vf "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,setsar=1" \
-c:v libx264 -crf 18 -preset slow -pix_fmt yuv420p \
-c:a aac -b:a 192k output_9x16.mp4For generated stills, use even dimensions for codecs that require chroma subsampling:
ffmpeg -y -loop 1 -framerate 30 -t 5 -i still.png \
-vf "scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1,format=yuv420p" \
-c:v libx264 -crf 18 -preset slow -an still_card.mp4Frame rate changes are a common source of judder, drift, duplicate frames, and broken captions. Inspect both r_frame_rate and avg_frame_rate; variable-frame-rate sources often have a confusing nominal rate.
Rules:
fps, trim, setpts, atrim, or asetpts, re-encode the filtered stream.Force constant 30 fps for a platform-specific version:
ffmpeg -y -i input.mp4 \
-vf "fps=30,scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,setsar=1" \
-c:v libx264 -crf 18 -preset slow -pix_fmt yuv420p \
-c:a aac -b:a 192k -ar 48000 output_30fps_vertical.mp4Trim accurately by re-encoding:
ffmpeg -y -ss 12.345 -to 27.890 -i input.mp4 \
-map 0:v:0 -map 0:a? \
-c:v libx264 -crf 18 -preset slow \
-c:a aac -b:a 192k clip_exact.mp4Fast keyframe-aligned trim by stream copy:
ffmpeg -y -ss 12 -to 28 -i input.mp4 -map 0 -c copy clip_fast.mp4If audio and video durations differ after finishing, inspect start times, PTS/DTS warnings, and whether a filtergraph dropped or duplicated frames. Do not hide sync defects with arbitrary audio padding unless the spec calls for it.
FFmpeg can set or transform color metadata, but metadata is not the same as a correct color conversion. Generated assets often lack reliable color tags. Mismatched full/limited range, Rec.709/Rec.2020 primaries, transfer functions, or HDR/SDR assumptions can produce washed-out or crushed exports.
Minimum practice:
Example of tagging an SDR H.264 delivery without claiming a color transform:
ffmpeg -y -i input.mp4 \
-c:v libx264 -crf 18 -preset slow -pix_fmt yuv420p \
-color_primaries bt709 -color_trc bt709 -colorspace bt709 -color_range tv \
-c:a aac -b:a 192k output_sdr_tagged.mp4If you must transform between matrices/ranges, use a color filter such as zscale or colorspace only after checking the installed filter options and validating with test frames. Do not use metadata flags alone as a conversion.
Finish audio to the delivery spec. If there is no spec:
Use standards language precisely:
Measure first:
ffmpeg -hide_banner -i input.mp4 \
-af loudnorm=I=-16:TP=-1.5:LRA=11:print_format=json \
-f null -For repeatable loudness normalization, use two-pass loudnorm: first measure, then feed measured values into the second pass. Example target shown is a common online-video heuristic, not a standard:
ffmpeg -hide_banner -i input.mp4 \
-af loudnorm=I=-16:TP=-1.5:LRA=11:print_format=json \
-f null -
ffmpeg -y -i input.mp4 \
-af "loudnorm=I=-16:TP=-1.5:LRA=11:measured_I=-18.7:measured_TP=-3.2:measured_LRA=8.4:measured_thresh=-29.1:offset=0.2:linear=true:print_format=summary,aformat=sample_rates=48000:channel_layouts=stereo" \
-c:v copy -c:a aac -b:a 192k output_loudnorm.mp4Do not normalize blindly when preserving a theatrical mix, music dynamics, stems, spatial audio, or client-approved master. Ask before downmixing 5.1/7.1/Atmos/spatial audio to stereo.
Downmix example. Inspect the probed channel layout first and adapt channel names; some 5.1 sources expose side channels as SL/SR rather than back channels BL/BR.
ffmpeg -y -i input.mov \
-map 0:v:0 -map 0:a:0 \
-c:v copy \
-af "pan=stereo|FL=0.707*FC+FL+0.707*BL+0.5*LFE|FR=0.707*FC+FR+0.707*BR+0.5*LFE,aformat=sample_rates=48000:channel_layouts=stereo" \
-c:a aac -b:a 192k output_stereo.mp4Review downmixes by ear. Formulaic downmix coefficients are a starting point, not a substitute for listening.
Make caption handling a delivery decision, not an afterthought.
Use sidecar captions when accessibility, searchability, user control, localization, or platform-native captions matter. Common sidecar formats include SRT and WebVTT; the target platform decides what is accepted.
Burn in captions when the platform does not support caption tracks, when the user requests open captions, or when short-form social readability is the priority. Burned captions are pixels: they cannot be turned off, searched, restyled by assistive technology, or corrected without re-rendering.
WCAG guidance treats captions as synchronized text for speech and needed non-speech audio. Do not treat auto-generated captions as final without review when accessibility or client compliance matters.
Add a compatible SRT sidecar as a subtitle stream when the container/player supports it:
ffmpeg -y -i input.mp4 -i captions.srt \
-map 0:v -map 0:a? -map 1:0 \
-c:v copy -c:a copy -c:s mov_text \
-metadata:s:s:0 language=eng output_with_captions.mp4Burn in SRT captions with libass-backed rendering:
ffmpeg -y -i input.mp4 \
-vf "subtitles=captions.srt:force_style='Fontname=Arial,Fontsize=42,Outline=2,Shadow=1,MarginV=80'" \
-c:v libx264 -crf 18 -preset slow -pix_fmt yuv420p \
-c:a copy output_burned_captions.mp4Before using subtitles, confirm the FFmpeg build has the filter and libass support:
ffmpeg -hide_banner -filters | findstr /i subtitlesCaption QA:
Generated image sequences are common in AI pipelines. Stabilize numbering, frame rate, pixel format, and alpha policy before encoding.
Encode a numbered PNG sequence:
ffmpeg -y -framerate 24 -start_number 1 -i "frames/frame_%05d.png" \
-c:v libx264 -crf 18 -preset slow -pix_fmt yuv420p output_from_sequence.mp4Preserve alpha for compositing review:
ffmpeg -y -framerate 24 -i "frames/frame_%05d.png" \
-c:v prores_ks -profile:v 4444 -pix_fmt yuva444p10le output_alpha.movExtract frames for QA:
ffmpeg -y -i output.mp4 -vf "fps=1,scale=320:-1" "qa/thumb_%04d.jpg"If a sequence has missing frames, duplicate numbers, mixed sizes, mixed color profiles, or inconsistent alpha, stop and repair the sequence before encoding.
GIF is often useful for previews but inefficient for photographic or long content. Prefer MP4/WebM for real delivery unless GIF is explicitly requested.
Use a palette for better GIF quality:
ffmpeg -y -ss 0 -t 4 -i input.mp4 \
-vf "fps=12,scale=640:-1:flags=lanczos,palettegen" palette.png
ffmpeg -y -ss 0 -t 4 -i input.mp4 -i palette.png \
-lavfi "fps=12,scale=640:-1:flags=lanczos[x];[x][1:v]paletteuse=dither=bayer:bayer_scale=5" \
preview.gifCreate quick social variants from a master, but re-check current platform rules before final delivery:
# 16:9 landscape
ffmpeg -y -i master.mov \
-vf "scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1" \
-c:v libx264 -crf 20 -preset medium -pix_fmt yuv420p \
-c:a aac -b:a 192k -movflags +faststart deliverable_16x9.mp4
# 9:16 vertical
ffmpeg -y -i master.mov \
-vf "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,setsar=1" \
-c:v libx264 -crf 20 -preset medium -pix_fmt yuv420p \
-c:a aac -b:a 192k -movflags +faststart deliverable_9x16.mp4
# 1:1 square
ffmpeg -y -i master.mov \
-vf "scale=1080:1080:force_original_aspect_ratio=increase,crop=1080:1080,setsar=1" \
-c:v libx264 -crf 20 -preset medium -pix_fmt yuv420p \
-c:a aac -b:a 192k -movflags +faststart deliverable_1x1.mp4Social and ad specs change. Treat current platform docs, client media plans, and ad-manager requirements as binding; treat examples here as starting points.
Concat by stream copy only when inputs are compatible in codec, time base, resolution, pixel format, frame rate, sample rate, and channel layout. Otherwise normalize each clip first or use the concat filter and re-encode.
Concat demuxer with compatible clips:
file 'clip01.mp4'
file 'clip02.mp4'
file 'clip03.mp4'ffmpeg -y -f concat -safe 0 -i concat_list.txt -c copy joined.mp4Normalize first, then concat:
ffmpeg -y -i clip01.mov -vf "scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,fps=30,setsar=1,format=yuv420p" -af "aformat=sample_rates=48000:channel_layouts=stereo" -c:v libx264 -crf 18 -preset slow -c:a aac -b:a 192k norm01.mp4Overlay watermark:
ffmpeg -y -i input.mp4 -i logo.png \
-filter_complex "[1:v]scale=180:-1[logo];[0:v][logo]overlay=W-w-40:H-h-40:format=auto" \
-c:v libx264 -crf 18 -preset slow -pix_fmt yuv420p \
-c:a copy output_watermarked.mp4Add a slate from an image while keeping program audio in sync by concatenating explicit slate silence before the program audio:
ffmpeg -y -loop 1 -t 3 -i slate.png -i program.mp4 -f lavfi -t 3 -i anullsrc=r=48000:cl=stereo \
-filter_complex "[0:v]scale=1920:1080,setsar=1,fps=30,format=yuv420p[v0];[1:v]scale=1920:1080,setsar=1,fps=30,format=yuv420p[v1];[2:a]aformat=sample_rates=48000:channel_layouts=stereo[a0];[1:a]aformat=sample_rates=48000:channel_layouts=stereo[a1];[v0][a0][v1][a1]concat=n=2:v=1:a=1[v][a]" \
-map "[v]" -map "[a]" \
-c:v libx264 -crf 18 -preset slow -c:a aac -b:a 192k output_with_slate.mp4If the program has no audio, use a video-only concat (concat=n=2:v=1:a=0) and omit the audio maps. Do not map program audio directly under a prepended slate unless that offset is intentional.
Use a manifest-driven pattern for batch work. Record input path, output path, requested spec, exact command, FFmpeg version, source hash, output hash, and QA status. Avoid handwritten one-off commands for dozens of variants.
Example manifest fields:
{
"job_id": "campaign-final-2026-07-11",
"ffmpeg_version": "record from ffmpeg -version",
"source": "master.mov",
"source_sha256": "computed before processing",
"deliverables": [
{
"name": "youtube_16x9",
"output": "deliverable_16x9.mp4",
"spec": "MP4 H.264 AAC 1920x1080 SDR",
"command": "exact command string",
"output_sha256": "computed after processing",
"qa": "pass with noted warnings"
}
]
}Use SHA-256 for file integrity manifests:
certutil -hashfile output.mp4 SHA256On macOS/Linux:
shasum -a 256 output.mp4Use FFmpeg frame hashes when verifying decoded frame equivalence or detecting accidental media changes:
ffmpeg -v error -i output.mp4 -f framehash output.framehash
ffmpeg -v error -i output.mp4 -f framemd5 output.framemd5File checksums change if metadata changes. Frame hashes can remain stable across some container metadata changes but will change after lossy re-encode, scaling, audio normalization, or pixel/sample-format conversion.
Run automated checks:
ffmpeg -v error -i output.mp4 -f null -
ffprobe -v error -show_format -show_streams -print_format json output.mp4 > output.ffprobe.jsonInspect:
For important deliverables, sample frames rather than trusting command success:
ffmpeg -y -i output.mp4 -vf "select='eq(n,0)+eq(n,300)+eq(n,600)',scale=480:-1" -vsync vfr "qa_sample_%03d.jpg"Stop and ask for approval or specialist review when:
Avoid legal advice. State the production risk and recommend client, counsel, platform, or accessibility review as appropriate.
Intent: prepare a vertical short from an AI video master and reviewed captions.
Inputs:
master.mov: generated 16:9 or mixed-aspect video with audio;captions.srt: reviewed caption file;Workflow:
ffmpeg -hide_banner -version > ffmpeg_version.txt
ffprobe -v error -show_format -show_streams -print_format json master.mov > master.ffprobe.json
ffmpeg -y -i master.mov \
-vf "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,setsar=1,subtitles=captions.srt:force_style='Fontname=Arial,Fontsize=44,Outline=2,Shadow=1,MarginV=110',format=yuv420p" \
-af "loudnorm=I=-16:TP=-1.5:LRA=11,aformat=sample_rates=48000:channel_layouts=stereo" \
-c:v libx264 -crf 20 -preset slow \
-c:a aac -b:a 192k -movflags +faststart \
final_9x16_burned.mp4
ffmpeg -v error -i final_9x16_burned.mp4 -f null -
ffprobe -v error -show_format -show_streams -print_format json final_9x16_burned.mp4 > final_9x16_burned.ffprobe.json
certutil -hashfile final_9x16_burned.mp4 SHA256Why it is structured this way:
Likely failure modes:
subtitles filter unavailable due to missing libass;Intent: convert an approved MOV master into an MP4 wrapper for review upload without re-encoding.
Workflow:
ffprobe -v error -show_format -show_streams -print_format json approved_master.mov > approved_master.ffprobe.json
ffmpeg -y -i approved_master.mov \
-map 0 -c copy -movflags +faststart review_upload.mp4
ffprobe -v error -show_format -show_streams -print_format json review_upload.mp4 > review_upload.ffprobe.json
certutil -hashfile approved_master.mov SHA256
certutil -hashfile review_upload.mp4 SHA256Decision notes:
Intent: produce three deterministic variants from one master while preserving a manifest trail.
Workflow:
ffprobe -v error -show_format -show_streams -print_format json master.mov > master.ffprobe.json
ffmpeg -y -i master.mov \
-vf "scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1,format=yuv420p" \
-c:v libx264 -crf 20 -preset medium -c:a aac -b:a 192k -movflags +faststart deliverable_16x9.mp4
ffmpeg -y -i master.mov \
-vf "scale=1080:1080:force_original_aspect_ratio=increase,crop=1080:1080,setsar=1,format=yuv420p" \
-c:v libx264 -crf 20 -preset medium -c:a aac -b:a 192k -movflags +faststart deliverable_1x1.mp4
ffmpeg -y -i master.mov \
-vf "scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,setsar=1,format=yuv420p" \
-c:v libx264 -crf 20 -preset medium -c:a aac -b:a 192k -movflags +faststart deliverable_9x16.mp4Variation:
Consequential facts in this skill are based on these sources, checked 2026-07-11:
© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts) in skills/production/runtime-assembly/ffmpeg-media-finishing of calesthio/generative-media-skills.
Open the folder on GitHubat commit 8c85352
Ffmpeg Media Finishing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ffmpeg Media Finishing this skillcalesthio/generative-media-skills | 197 | — | ~8.2k | Automated safety check: Pass | MIT | |
| Video Understandcalesthio/OpenMontage | 66k | — | ~841 | Automated safety check: Pass | AGPL-3.0 | |
| Karaoke CaptionsAI-Builder-Club/skills | 1.3k | — | ~850 | Automated safety check: Pass | None | |
| Ffmpegrendi-api/ffmpeg-cheatsheet | 1.7k | — | ~1.2k | Automated safety check: Pass | None | |
| AutoshortsUpload-Post/skill-autoshorts | 151 | — | ~5.3k | Automated safety check: Notes | MIT | |
| Stage EditOrkas-AI/Orkas-VideoStudio | 499 | — | ~2.4k | Automated safety check: Pass | MIT |
calesthio/OpenMontage
Understand video content locally using ffmpeg frame extraction and Whisper transcription.
AI-Builder-Club/skills
Generate TikTok/Shorts-style karaoke captions using MLX Whisper, ASS subtitles, and FFmpeg libass.
rendi-api/ffmpeg-cheatsheet
A skill your agent uses when the user asks for FFmpeg or FFprobe commands, video/audio conversion, trimming, resizing, padding, overlays, subtitles, thumbnails, GIFs, storyboards, slideshows…
Upload-Post/skill-autoshorts
Daily pipeline that picks one long video from a folder, transcribes it with Whisper, uses Gemini 3 Flash multimodal to find every viral short-form moment, cuts each candidate with FFmpeg, adds a…
Orkas-AI/Orkas-VideoStudio
Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…
openclaw/openclaw
Add subtitles, captions, narration cues, or zoom to a proof video or PR recording using repo-local capture helpers and a system ffmpeg renderer.
calesthio/generative-media-skills
A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.
calesthio/generative-media-skills
Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…
calesthio/generative-media-skills
Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…
calesthio/generative-media-skills
Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…
calesthio/generative-media-skills
Provider-independent quality assurance for AI-generated and AI-assisted media.
calesthio/generative-media-skills
Provider-independent production workflow for AI agents assembling generated or source media into HyperFrames HTML/CSS/JS videos.
Works with
Categories
Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables. Ffmpeg Media Finishing is an agent skill from calesthio/generative-media-skills. Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.
Ffmpeg Media Finishing fits situations like: finalizing images; image sequences; social/export variants; QA with ffmpeg/ffprobe.
Run `npx skills add calesthio/generative-media-skills --skill ffmpeg-media-finishing -a claude-code`. Or copy the skill folder (skills/production/runtime-assembly/ffmpeg-media-finishing in calesthio/generative-media-skills) into .claude/skills/ffmpeg-media-finishing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add calesthio/generative-media-skills --skill ffmpeg-media-finishing -a codex`. Or copy the skill folder (skills/production/runtime-assembly/ffmpeg-media-finishing in calesthio/generative-media-skills) into .agents/skills/ffmpeg-media-finishing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill ffmpeg-media-finishing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ffmpeg-media-finishing, .gemini/skills/ffmpeg-media-finishing, .github/skills/ffmpeg-media-finishing and .opencode/skills/ffmpeg-media-finishing in your project.
Going by SKILL.md and its folder, Ffmpeg Media Finishing needs Python for the scripts in its folder and the command-line tools its instructions call (ffmpeg, ffprobe and python). Our summary lists: Python 3.
SKILL.md names 6 domains. As links in the text: ffmpeg.org, w3.org, itu.int, tech.ebu.ch, support.google.com and ads.tiktok.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Ffmpeg Media Finishing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.2k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ffmpeg Media Finishing: Video Understand (calesthio/OpenMontage, 66k stars), Karaoke Captions (AI-Builder-Club/skills, 1.3k stars), Ffmpeg (rendi-api/ffmpeg-cheatsheet, 1.7k stars) and Autoshorts (Upload-Post/skill-autoshorts, 151 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.
Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.