Multimedia Accessibility
Owl-Listener/inclusive-design-skills
Design accessible video, audio, and multimedia content with captions, transcripts, and audio descriptions.
Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…
$ npx skills add calesthio/generative-media-skills --skill captions-media-accessibility -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install calesthio/generative-media-skills captions-media-accessibility --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/production/audio-craft/captions-media-accessibility .claude/skills/captions-media-accessibility && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "captions-media-accessibility" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/audio-craft/captions-media-accessibility into .claude/skills/captions-media-accessibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "captions-media-accessibility", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/calesthio/generative-media-skills/tree/main/skills/production/audio-craft/captions-media-accessibilityType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add calesthio/generative-media-skills --skill captions-media-accessibility -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install calesthio/generative-media-skills captions-media-accessibility --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/production/audio-craft/captions-media-accessibility .agents/skills/captions-media-accessibility && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "captions-media-accessibility" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/audio-craft/captions-media-accessibility into .agents/skills/captions-media-accessibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "captions-media-accessibility", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill captions-media-accessibility -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install calesthio/generative-media-skills captions-media-accessibility --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/production/audio-craft/captions-media-accessibility .cursor/skills/captions-media-accessibility && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "captions-media-accessibility" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/audio-craft/captions-media-accessibility into .cursor/skills/captions-media-accessibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "captions-media-accessibility", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/calesthio/generative-media-skills.git --path skills/production/audio-craft/captions-media-accessibility--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add calesthio/generative-media-skills --skill captions-media-accessibility -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install calesthio/generative-media-skills captions-media-accessibility --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/production/audio-craft/captions-media-accessibility .gemini/skills/captions-media-accessibility && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "captions-media-accessibility" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/audio-craft/captions-media-accessibility into .gemini/skills/captions-media-accessibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "captions-media-accessibility", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install calesthio/generative-media-skills captions-media-accessibilityInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add calesthio/generative-media-skills --skill captions-media-accessibility -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/production/audio-craft/captions-media-accessibility .github/skills/captions-media-accessibility && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "captions-media-accessibility" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/audio-craft/captions-media-accessibility into .github/skills/captions-media-accessibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "captions-media-accessibility", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill captions-media-accessibility -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install calesthio/generative-media-skills captions-media-accessibility --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/production/audio-craft/captions-media-accessibility .opencode/skills/captions-media-accessibility && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "captions-media-accessibility" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/audio-craft/captions-media-accessibility into .opencode/skills/captions-media-accessibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "captions-media-accessibility", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
captions-media-accessibilityProvider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…
Captions Media Accessibility is an agent skill from calesthio/generative-media-skills. Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video recuts, training media, and localized content. Use when planning, authoring, reviewing, localizing, burning in, exporting, or QAing captions, subtitles, SDH, transcripts, audio description, flashing/motion safety, caption readability, or accessible media handoff files.
Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `EVAL.md`, `scripts/validate_captions.py` and `tests/test_validate_captions.py`).
It sits in Media & Creative, covering Transcription, Accessibility and Plain language and style rules. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.
Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
w3.orgdcmp.orghelp.vimeo.comloc.govpartnerhelp.netflixstudios.comsection508.govlaw.cornell.edusupport.google.comdeveloper.mozilla.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Captions Media Accessibility loads about 6.9k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 3,173 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 3,173 words, ~6,908 tokens.
.claude/skills/captions-media-accessibility/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Treat captions and accessibility tracks as production assets, not polish. Build them from the script, audio mix, edit, localization plan, and platform delivery target. If the user asks for legal compliance, say what standards you are using and recommend review by qualified accessibility/legal counsel; do not promise legal compliance from a generated file alone.
Before writing or exporting captions, identify:
Use these as factual constraints when relevant, citing the source in your output or handoff notes when the claim matters.
WEBVTT, contains timed cues, and can be used for captions, subtitles, descriptions, chapters, or metadata depending on player support. Sources: W3C/MDN and Library of Congress WebVTT documentation, verified 2026-07-10.WEBVTT file signature, cue timings with period milliseconds, ordered cue start times, and end times greater than start times. Cue overlap is allowed by the format for captions/subtitles, but many caption delivery workflows still reject or flag overlaps. Source: W3C WebVTT Candidate Recommendation Draft, verified 2026-07-11.HH:MM:SS,mmm --> HH:MM:SS,mmm timing and UTF-8 for modern platform upload, even though legacy SRT files may use other encodings. Source: Library of Congress SRT format description plus platform guidance already cited here, verified 2026-07-11.These are craft defaults. Override them for a formal client style guide, broadcaster spec, language-specific subtitle convention, or platform validation error.
Create a caption master from these inputs:
For each caption cue:
[narrator], [offscreen voice], [woman 1], or client-approved labels.[alarm blares], [soft piano music], [audience laughs], [door unlocks]. Do not caption every incidental sound.Line breaks:
Timing:
Burn-in is typography over moving imagery. Design it like a legible UI:
Provide a transcript when the user asks for one, when the media is audio-only, when training/search/review workflows benefit from it, or when accessibility requirements call for a time-based media alternative.
Basic transcript:
Descriptive transcript:
Plan audio description before picture lock when possible. The cheapest and cleanest solution for many explainers, demos, and training videos is integrated description: write the main narration so it naturally says essential visual information.
Choose the method:
What to describe:
Run a sensory-safety pass for generated motion graphics, glitch edits, lightning, alarms, camera flashes, strobes, rapid cuts, high-contrast pattern flicker, red flashes, and aggressive kinetic typography.
Do not treat localization as "translate the SRT." Prepare a package:
Localization decisions:
Review with the final video, final audio, and exported files.
Content:
Timing and readability:
Accessibility and design:
Delivery:
WEBVTT and uses period milliseconds.This skill includes a standard-library Python 3.11+ validator at scripts/validate_captions.py. Use it when an agent needs a local, deterministic check of SRT or WebVTT sidecar structure before handoff, platform upload, or regression review.
What it validates:
WEBVTT signature/header separation before cue blocks.STYLE, REGION, and cue timing settings; unsupported, duplicate, malformed, or out-of-range settings fail clearly rather than being silently ignored.srt/vtt format.MM:SS.mmm or HH:MM:SS.mmm.Boundaries:
STYLE blocks or prove player-specific rendering behavior.Example CLI calls:
python skills/production/audio-craft/captions-media-accessibility/scripts/validate_captions.py captions_en.srt
python skills/production/audio-craft/captions-media-accessibility/scripts/validate_captions.py --format vtt captions_en.txt
python skills/production/audio-craft/captions-media-accessibility/scripts/validate_captions.py --max-lines 2 --max-chars-per-line 42 --cps-warning 20 captions_en.vtt
python skills/production/audio-craft/captions-media-accessibility/scripts/validate_captions.py --warnings-as-errors captions_en.srtExit codes:
0: no validation errors; warnings may be present unless --warnings-as-errors is set.2: validation errors, or warnings with --warnings-as-errors.3: operational or parse failure, such as unreadable file, invalid UTF-8, unknown format, missing WebVTT header, or malformed cue block.JSON output schema:
{
"tool": "validate_captions",
"version": "1.0.0",
"status": "passed",
"summary": {
"files": 1,
"cues": 2,
"errors": 0,
"warnings": 0
},
"files": [
{
"path": "captions_en.srt",
"format": "srt",
"cues": 2,
"errors": 0,
"warnings": 0
}
],
"findings": [
{
"severity": "warning",
"code": "characters_per_second",
"cue": 2,
"location": {
"file": "captions_en.srt",
"line": 6
},
"message": "Cue reads at 22.50 characters per second; warning threshold is 20"
}
]
}Stable finding fields are severity, code, cue, location, and message. location always contains file and may contain line and column. Known finding codes include invalid_utf8, file_read_failed, format_unknown, webvtt_header_missing, webvtt_header_invalid, webvtt_late_header_block, webvtt_timing_missing, srt_block_too_short, srt_index_invalid, timestamp_syntax, time_order, cue_order, cue_overlap, cue_text_missing, srt_index_sequence, max_lines, max_chars_per_line, characters_per_second, and no_cues.
Production intent: 20-second vertical ad for a running app, English source, uploaded to TikTok/Reels/YouTube Shorts, with always-visible creative captions plus an SRT master.
Approach:
Example SRT excerpt:
1
00:00:00,200 --> 00:00:02,000
[upbeat electronic music]
2
00:00:02,100 --> 00:00:04,300
MAYA: I stopped guessing
what my training needed.
3
00:00:04,500 --> 00:00:07,000
Now RunPilot adjusts my plan
after every workout.
4
00:00:07,200 --> 00:00:09,100
[watch chimes]
Today's run just got easier.Why it is structured this way: the SRT keeps speaker and sound context for accessibility, while the separate burned-in creative layer can use fewer words only because the sidecar remains complete.
Likely failure modes: platform UI covers the bottom line; ASR hears "RunPilot" as "run pilot"; the burned-in layer becomes the only delivered caption and omits [watch chimes].
Production intent: 90-second onboarding video showing how to reset a dashboard filter. The user wants accessible training media, not a cinematic story.
Accessible narration rewrite:
Original narration:
"Click here, then choose last quarter."
Integrated description rewrite:
"In the left sidebar, open Filters, then choose Last quarter from the Date range menu."Additional handoff notes:
Why it is structured this way: essential visual information is moved into the main audio, reducing the need for a separate audio-description track while improving the training script for everyone.
Production intent: English product-launch video localized into Spanish (Latin America), French (France), and Japanese; captions and subtitles are required, dub may be added later.
Handoff:
Files:
- launch_final_ref_23976.mp4
- launch_en-US_captions_SDH.vtt
- launch_en-US_transcript_descriptive.md
- launch_source_dialogue_glossary.csv
- launch_on_screen_text_markers.csv
Deliverables requested:
- es-419 translated subtitles, not SDH, WebVTT
- fr-FR translated subtitles, not SDH, WebVTT
- ja-JP translated subtitles, not SDH, WebVTT
- separate translator notes for UI text that should remain in English
Style:
- Product name "Northstar Studio" is never translated.
- Tone is confident but not slang-heavy.
- Preserve the joke in scene 4 by adapting the idiom, not literal word order.
- If later dubbed, regenerate captions from the final dubbed audio.Why it is structured this way: it separates same-language accessibility assets from translated subtitles, gives locale targets, and avoids freezing English line breaks into other languages.
Verified on 2026-07-10:
© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts) in skills/production/audio-craft/captions-media-accessibility of calesthio/generative-media-skills.
Open the folder on GitHubat commit 8c85352
Captions Media Accessibility next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Captions Media Accessibility this skillcalesthio/generative-media-skills | 197 | — | ~6.9k | Automated safety check: Pass | MIT | |
| Multimedia AccessibilityOwl-Listener/inclusive-design-skills | 105 | — | ~894 | Automated safety check: Pass | MIT | |
| Clipifylouisedesadeleer/clipify | 587 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Oracle Title ForgeSoul-Brews-Studio/arra-oracle-skills-cli | 123 | — | ~2k | Automated safety check: Pass | MIT | |
| Video Transcribewendy7756/AI-Video-Transcriber | 3.3k | — | ~937 | Automated safety check: Notes | Apache-2.0 | |
| YouTube Captions FetcherZeroPointRepo/youtube-skills | 1k | 1 repos | ~1.1k | Automated safety check: Pass | MIT |
Owl-Listener/inclusive-design-skills
Design accessible video, audio, and multimedia content with captions, transcripts, and audio descriptions.
louisedesadeleer/clipify
Find compelling moments in a video — funny dialogue OR repeated impact actions like axe chops, hits, throws, drumbeats — cut them as standalone clips, optionally reformat 16:9 ↔ 9:16, time-warp pans…
Soul-Brews-Studio/arra-oracle-skills-cli
Forge a title + subtitle (or reframe) for a BOOK, article, talk, or any technical piece, then hand off to a cover so it has LIFE and honesty — not clinical/dated.
wendy7756/AI-Video-Transcriber
Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.
ZeroPointRepo/youtube-skills
Fetches timestamped captions or plain-text transcripts for any YouTube video through the TranscriptAPI service, for reading, quoting, translating or accessibility.
content-designer/ux-writing-skill
Applies UX writing practice to interface copy such as buttons, errors, forms and onboarding, using four quality standards and accessibility guidance.
calesthio/generative-media-skills
A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.
calesthio/generative-media-skills
Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…
calesthio/generative-media-skills
Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…
calesthio/generative-media-skills
Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.
calesthio/generative-media-skills
Provider-independent quality assurance for AI-generated and AI-assisted media.
calesthio/generative-media-skills
Provider-independent production workflow for AI agents assembling generated or source media into HyperFrames HTML/CSS/JS videos.
Categories
Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…. Captions Media Accessibility is an agent skill from calesthio/generative-media-skills. Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video recuts, training media, and localized content.
Captions Media Accessibility fits situations like: audio description; flashing/motion safety; caption readability; accessible media handoff files.
Run `npx skills add calesthio/generative-media-skills --skill captions-media-accessibility -a claude-code`. Or copy the skill folder (skills/production/audio-craft/captions-media-accessibility in calesthio/generative-media-skills) into .claude/skills/captions-media-accessibility in your project. Claude Code loads it when a task matches its description.
Run `npx skills add calesthio/generative-media-skills --skill captions-media-accessibility -a codex`. Or copy the skill folder (skills/production/audio-craft/captions-media-accessibility in calesthio/generative-media-skills) into .agents/skills/captions-media-accessibility in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill captions-media-accessibility -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/captions-media-accessibility, .gemini/skills/captions-media-accessibility, .github/skills/captions-media-accessibility and .opencode/skills/captions-media-accessibility in your project.
Going by SKILL.md and its folder, Captions Media Accessibility needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md names 9 domains. As links in the text: w3.org, dcmp.org, help.vimeo.com, loc.gov, partnerhelp.netflixstudios.com, section508.gov, law.cornell.edu, support.google.com and developer.mozilla.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Captions Media Accessibility is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Captions Media Accessibility: Multimedia Accessibility (Owl-Listener/inclusive-design-skills, 105 stars), Clipify (louisedesadeleer/clipify, 587 stars), Oracle Title Forge (Soul-Brews-Studio/arra-oracle-skills-cli, 123 stars) and Video Transcribe (wendy7756/AI-Video-Transcriber, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.
Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.