Agent skill

Document To Narration

by jwynia in jwynia/agent-skills

Convert written documents to narrated video scripts with TTS audio and word-level timing.

MITAuto-check passedMedia & Creative

Install Document To Narration

skills CLI
$ npx skills add jwynia/agent-skills --skill document-to-narration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jwynia/agent-skills document-to-narration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jwynia/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/education/document-to-narration .claude/skills/document-to-narration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
document-to-narration
GitHub stars
170
Token cost
~3.5k tokens
SKILL.md length
970 words
Files
14 (incl. scripts, references)
Skills in repo
111
Repo updated
First seen
Licence
MIT

At a glance

Convert written documents to narrated video scripts with TTS audio and word-level timing.

  • Works in 4 steps: Setup (First Time Only) → Prepare Your Document → Run the Pipeline → …
  • Preparing essays
  • SKILL.md covers Core Principle, When to Use This Skill, Prerequisites and Complete Pipeline, plus 5 more sections
  • Runs Python, TypeScript and Shell scripts from its folder; calls python, deno and pip

What it does

Document To Narration is an agent skill from jwynia/agent-skills. Convert written documents to narrated video scripts with TTS audio and word-level timing. Use when preparing essays, blog posts, or articles for video narration. Outputs scene files, audio, and VTT with precise word timestamps. Keywords: narration, voiceover, TTS, scenes, audio, timing, video script, spoken.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts and reference files (for example `data/adaptation-rules.json`, `references/spoken-adaptation-guide.md` and `scripts/extract-scene-boundaries.py`). Compatibility notes: Requires Deno, Python 3.12 with venv, ffmpeg, whisper-cpp

It sits in Media & Creative, covering Text to speech and voice and Video scripts and shorts. It works with Whisper. The licence is MIT.

When your agent uses it

  • Preparing essays
  • Articles for video narration

Example prompts

  • “/document-to-narration”

Requirements

  • Python 3
  • Node.js
  • A Bash shell
  • Compatibility (from SKILL.md): Requires Deno, Python 3.12 with venv, ffmpeg, whisper-cpp

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Setup (First Time Only)
  2. Prepare Your Document
  3. Run the Pipeline
  4. Review Output

What it can do on your machine

Read from SKILL.md and the folder at commit e02ec7e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (Python, TypeScript and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • deno
    • pip
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • deno.land

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Deno, Python 3.12 with venv, ffmpeg, whisper-cpp

    From compatibility in the SKILL.md frontmatter.

Context cost

Document To Narration loads about 3.5k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 970 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jwynia/agent-skills at commit e02ec7e, republished under its MIT licence (© jwynia). 970 words, ~3,543 tokens.

Download SKILL.mdSave it as .claude/skills/document-to-narration/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
document-to-narration
description
Convert written documents to narrated video scripts with TTS audio and word-level timing. Use when preparing essays, blog posts, or articles for video narration. Outputs scene files, audio, and VTT with precise word timestamps. Keywords: narration, voiceover, TTS, scenes, audio, timing, video script, spoken.
compatibility
Requires Deno, Python 3.12 with venv, ffmpeg, whisper-cpp
license
MIT
metadata.author
jwynia
metadata.version
1.0
metadata.domain
video-production
metadata.type
generator
metadata.mode
generative

Document to Narration

Convert written documents into narrated video scripts with precise word-level timing.

Core Principle

The agent interprets; the document guides. Rather than rigid template-based splits, this skill uses agent judgment to find where the content naturally breathes, argues, and transitions. The document's argument flow determines scene breaks, not a predetermined structure.

When to Use This Skill

Use this skill when:

  • Converting a blog post or essay to video narration
  • Preparing content for TTS audio generation
  • Breaking long-form content into digestible scenes
  • Creating word-level synchronized captions for video

Do NOT use this skill when:

  • The content is already in scene/script format
  • You need real-time voice synthesis (this is batch processing)
  • Working with dialogue or multi-speaker content (single voice only)

Prerequisites

  • Deno installed (https://deno.land/)
  • Python 3.12 with venv support
  • ffmpeg for audio conversion
  • whisper-cpp (installed via @remotion/install-whisper-cpp)
  • TTS model at tts/model/ (not in git due to size - see Model Setup below)

Complete Pipeline

There are two approaches: per-scene (legacy) and full narration (recommended).

Generates a single audio file for consistent volume and pacing:

Document (.md)
    ↓ [agent interprets scene breaks]
Scene .txt files (01-scene-name.txt, 02-scene-name.txt, ...)
    ↓ [TTS via narrate-full.py - SINGLE PASS]
full-narration.wav (one consistent audio file)
    ↓ [Whisper via transcribe-full.py]
full-narration.json + full-narration.vtt (word-level timing)
    ↓ [extract-scene-boundaries.py]
Scene timing boundaries for video composition
Per-Scene Pipeline (Legacy)

Generates separate audio per scene - can cause volume inconsistencies:

Scene .txt files
    ↓ [TTS via narrate-scenes.py - MULTIPLE PASSES]
Scene .wav files (volume may vary between scenes)
    ↓ [concatenate]
Combined audio (may have clipping at boundaries)

Warning: Per-scene TTS generates audio with different volume levels and pacing. When concatenated, this causes audible jumps and clipping. Use the full narration pipeline instead.

Quick Start

bash
cd .claude/skills/document-to-narration
source tts/.venv/bin/activate

# 1. Split document into scenes (manual or scripted)
deno run --allow-read --allow-write scripts/split-to-scenes.ts input.md --output ./output/

# 2. Generate single audio file
python scripts/narrate-full.py ./output/scenes/

# 3. Transcribe with word-level timestamps
python scripts/transcribe-full.py ./output/full-narration.wav

# 4. Extract scene boundaries for video timing
python scripts/extract-scene-boundaries.py ./output/scenes/ ./output/full-narration.json --typescript
Legacy Per-Scene Pipeline
bash
# 1. Split document into scenes
deno run --allow-read --allow-write scripts/split-to-scenes.ts input.md --output ./output/

# 2. Generate audio per scene (may have volume inconsistencies)
source tts/.venv/bin/activate
python scripts/narrate-scenes.py ./output/scenes/

# 3. Transcribe (DEPRECATED: transcribe-scenes.ts requires whisper-cpp)
# Use transcribe-full.py instead after concatenating audio

Instructions

Step 1: Setup (First Time Only)
Create Python Virtual Environment
bash
cd .claude/skills/document-to-narration/tts
python3.12 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
TTS Model Setup

The fine-tuned voice model (~7.8GB) is not included in git due to size. Place your Qwen3-TTS model files in tts/model/:

tts/model/
├── config.json
├── generation_config.json
├── model.safetensors      # Main model weights
├── tokenizer_config.json
├── vocab.json
├── merges.txt
└── speech_tokenizer/
    └── ...
Install Whisper (if not already installed)

The @remotion/install-whisper-cpp package handles this:

typescript
import { installWhisperCpp, downloadWhisperModel } from '@remotion/install-whisper-cpp';

await installWhisperCpp({ to: './whisper-cpp', version: '1.5.5' });
await downloadWhisperModel({ model: 'medium', folder: './whisper-cpp' });
Step 2: Prepare Your Document

The skill works best with:

  • Markdown documents with clear heading structure (H1, H2)
  • Well-structured arguments with distinct sections
  • Content that reads naturally aloud
Step 3: Run the Pipeline
bash
deno run -A scripts/full-pipeline.ts /path/to/essay.md --output ./output/essay-name/
Step 4: Review Output
output/essay-name/
├── scenes/
│   ├── 01-opening-hook.txt      # Scene script
│   ├── 01-opening-hook.wav      # Generated audio
│   ├── 01-opening-hook.vtt      # Word-level captions
│   ├── 02-core-argument.txt
│   ├── 02-core-argument.wav
│   ├── 02-core-argument.vtt
│   └── ...
└── manifest.json                # Complete timing data

Scene Boundary Heuristics

The agent identifies scene breaks using these heuristics:

Strong Boundaries (Almost Always Break)
  • H2 heading changes
  • "Here's the thing" / "The point is" pivot statements
  • Major metaphor introduction
  • Explicit enumeration ("First...", "Second...")
  • Significant perspective shifts
Moderate Boundaries (Consider Breaking)
  • Long paragraph after short ones (or vice versa)
  • Example-to-principle transitions
  • "But" / "However" / "Meanwhile" at paragraph start
  • Question-then-answer patterns
Weak Boundaries (Usually Keep Together)
  • Paragraph-to-paragraph within same example
  • Sequential evidence for same point
  • Build-up to a punchline/reveal
Scene Length Guidance
  • Target: 100-300 words per scene (30-90 seconds of audio)
  • Minimum: 50 words (avoid micro-scenes)
  • Maximum: 500 words (avoid cognitive overload)

Anti-Patterns

The Paragraph Slicer

Pattern: Breaking at every paragraph or heading mechanically. Problem: Ignores argument flow. Scenes feel choppy and disconnected. Fix: Look for rhetorical units, not structural units. Multiple paragraphs often form one scene.

The Wall of Text

Pattern: Keeping entire sections as single scenes. Problem: Creates TTS audio that's too long. Loses natural breathing room. Fix: Target 100-300 words. Find the natural pause point within sections.

The Verbatim Transcriber

Pattern: Copying written text exactly without spoken adaptation. Problem: Written conventions don't work when spoken. Parentheticals, complex punctuation, and nested clauses confuse TTS and listeners. Fix: Apply adaptation rules. Read it aloud mentally.

The Over-Adapter

Pattern: Rewriting content so heavily it loses the author's voice. Problem: The result doesn't sound like the original author. Fix: Preserve voice, adjust mechanics. If the author uses rhetorical questions, keep them.

Available Scripts

scripts/split-to-scenes.ts

Parse a markdown document and output scene text files.

bash
deno run --allow-read --allow-write scripts/split-to-scenes.ts input.md --output ./output/
deno run --allow-read --allow-write scripts/split-to-scenes.ts input.md --output ./output/ --adapt
deno run --allow-read scripts/split-to-scenes.ts input.md --dry-run

Options:

  • --output - Directory for scene files (created if doesn't exist)
  • --adapt - Apply spoken adaptation rules
  • --dry-run - Preview scene breaks without writing files

Output: Numbered .txt files and initial manifest.json

Show full SKILL.md (386 more words)Show less

Generate a single TTS audio file from all scene files. Produces consistent volume and pacing.

bash
python scripts/narrate-full.py ./output/scenes/
python scripts/narrate-full.py ./output/scenes/ --force
python scripts/narrate-full.py ./output/scenes/ --speaker jwynia
python scripts/narrate-full.py ./output/scenes/ --output ./custom/path/audio.wav

Options:

  • --force - Regenerate even if output exists
  • --speaker - Speaker name (default: jwynia)
  • --output - Custom output path (default: ../full-narration.wav)

Output: Single full-narration.wav in parent directory of scenes

scripts/narrate-scenes.py (Legacy)

Generate TTS audio for each scene file separately. Not recommended - can cause volume inconsistencies when concatenated.

bash
python scripts/narrate-scenes.py ./output/scenes/
python scripts/narrate-scenes.py ./output/scenes/ --force
python scripts/narrate-scenes.py ./output/scenes/ --speaker jwynia

Options:

  • --force - Regenerate even if output exists
  • --speaker - Speaker name (default: jwynia)

Output: .wav files alongside each .txt file

Transcribe audio with word-level timestamps using Python's openai-whisper.

bash
python scripts/transcribe-full.py ./output/full-narration.wav
python scripts/transcribe-full.py ./output/full-narration.wav --model large-v3
python scripts/transcribe-full.py ./output/full-narration.wav --output-dir ./captions/

Options:

  • --model - Whisper model: tiny, base, small, medium, large, large-v2, large-v3 (default: medium)
  • --output-dir - Output directory (default: same as audio file)

Output:

  • .vtt file with word-level timestamps
  • .json file with captions array for Remotion

Dependencies: Requires openai-whisper in Python environment:

bash
pip install openai-whisper
scripts/extract-scene-boundaries.py

Extract scene timing boundaries from transcript by matching scene opening phrases.

bash
# Human-readable table
python scripts/extract-scene-boundaries.py ./output/scenes/ ./output/full-narration.json

# JSON output
python scripts/extract-scene-boundaries.py ./output/scenes/ ./output/full-narration.json --json

# TypeScript for Video.tsx
python scripts/extract-scene-boundaries.py ./output/scenes/ ./output/full-narration.json --typescript

Options:

  • --json - Output as JSON array
  • --typescript - Output as TypeScript code for Video.tsx scenes array

Output: Scene numbers, slugs, start times, and durations

scripts/transcribe-scenes.ts (Deprecated)

Deprecated: Requires whisper-cpp binary which may not be installed. Use transcribe-full.py instead.

Transcribe per-scene audio files using whisper-cpp.

bash
deno run --allow-read --allow-write --allow-run scripts/transcribe-scenes.ts ./output/scenes/

Output: .vtt files with word-level timestamps

scripts/full-pipeline.ts

Orchestrate the complete pipeline.

bash
deno run -A scripts/full-pipeline.ts input.md --output ./output/project-name/

Options:

  • --output - Output directory (required)
  • --adapt - Apply spoken adaptation
  • --skip-tts - Skip audio generation (text only)
  • --skip-transcribe - Skip Whisper transcription

Output Format

manifest.json
json
{
  "source": "appliance-vs-trade-tool-draft.md",
  "created_at": "2024-01-15T10:30:00Z",
  "total_scenes": 9,
  "total_duration_seconds": 420,
  "scenes": [
    {
      "number": 1,
      "slug": "popcorn-opening",
      "word_count": 185,
      "audio_duration_seconds": 55.2,
      "files": {
        "text": "scenes/01-popcorn-opening.txt",
        "audio": "scenes/01-popcorn-opening.wav",
        "captions": "scenes/01-popcorn-opening.vtt"
      },
      "captions": [
        { "text": "Two", "startMs": 0, "endMs": 180, "confidence": 0.98 },
        { "text": "people", "startMs": 180, "endMs": 450, "confidence": 0.97 }
      ]
    }
  ]
}
VTT Format
vtt
WEBVTT

00:00.000 --> 00:00.180
Two

00:00.180 --> 00:00.450
people

00:00.450 --> 00:00.720
walk

00:00.720 --> 00:01.100
into

Spoken Adaptation

When --adapt is enabled, the skill transforms written conventions to spoken equivalents:

WrittenSpoken
Parenthetical asidesEm-dash or separate sentence
"e.g.""for example"
"i.e.""that is"
Long nested clausesSplit into multiple sentences
SemicolonsPeriods
*emphasis*Context-appropriate stress

Preserve:

  • Author's voice and tone
  • Rhetorical questions
  • Deliberate repetition
  • Key phrases and memorable formulations

Integration

With remotion-designer
  • Pass manifest scene list to remotion-designer
  • Each scene becomes a visual design unit
  • Word-level timing drives text animation
With Remotion Compositions
tsx
import { Audio, useCurrentFrame, Sequence } from 'remotion';
import manifest from './output/manifest.json';

// Use scene durations for Sequence timing
{manifest.scenes.map((scene, i) => (
  <Sequence
    from={accumulatedFrames}
    durationInFrames={scene.audio_duration_seconds * fps}
  >
    <Audio src={staticFile(scene.files.audio)} />
    <CaptionRenderer captions={scene.captions} />
  </Sequence>
))}

Technical Notes

WAV Format Conversion

Whisper requires 16kHz mono WAV. The pipeline handles conversion automatically:

bash
ffmpeg -i input.wav -ar 16000 -ac 1 output_16khz.wav
TTS Model

The fine-tuned voice model (~7.8GB) is bundled at tts/model/. Uses Qwen3-TTS with custom speaker embedding.

Performance
  • TTS: ~5-30 seconds per sentence (Apple Silicon MPS or NVIDIA CUDA)
  • Whisper: ~0.5-2x realtime depending on model size
  • Full essay (~2000 words): ~10-20 minutes total processing

What This Skill Does NOT Do

  • Generate video visuals (use remotion-designer)
  • Real-time voice synthesis
  • Multi-speaker dialogue
  • Edit or improve the content's argument
  • Make editorial changes beyond mechanical spoken adaptation

© jwynia, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (scripts, references) in skills/education/document-to-narration of jwynia/agent-skills.

  • SKILL.md
  • .gitignore
  • data/adaptation-rules.json
  • references/spoken-adaptation-guide.md
  • scripts/extract-scene-boundaries.py
  • scripts/full-pipeline.ts
  • scripts/narrate-full.py
  • scripts/narrate-scenes.py
  • scripts/split-to-scenes.ts
  • scripts/transcribe-full.py
  • scripts/transcribe-scenes.ts
  • tts/model/.gitkeep
  • tts/requirements.txt
  • tts/setup-venv.sh

Open the folder on GitHubat commit e02ec7e

Compare with similar skills

Document To Narration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Document To Narration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Document To Narration this skilljwynia/agent-skills170—~3.5kAutomated safety check: PassMIT
ShortsAgriciDaniel/claude-shorts219—~3.2kAutomated safety check: NotesMIT
ShowtimeFavioVazquez/showtime220—~3kAutomated safety check: PassMIT
TSX Vertical Shorts Builderhassancs91/claude-faceless-shorts-creator271—~1.9kAutomated safety check: NotesMIT
Video Scriptzenstory-ai/video-recap-skills561—~2.4kAutomated safety check: PassMIT
Roadshow Video Generatoranbeime/skill7.8k—~1.4kAutomated safety check: NotesNone

Similar skills

  • Shorts

    AgriciDaniel/claude-shorts

    Interactive longform-to-shortform video creator. An agent skill from AgriciDaniel/claude-shorts.

    219 GitHub stars~3.2k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Showtime

    FavioVazquez/showtime

    A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…

    220 GitHub stars~3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • TSX Vertical Shorts Builder

    hassancs91/claude-faceless-shorts-creator

    Builds a roughly 40-second vertical short end to end from a topic, scripting beats, rendering TSX compositions and layering voice and captions.

    271 GitHub stars~1.9k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Video Script

    zenstory-ai/video-recap-skills

    对已完成分析的视频进行导演与剪辑策划,再写带时间戳的中文解说并校验;也处理已有短片的 宣发标题、花字修订和外部文案回填。普通策划输入 workdir 的 agentnarrationbrief.md 与 vlmanalysis.json;文案返修输入当前成片的工程与内容证据。策划输出 recapstoryplan.json、visualaudioboard.json、 可选…

    561 GitHub stars~2.4k tokensUpdated 6 days ago
    Media & CreativeAuto-check passed
  • Turns a document into a narrated roadshow video through ten staged roles, from document analysis and slide planning to audio, subtitles and final composition.

    7.8k GitHub stars~1.4k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Bggg Tiktok Readvideo

    binggandata/bggg-skills

    把 TikTok、Reels、YouTube Shorts、UGC 广告、本地 MP4/MOV/WebM 等视频拆成 Codex 可读的视频上下文。

    605 GitHub stars~1.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from jwynia/agent-skills

All 111 skills in this repo
  • Devcontainer

    jwynia/agent-skills

    Diagnose devcontainer configuration problems and guide development environment setup.

    170 GitHub stars~1.2k tokensUpdated 7 mo ago
    Auto-check: notes
  • Frontend Design

    jwynia/agent-skills

    Create distinctive, production-grade frontend interfaces with high design quality.

    170 GitHub stars~3.2k tokensUpdated 7 mo ago
    Auto-check passed
  • Gitea Workflow

    jwynia/agent-skills

    Orchestrate agile development workflows for Gitea repositories using the tea CLI.

    170 GitHub stars~3.8k tokensUpdated 7 mo ago
    Auto-check passed
  • Godot Asset Generator

    jwynia/agent-skills

    Generate game assets using AI image generation APIs (DALL-E, Replicate, fal.ai) and prepare them for Godot.

    170 GitHub stars~3.8k tokensUpdated 7 mo ago
    Auto-check passed
  • Mastra Hono

    jwynia/agent-skills

    Develop AI agents, tools, and workflows with Mastra v1 Beta and Hono servers.

    170 GitHub stars~2.9k tokensUpdated 7 mo ago
    Auto-check passed
  • PPTX Generator

    jwynia/agent-skills

    Create and manipulate PowerPoint PPTX files programmatically.

    170 GitHub stars~3.1k tokensUpdated 7 mo ago
    Auto-check passed

Works with

Questions about Document To Narration

What does Document To Narration do?

Convert written documents to narrated video scripts with TTS audio and word-level timing. Document To Narration is an agent skill from jwynia/agent-skills. Convert written documents to narrated video scripts with TTS audio and word-level timing.

When should I use Document To Narration?

Document To Narration fits situations like: preparing essays; articles for video narration.

How do I install Document To Narration in Claude Code?

Run `npx skills add jwynia/agent-skills --skill document-to-narration -a claude-code`. Or copy the skill folder (skills/education/document-to-narration in jwynia/agent-skills) into .claude/skills/document-to-narration in your project. Claude Code loads it when a task matches its description.

How do I install Document To Narration in Codex?

Run `npx skills add jwynia/agent-skills --skill document-to-narration -a codex`. Or copy the skill folder (skills/education/document-to-narration in jwynia/agent-skills) into .agents/skills/document-to-narration in your project. Codex loads it when a task matches its description.

Can I use Document To Narration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jwynia/agent-skills --skill document-to-narration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/document-to-narration, .gemini/skills/document-to-narration, .github/skills/document-to-narration and .opencode/skills/document-to-narration in your project.

What does Document To Narration need to run?

Going by SKILL.md and its folder, Document To Narration needs Python, TypeScript and a shell for the scripts in its folder and the command-line tools its instructions call (python, deno, pip and ffmpeg). Our summary lists: Python 3; Node.js; A Bash shell. Compatibility (from SKILL.md): Requires Deno, Python 3.12 with venv, ffmpeg, whisper-cpp.

Does Document To Narration access the network?

SKILL.md names 1 domain. As links in the text: deno.land. This is read from the text; nothing was executed.

Is Document To Narration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Document To Narration use?

Document To Narration is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Document To Narration use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 908 tokens, read only when the agent opens those files.

What are the alternatives to Document To Narration?

Skills that share tags, products or a category with Document To Narration: Shorts (AgriciDaniel/claude-shorts, 219 stars), Showtime (FavioVazquez/showtime, 220 stars), TSX Vertical Shorts Builder (hassancs91/claude-faceless-shorts-creator, 271 stars) and Video Script (zenstory-ai/video-recap-skills, 561 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Document To Narration?

jwynia (a GitHub user) maintains it in jwynia/agent-skills, which has 170 GitHub stars. The repository holds 111 skills in this directory. The repository was last updated on February 24, 2026.

Source: jwynia/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.