Agent skill

AI Presenter Video

by NousResearch in NousResearch/hermes-agent

Produces a presenter-led video from a topic or script plus one authorized presenter image, with captions, lip-sync checks and acceptance reports.

MITAuto-check passedMedia & Creative

Install AI Presenter Video

skills CLI
$ npx skills add NousResearch/hermes-agent --skill ai-presenter-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NousResearch/hermes-agent ai-presenter-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NousResearch/hermes-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/optional-skills/creative/ai-presenter-video .claude/skills/ai-presenter-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-presenter-video
GitHub stars
253k
Token cost
~2.3k tokens
SKILL.md length
956 words
Files
9 (incl. scripts, references, assets)
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

Produces a presenter-led video from a topic or script plus one authorized presenter image, with captions, lip-sync checks and acceptance reports.

  • Works in 6 steps: Start or resume a job. New job → Manual input review. Actually look at… → Lock content and audio — read… → …
  • Making a presenter-led explainer video from a script and one portrait image
  • SKILL.md covers Hermes adaptations (read first), Workflow, Operating rules (non-negotiable) and Defaults for minimal input, plus 3 more sections
  • Runs Python and Shell scripts from its folder; calls python3 and bash

What it does

The skill takes a topic or finished script and one image of an authorized adult presenter and produces a presenter-led video. The pipeline covers locked narration, avatar generation with lip-sync checks, captions, deterministic editing, loudness-normalized master and share encodes, and machine and visual acceptance reports. It also handles continuing, revising, captioning, lip-sync repair and re-exporting an existing job.

The workflow is provider-neutral: the agent picks tools from what the session has, such as FAL video and image models, text-to-speech, whisper-style speech recognition for word timestamps and ffmpeg for everything deterministic. Visual review samples frames and a contact sheet for identity, mouth timing, hands, blinking and continuity, while numeric checks come from ffprobe output. The scripts init_job.py, preflight.py and finalize_delivery.sh run without network or credentials, and assets/job.template.json holds the job file layout.

Because remote avatar and voice generation is billable, the agent must state the uploaded assets, requested seconds, known cost, pilot size and retry details before the first paid call. The excerpt is cut off after that rule, and the references cover editing, generation and QA recovery.

When your agent uses it

  • Making a presenter-led explainer video from a script and one portrait image
  • Repairing lip-sync or captions on an existing presenter video job
  • Re-exporting a presenter video as loudness-normalized master and share files
  • Reviewing a generated avatar video for identity and mouth-timing problems

Example prompts

  • “Make a presenter video from this script and the portrait in ./assets/presenter.png.”
  • “Repair the lip-sync on the existing presenter-video job and re-export it.”
  • “Add captions to the finished avatar video and run the acceptance checks.”

Requirements

  • One image of an authorized adult presenter
  • Access to billable avatar and voice generation, such as FAL models and text-to-speech
  • ffmpeg and ffprobe
  • Python for the job scripts

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Start or resume a job. New job
  2. Manual input review. Actually look at the presenter image
  3. Lock content and audio — read references/generation.md. Script →
  4. Plan and generate the presenter — read references/generation.md.
  5. Edit — read references/editing.md. Deterministic timeline driven by
  6. Verify and deliver — read references/qa-recovery.md, render, then

What it can do on your machine

Read from SKILL.md and the folder at commit b30f95a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Presenter Video loads about 2.3k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 19 tokens; SKILL.md has 956 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NousResearch/hermes-agent at commit b30f95a, republished under its MIT licence (© NousResearch). 956 words, ~2,310 tokens.

Download SKILL.mdSave it as .claude/skills/ai-presenter-video/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
ai-presenter-video
description
Make a verified AI presenter video from script + image.
version
1.0.0
author
cclank (https://github.com/cclank/lanshu-create-ai-presenter-video), ported by Hermes Agent
license
MIT
platforms
linux, macos
required_commands
ffmpeg, ffprobe, python3

AI Presenter Video

Turn a topic (or finished script) plus ONE authorized adult presenter image into a complete, publish-ready presenter-led video: locked narration, avatar generation with lip-sync QA, captions, deterministic editing, loudness-normalized master/share encodes, and machine + visual acceptance reports.

Use this skill for new presenter videos AND for continuing, revising, captioning, lip-sync-repairing, or re-exporting an existing presenter-video job. The workflow is provider-neutral: pick generation capabilities from what is actually available in the session (FAL video/image models via image_generate and the video-gen plugin, TTS via text_to_speech, ASR via the whisper/STT tooling, ffmpeg for everything deterministic).

Ported from cclank/lanshu-create-ai-presenter-video (MIT). Upstream body kept substantively verbatim in references/; Hermes adaptations live in this hub file. Scripts are deterministic (no network, no credentials).

Hermes adaptations (read first)

  • Skill dir resolution — upstream hardcoded its own agent's skills path. In Hermes the loader expands ${HERMES_SKILL_DIR} to this skill's installed directory, so every command below uses that token directly:

    bash
    SKILL_DIR="${HERMES_SKILL_DIR}"

    Shell variables do not persist between tool calls — re-paste the assignment (or the expanded path) in each terminal call that uses it.

  • Capability mapping — where the references say "a voice generation capability", use text_to_speech (OpenAI/Edge/ElevenLabs per user config); "presenter/avatar generation" → FAL image-to-video families (Kling, Wan, MiniMax H3 etc.) through the configured video tooling, or an avatar/lipsync endpoint the user has access to; "word-timestamp ASR" → whisper via the STT tooling or faster-whisper in a venv; "deterministic compositor" → ffmpeg filtergraphs, or the hyperframes skill when installed (the editing reference has a HyperFrames section that maps directly onto it).

  • Visual QA — do the "normal-speed visual review" steps with vision_analyze on the generated contact sheet plus sampled frames (identity, mouth timing, hands, blinking, continuity). Numeric checks come from the scripts' ffprobe output.

  • Paid-generation consent — remote avatar/TTS generation is billable. Follow the upstream operating rules: before the first paid call state the uploaded assets, requested seconds, known cost, pilot size, and retry ceiling, and get the user's explicit go-ahead. Never upload the presenter image to a remote provider before remote_upload_approved is true in job.json.

  • Consent flags live under input — rights_confirmed, adult_presenter_confirmed, remote_upload_approved, and voice_clone_approved sit inside the input object of job.json (init flags set them; hand-editing must target input.*, not the job root). manual_input_review.* sits at the root. preflight.py distinguishes errors (block everything) from remote_blockers (block only remote generation) — local script/audio work may proceed while remote is blocked.

Workflow

  1. Start or resume a job. New job:

    bash
    python3 "$SKILL_DIR/scripts/init_job.py" \
      --job-dir ~/Videos/my-presenter-video \
      --presenter-image /path/to/presenter.png \
      --topic "explain context engineering in one minute" \
      --duration 60 --aspect 9:16 \
      --rights-confirmed --adult-presenter-confirmed

    Use --script for an existing script file; other flags: --voice-sample, --supporting-media, --width, --height, --fps, --watermark, --cta. For an existing job, read job.json + QA reports and resume from the earliest unfinished state — never regenerate accepted work.

  2. Manual input review. Actually look at the presenter image (vision_analyze) and listen to any voice sample; record findings by setting the manual_input_review booleans in job.json, e.g.:

    bash
    python3 - <<'PY'
    import json
    p = "~/Videos/my-presenter-video/job.json"  # expand ~ or use an absolute path
    import os; p = os.path.expanduser(p)
    j = json.load(open(p))
    j["manual_input_review"].update(image_viewed=True, single_clear_face=True,
                                    image_has_no_unwanted_text=True)
    json.dump(j, open(p, "w"), indent=2)
    PY

    Then gate:

    bash
    python3 "$SKILL_DIR/scripts/preflight.py" ~/Videos/my-presenter-video/job.json

    Proceed only when ok: true; do remote generation only when remote_ready: true. Note: preflight also updates job.json in place (records the report path) — re-read it after running rather than editing a stale copy.

  3. Lock content and audio — read references/generation.md. Script → full narration via text_to_speech → ASR-verify the narration against the script → record real durations. The locked audio is the master clock for everything downstream.

  4. Plan and generate the presenter — read references/generation.md. Short low-cost pilot first; full run only after the pilot passes identity and mouth-timing review.

  5. Edit — read references/editing.md. Deterministic timeline driven by the locked audio; captions and keyword callouts only after audio and media are final.

  6. Verify and deliver — read references/qa-recovery.md, render, then:

    bash
    bash "$SKILL_DIR/scripts/finalize_delivery.sh" \
      ~/Videos/my-presenter-video/renders/rendered.mp4 \
      ~/Videos/my-presenter-video/outputs my-video

    The finalizer preserves aspect ratio, runs two-pass loudness normalization (program ≈ −16 LUFS), produces master + share encodes, decode-verifies both, writes a delivery report JSON, and emits a nine-frame contact sheet. Inspect the contact sheet with vision_analyze before claiming completion.

Show full SKILL.md (343 more words)Show less

Operating rules (non-negotiable)

  • Confirm image rights, adult status, remote-upload approval, and voice-cloning authorization before the relevant remote action.
  • Never infer or clone a real person's voice from an image; use an authorized sample or a stock TTS voice.
  • Lock the complete narration before presenter generation, caption timing, or final scene boundaries.
  • Mute video sources in the final composition; only the approved narration and intentional mix tracks carry audio.
  • Preserve provider request bodies and task IDs (minus credentials/expiring URLs). Poll interrupted work before resubmitting — avoid double billing.
  • Stop after three rejected paid candidates and summarize the failure mode.
  • Do not claim completion until the final files fully decode and the contact sheet or full playback has been reviewed.

Defaults for minimal input

9:16, 1080×1920, 30fps; topic-derived videos target 45–75s; stock voice when no authorized sample; presenter-led layout with hook → 2–4 beats → close; no music/CTA unless requested; language inferred from the request.

Reference routing

  • references/generation.md — intake, content, voice, capability selection, presenter prompts, paid generation, provider changes.
  • references/editing.md — timeline contract, openings/closes, captions, keyword-callout presets, HyperFrames composition, exports.
  • references/qa-recovery.md — technical acceptance, visual acceptance, and recovery for lip-sync/identity/hands/exposure/freeze/caption/audio faults.

Pitfalls

  • preflight.py requires ffprobe; on a bare box install ffmpeg first.
  • The consent booleans set by init flags land under input.*; editing them at the job-json root silently does nothing (preflight keeps blocking).
  • finalize_delivery.sh needs bash + jq + awk and a fully decodable input — a truncated render fails the decode check by design, not by accident.
  • Long avatar clips drift: prefer one continuous presenter source sliced on the audio timeline over many regenerated chapter clips (identity drift across regenerations is the #1 visual-QA failure).
  • FAL i2v endpoints cap duration (typically 5–15s); plan chapter-level presenter segments accordingly and reuse the pilot's seed/params for consistency where the endpoint supports it.

Verification

Validated hands-on (Aug 2026): init_job.py → job.json with correct state machine; preflight.py correctly blocked on unreviewed inputs, flipped to ok: true after review booleans, and kept remote_ready: false until input.remote_upload_approved; finalize_delivery.sh on a synthetic 5s 1080×1920 render produced decode-verified master (631kbit/s) + share encodes, delivery-report JSON, and a 9-frame contact sheet, exit 0.

© NousResearch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references, assets) in optional-skills/creative/ai-presenter-video of NousResearch/hermes-agent.

  • SKILL.md
  • LICENSE
  • assets/job.template.json
  • references/editing.md
  • references/generation.md
  • references/qa-recovery.md
  • scripts/finalize_delivery.sh
  • scripts/init_job.py
  • scripts/preflight.py

Open the folder on GitHubat commit b30f95a

Compare with similar skills

AI Presenter Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Presenter Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Presenter Video this skillNousResearch/hermes-agent253k—~2.3kAutomated safety check: PassMIT
Super Video MakerBomx/super-video-maker-skill310—~11kAutomated safety check: NotesNone
Hearyourvoicekillernay/HearYourVOICE140—~10kAutomated safety check: NotesMIT
AI Video GenaAAaqwq/AGI-Super-Team1051 repos~819Automated safety check: NotesMIT
Vox DirectorAlisa0808/vox-director2.2k—~5.6kAutomated safety check: PassMIT
Video Productionspeechlab0210/video-production-skill105—~4.1kAutomated safety check: NotesMIT

Similar skills

  • Super Video Maker

    Bomx/super-video-maker-skill

    End-to-end AI video production skill for agentic frameworks.

    310 GitHub stars~11k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Hearyourvoice

    killernay/HearYourVOICE

    The repeatable workflow for short Thai documentary/explainer videos — one topic in, one finished MP4 out.

    140 GitHub stars~10k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • AI Video Gen

    aAAaqwq/AGI-Super-Team

    End-to-end AI video generation - create videos from text prompts using image generation, video synthesis, voice-over, and editing.

    105 GitHub starsUsed in 1 repo~819 tokens
    Media & CreativeAuto-check: notes
  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Video Production

    speechlab0210/video-production-skill

    AI educational video production pipeline. An agent skill from speechlab0210/video-production-skill.

    105 GitHub stars~4.1k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Muapi Director

    Anil-matcha/vox-ai-motion-graphics-generator

    Turn ONE topic, talking-head video, or photo into a finished Vox-style paper-collage explainer / ad video on the MuAPI platform (api.muapi.ai) + local ffmpeg — script, collage keyframes, motion…

    246 GitHub stars~679 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from NousResearch/hermes-agent

All 31 skills in this repo
  • Word DOCX Toolkit

    NousResearch/hermes-agent

    Creates, reads, edits and templates Word .docx files with python-docx scripts, including tracked changes, comments, tables of contents and health checks.

    253k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Grounded Citations

    NousResearch/hermes-agent

    Attaches a numbered, URL-backed citation to every outside fact in an answer or document, rejecting quotes that aren't real.

    253k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • PDF

    NousResearch/hermes-agent

    PDF files: create, read, merge, fill, OCR, edit text. An agent skill from NousResearch/hermes-agent.

    253k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Scrollcraft

    NousResearch/hermes-agent

    Premium scroll-driven landing pages; scroll = timeline. An agent skill from NousResearch/hermes-agent.

    253k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • XLSX

    NousResearch/hermes-agent

    Create, read, edit Excel .xlsx workbooks and CSVs. An agent skill from NousResearch/hermes-agent.

    253k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Powerpoint

    NousResearch/hermes-agent

    Create, read, edit .pptx decks with python-pptx. An agent skill from NousResearch/hermes-agent.

    253k GitHub stars~2.7k tokensUpdated today
    Auto-check passed

Questions about AI Presenter Video

What does AI Presenter Video do?

Produces a presenter-led video from a topic or script plus one authorized presenter image, with captions, lip-sync checks and acceptance reports. The skill takes a topic or finished script and one image of an authorized adult presenter and produces a presenter-led video. The pipeline covers locked narration, avatar generation with lip-sync checks, captions, deterministic editing, loudness-normalized master and share encodes, and machine and visual acceptance reports.

When should I use AI Presenter Video?

AI Presenter Video fits situations like: making a presenter-led explainer video from a script and one portrait image; repairing lip-sync or captions on an existing presenter video job; re-exporting a presenter video as loudness-normalized master and share files; reviewing a generated avatar video for identity and mouth-timing problems.

How do I install AI Presenter Video in Claude Code?

Run `npx skills add NousResearch/hermes-agent --skill ai-presenter-video -a claude-code`. Or copy the skill folder (optional-skills/creative/ai-presenter-video in NousResearch/hermes-agent) into .claude/skills/ai-presenter-video in your project. Claude Code loads it when a task matches its description.

How do I install AI Presenter Video in Codex?

Run `npx skills add NousResearch/hermes-agent --skill ai-presenter-video -a codex`. Or copy the skill folder (optional-skills/creative/ai-presenter-video in NousResearch/hermes-agent) into .agents/skills/ai-presenter-video in your project. Codex loads it when a task matches its description.

Can I use AI Presenter Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NousResearch/hermes-agent --skill ai-presenter-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-presenter-video, .gemini/skills/ai-presenter-video, .github/skills/ai-presenter-video and .opencode/skills/ai-presenter-video in your project.

What does AI Presenter Video need to run?

Going by SKILL.md and its folder, AI Presenter Video needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python3 and bash). Our summary lists: One image of an authorized adult presenter; Access to billable avatar and voice generation, such as FAL models and text-to-speech; ffmpeg and ffprobe; Python for the job scripts.

Does AI Presenter Video access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Presenter Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does AI Presenter Video use?

AI Presenter Video is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Presenter Video use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to AI Presenter Video?

Skills that share tags, products or a category with AI Presenter Video: Super Video Maker (Bomx/super-video-maker-skill, 310 stars), Hearyourvoice (killernay/HearYourVOICE, 140 stars), AI Video Gen (aAAaqwq/AGI-Super-Team, 105 stars) and Vox Director (Alisa0808/vox-director, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Presenter Video?

NousResearch (a GitHub organization) maintains it in NousResearch/hermes-agent, which has 252,613 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 11, 2026.

Source: NousResearch/hermes-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.