Agent skill

Proof Video

by openclaw in openclaw/openclaw

Add subtitles, captions, narration cues, or zoom to a proof video or PR recording using repo-local capture helpers and a system ffmpeg renderer.

MITAuto-check passedMedia & Creative

Install Proof Video

skills CLI
$ npx skills add openclaw/openclaw --skill proof-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openclaw/openclaw proof-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openclaw/openclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/proof-video .claude/skills/proof-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
proof-video
GitHub stars
392k
Token cost
~2.4k tokens
SKILL.md length
774 words
Files
1
Skills in repo
93
Repo updated
First seen
Licence
MIT

At a glance

Add subtitles, captions, narration cues, or zoom to a proof video or PR recording using repo-local capture helpers and a system ffmpeg renderer.

  • Tasks that involve Transcription
  • SKILL.md covers Standalone recording, Render, Cue sidecar v1 and native… and Verified capture mode and…
  • Calls node, brew and apt
  • Tasks that involve Text to speech and voice

What it does

Proof Video is an agent skill from openclaw/openclaw. Add subtitles, captions, narration cues, or zoom to a proof video or PR recording using repo-local capture helpers and a system ffmpeg renderer.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription, Text to speech and voice and Video production. It works with FFmpeg. The repository describes itself as: The AI that really does things. Any OS. Any Platform. The lobster way. 🦞. The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Text to speech and voice
  • Tasks that involve Video production

Example prompts

  • “/proof-video”

What it can do on your machine

Read from SKILL.md and the folder at commit 1eb5970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node
    • brew
    • apt

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Proof Video loads about 2.4k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 774 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openclaw/openclaw at commit 1eb5970, republished under its MIT licence (© openclaw). 774 words, ~2,439 tokens.

Download SKILL.mdSave it as .claude/skills/proof-video/SKILL.md (or your agent's skills folder).
name
proof-video
description
Add subtitles, captions, narration cues, or zoom to a proof video or PR recording using repo-local capture helpers and a system ffmpeg renderer.

Proof video

Use this developer skill for captioned Control UI recordings and native captures. Captions are burned into an MP4; they are visual narration, not synthesized audio. Keep the raw video as evidence, attach the polished MP4, and never commit media or .artifacts/ output. Inspect the entire capture for unrelated or private content before publishing under the PR's existing authorization.

Standalone recording

From the repository root, save this script as .artifacts/proof-video/record.mts (create the parent directory first), then run node --import tsx .artifacts/proof-video/record.mts. It starts an isolated mocked Control UI, records five captions, a real toolbar zoom, fast-forwarded startup, and an eight-second pause at 4× speed. Every invocation allocates a fresh evidence directory.

ts
import { chromium } from "playwright";
import { createControlUiE2eArtifactDir } from "../../ui/src/test-helpers/control-ui-e2e-artifacts.ts";
import {
  installMockGateway,
  resolvePlaywrightChromiumExecutablePath,
  startControlUiE2eServer,
} from "../../ui/src/test-helpers/control-ui-e2e.ts";
import { startProofRecording } from "../../ui/src/test-helpers/proof-video.ts";

const dir = createControlUiE2eArtifactDir("captioned-toolbar", ".artifacts/proof-video");
const server = await startControlUiE2eServer(undefined, { source: true });
const browser = await chromium.launch({
  executablePath: resolvePlaywrightChromiumExecutablePath(chromium.executablePath()),
  args: ["--force-device-scale-factor=2"],
});
try {
  const recording = await startProofRecording(browser, {
    dir,
    contextOptions: { colorScheme: "dark", locale: "en-US", reducedMotion: "reduce" },
  });
  const { page } = recording;
  const sessions = ["Main", "Weekly planning", "Reading list"].map((label, index) => ({
    key: index === 0 ? "agent:main:main" : `agent:main:proof-${index}`,
    label,
    displayName: label,
    kind: "direct",
    status: "done",
    hasActiveRun: false,
    updatedAt: Date.parse("2026-09-01T12:00:00Z") - index * 60000,
    contextTokens: 200000,
    totalTokens: 0,
    model: "gpt-5.6-luna",
    modelProvider: "openai",
  }));
  const gateway = await installMockGateway(page, {
    sessionKey: "agent:main:main",
    sessionArchiveFiltering: true,
    methodResponses: {
      "sessions.list": {
        sessions,
        count: sessions.length,
        totalCount: sessions.length,
        hasMore: false,
        offset: 0,
        limitApplied: 50,
        nextOffset: null,
        path: "",
        ts: sessions[0].updatedAt,
        defaults: { contextTokens: 200000, model: "gpt-5.6-luna", modelProvider: "openai" },
      },
    },
  });
  recording.caption("Open the mocked Control UI.");
  recording.speed(8);
  await page.goto(`${server.baseUrl}chat`);
  await gateway.waitForRequest("sessions.list");
  await page.getByRole("link", { name: "Weekly planning", exact: true }).waitFor();
  const toolbar = page.locator(".sidebar-session-toolbar");
  await toolbar.getByText("Sessions", { exact: true }).waitFor();
  recording.speed(1);
  await page.waitForTimeout(1500);
  recording.caption("Find your conversations in the session list.");
  await page.waitForTimeout(3000);
  await recording.zoom(toolbar);
  recording.caption("Zoom in on the session filter controls.");
  await toolbar.getByRole("button", { name: "Filter & sort" }).click();
  await page.locator(".sidebar-session-sort-menu").waitFor({ state: "visible" });
  await page.waitForTimeout(4000);
  recording.zoomOut();
  recording.caption("Fast-forward an eight-second pause.");
  recording.speed(4);
  await page.waitForTimeout(8000);
  recording.speed(1);
  recording.caption("Return to the full session list.");
  await page.getByRole("menuitemradio", { name: "All", exact: true }).click();
  await page.waitForFunction(() =>
    document
      .querySelector(".sidebar-session-sort")
      ?.classList.contains("sidebar-session-sort--filtered"),
  );
  await page.waitForTimeout(3000);
  console.log(JSON.stringify(await recording.finish(), null, 2));
} finally {
  await browser.close();
  await server.close();
}

Use caption(text) to replace the current caption, clearCaption() to end it, zoom(locatorOrRect, { factor: 2, padding: 24, ease: 0.5 }) to start or replace a zoom, and zoomOut() to end it. Rectangles use viewport CSS pixels. Padding expands and clamps the target rectangle; its center determines the zoom focus. speed(factor) changes playback speed until the next call; speed(1) restores normal playback. finish() closes open cues at the finalized video's end, closes the recording context, rasterizes captions, writes cues.json, and returns { dir, videoPath, cuesPath }. The caller closes the browser and server.

The helper is test tooling at ui/src/test-helpers/proof-video.ts; production code must not import it. contextOptions carries ordinary Playwright context options, while the helper owns the viewport, recording size, and device scale. Keep narration to one or two short lines; oversized text fails instead of silently clipping. Caption PNGs are transparent viewport-width, 120-pixel-high strips with a centered dark rounded card and 30-pixel system text.

Render

Use the exact directory printed by the recording:

bash
node --import tsx scripts/render-proof-video.mts --cues .artifacts/proof-video/<capture>/cues.json

The renderer prints the output path and duration. Optional --input <video> overrides the raw input and --out <file.mp4> chooses the output. Existing output files are refused, preserving raw evidence and earlier renders. It uses trusted system ffmpeg and ffprobe, never Playwright's bundled encoder. Install a full system build with brew install ffmpeg or apt install ffmpeg.

Speed segments split the raw timeline with trim/setpts/concat, then normalize it to the input frame rate. Caption windows and zoom transitions are remapped to output time. Zoom uses zoompan with d=1 and matching input FPS; a 100-frame, 4-second, 25-FPS synthetic capture retained exactly 100 frames and 4 seconds, with a visibly and numerically different zoom hold. Dynamic crop width/height expressions do not animate: ffmpeg evaluates them only at initialization. Output is viewport-sized H.264 MP4, yuv420p, CRF 23, with faststart and no audio.

Show full SKILL.md (365 more words)Show less

Cue sidecar v1 and native captures

Save a hand-written cues.json beside an untouched native .mov or .webm. Use scale for actual raw pixels per viewport CSS pixel, not display metadata. Input dimensions must equal viewport × scale and frame rate must be constant. All times below are seconds in the RAW video. Caption image paths are relative to the sidecar directory. Zoom cues cannot overlap other zoom cues; speed cues cannot overlap other speed cues. Adjacent cues are allowed.

json
{
  "version": 1,
  "video": "capture.mov",
  "viewport": { "width": 1280, "height": 800 },
  "scale": 2,
  "captionHeight": 120,
  "cues": [
    {
      "kind": "caption",
      "start": 1,
      "end": 4,
      "text": "Inspect the session controls.",
      "image": "captions/0.png"
    },
    {
      "kind": "zoom",
      "start": 1,
      "end": 4,
      "rect": { "x": 20, "y": 150, "width": 280, "height": 80 },
      "factor": 2,
      "ease": 0.5
    },
    { "kind": "speed", "start": 5, "end": 9, "factor": 4 }
  ]
}
bash
node --import tsx scripts/render-proof-video.mts --cues <dir>/cues.json --rasterize

--rasterize creates missing caption PNGs with fresh Playwright Chromium pages. The exported rasterizeCaptions(browser, cues, dir) also supports programmatic native-capture workflows. Neither path needs libass or drawtext: Chromium rasterizes text and ffmpeg overlays PNGs after the zoom.

Verified capture mode and limitations

On macOS with Playwright 1.62.1 and Chromium 151.0.7922.34, context deviceScaleFactor: 2 produced a 2560×1600 WebM containing only 1280×800 page content plus gray padding. Launching with --force-device-scale-factor=2 and leaving context emulation at 1 recorded true 2× content: a 100-CSS-pixel marker measured 200×200 at raw coordinates (2360, 1400). The helper defaults to deviceScaleFactor: 2 for recording size and sidecar scale; the launch flag in the template is required. This parameter does not enable context emulation. For verified 1× capture, launch without the flag and pass deviceScaleFactor: 1. The renderer uses Lanczos pre-upscaling for 1× sources, but cannot restore lost detail: a 2× zoom on those recordings is softer.

The cue clock starts at context creation using performance.now(). Playwright's video clock can differ by roughly 100 ms, and startup/load can increase the skew; inspect action/caption alignment in the finalized recording. Avoid clipped, locator, or full-page screenshots during recording: Chromium can corrupt the screencast surface. Caption rasterization happens only after recording closes, in a fresh context. See the Control UI E2E skill's screenshot guidance.

GitHub's player ignores subtitle tracks, so captions must be burned in. The renderer requires crop, scale, overlay, setpts, trim, concat, fps, and zoompan in the system ffmpeg build. It does not use subtitles= or drawtext=. Variable frame-rate native captures need conversion to a separate constant-frame-rate working copy first; retain the original. Frame-rate quantization can shift speed-segment boundaries by about one output frame. Check output duration and extract caption, zoom-hold, and sped-segment frames before attaching the MP4.

© openclaw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/proof-video of openclaw/openclaw.

Open the folder on GitHubat commit 1eb5970

Compare with similar skills

Proof Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Proof Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Proof Video this skillopenclaw/openclaw392k—~2.4kAutomated safety check: PassMIT
Stage EditOrkas-AI/Orkas-VideoStudio498—~2.4kAutomated safety check: PassMIT
Explain Videolimin112/min-skill362—~2.4kAutomated safety check: PassNone
Qiaomu Cutjoeseesun/qiaomu-cut-skill369—~6.8kAutomated safety check: NotesMIT
ShowtimeFavioVazquez/showtime178—~3kAutomated safety check: PassMIT
Content To Videoarchitectds/modeldock117—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Stage Edit

    Orkas-AI/Orkas-VideoStudio

    Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

    498 GitHub stars~2.4k tokensUpdated 16 days ago
    Media & CreativeAuto-check passed
  • Explain Video

    limin112/min-skill

    Build a narrated explainer video from a concept — discussion → structure → HTML slide deck → narration script → TTS voice → subtitles → background music → Playwright screen recording → ffmpeg…

    362 GitHub stars~2.4k tokensUpdated 15 days ago
    Media & CreativeAuto-check passed
  • Qiaomu Cut

    joeseesun/qiaomu-cut-skill

    把一句话需求转成可复现、可验收视频工程的乔木智能剪辑导演。Use when the user asks to create, plan, edit, remix, explain, narrate, subtitle, animate, composite, or render a video—including one-line requests such as “制作一个科普视频:介绍…

    369 GitHub stars~6.8k tokensUpdated 10 days ago
    Media & CreativeAuto-check: notes
  • Showtime

    FavioVazquez/showtime

    A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…

    178 GitHub stars~3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Content To Video

    architectds/modeldock

    Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video.

    117 GitHub stars~2.4k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Vox Explainer

    CK42BB/vox-explainer-skill

    End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt.

    110 GitHub stars~2.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed

More from openclaw/openclaw

All 93 skills in this repo
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Tmux

    openclaw/openclaw

    Control tmux sessions/panes for interactive CLIs: list, capture output, send keys, paste text, monitor prompts.

    392k GitHub starsUsed in 2 repos~640 tokens
    Auto-check passed
  • Feishu Doc

    openclaw/openclaw

    Feishu document read/write workflows. An agent skill from openclaw/openclaw.

    392k GitHub stars~516 tokensUpdated today
    Auto-check passed
  • Openclaw PR Maintainer

    openclaw/openclaw

    Review, triage, repair, or land OpenClaw issues and pull requests with current-source evidence and the native maintainer workflow.

    392k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Browser Automation

    openclaw/openclaw

    A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

    392k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Clawsweeper

    openclaw/openclaw

    A skill your agent uses for all ClawSweeper work: OpenClaw issue/PR sweep reports, repair jobs, cloud fix PRs, @clawsweeper maintainer mention commands, trusted ClawSweeper-reviewed…

    392k GitHub stars~3k tokensUpdated today
    Auto-check passed

Works with

Questions about Proof Video

What does Proof Video do?

Add subtitles, captions, narration cues, or zoom to a proof video or PR recording using repo-local capture helpers and a system ffmpeg renderer. Proof Video is an agent skill from openclaw/openclaw. Add subtitles, captions, narration cues, or zoom to a proof video or PR recording using repo-local capture helpers and a system ffmpeg renderer.

When should I use Proof Video?

Proof Video fits situations like: tasks that involve Transcription; tasks that involve Text to speech and voice; tasks that involve Video production.

How do I install Proof Video in Claude Code?

Run `npx skills add openclaw/openclaw --skill proof-video -a claude-code`. Or copy the skill folder (.agents/skills/proof-video in openclaw/openclaw) into .claude/skills/proof-video in your project. Claude Code loads it when a task matches its description.

How do I install Proof Video in Codex?

Run `npx skills add openclaw/openclaw --skill proof-video -a codex`. Or copy the skill folder (.agents/skills/proof-video in openclaw/openclaw) into .agents/skills/proof-video in your project. Codex loads it when a task matches its description.

Can I use Proof Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openclaw/openclaw --skill proof-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/proof-video, .gemini/skills/proof-video, .github/skills/proof-video and .opencode/skills/proof-video in your project.

What does Proof Video need to run?

Going by SKILL.md and its folder, Proof Video needs the command-line tools its instructions call (node, brew and apt).

Does Proof Video access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Proof Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Proof Video use?

Proof Video is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Proof Video use?

About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Proof Video?

Skills that share tags, products or a category with Proof Video: Stage Edit (Orkas-AI/Orkas-VideoStudio, 498 stars), Explain Video (limin112/min-skill, 362 stars), Qiaomu Cut (joeseesun/qiaomu-cut-skill, 369 stars) and Showtime (FavioVazquez/showtime, 178 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Proof Video?

openclaw (a GitHub organization) maintains it in openclaw/openclaw, which has 391,610 GitHub stars. The repository holds 93 skills in this directory. The repository was last updated on October 8, 2026.

Source: openclaw/openclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.