Agent skill

Video Description Oversight

by calesthio in calesthio/generative-media-skills

Provider-independent governance workflow for verifying and correcting human- or model-generated video descriptions.

MITAuto-check passedResearch & Science

Install Video Description Oversight

skills CLI
$ npx skills add calesthio/generative-media-skills --skill video-description-oversight -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/generative-media-skills video-description-oversight --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/production/governance-delivery/video-description-oversight .claude/skills/video-description-oversight && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-description-oversight
GitHub stars
197
Token cost
~3.5k tokens
SKILL.md length
1,536 words
Files
2
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

Provider-independent governance workflow for verifying and correcting human- or model-generated video descriptions.

  • Works in 6 steps: Were all valid critique findings… → Did revision introduce new… → Does it remain objective and temporally… → …
  • Pre-caption critique and post-caption revision
  • SKILL.md covers Evidence stance, Activation and boundaries, Inputs and evidence package and Triage the pre-caption, plus 12 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Video Description Oversight is an agent skill from calesthio/generative-media-skills. Provider-independent governance workflow for verifying and correcting human- or model-generated video descriptions. Use for pre-caption critique and post-caption revision, aspect-by-aspect factual review, critique precision/recall/constructiveness, second-stage peer review, calibration, appeals, versioned triplets, provenance, and acceptance reporting; not for pixel/audio QA, accessibility captioning, model training, or creative shot direction.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EVAL.md`).

It sits in Research & Science, covering Performance reviews, Fine-tuning and Peer review. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.

When your agent uses it

  • Pre-caption critique and post-caption revision
  • Aspect-by-aspect factual review
  • Critique precision/recall/constructiveness
  • Second-stage peer review

Example prompts

  • “/video-description-oversight”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Were all valid critique findings addressed?
  2. Did revision introduce new hallucinations, omissions, ambiguity, or writing problems?
  3. Does it remain objective and temporally ordered?
  4. Does terminology match the approved glossary?
  5. Are uncertainty and intentionally omitted aspects preserved?
  6. Does the post-caption suit its declared use and privacy/rights scope?

What it can do on your machine

Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org
    • huggingface.co
    • linzhiqiu.github.io
    • github.com
    • doi.org
    • spec.c2pa.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Description Oversight loads about 3.5k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 1,536 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 1,536 words, ~3,450 tokens.

Download SKILL.mdSave it as .claude/skills/video-description-oversight/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
video-description-oversight
description
Provider-independent governance workflow for verifying and correcting human- or model-generated video descriptions. Use for pre-caption critique and post-caption revision, aspect-by-aspect factual review, critique precision/recall/constructiveness, second-stage peer review, calibration, appeals, versioned triplets, provenance, and acceptance reporting; not for pixel/audio QA, accessibility captioning, model training, or creative shot direction.

Video description oversight

Use this skill when language about a video must be accepted as evidence-quality metadata rather than plausible prose. It governs a correction loop:

text
source video + specification + pre-caption
  -> evidence-backed critique
  -> revised post-caption
  -> independent acceptance review
  -> versioned triplet and decision record

The reviewer verifies what language says about media. Technical file integrity, visual artifact QA, accessibility captions, creative intent, and model post-training belong elsewhere.

Evidence stance

  • Documented fact: a finding supported by the source video, approved specification, or cited standard/research.
  • Reviewer observation: a timestamped comparison between video and description.
  • Production heuristic: an operational quality-control choice that must fit the project's risk and workforce.

The workflow is informed by Lin et al.'s CHAI framework, which reports that critiques used to revise precise video captions are most useful when accurate, complete, and constructive. Sources were verified 2026-07-14. CHAI's performance numbers are first-party research results on its data and annotator program, not universal service-level guarantees.

Activation and boundaries

Use this skill for descriptions used in:

  • search, archive, edit logging, or reference analysis;
  • training/evaluation datasets;
  • generation-prompt handoff;
  • cinematography-aware video understanding;
  • regulated or high-risk metadata where language errors matter.

Route elsewhere for:

  • codec, color, audio, anatomy, artifact, or final-media release QA;
  • SDH/subtitle timing and accessibility compliance;
  • deciding the desired shot or creative treatment;
  • training/fine-tuning reward or caption models;
  • provider-specific video analysis API calls;
  • payroll, labor classification, or workforce procurement.

The workflow requires a description specification. If none exists, establish one through precise-video-description before grading completeness or terminology.

Inputs and evidence package

Collect:

  • immutable source asset/version, checksum, duration/timebase, rights and privacy classification;
  • pre-caption text, author/model/version, prompt/instruction, timestamp, and provenance;
  • approved description scope and five-aspect requirement: Subject, Scene, Motion, Spatial, Camera;
  • project glossary/version, reference examples, uncertainty policy, and domain limitations;
  • target use and severity policy;
  • reviewer identity/pseudonymous ID, expertise, language, calibration status, and conflicts;
  • prior critiques, revisions, appeals, and accepted version if any.

Do not let a reviewer critique video they cannot access. A text-only “blind critique” cannot establish visual accuracy.

Triage the pre-caption

Before detailed review, decide:

  • reviewable: coherent enough for correction;
  • requires domain expert: specialized medicine, science, culture, language, or cinematography exceeds reviewer competence;
  • requires new source: video is corrupted, too low quality, incomplete, or mismatched;
  • reject and regenerate/rewrite: caption is irrelevant or so defective that itemized correction is less reliable than a new draft;
  • blocked: rights, privacy, or specification is missing.

Record the reason. Do not force a critique workflow onto an unusable source/caption pair.

Review by aspect and time

Use the project's precise-description contract rather than intuition.

Subject

Check entity count, stable identifiers, visible attributes, pose, relationships, entry/exit, and unverified identity/demographic inference.

Scene

Check setting, time/weather evidence, overlays versus physical objects, transitions, and unsupported mood/theme claims.

Motion

Check action verbs, actors/targets, temporal order, simultaneous activity, contact, direction, and causal overstatement.

Spatial

Check shot size, frame position, depth, overlap/occlusion, start/end framing, and subject-relative versus frame-relative direction.

Camera

Check translation versus rotation, zoom versus movement, angle/height/roll, focus changes, steadiness, playback effects, edit transitions, and fabricated technical settings.

Review at normal speed and frame-level around disputed events. Each finding must include a locator or bounded interval where possible.

Write a correctional critique

Every critique must pass three independent gates:

Precision

Every finding is real and supported by visible evidence/specification. Do not add a plausible correction that the video does not establish.

Recall

All consequential errors and omissions within the agreed scope are covered. Recall is not maximal verbosity; trivial details outside the contract can remain omitted.

Constructiveness

Each finding says how to repair the caption: replace, add, delete, reorder, qualify, or mark uncertain. “This is wrong” is not a usable correction.

Recommended finding shape:

json
{
  "aspect": "camera",
  "time_range": ["00:02.100", "00:04.800"],
  "error_type": "incorrect-term",
  "caption_claim": "the camera zooms in",
  "evidence": "foreground/background parallax increases while field of view appears stable",
  "correction": "replace with 'the camera moves forward'",
  "severity": "major",
  "reviewer_id": "reviewer-17",
  "spec_version": "video-language-2.1"
}

This is an example schema. It must not contain invented confidence scores. Separate observed evidence from the proposed wording.

Handle accurate captions

Do not invent feedback to prove that review occurred. If no corrections are required, record an explicit sentinel such as:

text
The description matches the reviewed source and specification; no edits are required.

An “accurate” decision still needs review scope, asset/version, reviewer, date, glossary/specification version, and any unreviewed lanes.

Produce and verify the post-caption

The author/model revises from the accepted critique. Preserve all versions; never overwrite the pre-caption.

Second-stage acceptance asks:

  1. Were all valid critique findings addressed?
  2. Did revision introduce new hallucinations, omissions, ambiguity, or writing problems?
  3. Does it remain objective and temporally ordered?
  4. Does terminology match the approved glossary?
  5. Are uncertainty and intentionally omitted aspects preserved?
  6. Does the post-caption suit its declared use and privacy/rights scope?

The accepting reviewer should be different from the first reviewer for high-risk or dataset use. Lower-risk internal work may use calibrated spot review according to a documented policy.

Severity and disposition

Define severity from downstream consequence, not word count:

  • Critical: false identity, safety/regulatory claim, privacy leak, materially deceptive event, unusable dataset target, or unsupported statement that could cause harm.
  • Major: wrong subject/action/camera/spatial relation, consequential omission, temporal inversion, or terminology that changes meaning.
  • Minor: localized wording or non-consequential omission within scope.
  • Observation: non-blocking note or limitation.

Possible dispositions: accepted, accepted_with_limitations, revise, reject, escalate, blocked.

Do not mechanically waive privacy, consent, safety, or material factual errors.

Peer review, appeals, and calibration

Use gold examples and periodic blind calibration. Track agreement by aspect because a team may agree on Subject/Scene while failing on Spatial/Camera.

When reviewers disagree:

  1. identify whether the conflict is evidence, terminology, scope, or specification ambiguity;
  2. replay the disputed interval and apply the glossary decision rule;
  3. escalate to an appropriate senior/domain reviewer if unresolved;
  4. record the decision and update training material/specification if the ambiguity will recur;
  5. allow a documented appeal without erasing the original review.

Do not impose a universal kappa threshold, daily review quota, compensation scheme, or expertise ladder. CHAI used a highly trained professional pipeline; teams with different content or reviewers must establish their own validated calibration criteria.

Show full SKILL.md (597 more words)Show less

Versioned outputs and provenance

Preserve:

  • source asset ID/hash/version;
  • specification/glossary version;
  • pre-caption, critique findings, post-caption;
  • author/model/prompt/version and reviewer IDs;
  • timestamps, iteration number, status, acceptance/appeal decisions;
  • domain, language, intended use, rights/privacy restrictions;
  • known unreviewed aspects and residual risks.

The (pre-caption, critique, post-caption) triplet can become a valuable dataset asset, but production permission does not automatically grant model-training permission. Create a datasheet and confirm source-video, caption, reviewer, and derivative-data rights before reuse.

C2PA or signed records can support provenance; they do not prove that a caption is true.

Privacy, workforce, and domain limitations

Detailed video descriptions may expose identity, location, private behavior, screens, health/financial details, or copyrighted story content. Minimize access, use pseudonymous reviewer IDs, define retention/deletion, separate public output from private review metadata, and audit exports.

Match reviewers to domain/language complexity. Cinematography expertise does not imply medical or cultural expertise; language fluency does not guarantee camera-motion discrimination. CHAI's findings came from trained professional creators and predominantly professional video domains. Do not assume equivalent results from untrained crowdworkers, multilingual auto-translation, long-form footage, or specialized domains.

Quality dashboard

Track trends rather than optimizing one number:

  • corrections and severity by aspect;
  • false-positive critique rate;
  • omissions found at second-stage review;
  • revision iterations and rejection reasons;
  • appeal rate and root cause;
  • agreement/calibration by aspect and domain;
  • source model/author defect patterns;
  • specification ambiguities and glossary changes;
  • review time and fatigue indicators without intrusive surveillance.

High critique volume can indicate either poor sources or overcritical reviewers. Audit evidence before drawing conclusions.

Example 1: camera-term correction

This is a complete example, not a mandatory formula.

Source: six-second kitchen shot. Pre-caption claim: “The camera zooms in on the woman as she glances to the right.”

Evidence review: background/foreground parallax changes, supporting forward camera translation; the woman looks toward frame-right, but whether that is her left/right is not needed. Focus remains on her eye.

Critique: “Camera, 00:02.0-00:05.0: replace ‘zooms in’ with ‘moves forward’ because the viewpoint translates and parallax changes rather than only the field of view. Spatial/Motion: keep ‘looks toward frame-right’; do not convert it to the subject's right without evidence. The remaining subject and scene description is accurate.”

Post-caption: “In a dim kitchen, the camera moves smoothly forward toward a woman as she raises her gaze toward frame-right; focus remains on her near eye.”

Acceptance: second reviewer confirms correction, no new claims, and records the exact source/spec versions.

Failure to avoid: “The caption is wrong about camera motion.” It lacks the supported replacement and is non-constructive.

Example 2: incomplete generated game description

This is a complete example, not a mandatory formula.

Pre-caption: “A knight runs through a level while the camera follows.”

Reviewed scope: all five aspects for dataset metadata.

Findings:

  • Subject: “knight” overstates identity; visible evidence supports “small armored character.”
  • Scene: missing elevated platforms and score overlay.
  • Motion: add frame-right travel and stop before the cut.
  • Spatial: add side view, left-third hold, and platforms moving frame-left; world/camera displacement remains uncertain.
  • Camera: “follows” is plausible from stable framing but should be qualified as “consistent with lateral tracking”; add final hard cut to a closer view.

Constructive critique: provides each replacement/addition and preserves uncertainty. Post-caption: incorporates all five without claiming world speed or exact lens. Disposition: accepted after second review.

Failure to avoid: adding speculative game title, character name, player intent, or “dynamic exciting atmosphere.”

Sources

Verified 2026-07-14:

© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/production/governance-delivery/video-description-oversight of calesthio/generative-media-skills.

  • SKILL.md
  • EVAL.md

Open the folder on GitHubat commit 8c85352

Compare with similar skills

Video Description Oversight next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Description Oversight compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Description Oversight this skillcalesthio/generative-media-skills197—~3.5kAutomated safety check: PassMIT
Scientific Writergaasher/Agent-Loop-Skills174—~3.3kAutomated safety check: PassMIT
External Model Validationaipoch/medical-research-skills1.9k—~3.2kAutomated safety check: PassMIT
Jeg Rebuttalbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.2kAutomated safety check: PassMIT
Quaxnstarman/quax143—~5.5kAutomated safety check: PassApache-2.0
Paper To HTMLysyecust/lecture-to-notes273—~1.3kAutomated safety check: PassCustom licence

Similar skills

  • Scientific Writer

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has a scientific draft (with its dataset, figures, and optional analysis code) and wants it iteratively revised until it clears a quality bar.

    174 GitHub stars~3.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • External Model Validation

    aipoch/medical-research-skills

    A skill your agent uses when validating an existing prognostic risk signature on an external bulk expression cohort with survival outcomes, producing risk scores, Kaplan-Meier curves, risk…

    1.9k GitHub stars~3.2k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Jeg Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding after a Journal of Economic Growth revise-and-resubmit to organize responses about growth mechanisms, model assumptions, empirical identification, calibration…

    1.2k GitHub stars~1.2k tokensUpdated 14 days ago
    Business, Finance & HRAuto-check passed
  • Quax

    nstarman/quax

    A skill your agent uses when writing, reviewing, or debugging JAX code that involves quax — custom array-ish objects (physical units, LoRA, sparse, symbolic zero, named axes), quax.quaxify…

    143 GitHub stars~5.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Paper To HTML

    ysyecust/lecture-to-notes

    Generate a self-contained, beautifully styled HTML analysis of an academic paper.

    273 GitHub stars~1.3k tokensUpdated 8 days ago
    Research & ScienceAuto-check passed
  • Resolving Clinical Context

    maziyarpanahi/openmed

    Assign negation, temporality, and uncertainty (the ConText axes) to clinical entities extracted by OpenMed, so "denies chest pain" is not counted as chest pain and "history of MI" is not counted as…

    5.5k GitHub stars~1.8k tokensUpdated today
    Research & ScienceAuto-check passed

More from calesthio/generative-media-skills

All 26 skills in this repo
  • 3D Asset Production

    calesthio/generative-media-skills

    A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.

    197 GitHub stars~9.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Audio Mixing Mastering

    calesthio/generative-media-skills

    Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…

    197 GitHub stars~7k tokensUpdated 2 mo ago
    Auto-check passed
  • Captions Media Accessibility

    calesthio/generative-media-skills

    Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…

    197 GitHub stars~6.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Comfyui Media Workflows

    calesthio/generative-media-skills

    Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…

    197 GitHub stars~8.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Ffmpeg Media Finishing

    calesthio/generative-media-skills

    Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.

    197 GitHub stars~8.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Generated Media QA

    calesthio/generative-media-skills

    Provider-independent quality assurance for AI-generated and AI-assisted media.

    197 GitHub stars~8k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Video Description Oversight

What does Video Description Oversight do?

Provider-independent governance workflow for verifying and correcting human- or model-generated video descriptions. Video Description Oversight is an agent skill from calesthio/generative-media-skills. Provider-independent governance workflow for verifying and correcting human- or model-generated video descriptions.

When should I use Video Description Oversight?

Video Description Oversight fits situations like: pre-caption critique and post-caption revision; aspect-by-aspect factual review; critique precision/recall/constructiveness; second-stage peer review.

How do I install Video Description Oversight in Claude Code?

Run `npx skills add calesthio/generative-media-skills --skill video-description-oversight -a claude-code`. Or copy the skill folder (skills/production/governance-delivery/video-description-oversight in calesthio/generative-media-skills) into .claude/skills/video-description-oversight in your project. Claude Code loads it when a task matches its description.

How do I install Video Description Oversight in Codex?

Run `npx skills add calesthio/generative-media-skills --skill video-description-oversight -a codex`. Or copy the skill folder (skills/production/governance-delivery/video-description-oversight in calesthio/generative-media-skills) into .agents/skills/video-description-oversight in your project. Codex loads it when a task matches its description.

Can I use Video Description Oversight in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill video-description-oversight -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-description-oversight, .gemini/skills/video-description-oversight, .github/skills/video-description-oversight and .opencode/skills/video-description-oversight in your project.

What does Video Description Oversight need to run?

SKILL.md names no scripts, command-line tools or credentials: Video Description Oversight is instructions for the agent only.

Does Video Description Oversight access the network?

SKILL.md names 6 domains. As links in the text: arxiv.org, huggingface.co, linzhiqiu.github.io, github.com, doi.org and spec.c2pa.org. This is read from the text; nothing was executed.

Is Video Description Oversight safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Video Description Oversight use?

Video Description Oversight is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Description Oversight use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video Description Oversight?

Skills that share tags, products or a category with Video Description Oversight: Scientific Writer (gaasher/Agent-Loop-Skills, 174 stars), External Model Validation (aipoch/medical-research-skills, 1.9k stars), Jeg Rebuttal (brycewang-stanford/Awesome-Journal-Skills, 1.2k stars) and Quax (nstarman/quax, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Description Oversight?

calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.

Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.