Agent skill

Cdaf

by UditAkhourii in UditAkhourii/cdaf

Read CDAF sidecar files (.cdaf) instead of processing video with vision.

MITAuto-check passedAI & LLM Engineering

Install Cdaf

skills CLI
$ npx skills add UditAkhourii/cdaf --skill cdaf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install UditAkhourii/cdaf cdaf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/UditAkhourii/cdaf.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/claude-code/cdaf .claude/skills/cdaf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cdaf
GitHub stars
121
Token cost
~1.9k tokens
SKILL.md length
987 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Read CDAF sidecar files (.cdaf) instead of processing video with vision.

  • Works in 4 steps: Check for a sidecar: same directory,… → Verify freshness before trusting it (see… → If fresh: read the sidecar instead of… → …
  • A task involves understanding
  • SKILL.md covers The rule, Verifying freshness, Reading a sidecar and How much to trust a fresh…, plus 3 more sections
  • Calls ffmpeg, curl and pip; needs GEMINI_API_KEY

What it does

Cdaf is an agent skill from UditAkhourii/cdaf. Read CDAF sidecar files (.cdaf) instead of processing video with vision. Use whenever a task involves understanding, summarizing, searching, editing, or selecting from video files (b-roll, raw clips, footage libraries) — check for a .cdaf sidecar FIRST before analyzing any video directly.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with Google Gemini and Remotion. The repository describes itself as: CDAF (Cached Descriptive Asset Files) - open sidecar format for video so AI agents stop re-analyzing the same footage. Spec, CLI, agent skill, reproducible benchmark. The licence is MIT.

When your agent uses it

  • A task involves understanding
  • Selecting from video files (b-roll
  • Footage libraries) — check for a .cdaf sidecar FIRST before analyzing any video directly

Example prompts

  • “/cdaf”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Check for a sidecar: same directory, same basename, .cdaf extension.
  2. Verify freshness before trusting it (see below). A stale sidecar describes an
  3. If fresh: read the sidecar instead of processing the video. Use it as the
  4. If missing or stale: generate one (see below) so the cost is paid once. If you

What it can do on your machine

Read from SKILL.md and the folder at commit e562064. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ffmpeg
    • curl
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • aistudio.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cdaf loads about 1.9k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 987 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from UditAkhourii/cdaf at commit e562064, republished under its MIT licence (© UditAkhourii). 987 words, ~1,907 tokens.

Download SKILL.mdSave it as .claude/skills/cdaf/SKILL.md (or your agent's skills folder).
name
cdaf
description
Read CDAF sidecar files (.cdaf) instead of processing video with vision. Use whenever a task involves understanding, summarizing, searching, editing, or selecting from video files (b-roll, raw clips, footage libraries) — check for a .cdaf sidecar FIRST before analyzing any video directly.

CDAF — Cached Descriptive Asset Files

A .cdaf file is a timestamped, pre-computed description of a video, sitting next to it with the same basename (clip.mp4 → clip.cdaf). Reading it costs a few hundred text tokens; analyzing the video directly costs orders of magnitude more (~263 tokens per second of footage on Gemini-class models). Always prefer the sidecar when it is fresh.

Format spec and tooling: https://github.com/UditAkhourii/cdaf

The rule

Before analyzing ANY video file (.mp4, .mov, .mkv, .webm, .avi, .m4v):

  1. Check for a sidecar: same directory, same basename, .cdaf extension.
  2. Verify freshness before trusting it (see below). A stale sidecar describes an older version of the video — using it is worse than not having one.
  3. If fresh: read the sidecar instead of processing the video. Use it as the account of what the video contains — within the limits below.
  4. If missing or stale: generate one (see below) so the cost is paid once. If you cannot generate, fall back to direct video analysis.

Verifying freshness

The sidecar header carries bytes (file size) and sha256 (content hash) of the exact video it describes. Verification never needs an API key or network access.

  • Cheap check (usually enough): compare the video's current file size to the header's bytes value. Different size → provably stale.
  • Strict check: cdaf validate <video> (exit 0 = fresh), or hash the file yourself and compare to the header's sha256:
    • PowerShell: (Get-FileHash clip.mp4 -Algorithm SHA256).Hash.ToLower()
    • POSIX: sha256sum clip.mp4 / shasum -a 256 clip.mp4
  • Use the strict check when the decision is expensive to get wrong (publishing, final edits); the cheap check suffices for exploration.

Reading a sidecar

It is plain UTF-8 text — use the Read tool directly, or cdaf read <video> (which verifies the hash automatically and refuses to print a stale sidecar).

Format: a key: value header between --- CDAF/1.0 and ---, then markdown:

  • ## Summary — what the clip is
  • ## Segments — [MM:SS.d-MM:SS.d] description lines covering the whole video (see the trust note below before cutting on these timestamps)
  • ## Transcript — spoken words with timestamps (or (no speech))
  • ## On-screen Text — visible text with timestamps (or (none))
  • ## Tags — retrieval keywords

How much to trust a fresh sidecar

Freshness proves the sidecar describes these exact bytes. It does not prove the description is complete or correct — it is generated text, the record of one model's pass. Three limits matter when a mistake is expensive:

  • Timestamps are approximate. Boundaries inferred by a model drift (over a second on measured clips), miss real cuts, and occasionally mark cuts that do not exist. They are fine for locating, ranking, and rough trims. Before cutting on them, verify against the container:

    bash
    ffmpeg -v error -i clip.mp4 -vf "select='gt(scene,0.1)',metadata=print:file=-" -f null -
  • Descriptions can add, not just omit. Shown a whole video at once, a model may narrate the outcome a clip implies but never shows — reporting that a task was completed when the footage only shows it being started. Such entries are fluent, specific, and indistinguishable from correct ones. Treat any claim that something was finished, fixed, repaired, or achieved as unverified: check the frames before relying on it, and say the sidecar is your source when you report it. An omission is a visible gap; an addition reads exactly like a fact.

  • Fine visual state is unreliable. Strike-through and other mark-up on a list (crossed out, checked, highlighted), small or stylised text, and subtle motion are often missed or reported at chance. Incidental background text on packaging or labels has produced confident phantom brand names. If such a detail carries the meaning of the shot, look at the frame.

Provenance keys. When a producer records how the body was made, x- keys tell you which parts were measured rather than inferred:

KeyMeans
x-shot-source: ffmpeg-scene-detect@<threshold>Boundaries came from the container and are frame-exact — no need for the ffmpeg check above.
x-shot-isolation: per-shotEach segment was described without sight of the others, which suppresses invented continuity between shots.
x-transcript-timing: measured-rmsTranscript times were measured from audio. none means no measurable speech.

Absent these keys, assume the weaker case: inferred boundaries, whole-video context, and guessed transcript times.

None of this argues for re-watching by default — that would forfeit the entire saving. Verify the specific claim your decision rests on, not the whole clip.

Show full SKILL.md (295 more words)Show less

Generating sidecars

Two providers. Both write the same v1.0 format and either output passes cdaf validate.

Local model — no API key, no cost, footage stays on the machine
bash
cdaf generate <video> --local          # or --provider local

Needs ffmpeg and an OpenAI-compatible endpoint serving a model with a vision encoder (default http://127.0.0.1:8090/v1, override with --base-url / --model, or the CDAF_BASE_URL / CDAF_LOCAL_MODEL env vars). An audio encoder, where the model has one, is used for the transcript. Check the endpoint is up before offering this route:

bash
curl -s localhost:8090/props   # llama-server: reports which modalities are loaded

Slower per clip than the API, but free and private, and cost scales per shot rather than per second of footage — so long clips are far cheaper here. Set CDAF_PROVIDER=local to make it the default.

Gemini API
bash
cdaf generate <video-or-directory>      # skips sidecars that are already fresh
cdaf generate <video> --force           # regenerate even if fresh
cdaf generate ./footage --detail rich   # brief | standard | rich

Needs Python >= 3.10 and GEMINI_API_KEY (free tier: https://aistudio.google.com/apikey). Install the CLI once:

bash
pip install "cdaf[generate] @ git+https://github.com/UditAkhourii/cdaf.git#subdirectory=cli"

Faster per clip and handles whole directories, but calls a paid API. Ask the user before batch-generating a large library, and tell them roughly how many videos you are about to process.

Working across a footage library

  • Survey coverage: cdaf status <dir> lists every video as FRESH/STALE/MISSING.
  • To find footage matching a need ("sunset city shots"), grep the .cdaf files — never open the videos: search *.cdaf for the relevant keywords, then rank by the Segments detail.
  • Batch-fill gaps: cdaf generate <dir> (fresh sidecars are skipped automatically).

What NOT to do

  • Do not treat a sidecar as fresh without at least the size check.
  • Do not invent visual details beyond what the sidecar states; if the task needs information the sidecar lacks (exact colors, a specific frame), say so and fall back to targeted direct analysis of just the needed timestamp range.
  • Do not edit .cdaf files by hand to "update" them — the header hash would then describe a video the body no longer matches. Regenerate with cdaf generate <video> --force instead.

© UditAkhourii, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/claude-code/cdaf of UditAkhourii/cdaf.

Open the folder on GitHubat commit e562064

Compare with similar skills

Cdaf next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cdaf compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cdaf this skillUditAkhourii/cdaf121—~1.9kAutomated safety check: PassMIT
ModLens Image Vision Bridgeliustack/modlens4.2k—~1.3kAutomated safety check: NotesMIT
Gemini API Devgoogle-gemini/gemini-skills4.3k—~5.1kAutomated safety check: PassApache-2.0
Gemini Live APIgoogle/skills21k—~2.5kAutomated safety check: NotesApache-2.0
Gemini Video Understandingeinverne/dotfiles121—~2.6kAutomated safety check: NotesMIT
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0

Similar skills

  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.2k GitHub stars~1.3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Gemini API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice…

    4.3k GitHub stars~5.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Gemini Live API

    google/skills

    Official

    Generates a Gemini LiveAPI client service class in the user's chosen programming language.

    21k GitHub stars~2.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.

    121 GitHub stars~2.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Evals Context

    zgsm-ai/costrict

    Provides context about the CoStrict evals system structure in this monorepo.

    4.4k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed

Questions about Cdaf

What does Cdaf do?

Read CDAF sidecar files (.cdaf) instead of processing video with vision. Cdaf is an agent skill from UditAkhourii/cdaf.cdaf) instead of processing video with vision.

When should I use Cdaf?

Cdaf fits situations like: A task involves understanding; selecting from video files (b-roll; footage libraries) — check for a .cdaf sidecar FIRST before analyzing any video directly.

How do I install Cdaf in Claude Code?

Run `npx skills add UditAkhourii/cdaf --skill cdaf -a claude-code`. Or copy the skill folder (skills/claude-code/cdaf in UditAkhourii/cdaf) into .claude/skills/cdaf in your project. Claude Code loads it when a task matches its description.

How do I install Cdaf in Codex?

Run `npx skills add UditAkhourii/cdaf --skill cdaf -a codex`. Or copy the skill folder (skills/claude-code/cdaf in UditAkhourii/cdaf) into .agents/skills/cdaf in your project. Codex loads it when a task matches its description.

Can I use Cdaf in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add UditAkhourii/cdaf --skill cdaf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cdaf, .gemini/skills/cdaf, .github/skills/cdaf and .opencode/skills/cdaf in your project.

What does Cdaf need to run?

Going by SKILL.md and its folder, Cdaf needs the command-line tools its instructions call (ffmpeg, curl and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY.

Does Cdaf access the network?

SKILL.md names 1 domain. As links in the text: aistudio.google.com. This is read from the text; nothing was executed.

Is Cdaf safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cdaf use?

Cdaf is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cdaf use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cdaf?

Skills that share tags, products or a category with Cdaf: ModLens Image Vision Bridge (liustack/modlens, 4.2k stars), Gemini API Dev (google-gemini/gemini-skills, 4.3k stars), Gemini Live API (google/skills, 21k stars) and Gemini Video Understanding (einverne/dotfiles, 121 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cdaf?

UditAkhourii (a GitHub user) maintains it in UditAkhourii/cdaf, which has 121 GitHub stars. The repository was last updated on August 28, 2026.

Source: UditAkhourii/cdaf on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.