Agent skill

VideoDB Video and Audio Workflows

by video-db in video-db/skills

Ingests video, audio, live streams and desktop captures, indexes them for timestamped search, then edits, subtitles, dubs and clips the results.

MITAuto-check: notesMedia & Creative

Install VideoDB Video and Audio Workflows

skills CLI
$ npx skills add video-db/skills --skill videodb -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install video-db/skills videodb --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/video-db/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/python .claude/skills/videodb && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
videodb
GitHub stars
127
Token cost
~5.5k tokens
SKILL.md length
1,608 words
Files
24 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Ingests video, audio, live streams and desktop captures, indexes them for timestamped search, then edits, subtitles, dubs and clips the results.

  • Works in 5 steps: Desktop Perception → Video ingest + stream → Understand + index + retrieve… → …
  • Searching a folder of recordings for every scene that shows a given object
  • SKILL.md covers 1) Desktop Perception, 2) Video ingest + stream, 3) Understand + index +… and 4) Timeline editing + generation, plus 10 more sections
  • Calls python, pip and npx; reaches console.videodb.io and mcp.videodb.io; needs VIDEO_DB_API_KEY

What it does

VideoDB gives an agent a way to see, understand and act on video and audio. It can take a local file, a URL or an RTSP feed, record the desktop with screen, microphone and system audio, and hand back a playable stream link together with live context from whatever it is watching.

For understanding, analyzers cover speech, scenes, objects, on-screen text, brands and activity, and the output feeds indexes you can search, filter, count and group, returning matching moments with timestamps, playable evidence and automatic clips. For editing, it transcodes (codec, bitrate, frame rate, resolution, aspect ratio), generates, translates and burns in subtitles, overlays text, images and branding, adds music, voiceover and dubbing, and can generate image, audio and video assets.

Live streams and desktop sessions can trigger real-time alerts, and a session ends with a summary and a searchable timeline. The skill is limited to Read, Grep, Glob and python shell commands, and ships a reference folder covering capture, editing, indexing, search, real-time streams and a sandbox.

When your agent uses it

  • Searching a folder of recordings for every scene that shows a given object
  • Adding subtitles, translations or dubbing to a video
  • Recording a desktop session and getting a summary when it ends
  • Watching an RTSP camera feed and raising alerts on events
  • Converting a clip to a different codec, frame rate or aspect ratio

Example prompts

  • “Ingest ./demo/walkthrough.mp4 and give me a playable stream link.”
  • “Index every video in ./recordings and list the timestamps where a whiteboard appears.”
  • “Generate English and Spanish subtitles for talk.mp4 and burn them in.”
  • “Record my screen session and write up a summary when I stop.”

Requirements

  • Python
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash(python:*)

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Desktop Perception
  2. Video ingest + stream
  3. Understand + index + retrieve (timestamps + evidence)
  4. Timeline editing + generation
  5. Live streams (RTSP) + monitoring

What it can do on your machine

Read from SKILL.md and the folder at commit b47b587. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash(python:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • console.videodb.io
    • mcp.videodb.io
    • youtube.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VIDEO_DB_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

VideoDB Video and Audio Workflows loads about 5.5k tokens when it runs. Until then it costs about 173 tokens; SKILL.md has 1,608 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~173
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:67
    Python code. This ensures `load_dotenv(".env")` finds the correct `.env` file.
  • NoteMentions a .env fileSKILL.md:71
    load_dotenv(".env")
  • NoteMentions a .env fileSKILL.md:79
    2. Project's `.env` file in current directory
  • NoteMentions a .env fileSKILL.md:90
    load_dotenv(".env")
  • NoteMentions a .env fileSKILL.md:122
    - **Project `.env` file**: Save `VIDEO_DB_API_KEY=your-key` in the project's `.env` file

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from video-db/skills at commit b47b587, republished under its MIT licence (© video-db). 1,608 words, ~5,454 tokens.

Download SKILL.mdSave it as .claude/skills/videodb/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.
name
videodb
description
See, Understand, Act on video and audio. See- ingest from local files, URLs, RTSP/live feeds, or live record desktop; return realtime context and playable stream links. Understand- run analyzers over speech, scenes, objects, OCR, brands and activity; build searchable indexes; then search moments, ask questions about a video, filter and aggregate results with timestamps and auto-clips. Act- transcode and normalize (codec, fps, resolution, aspect ratio), perform timeline edits (subtitles, text/image overlays, branding, audio overlays, dubbing, translation), generate media assets (image, audio, video), and create real time alerts for events from live streams or desktop capture.
allowed-tools
Read, Grep, Glob, Bash(python:*)
argument-hint
[task description]

VideoDB Skill

Perception + memory + actions for video, live streams, and desktop sessions.

Use this skill when you need to:

1) Desktop Perception

  • Start/stop a desktop session capturing screen, mic, and system audio
  • Stream live context and store episodic session memory
  • Run real-time alerts/triggers on what’s spoken and what's happening on screen
  • Produce session summaries, a searchable timeline, and playable evidence links

2) Video ingest + stream

  • Ingest a file or URL and return a playable web stream link
  • Transcode/normalize: codec, bitrate, fps, resolution, aspect ratio

3) Understand + index + retrieve (timestamps + evidence)

  • Understand: run analyzers over speech, scenes, objects, OCR, brands, activity
  • Index: turn analyzer artifacts into semantic, filterable, and aggregatable indexes
  • Retrieve: search moments, ask questions, filter exactly, count and group — with timestamps and playable evidence
  • Auto-create clips from results

4) Timeline editing + generation

  • Subtitles: generate, translate, burn-in
  • Overlays: text/image/branding, motion captions
  • Audio: background music, voiceover, dubbing
  • Programmatic composition and exports via timeline operations

5) Live streams (RTSP) + monitoring

  • Connect RTSP/live feeds
  • Run real-time visual and spoken understanding and emit events/alerts for monitoring workflows

Common inputs

  • Local file path, public URL, or RTSP URL
  • Desktop capture request: start / stop / summarize session
  • Desired operations: get context for understanding, transcode spec, index spec, search query, clip ranges, timeline edits, alert rules

Common outputs

  • Stream URL — make it playable: https://console.videodb.io/player?url={STREAM_URL}
  • Search results with timestamps and evidence links
  • Generated assets: subtitles, audio, images, clips
  • Event/alert payloads for live streams
  • Desktop session summaries and memory entries

Canonical prompts (examples)

  • “Start desktop capture and alert when a password field appears.”
  • “Record my session and produce an actionable summary when it ends.”
  • “Ingest this file and return a playable stream link.”
  • “Index this folder and find every scene with people, return timestamps.”
  • “Generate subtitles, burn them in, and add light background music.”
  • “Connect this RTSP URL and alert when a person enters the zone.”

Running Python code

CRITICAL: Always cd to the user's project directory before running Python code. This ensures load_dotenv(".env") finds the correct .env file.

python
from dotenv import load_dotenv
load_dotenv(".env")

import videodb
conn = videodb.connect()

This reads VIDEO_DB_API_KEY from:

  1. Environment (if already exported)
  2. Project's .env file in current directory

If the key is missing, videodb.connect() raises AuthenticationError automatically.

Do NOT write a script file when a short inline command works.

When writing inline Python (python -c "..."), always use properly formatted code — use semicolons to separate statements and keep it readable. For anything longer than ~3 statements, use a heredoc instead:

bash
python << 'EOF'
from dotenv import load_dotenv
load_dotenv(".env")

import videodb
conn = videodb.connect()
coll = conn.get_collection()
print(f"Videos: {len(coll.get_videos())}")
EOF

Setup

When the user asks to "setup videodb" or similar:

1. Install SDK
bash
pip install "videodb[capture]>=0.5.0" python-dotenv

If videodb[capture] fails on Linux, install without the capture extra:

bash
pip install "videodb>=0.5.0" python-dotenv

The >=0.5.0 pin matters — the understand/index/ask/aggregate APIs do not exist in earlier versions.

2. Configure API key

The user must set VIDEO_DB_API_KEY using either method:

  • Export in terminal (recommended): export VIDEO_DB_API_KEY=your-key
  • Project .env file: Save VIDEO_DB_API_KEY=your-key in the project's .env file

Get a free API key at https://console.videodb.io (50 free uploads, no credit card).

Do NOT read, write, or handle the API key yourself. Always let the user set it.

3. Authorize the hosted MCP server (Claude Code plugin only)

The plugin bundles the hosted MCP server at https://mcp.videodb.io/mcp, which is authorized separately from the SDK key above. Tell the user to run /mcp, select videodb, and complete browser authorization with their VideoDB account. Verify with "list my VideoDB collections".

The MCP tools need no local Python install and no API key. Skip this step when the skill was installed with npx skills add or is running outside Claude Code — there is no MCP server in that case, and the SDK path above is the only one available.

Quick Reference

Upload media
python
# URL
video = coll.upload(url="https://example.com/video.mp4")

# YouTube
video = coll.upload(url="https://www.youtube.com/watch?v=VIDEO_ID")

# Local file
video = coll.upload(file_path="/path/to/video.mp4")
Understand → index → retrieve (default path)

Three stages. Run analyzers to produce artifacts, index each artifact, then retrieve.

python
import time

# 1. Understand. Naming each analyzer keeps `analyzer.name` meaningful downstream.
understanding = video.understand(
    analyzers=[
        {"type": "spoken_words", "name": "transcript"},
        {"type": "vlm", "name": "scene",
         "config": {"prompt": "Describe the scene and any on-screen text."}},
    ],
    segmentation={"type": "shot", "threshold": 30},
)

# A run with a failed or skipped analyzer ends `partial`, which the SDK does not
# treat as terminal — wait_until_complete() would poll to TimeoutError. Poll the
# analyzers instead. The `analyzers and` guard is load-bearing: a refresh can
# transiently return an empty list, and all([]) is True, which would exit the
# loop while the run is still going.
deadline = time.time() + 3600
while time.time() < deadline:
    analyzers = understanding.refresh().list_analyzers()
    if analyzers and all(a.is_complete for a in analyzers):
        break
    time.sleep(15)

# 2. Index each artifact that succeeded.
for analyzer in understanding.list_analyzers():
    if analyzer.is_successful:
        video.index(source=analyzer, name=analyzer.name).wait_until_complete()

# 3. Retrieve.
response = video.search("discussion about pricing")
for shot in response.shots:
    print(f"[{shot.start:.1f}s - {shot.end:.1f}s] {shot.text}")
if response.response_type in ("shots", "deepsearch") and len(response):
    stream_url = response.compile()   # raises SearchError otherwise

Analyzer types: spoken_words (→ artifact transcript), vlm (→ scene), object_detection (→ objects), ocr, brand_detection (→ brands), activity_recognition (→ activity), location_detection (→ location), faces, audio_event_detection. They are plain strings — there is no SDK enum.

See reference/indexing.md for segmentation, sampling, field configuration, and cost tuning.

Retrieval

search(query) is the default — it plans the retrieval and picks the indexes itself. Reach past it when you need something specific:

python
# Target a specific index, with a relevance floor
video.semantic_search("a customer holding the product", index_names=["scene"], score_threshold=0.7)

# Exact filtering, no natural-language interpretation
video.query(index_name="objects",
            filter=[{"field": "frames.detections.label", "op": "contains", "value": "car"}])

# Counts and facets — returns the raw server payload, not a SearchResult
video.aggregate(index_name="objects", group_by="frames.detections.label", metric="count")

# A written answer plus the moments it came from
answer = video.ask("What did they say about pricing?", include_sources=True)

All five exist on Collection too, fanning out across every indexed video. See reference/search.md.

search() now returns SearchResponse, not SearchResult. get_shots(), compile(), play(), and iteration all work, but there is no .stream_url on it — use .compile().

Transcript + subtitle
python
# force=True skips the error if the video is already indexed
video.index_spoken_words(force=True)
text = video.get_transcript_text()
stream_url = video.add_subtitle()

index_spoken_words() is the correct call here even on 0.5.0 — add_subtitle() and CaptionAsset(src="auto") read the v1 spoken-word index. A v2 spoken_words artifact does not substitute for it. This is the one place v1 indexing is still the right answer.

Legacy indexing (existing codebases)
python
# v1 API — still supported in 0.5.0, not deprecated. New code should use the v2 path above.
from videodb import IndexType

video.index_spoken_words(force=True)
scene_index_id = video.index_scenes(prompt="Describe the visual content.")
results = video.legacy_search(
    "person writing on a whiteboard",
    index_type=IndexType.scene,
    scene_index_id=scene_index_id,
)

Recognise this pattern in existing repos and leave it alone unless asked to migrate — it still works. See reference/migration.md to port it, or reference/legacy/search.md to maintain it.

Timeline editing

Use the Editor API to compose videos, images, audio, and text. See reference/editor.md for full workflow.

python
from videodb.editor import Timeline, Track, Clip, VideoAsset, ImageAsset, AudioAsset, Fit

timeline = Timeline(conn)
timeline.resolution = "1280x720"

video_track = Track()
video_track.add_clip(0, Clip(asset=VideoAsset(id=video.id, start=10), duration=20))

audio_track = Track()
audio_track.add_clip(0, Clip(asset=AudioAsset(id=music.id, volume=0.2), duration=20))

timeline.add_track(video_track)
timeline.add_track(audio_track)
stream_url = timeline.generate_stream()
Transcode video (resolution / quality change)
python
from videodb import TranscodeMode, VideoConfig, AudioConfig

# Change resolution, quality, or aspect ratio server-side
job_id = conn.transcode(
    source="https://example.com/video.mp4",
    callback_url="https://example.com/webhook",
    mode=TranscodeMode.economy,
    video_config=VideoConfig(resolution=720, quality=23, aspect_ratio="16:9"),
    audio_config=AudioConfig(mute=False),
)
Reframe aspect ratio (for social platforms)

Warning: reframe() is a slow server-side operation. For long videos it can take several minutes and may time out. Best practices:

  • Always limit to a short segment using start/end when possible
  • For full-length videos, use callback_url for async processing
  • Trim the video on a Timeline first, then reframe the shorter result
python
from videodb import ReframeMode

# Always prefer reframing a short segment:
reframed = video.reframe(start=0, end=60, target="vertical", mode=ReframeMode.smart)

# Async reframe for full-length videos (returns None, result via webhook):
video.reframe(target="vertical", callback_url="https://example.com/webhook")

# Presets: "vertical" (9:16), "square" (1:1), "landscape" (16:9)
reframed = video.reframe(start=0, end=60, target="square")

# Custom dimensions
reframed = video.reframe(start=0, end=60, target={"width": 1280, "height": 720})
Generative media
python
image = coll.generate_image(
    prompt="a sunset over mountains",
    aspect_ratio="16:9",
)
Sandbox Compute (self-hosted / open-weight models)

Run open-weight models (Gemma, Qwen, Whisper, OmniVoice, FLUX, RT-DETR) by creating a sandbox and passing sandbox_id to a supported job. Requires videodb>=0.5.1.

python
from videodb import SandboxTier, SandboxModel

# 1. Create a sandbox sized for the largest model, then wait until active.
sandbox = conn.create_sandbox(
    tier=SandboxTier.medium,
    models=[SandboxModel.GEMMA_4_31B.value],   # exact ID, NO -FP8 suffix
)
sandbox.wait_for_ready(timeout=300, interval=5)

# 2. Understanding: set config.model + config.sandbox_id on the analyzer.
understanding = video.understand(analyzers=[{
    "type": "vlm", "name": "scene",
    "config": {"model": "google/gemma-4-31B-it", "sandbox_id": sandbox.id,
               "prompt": "Describe the scene."},
}])

# 2b. Generation: pass model_name + sandbox_id (jobs return GenerationJob → .wait()).
response = coll.generate_text(prompt="Summarize this.", model_name="Qwen/Qwen3.5-9B",
                             sandbox_id=sandbox.id, max_tokens=300)
job = coll.generate_image(prompt="a city at sunset", model_name="black-forest-labs/FLUX.1-dev",
                          sandbox_id=sandbox.id)
image = job.wait(timeout=900, interval=5)

# 3. Stop when done — provisioning/active/alert all count toward the tier limit.
sandbox.stop(); sandbox.wait_for_stop()

Model IDs must match the catalog exactly (no -FP8 suffix) or create_sandbox raises Unsupported sandbox model. See reference/sandbox.md for the full model catalog, tiers, categories, pricing, and pitfalls.

Error handling

python
from videodb.exceptions import AuthenticationError, InvalidRequestError

try:
    conn = videodb.connect()
except AuthenticationError:
    print("Check your VIDEO_DB_API_KEY")

try:
    video = coll.upload(url="https://example.com/video.mp4")
except InvalidRequestError as e:
    print(f"Upload failed: {e}")
Show full SKILL.md (695 more words)Show less
Common pitfalls
ScenarioError messageSolution
Search result has no stream URLAttributeError: 'SearchResponse' object has no attribute 'stream_url'search() returns SearchResponse in 0.5.0. Use results.compile()
search(score_threshold=) searches the wrong indexesno error, unexpected resultsscore_threshold does not route to legacy. Use semantic_search(score_threshold=), or legacy_search() for v1 indexes
Semantic index on object detectionuse_for includes semantic but no scene has embeddable textObject artifacts have no top-level text. Omit use_for (it degrades automatically) or pass ["query", "aggregate"]
Indexing a field that does not existfields.filter names not present in any scene's dataThe error lists the available field names — read it. Or check index.field_schema
Search finds no matchesv2 returns an empty SearchResponse; only legacy_search() raises InvalidRequestError: No results foundCheck len(response). Wrap only legacy calls in try/except
Indexing an already-indexed video (v1)Spoken word index for video already existsUse video.index_spoken_words(force=True) to skip if already indexed
Reframe times outBlocks indefinitely on long videosUse start/end to limit segment, or pass callback_url for async
Negative timestamps on TimelineSilently produces broken streamAlways validate start >= 0 before creating VideoAsset
generate_video() / create_collection() failsOperation not allowed or maximum limitPlan-gated features — inform the user about plan limits

Additional docs

Reference documentation is in ${CLAUDE_SKILL_DIR}/reference/. Read files there with that prefix; the links below are relative to this SKILL.md.

Legacy v1 indexing and search. These APIs still work and are not deprecated, but read these only when maintaining existing v1 code:

Screen Recording (Desktop Capture)

Use ws_listener.py to capture WebSocket events during recording sessions. Desktop capture supports macOS only.

${CLAUDE_SKILL_DIR} is this skill's install directory, set by Claude Code. On agents that do not set it, substitute the directory holding this SKILL.md.

Quick Start
  1. Start listener: python "${CLAUDE_SKILL_DIR}/scripts/ws_listener.py" --cwd=<PROJECT_ROOT> &
  2. Get WebSocket ID: cat /tmp/videodb_ws_id
  3. Run capture code (see reference/capture.md for full workflow)
  4. Events written to: /tmp/videodb_events.jsonl
Query Events
python
import json
events = [json.loads(l) for l in open("/tmp/videodb_events.jsonl")]

# Get all transcripts
transcripts = [e["data"]["text"] for e in events if e.get("channel") == "transcript"]

# Get visual descriptions from last 5 minutes
import time
cutoff = time.time() - 300
recent_visual = [e for e in events 
                 if e.get("channel") == "visual_index" and e["unix_ts"] > cutoff]
Utility Scripts
  • ${CLAUDE_SKILL_DIR}/scripts/ws_listener.py - WebSocket event listener (dumps to JSONL)

For complete capture workflow, see reference/capture.md.

Do not use ffmpeg, moviepy, or local encoding tools when VideoDB supports the operation. The following are all handled server-side by VideoDB — trimming, combining clips, overlaying audio or music, adding subtitles, text/image overlays, transcoding, resolution changes, aspect-ratio conversion, resizing for platform requirements, transcription, volume control, fade transitions, and media generation. Only fall back to local tools for operations listed under Limitations in reference/editor.md (speed changes, crop/zoom, colour grading, keyframe animation).

When to use what
ProblemVideoDB solution
Make a video searchablevideo.understand(analyzers=[...]) then video.index(source=analyzer)
Find moments by what was said or shownvideo.search(query), or semantic_search(index_names=[...]) to target an index
Answer a question about a videovideo.ask(question, include_sources=True)
Count or group what appears in a videovideo.aggregate(index_name=..., group_by=..., metric="count")
Filter moments on exact field valuesvideo.query(index_name=..., filter={...})
Platform rejects video aspect ratio or resolutionvideo.reframe() or conn.transcode() with VideoConfig
Need to resize video for Twitter/Instagram/TikTokvideo.reframe(target="vertical") or target="square"
Need to change resolution (e.g. 1080p → 720p)conn.transcode() with VideoConfig(resolution=720)
Need to overlay audio/music on videoAudioAsset on an Editor Timeline with volume control
Need to add subtitlesvideo.add_subtitle() or CaptionAsset on Editor Timeline
Need to combine/trim clipsVideoAsset on an Editor Timeline
Need to compose images with voiceoverImageAsset + AudioAsset on separate Editor tracks
Need to generate voiceover, music, or SFXcoll.generate_voice(), generate_music(), generate_sound_effect()

© video-db, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 23 other files (scripts) in python of video-db/skills.

  • SKILL.md
  • .env.example
  • reference/api-reference.md
  • reference/capture-reference.md
  • reference/capture.md
  • reference/editor-reference.md
  • reference/editor.md
  • reference/generative.md
  • reference/indexing-reference.md
  • reference/indexing.md
  • reference/legacy/index-reference.md
  • reference/legacy/index.md
  • reference/legacy/search.md
  • reference/migration.md
  • reference/rtstream-reference.md
  • reference/rtstream.md
  • reference/sandbox-reference.md
  • reference/sandbox.md
  • reference/search-reference.md
  • … and 5 more

Open the folder on GitHubat commit b47b587

Compare with similar skills

VideoDB Video and Audio Workflows next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

VideoDB Video and Audio Workflows compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
VideoDB Video and Audio Workflows this skillvideo-db/skills127—~5.5kAutomated safety check: NotesMIT
Cassette Video EditCassette-Editor/oh-my-cassette1581 repos~3.4kAutomated safety check: PassMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
CLI Anything VideocaptionerHKUDS/CLI-Anything52k—~1.2kAutomated safety check: PassApache-2.0
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo129k—~2.1kAutomated safety check: WarnMIT
Vox DirectorAlisa0808/vox-director2.2k—~5.6kAutomated safety check: PassMIT

Similar skills

  • Cassette Video Edit

    Cassette-Editor/oh-my-cassette

    Edit, trim, cut, caption, subtitle, reframe, combine, add background music to, or export video, audio, and image files through Cassette.

    158 GitHub starsUsed in 1 repo~3.4k tokens
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated 11 days ago
    Media & CreativeAuto-check: notes
  • CLI Anything Videocaptioner

    HKUDS/CLI-Anything

    AI-powered video captioning — transcribe speech, optimize/translate subtitles, and burn them into video via the stable VideoCaptioner backend.

    52k GitHub stars~1.2k tokensUpdated 17 days ago
    Media & CreativeAuto-check passed
  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    276k GitHub starsUsed in 3 repos~3.5k tokens
    Media & CreativeAuto-check: notes

Works with

Questions about VideoDB Video and Audio Workflows

What does VideoDB Video and Audio Workflows do?

Ingests video, audio, live streams and desktop captures, indexes them for timestamped search, then edits, subtitles, dubs and clips the results. VideoDB gives an agent a way to see, understand and act on video and audio. It can take a local file, a URL or an RTSP feed, record the desktop with screen, microphone and system audio, and hand back a playable stream link together with live context from whatever it is watching.

When should I use VideoDB Video and Audio Workflows?

VideoDB Video and Audio Workflows fits situations like: searching a folder of recordings for every scene that shows a given object; adding subtitles, translations or dubbing to a video; recording a desktop session and getting a summary when it ends; watching an RTSP camera feed and raising alerts on events.

How do I install VideoDB Video and Audio Workflows in Claude Code?

Run `npx skills add video-db/skills --skill videodb -a claude-code`. Or copy the skill folder (python in video-db/skills) into .claude/skills/videodb in your project. Claude Code loads it when a task matches its description.

How do I install VideoDB Video and Audio Workflows in Codex?

Run `npx skills add video-db/skills --skill videodb -a codex`. Or copy the skill folder (python in video-db/skills) into .agents/skills/videodb in your project. Codex loads it when a task matches its description.

Can I use VideoDB Video and Audio Workflows in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add video-db/skills --skill videodb -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/videodb, .gemini/skills/videodb, .github/skills/videodb and .opencode/skills/videodb in your project.

What does VideoDB Video and Audio Workflows need to run?

Going by SKILL.md and its folder, VideoDB Video and Audio Workflows needs the command-line tools its instructions call (python, pip and npx) and credentials named VIDEO_DB_API_KEY. Our summary lists: Python. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash(python:*).

Does VideoDB Video and Audio Workflows access the network?

SKILL.md names 3 domains. In commands or code: console.videodb.io, mcp.videodb.io and youtube.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is VideoDB Video and Audio Workflows safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does VideoDB Video and Audio Workflows use?

VideoDB Video and Audio Workflows is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does VideoDB Video and Audio Workflows use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to VideoDB Video and Audio Workflows?

Skills that share tags, products or a category with VideoDB Video and Audio Workflows: Cassette Video Edit (Cassette-Editor/oh-my-cassette, 158 stars), Edu Math Video (wy51ai/edulab, 1.4k stars), CLI Anything Videocaptioner (HKUDS/CLI-Anything, 52k stars) and MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 129k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains VideoDB Video and Audio Workflows?

video-db (a GitHub organization) maintains it in video-db/skills, which has 127 GitHub stars. The repository was last updated on September 11, 2026.

Source: video-db/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.