Agent skill

Video Ingest

by putervision in putervision/state-memory-mcp

Teaches the agent to process, ingest, analyze, and compare WebM, MP4, and GIF video recordings using vision-memory-mcp and state-memory-mcp.

MITAuto-check passedAgent Workflows

Install Video Ingest

skills CLI
$ npx skills add putervision/state-memory-mcp --skill video-ingest -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install putervision/state-memory-mcp video-ingest --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/putervision/state-memory-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/video-ingest .claude/skills/video-ingest && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-ingest
GitHub stars
111
Token cost
~902 tokens
SKILL.md length
334 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Teaches the agent to process, ingest, analyze, and compare WebM, MP4, and GIF video recordings using vision-memory-mcp and state-memory-mcp.

  • Works in 5 steps: When to Use Video Ingestion → Ingestion Strategies & Parameter Tuning → Mandatory Dual-MCP Evidence Workflow → …
  • Tasks that involve MCP servers
  • SKILL.md covers 1. When to Use Video Ingestion, 2. Ingestion Strategies &…, 3. Mandatory Dual-MCP Evidence… and 4. CLI Quick Reference, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Video Ingest is an agent skill from putervision/state-memory-mcp. Teaches the agent to process, ingest, analyze, and compare WebM, MP4, and GIF video recordings using vision-memory-mcp and state-memory-mcp.

Its SKILL.md is about 900 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol. The repository describes itself as: Persistent, branch-aware workflow state memory MCP server for AI coding assistants. Tracks tasks, accepted decisions, and active blockers to prevent session context bloat and… The licence is MIT.

When your agent uses it

  • Tasks that involve MCP servers

Example prompts

  • “Use the video-ingest skill to teach the agent to process, ingest, analyze, and compare WebM, MP4, and GIF video recordings using vision-memory-mcp…”
  • “/video-ingest”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. When to Use Video Ingestion
  2. Ingestion Strategies & Parameter Tuning
  3. Mandatory Dual-MCP Evidence Workflow
  4. CLI Quick Reference
  5. Consolidated Video MCP Tools Reference

What it can do on your machine

Read from SKILL.md and the folder at commit 0f60ae4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Ingest loads about 902 tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 334 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~902

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from putervision/state-memory-mcp at commit 0f60ae4, republished under its MIT licence (© putervision). 334 words, ~902 tokens.

Download SKILL.mdSave it as .claude/skills/video-ingest/SKILL.md (or your agent's skills folder).
name
video-ingest
description
Teaches the agent to process, ingest, analyze, and compare WebM, MP4, and GIF video recordings using vision-memory-mcp and state-memory-mcp.

Video Frame Digesting & Temporal Memory Skill (video-ingest)

This skill provides step-by-step guidance, best practices, and operational patterns for digesting WebM, MP4, and GIF video recordings using @putervision/vision-memory-mcp.


1. When to Use Video Ingestion

Use video ingestion whenever you encounter:

  • E2E Playwright / Cypress / Selenium Test Artifacts: Recorded .webm screenchunks or .mp4 test run videos.
  • Bug Reproduction Videos: User-uploaded screen recordings demonstrating UI glitches or crashes.
  • UI Walkthrough Recordings: Demonstrations of complex multi-step user workflows.
  • Visual Regression Diagnostics: Comparing a passing baseline video run against a failing test run.

2. Ingestion Strategies & Parameter Tuning

ScenarioRecommended ParametersWhy
Action Event Timestamps (Highest Precision)action_timestamps: [1.2, 3.5, 7.0]Samples keyframes at exact interaction timestamps (clicks, types, navigation events) from test runners or state-memory logs.
Dynamic UI / Animationsscene_threshold: 0.3, fps: 1Combines scene-change detection (gt(scene,0.3)) with 1 fps background sampling to capture major screen transitions without frame bloat.
High-Speed Test Runsfps: 2 or fps: 5Increases frame rate sampling for rapidly switching UI test steps.
Long Screen Recordingsfps: 0.5, scene_threshold: 0.4Lowers sampling rate to conserve storage while extracting unique keyframe states.

3. Mandatory Dual-MCP Evidence Workflow

When diagnosing bugs or linking test runs to task nodes:

  1. Ingest Video: Call manage_video (action: "ingest") with file path and action timestamps.
  2. Extract Evidence Payload: Read the returned evidence_payload containing source_video_id, frame_range, and timestamps_ms.
  3. Build Evidence Pack: Call create_evidence_pack linking keyframe_state_ids with state-memory-mcp task or blocker node IDs.
  4. Compare Trajectories (On Failure): Call compare_states(video_a_id: "...", video_b_id: "...") to pinpoint exact frame divergence points.

4. CLI Quick Reference

bash
# Ingest WebM / MP4 video recording with custom category
vision-memory-mcp video ingest ./recording.webm --category e2e_test

# Ingest with explicit action timestamps
vision-memory-mcp video ingest ./bug.mp4 --category bug_repro --action-timestamps 1.2,3.5,8.0

# Inspect chronological timeline & keyframes for a video ID
vision-memory-mcp video inspect vid_d7fc21c0ad52f861

# List all ingested video memory records
vision-memory-mcp video list

5. Consolidated Video MCP Tools Reference

  • manage_video (action: "ingest"): Ingests video file/base64, extracts keyframes, dHash deduplicates, and generates CLIP vector embeddings.
  • manage_video (action: "timeline"): Fetches step-by-step keyframes, exact timestamps (timestamp_ms), OCR snippets, and grounded target handles.
  • manage_video (action: "search"): Searches video memory by description query, category, tags, or file path.
  • compare_states (video_a_id, video_b_id): Calculates similarity score between two recordings and pinpoints exact timestamp divergence.
  • create_evidence_pack: Produces an immutable, cryptographically hashable evidence pack payload linking keyframes to task graph nodes.

© putervision, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/video-ingest of putervision/state-memory-mcp.

Open the folder on GitHubat commit 0f60ae4

Compare with similar skills

Video Ingest next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Ingest compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Ingest this skillputervision/state-memory-mcp111—~902Automated safety check: PassMIT
Agent QA Authoringvostride/agent-qa902—~569Automated safety check: PassCustom licence
Agent QA Result Triagevostride/agent-qa902—~394Automated safety check: PassCustom licence
Lightpandalightpanda-io/agent-skill101—~6kAutomated safety check: PassApache-2.0
Opik Comparecomet-ml/opik-mcp219—~2.6kAutomated safety check: NotesApache-2.0
Agent QA Debug Fixvostride/agent-qa902—~398Automated safety check: PassCustom licence

Similar skills

  • Agent QA Authoring

    vostride/agent-qa

    A skill your agent uses when creating, editing, validating, or running agent-qa tests, suites, or hooks.

    902 GitHub stars~569 tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Agent QA Result Triage

    vostride/agent-qa

    A skill your agent uses when investigating failed agent-qa runs, inspecting artifacts, classifying failures, or comparing recent runs.

    902 GitHub stars~394 tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Lightpanda

    lightpanda-io/agent-skill

    Lightpanda browser, drop-in replacement for Chrome-based browsing in any AI agent - faster and lighter for tasks without graphical rendering like data retrieval.

    101 GitHub stars~6k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Opik Compare

    comet-ml/opik-mcp

    Run a candidate against the baseline over an Opik test suite and read the numbers back — which cases broke, which got fixed, the per-metric deltas, worst rows, and whether the two runs are…

    219 GitHub stars~2.6k tokensUpdated yesterday
    Agent WorkflowsAuto-check: notes
  • Agent QA Debug Fix

    vostride/agent-qa

    Use after an agent-qa run has failed and you need to debug, patch, and verify the issue using MCP evidence, logs, artifacts, and local code changes instead of generated fix suggestions.

    902 GitHub stars~398 tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Unifi MCP Tool Builder

    enuno/unifi-mcp-server

    Specialized guide for adding new MCP tools to the UniFi MCP Server following project standards, UniFi API patterns, and test-driven development practices.

    282 GitHub stars~5.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from putervision/state-memory-mcp

  • Behavior MCP

    putervision/state-memory-mcp

    Teaches the agent to use the Behavior MCP server for ~60Hz in-browser behavior trees, triggers, and recordings.

    111 GitHub stars~853 tokensUpdated 5 days ago
    Auto-check passed
  • State Memory MCP

    putervision/state-memory-mcp

    Teaches the agent to use the state-memory-mcp MCP server to track workflow state, tasks, decisions, blockers, artifacts, plans, milestones, and their semantic relationships in a persistent graph…

    111 GitHub stars~1.9k tokensUpdated 5 days ago
    Auto-check passed
  • Webcrypt MCP

    putervision/state-memory-mcp

    Teaches the agent to use the WebCrypt MCP server for AES-256-GCM symmetric encryption, RSA-4096 hybrid encryption, key generation, digital signatures, hashing, and post-quantum cryptography.

    111 GitHub stars~847 tokensUpdated 5 days ago
    Auto-check passed
  • Agent Reasoning MCP

    putervision/state-memory-mcp

    Teaches the agent to use the Strategic Agent Reasoning MCP server for BDI goals, utility scoring, risk evaluation, and replanning.

    111 GitHub stars~889 tokensUpdated 5 days ago
    Auto-check passed
  • Vision Memory MCP

    putervision/state-memory-mcp

    Teaches the agent to use the Visual Memory MCP server to cache webpage and application screenshots, matching layout states and avoiding redundant LLM vision calls.

    111 GitHub stars~1.4k tokensUpdated 5 days ago
    Auto-check: notes
  • World Model MCP

    putervision/state-memory-mcp

    Teaches the agent to use the Spatial World Model MCP server to track entities, 3D/2D positions, spatial relationships, object permanence, and movement simulation.

    111 GitHub stars~804 tokensUpdated 5 days ago
    Auto-check passed

Questions about Video Ingest

What does Video Ingest do?

Teaches the agent to process, ingest, analyze, and compare WebM, MP4, and GIF video recordings using vision-memory-mcp and state-memory-mcp. Video Ingest is an agent skill from putervision/state-memory-mcp. Teaches the agent to process, ingest, analyze, and compare WebM, MP4, and GIF video recordings using vision-memory-mcp and state-memory-mcp.

When should I use Video Ingest?

Video Ingest fits situations like: tasks that involve MCP servers.

How do I install Video Ingest in Claude Code?

Run `npx skills add putervision/state-memory-mcp --skill video-ingest -a claude-code`. Or copy the skill folder (.agents/skills/video-ingest in putervision/state-memory-mcp) into .claude/skills/video-ingest in your project. Claude Code loads it when a task matches its description.

How do I install Video Ingest in Codex?

Run `npx skills add putervision/state-memory-mcp --skill video-ingest -a codex`. Or copy the skill folder (.agents/skills/video-ingest in putervision/state-memory-mcp) into .agents/skills/video-ingest in your project. Codex loads it when a task matches its description.

Can I use Video Ingest in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add putervision/state-memory-mcp --skill video-ingest -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-ingest, .gemini/skills/video-ingest, .github/skills/video-ingest and .opencode/skills/video-ingest in your project.

What does Video Ingest need to run?

SKILL.md names no scripts, command-line tools or credentials: Video Ingest is instructions for the agent only.

Does Video Ingest access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Video Ingest safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Video Ingest use?

Video Ingest is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Ingest use?

About 902 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video Ingest?

Skills that share tags, products or a category with Video Ingest: Agent QA Authoring (vostride/agent-qa, 902 stars), Agent QA Result Triage (vostride/agent-qa, 902 stars), Lightpanda (lightpanda-io/agent-skill, 101 stars) and Opik Compare (comet-ml/opik-mcp, 219 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Ingest?

putervision (a GitHub organization) maintains it in putervision/state-memory-mcp, which has 111 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 3, 2026.

Source: putervision/state-memory-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.