Agent skill

Vision Memory MCP

by putervision in putervision/state-memory-mcp

Teaches the agent to use the Visual Memory MCP server to cache webpage and application screenshots, matching layout states and avoiding redundant LLM vision calls.

MITAuto-check: notesAgent Workflows

Install Vision Memory MCP

skills CLI
$ npx skills add putervision/state-memory-mcp --skill vision-memory-mcp -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install putervision/state-memory-mcp vision-memory-mcp --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/putervision/state-memory-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/vision-memory-mcp .claude/skills/vision-memory-mcp && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vision-memory-mcp
GitHub stars
111
Token cost
~1.4k tokens
SKILL.md length
571 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Teaches the agent to use the Visual Memory MCP server to cache webpage and application screenshots, matching layout states and avoiding redundant LLM vision calls.

  • Works in 4 steps: Priority Order & Checklist → Complete Tool Reference → Agent Permissions & Auto-Run Configuration → …
  • Tasks that involve MCP servers
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Vision Memory MCP is an agent skill from putervision/state-memory-mcp. Teaches the agent to use the Visual Memory MCP server to cache webpage and application screenshots, matching layout states and avoiding redundant LLM vision calls.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol. The repository describes itself as: Persistent, branch-aware workflow state memory MCP server for AI coding assistants. Tracks tasks, accepted decisions, and active blockers to prevent session context bloat and… The licence is MIT.

When your agent uses it

  • Tasks that involve MCP servers

Example prompts

  • “Use the vision-memory-mcp skill to teach the agent to use the Visual Memory MCP server to cache webpage and application screenshots, matching layout…”
  • “/vision-memory-mcp”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Priority Order & Checklist
  2. Complete Tool Reference
  3. Agent Permissions & Auto-Run Configuration
  4. CLI Commands Reference

What it can do on your machine

Read from SKILL.md and the folder at commit 0f60ae4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vision Memory MCP loads about 1.4k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 571 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:49
    rkspace .vision-memory-mcp/, .gitignore, .env, and IDE agent rules.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from putervision/state-memory-mcp at commit 0f60ae4, republished under its MIT licence (© putervision). 571 words, ~1,387 tokens.

Download SKILL.mdSave it as .claude/skills/vision-memory-mcp/SKILL.md (or your agent's skills folder).
name
vision-memory-mcp
description
Teaches the agent to use the Visual Memory MCP server to cache webpage and application screenshots, matching layout states and avoiding redundant LLM vision calls.

Visual Memory (vision-memory-mcp)

This project utilizes vision-memory-mcp to cache visual states, record layout transitions, and avoid repetitive LLM vision calls.

1. Priority Order & Checklist

Whenever you capture a screenshot, examine a webpage, or need to verify a visual state, you MUST run this sequence:

  1. Orient: Call get_session_context to load active transitions and recent visual states.
  2. Search (Optional): Call recall_memory to search past visual states by text query or image query.
  3. Ingest/Verify: Call analyze_screenshot with the base64 screenshot before calling any vision LLMs.
  4. Cache Hit: If is_known is true, read the returned description and do NOT call your vision LLM.
  5. Cache Miss: If is_known is false, inspect the image with your vision model, summarize the layout, and register it back by calling analyze_screenshot with both the screenshot and description parameters.
  6. Log Transitions: Right after taking any UI action (click, type, navigate, scroll), call record_outcome to build the navigation graph.
  7. Snapshotting: Call manage_snapshot (action: "save") when reaching milestones, and manage_snapshot (action: "diff") to check for visual regressions.
2. Complete Tool Reference
Tool NameKey InputsDescription
analyze_screenshotscreenshot? (base64), file_path?, description?, items?Main ingestion (single or batch) and visual state retrieval tool.
recall_memoryquery?, screenshot?, file_path?, strategy?, limit?Search visual memory by text query or image query (read-only).
record_outcomefrom_state_id, to_state_id?, action, action_type? ('blocker' | 'click' | etc.)Record UI action transitions or log visual blockers for state-memory.
get_navigation_pathsfrom_state_id?, to_state_id?, to_description?, max_hops?Find historical path or instructions between states.
predict_next_actioncurrent_state_id, goal_description?, goal_state_id?Predict best next UI action and grounded element handles (target_selector, target_coords).
compare_statesstate_a_id & state_b_id OR video_a_id & video_b_idCompare two states visually (has_layout_change) or compare video runs.
get_session_contextinclude_recent?, include_frequent?Get recent/frequent states, transition graphs, disk stats, cache metrics, and version info.
manage_snapshotaction ('save' | 'diff' | 'export' | 'restore'), name?, archive_json?Unified snapshot management for visual checkpoints and regression detection.
manage_visual_specaction ('set' | 'verify' | 'list'), name?, screenshot?, tolerance?Register and verify visual design contract baselines (Visual SDD).
manage_videoaction ('ingest' | 'search' | 'timeline'), file_path?, query?, video_id?Ingest WebM/MP4 recordings, search video keyframes, or retrieve timelines.
create_evidence_packkeyframe_state_ids, source_video_id?, linked_state_memory_nodes?Package immutable evidence packs linking video keyframes to state-memory DAGs.
export_trajectoriesformat? ('json' | 'llava' | 'qwen2_vl' | 'joint'), trace_id?Export multimodal trajectories for model fine-tuning or joint workflow exports.
undo_visual_mutationtype? ('state' | 'transition' | 'any')Revert the last visual state ingestion or transition edge addition.
forget_statestate_idPurge a specific state and vector embedding for privacy.
wait_for_visual_statetarget_state_id, timeout_ms?Poll for target visual state until present or timeout occurs.
Show full SKILL.md (162 more words)Show less
3. Agent Permissions & Auto-Run Configuration

To bypass confirmation dialogs when running CLI cache commands or reading/writing brain images, add these allows to your configuration:

  • Google Antigravity (~/.gemini/config/config.json): Add these rules to your "globalPermissionGrants" -> "allow" list:
    • "command(vision-memory-mcp)" (Allows running any query/ingest command prefix)
    • "read_file(.*\\.gemini/antigravity/brain/.*)" (Allows reading brain screenshots)
    • "write_file(.*\\.gemini/antigravity/brain/.*)" (Allows saving brain snapshots)
4. CLI Commands Reference

Run these commands in the terminal for management and analytics:

  • vision-memory-mcp init [-y|--yes]: Scaffold workspace .vision-memory-mcp/, .gitignore, .env, and IDE agent rules.
  • vision-memory-mcp init-global: Re-initialize across all projects registered in ~/.vision-memory-mcp/projects.json.
  • vision-memory-mcp doctor: Health check storage writability, sharp bindings, Node runtime, and sub-directory Git repos.
  • vision-memory-mcp audit: Audit sub-directory Git repos, submodules, database locations, and total visual states.
  • vision-memory-mcp inspect: Display stored visual states in an ASCII table.
  • vision-memory-mcp metrics: Calculate cache hit rate, token savings, and ROI.
  • vision-memory-mcp view: Open an interactive force-directed graph visualizer in the browser.
  • vision-memory-mcp export --format [json\|mermaid\|html] --out [file]: Export the memory graph.
  • vision-memory-mcp prune: Purge expired or low-access states.

© putervision, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/vision-memory-mcp of putervision/state-memory-mcp.

Open the folder on GitHubat commit 0f60ae4

Compare with similar skills

Vision Memory MCP next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vision Memory MCP compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vision Memory MCP this skillputervision/state-memory-mcp111—~1.4kAutomated safety check: NotesMIT
MCP Server Builderanthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
MCP Integration for Pluginsanthropics/claude-plugins-official38k11 repos~3.1kAutomated safety check: PassApache-2.0
Fastmcp Client CLIPrefectHQ/fastmcp28k1 repos~823Automated safety check: PassApache-2.0
Crush Configurationcharmbracelet/crush29k—~3.7kAutomated safety check: PassCustom licence

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • MCP Integration for Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to bundle Model Context Protocol servers in a Claude Code plugin, covering config files, stdio, SSE, HTTP and WebSocket server types, and authentication.

    38k GitHub starsUsed in 11 repos~3.1k tokens
    Agent WorkflowsAuto-check passed
  • Fastmcp Client CLI

    PrefectHQ/fastmcp

    Query and invoke tools on MCP servers using fastmcp list and fastmcp call.

    28k GitHub starsUsed in 1 repo~823 tokens
    Agent WorkflowsAuto-check passed
  • Crush Configuration

    charmbracelet/crush

    Explains how to configure the Crush coding agent with crushrc or crush.json, covering providers, models, LSPs, MCP servers, hooks, permissions and config precedence.

    29k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Context Mode Output Sandbox

    mksglu/context-mode

    Routes large command, file, API and browser output through context-mode tools so only the needed result enters the agent's context, instead of dumping it via Bash.

    26k GitHub stars~4.1k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from putervision/state-memory-mcp

  • Behavior MCP

    putervision/state-memory-mcp

    Teaches the agent to use the Behavior MCP server for ~60Hz in-browser behavior trees, triggers, and recordings.

    111 GitHub stars~853 tokensUpdated 4 days ago
    Auto-check passed
  • State Memory MCP

    putervision/state-memory-mcp

    Teaches the agent to use the state-memory-mcp MCP server to track workflow state, tasks, decisions, blockers, artifacts, plans, milestones, and their semantic relationships in a persistent graph…

    111 GitHub stars~1.9k tokensUpdated 4 days ago
    Auto-check passed
  • Video Ingest

    putervision/state-memory-mcp

    Teaches the agent to process, ingest, analyze, and compare WebM, MP4, and GIF video recordings using vision-memory-mcp and state-memory-mcp.

    111 GitHub stars~902 tokensUpdated 4 days ago
    Auto-check passed
  • Webcrypt MCP

    putervision/state-memory-mcp

    Teaches the agent to use the WebCrypt MCP server for AES-256-GCM symmetric encryption, RSA-4096 hybrid encryption, key generation, digital signatures, hashing, and post-quantum cryptography.

    111 GitHub stars~847 tokensUpdated 4 days ago
    Auto-check passed
  • Agent Reasoning MCP

    putervision/state-memory-mcp

    Teaches the agent to use the Strategic Agent Reasoning MCP server for BDI goals, utility scoring, risk evaluation, and replanning.

    111 GitHub stars~889 tokensUpdated 4 days ago
    Auto-check passed
  • World Model MCP

    putervision/state-memory-mcp

    Teaches the agent to use the Spatial World Model MCP server to track entities, 3D/2D positions, spatial relationships, object permanence, and movement simulation.

    111 GitHub stars~804 tokensUpdated 4 days ago
    Auto-check passed

Categories

Questions about Vision Memory MCP

What does Vision Memory MCP do?

Teaches the agent to use the Visual Memory MCP server to cache webpage and application screenshots, matching layout states and avoiding redundant LLM vision calls. Vision Memory MCP is an agent skill from putervision/state-memory-mcp. Teaches the agent to use the Visual Memory MCP server to cache webpage and application screenshots, matching layout states and avoiding redundant LLM vision calls.

When should I use Vision Memory MCP?

Vision Memory MCP fits situations like: tasks that involve MCP servers.

How do I install Vision Memory MCP in Claude Code?

Run `npx skills add putervision/state-memory-mcp --skill vision-memory-mcp -a claude-code`. Or copy the skill folder (.agents/skills/vision-memory-mcp in putervision/state-memory-mcp) into .claude/skills/vision-memory-mcp in your project. Claude Code loads it when a task matches its description.

How do I install Vision Memory MCP in Codex?

Run `npx skills add putervision/state-memory-mcp --skill vision-memory-mcp -a codex`. Or copy the skill folder (.agents/skills/vision-memory-mcp in putervision/state-memory-mcp) into .agents/skills/vision-memory-mcp in your project. Codex loads it when a task matches its description.

Can I use Vision Memory MCP in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add putervision/state-memory-mcp --skill vision-memory-mcp -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vision-memory-mcp, .gemini/skills/vision-memory-mcp, .github/skills/vision-memory-mcp and .opencode/skills/vision-memory-mcp in your project.

What does Vision Memory MCP need to run?

SKILL.md names no scripts, command-line tools or credentials: Vision Memory MCP is instructions for the agent only.

Does Vision Memory MCP access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vision Memory MCP safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Vision Memory MCP use?

Vision Memory MCP is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vision Memory MCP use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vision Memory MCP?

Skills that share tags, products or a category with Vision Memory MCP: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), MCP Integration for Plugins (anthropics/claude-plugins-official, 38k stars) and Fastmcp Client CLI (PrefectHQ/fastmcp, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vision Memory MCP?

putervision (a GitHub organization) maintains it in putervision/state-memory-mcp, which has 111 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 3, 2026.

Source: putervision/state-memory-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.