Agent skill

Luma Vision Image Analysis

by JochenYang in JochenYang/luma-mcp

Calls an external vision model through vision.js to analyze an image when the user explicitly invokes /skill luma-vision, for agents whose own model cannot see images.

MITAuto-check passedAI & LLM Engineering

SKILL.md written in Chinese; this summary is our English description.

Install Luma Vision Image Analysis

skills CLI
$ npx skills add JochenYang/luma-mcp --skill luma-vision -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JochenYang/luma-mcp luma-vision --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JochenYang/luma-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/vision-skill .claude/skills/luma-vision && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
luma-vision
GitHub stars
116
Token cost
~295 tokens
SKILL.md length
89 words
Files
2 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Calls an external vision model through vision.js to analyze an image when the user explicitly invokes /skill luma-vision, for agents whose own model cannot see images.

  • Analyzing a screenshot, error message or UI through an explicit /skill luma-vision command
  • SKILL.md covers 说明, 何时触发, 何时不触发 and 行为, plus 3 more sections
  • Runs JavaScript scripts from its folder; calls node; needs CUSTOM_API_KEY
  • Giving a text-only coding model the ability to understand an attached image

What it does

The skill only activates when the current message starts with /skill luma-vision and carries an image; it does not fire for a plainly pasted image or for an earlier attached image in the conversation. It exists because a model can declare image_in in its config without actually understanding image content, and this gives such a text-only model a path to real image analysis instead.

It runs node scripts/vision.js with the image source and a question, never substituting ReadMediaFile. Four image source forms are supported: a local path, an HTTP or HTTPS URL, a data URI for inline images and an @-prefixed path such as @clipboard.png, which has the @ stripped and is treated as local. The API endpoint, model name and key are read from CUSTOM_BASE_URL, CUSTOM_MODEL_NAME and CUSTOM_API_KEY in the environment. The SKILL.md is in Chinese.

When your agent uses it

  • Analyzing a screenshot, error message or UI through an explicit /skill luma-vision command
  • Giving a text-only coding model the ability to understand an attached image
  • Reading an image from a local path, a URL or a data URI with an external vision model

Example prompts

  • “/skill luma-vision ./error-screenshot.png what does this stack trace say?”
  • “/skill luma-vision https://example.com/diagram.png explain this architecture diagram.”
  • “/skill luma-vision @clipboard.png describe what's wrong with this UI.”

Requirements

  • Node, to run scripts/vision.js
  • CUSTOM_BASE_URL, CUSTOM_MODEL_NAME and CUSTOM_API_KEY set for a vision model provider

What it can do on your machine

Read from SKILL.md and the folder at commit e686dd8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CUSTOM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Luma Vision Image Analysis loads about 295 tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 89 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~295

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from JochenYang/luma-mcp at commit e686dd8, republished under its MIT licence (© JochenYang). 89 words, ~295 tokens.

Download SKILL.mdSave it as .claude/skills/luma-vision/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
luma-vision
description
多模型视觉理解能力。通过 /skill luma-vision 激活时,执行 vision.js 脚本调用外部视觉模型分析图片。

Luma Vision Skill — 多模型视觉理解

说明

所有模型要发送图片都必须在 config.toml 中声明 image_in 能力,否则前端会拦截图片提交。 但纯文本模型即使声明了 image_in 也无法真正理解图片内容。

本 skill 只在用户主动通过 /skill luma-vision 命令发送图片时激活。 如果用户直接粘贴图片(没有 /skill),不要执行本 skill 的逻辑。

何时触发

  • 仅限 用户当前消息以 /skill luma-vision 开头并附带图片
  • 历史消息中的 Attached image file: 内容不触发本 skill

何时不触发

  • 用户直接粘贴图片发送(没有 /skill 前缀)
  • 多模态模型收到图片时,直接用原生看图能力,不要走脚本

行为

直接执行 vision.js 脚本,不要用 ReadMediaFile:

bash
node "<skill_dir>/scripts/vision.js" "<图片来源>" "<问题描述>"

支持的图片来源

脚本兼容三种来源,适配不同 agent 的传图方式:

类型示例适用场景
本地路径./image.png、D:\photos\photo.jpgKimi Code skill 传的缓存路径
HTTP(S) URLhttps://example.com/image.png网页图片
Data URIdata:image/png;base64,...Claude Code 等直接内联传入
@ 前缀路径@clipboard.png自动剥离 @ 后按本地路径处理

环境变量

在系统环境变量中配置,脚本会自动读取:

变量说明
CUSTOM_BASE_URLAPI 地址
CUSTOM_MODEL_NAME模型名称
CUSTOM_API_KEYAPI Key

注意事项

  • 脚本路径:<skill_dir>/scripts/vision.js
  • 不要用 ReadMediaFile 代替

© JochenYang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in vision-skill of JochenYang/luma-mcp.

  • SKILL.md
  • scripts/vision.js

Open the folder on GitHubat commit e686dd8

Compare with similar skills

Luma Vision Image Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Luma Vision Image Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Luma Vision Image Analysis this skillJochenYang/luma-mcp116—~295Automated safety check: PassMIT
Pi AgentK-Dense-AI/scientific-agent-skills48k1 repos~2.1kAutomated safety check: PassMIT
Openspec AwareChorus-AIDLC/Chorus1.2k—~7.3kAutomated safety check: PassAGPL-3.0
Openspec Aware ChorusChorus-AIDLC/Chorus1.2k—~7.5kAutomated safety check: PassAGPL-3.0
Openspec AwareChorus-AIDLC/Chorus1.2k—~7.2kAutomated safety check: NotesAGPL-3.0
AutoRAG Setup and RepairMarker-Inc-Korea/AutoRAG5.1k—~5.5kAutomated safety check: PassMIT

Similar skills

  • Pi Agent

    K-Dense-AI/scientific-agent-skills

    Builds with and operates Pi, the minimal terminal coding harness.

    48k GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Openspec Aware

    Chorus-AIDLC/Chorus

    OpenSpec-mode authoring for Chorus PM workflows in Codex. An agent skill from Chorus-AIDLC/Chorus.

    1.2k GitHub stars~7.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Openspec Aware Chorus

    Chorus-AIDLC/Chorus

    OpenSpec-mode authoring for Chorus PM workflows on dsh — the default whenever OpenSpec is usable.

    1.2k GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Openspec Aware

    Chorus-AIDLC/Chorus

    OpenSpec-mode authoring for Chorus PM workflows in Hermes. An agent skill from Chorus-AIDLC/Chorus.

    1.2k GitHub stars~7.2k tokensUpdated today
    Agent WorkflowsAuto-check: notes
  • AutoRAG Setup and Repair

    Marker-Inc-Korea/AutoRAG

    Installs, configures, and repairs AutoRAG's search model, approved folders, indexes, and datasources, and registers its Lite MCP server.

    5.1k GitHub stars~5.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Agent Framework

    jihadkhawaja/Egroo

    Build, extend, and debug AI agents in Egroo using the Microsoft Agent Framework (C .NET).

    178 GitHub stars~1.9k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

Questions about Luma Vision Image Analysis

What does Luma Vision Image Analysis do?

Calls an external vision model through vision.js to analyze an image when the user explicitly invokes /skill luma-vision, for agents whose own model cannot see images. The skill only activates when the current message starts with /skill luma-vision and carries an image; it does not fire for a plainly pasted image or for an earlier attached image in the conversation. It exists because a model can declare image_in in its config without actually understanding image content, and this gives such a text-only model a path to real image analysis instead.

When should I use Luma Vision Image Analysis?

Luma Vision Image Analysis fits situations like: analyzing a screenshot, error message or UI through an explicit /skill luma-vision command; giving a text-only coding model the ability to understand an attached image; reading an image from a local path, a URL or a data URI with an external vision model.

How do I install Luma Vision Image Analysis in Claude Code?

Run `npx skills add JochenYang/luma-mcp --skill luma-vision -a claude-code`. Or copy the skill folder (vision-skill in JochenYang/luma-mcp) into .claude/skills/luma-vision in your project. Claude Code loads it when a task matches its description.

How do I install Luma Vision Image Analysis in Codex?

Run `npx skills add JochenYang/luma-mcp --skill luma-vision -a codex`. Or copy the skill folder (vision-skill in JochenYang/luma-mcp) into .agents/skills/luma-vision in your project. Codex loads it when a task matches its description.

Can I use Luma Vision Image Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JochenYang/luma-mcp --skill luma-vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/luma-vision, .gemini/skills/luma-vision, .github/skills/luma-vision and .opencode/skills/luma-vision in your project.

What does Luma Vision Image Analysis need to run?

Going by SKILL.md and its folder, Luma Vision Image Analysis needs JavaScript for the scripts in its folder, the command-line tools its instructions call (node) and credentials named CUSTOM_API_KEY. Our summary lists: Node, to run scripts/vision.js; CUSTOM_BASE_URL, CUSTOM_MODEL_NAME and CUSTOM_API_KEY set for a vision model provider.

Does Luma Vision Image Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Luma Vision Image Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Luma Vision Image Analysis use?

Luma Vision Image Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Luma Vision Image Analysis use?

About 295 tokens (SKILL.md is roughly 1.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Luma Vision Image Analysis?

Skills that share tags, products or a category with Luma Vision Image Analysis: Pi Agent (K-Dense-AI/scientific-agent-skills, 48k stars), Openspec Aware (Chorus-AIDLC/Chorus, 1.2k stars), Openspec Aware Chorus (Chorus-AIDLC/Chorus, 1.2k stars) and Openspec Aware (Chorus-AIDLC/Chorus, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Luma Vision Image Analysis?

JochenYang (a GitHub user) maintains it in JochenYang/luma-mcp, which has 116 GitHub stars. The repository was last updated on August 9, 2026.

Source: JochenYang/luma-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.