Segment Anything Model Guide
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual…
$ npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install linkfox-ai/linkfox-skills linkfox-multimodal-recognize-image --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/linkfox-ai/linkfox-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/linkfox-multimodal-recognize-image .claude/skills/linkfox-multimodal-recognize-image && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "linkfox-multimodal-recognize-image" agent skill from https://github.com/linkfox-ai/linkfox-skills/tree/main/skills/linkfox-multimodal-recognize-image into .claude/skills/linkfox-multimodal-recognize-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linkfox-multimodal-recognize-image", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/linkfox-ai/linkfox-skills/tree/main/skills/linkfox-multimodal-recognize-imageType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install linkfox-ai/linkfox-skills linkfox-multimodal-recognize-image --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/linkfox-ai/linkfox-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/linkfox-multimodal-recognize-image .agents/skills/linkfox-multimodal-recognize-image && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "linkfox-multimodal-recognize-image" agent skill from https://github.com/linkfox-ai/linkfox-skills/tree/main/skills/linkfox-multimodal-recognize-image into .agents/skills/linkfox-multimodal-recognize-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linkfox-multimodal-recognize-image", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install linkfox-ai/linkfox-skills linkfox-multimodal-recognize-image --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/linkfox-ai/linkfox-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/linkfox-multimodal-recognize-image .cursor/skills/linkfox-multimodal-recognize-image && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "linkfox-multimodal-recognize-image" agent skill from https://github.com/linkfox-ai/linkfox-skills/tree/main/skills/linkfox-multimodal-recognize-image into .cursor/skills/linkfox-multimodal-recognize-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linkfox-multimodal-recognize-image", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/linkfox-ai/linkfox-skills.git --path skills/linkfox-multimodal-recognize-image--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install linkfox-ai/linkfox-skills linkfox-multimodal-recognize-image --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/linkfox-ai/linkfox-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/linkfox-multimodal-recognize-image .gemini/skills/linkfox-multimodal-recognize-image && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "linkfox-multimodal-recognize-image" agent skill from https://github.com/linkfox-ai/linkfox-skills/tree/main/skills/linkfox-multimodal-recognize-image into .gemini/skills/linkfox-multimodal-recognize-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linkfox-multimodal-recognize-image", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install linkfox-ai/linkfox-skills linkfox-multimodal-recognize-imageInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/linkfox-ai/linkfox-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/linkfox-multimodal-recognize-image .github/skills/linkfox-multimodal-recognize-image && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "linkfox-multimodal-recognize-image" agent skill from https://github.com/linkfox-ai/linkfox-skills/tree/main/skills/linkfox-multimodal-recognize-image into .github/skills/linkfox-multimodal-recognize-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linkfox-multimodal-recognize-image", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install linkfox-ai/linkfox-skills linkfox-multimodal-recognize-image --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/linkfox-ai/linkfox-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/linkfox-multimodal-recognize-image .opencode/skills/linkfox-multimodal-recognize-image && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "linkfox-multimodal-recognize-image" agent skill from https://github.com/linkfox-ai/linkfox-skills/tree/main/skills/linkfox-multimodal-recognize-image into .opencode/skills/linkfox-multimodal-recognize-image/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "linkfox-multimodal-recognize-image", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
linkfox-multimodal-recognize-image基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual…
Linkfox Multimodal Recognize Image is an agent skill from linkfox-ai/linkfox-skills. 基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual Q&A时触发此技能。当用户提到图片识别、图片分析、图片描述、识别图片内容、分析产品图、从图片中读取文字、描述图片、提取视觉内容或理解照片内容时触发。当用户提供图片URL并就其视觉内容提问时,即使未明确说"图片识别",也应触发此技能。
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/api.md`, `references/onboarding.md` and `scripts/multimodal_recognize_image.py`).
It sits in AI & LLM Engineering, covering Computer vision. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 38fef04. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
skill.linkfox.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
LINKFOX_AGENT_API_KEYLINKFOXAGENT_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Linkfox Multimodal Recognize Image loads about 1.7k tokens when it runs, and up to ~2.9k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 826 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from linkfox-ai/linkfox-skills at commit 38fef04, republished under its MIT licence (© linkfox-ai). 826 words, ~1,687 tokens.
.claude/skills/linkfox-multimodal-recognize-image/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.This skill guides you on how to use the multimodal image recognition API to analyze images from URLs and extract meaningful information based on user intent.
The Image Recognition tool accepts an image URL and an optional natural-language requirement describing what the user wants to know about the image. The backend uses a multimodal AI model to interpret the visual content and return a textual description or analysis.
Supported formats: JPG, JPEG, PNG, GIF, WebP, BMP.
How it works: You provide a publicly accessible image URL and a requirement (what you want to learn from the image). The service downloads the image, runs multimodal analysis, and returns a text-based result.
| Parameter | Required | Description |
|---|---|---|
| imageUrl | Yes | A publicly accessible URL pointing to the image. Must be JPG, JPEG, PNG, GIF, WebP, or BMP. Maximum 1000 characters. |
| requirement | No | A natural-language description of what to identify or analyze in the image. Defaults to "Describe the content of this image" when omitted. Maximum 1000 characters. |
This tool requires a publicly accessible image URL. If the user provides a local image file path (e.g., C:\Users\...\photo.png, /home/.../image.jpg), you must upload it first to obtain a public URL.
Run the upload script:
python scripts/upload_image.py /path/to/local/image.pngThe script will return a public URL (valid for 24 hours) that can be used as the image URL parameter.
1. General Image Description
imageUrl to the provided URL, leave requirement as default.2. Product Image Analysis
requirement to: "This is an Amazon product listing image. Identify the product, key features, and selling points visible in the image."3. Text Extraction from an Image
requirement to: "Extract all visible text from this image, preserving layout where possible."4. A+ Page Image Review
requirement to: "This is an Amazon A+ product description image. Describe the visual content, key messaging, and branding elements."5. Comparison / Detail Inspection
requirement to: "Identify and describe any differences between the product and its packaging shown in the image."POST /multimodal/recognizeImage(完整参数/响应/错误码见 references/api.md)python scripts/multimodal_recognize_image.py '<JSON 参数>' [--inline]输出策略(脚本默认行为):
<cwd>/linkfox/<YYYY-MM-DD>/<session>/data/linkfox-multimodal-recognize-image-<timestamp>.json(<cwd> 为脚本执行时的工作目录,在 Claude Code 里即当前项目目录;<session> 取自环境变量 SESSION_ID,按用户任务自动聚合;禁止写入 /tmp,当前目录不可写则报错)total/costToken、最大列表字段的长度 + 前 3 条样本)--inline 强制全量打印到 stdout(同样落盘)读数据建议:先看摘要判断是否足够;需要具体字段时优先用 jq或ConvertFrom-Json 从保存的 json 文件按需抽取,避免整份 JSON 进入上下文。
发生以下异常情况时,采用 references/onboarding.md 引导解决问题:
LINKFOX_AGENT_API_KEY,也未配置 LINKFOXAGENT_API_KEY。costToken value from the response.Applicable -- Image analysis tasks:
| User Says | Scenario |
|---|---|
| "What's in this image/picture/photo" | General image description |
| "Analyze this product image" | Product visual analysis |
| "Read the text in this image" | OCR / text extraction |
| "Describe the A+ page images" | E-commerce content review |
| "What does this screenshot show" | Screenshot interpretation |
| "Identify objects in this photo" | Object detection / listing |
Not applicable -- Needs beyond image recognition: ``
按动态规则计费:消耗算力 = 输入消耗的算力 + 输出结果消耗的算力。
重要:本技能的服务按倍数动态计算,可能一次性消耗大量算力,必须提醒用户,由用户决定是否继续。
Feedback:
Auto-detect and report feedback via the Feedback API when any of the following apply:
Call the feedback API as specified in references/api.md. Do not interrupt the user's flow.
For more high-quality, professional cross-border e-commerce skills, set LinkFox Skills.
© linkfox-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/linkfox-multimodal-recognize-image of linkfox-ai/linkfox-skills.
Open the folder on GitHubat commit 38fef04
Linkfox Multimodal Recognize Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Linkfox Multimodal Recognize Image this skilllinkfox-ai/linkfox-skills | 107 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3.3k | Automated safety check: Pass | MIT | |
| CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs | 13k | 7 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Yolo Master AgentTencent/YOLO-Master | 745 | — | ~755 | Automated safety check: Pass | AGPL-3.0 | |
| Video Understandjjyaoao/HelloAgents | 3.2k | 1 repos | ~6.2k | Automated safety check: Pass | MIT | |
| LLaVA Vision-Language ModelOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Tencent/YOLO-Master
A skill your agent uses when the user wants to run a YOLO-Master task (train/val/predict/track/export/benchmark) or use the Agent Skill dispatcher.
jjyaoao/HelloAgents
Implement specialized video understanding capabilities using the z-ai-web-dev-sdk.
Orchestra-Research/AI-Research-SKILLs
Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.
edwardsanchez/MotionEyes
Pixel-based motion and UI change analysis from frame sequences or screenshots using computer vision and visual comparison.
linkfox-ai/linkfox-skills
1688平台以图搜图,通过商品图片精准检索外观相似或同款的1688货源,返回标题、价格、起批量、月销量、复购率、交易评分等核心数据。当用户提到1688以图搜图、1688找货源、以图找同款、跨境找工厂、1688识图、图片找货源、找相似货源、image search 1688、find supplier by…
linkfox-ai/linkfox-skills
亚马逊ABA(品牌分析)搜索词数据的查询与分析,涵盖15个站点近3年的周维度数据。当用户提到ABA数据、亚马逊搜索词分析、关键词挖掘、搜索排名趋势、市场机会分析、季节性关键词、高点击低转化分析、蓝海词发现、竞品关键词分析、ABA data, search term report, keyword mining, search ranking trends, blue ocean…
linkfox-ai/linkfox-skills
通过亚马逊前台的 Alexa 购物助手发起自然语言问答,获取与问题相关的导购回答、推荐商品分组、ASIN 列表,以及可继续追问的问题。每次调用仅支持 1 条 prompt,如需追问须由 agent 总结上下文后拼接新问题发起新请求。可用 url 补充亚马逊页面上下文。当用户提到亚马逊 Alexa、Alexa 购物助手、亚马逊智能助手、AI…
linkfox-ai/linkfox-skills
亚马逊反向选品:基于历史商业洞察报告沉淀的指标数据池,按 30+ 项商业维度(市场规模与增长、价格区间与档位份额、竞争密度与头部集中度、人群画像如年龄/性别/收入、评论卖点与痛点等)反向筛选亚马逊赛道与关键词。当用户提到反向选品、指标筛选、细分市场反查、蓝海赛道挖掘、低竞争赛道、新人友好赛道、品牌分散市场、痛点切入、卖点反查、定价档位机会、人群画像选品、Amazon niche reverse…
linkfox-ai/linkfox-skills
通过ASIN获取亚马逊商品详细信息,包括标题、图片、五点描述、规格参数、A+页面、价格、评分评论、变体等;可在取得原始HTML时尝试提取Item Highlights(商品亮点)。当用户提到亚马逊商品详情、ASIN查询、商品页面数据、Listing分析、五点描述提取、Item…
linkfox-ai/linkfox-skills
按ASIN获取并分析亚马逊商品评论,支持15个站点(含美国站),按星级筛选评论。当用户提到亚马逊评论、美国站评论、商品评价、买家投诉、差评、好评、星级评分、评论分析、评论情感、产品改良建议、Vine评论、已验证购买评论、竞品评论研究、Amazon reviews, US reviews, Amazon.com reviews, product feedback, negative review…
Categories
基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual…. Linkfox Multimodal Recognize Image is an agent skill from linkfox-ai/linkfox-skills.
Linkfox Multimodal Recognize Image fits situations like: tasks that involve Computer vision.
Run `npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a claude-code`. Or copy the skill folder (skills/linkfox-multimodal-recognize-image in linkfox-ai/linkfox-skills) into .claude/skills/linkfox-multimodal-recognize-image in your project. Claude Code loads it when a task matches its description.
Run `npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a codex`. Or copy the skill folder (skills/linkfox-multimodal-recognize-image in linkfox-ai/linkfox-skills) into .agents/skills/linkfox-multimodal-recognize-image in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/linkfox-multimodal-recognize-image, .gemini/skills/linkfox-multimodal-recognize-image, .github/skills/linkfox-multimodal-recognize-image and .opencode/skills/linkfox-multimodal-recognize-image in your project.
Going by SKILL.md and its folder, Linkfox Multimodal Recognize Image needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named LINKFOX_AGENT_API_KEY and LINKFOXAGENT_API_KEY. Our summary lists: Python 3; A credential in LINKFOX_AGENT_API_KEY; A credential in LINKFOXAGENT_API_KEY.
SKILL.md names 1 domain. As links in the text: skill.linkfox.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Linkfox Multimodal Recognize Image is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Linkfox Multimodal Recognize Image: Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Yolo Master Agent (Tencent/YOLO-Master, 745 stars) and Video Understand (jjyaoao/HelloAgents, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
linkfox-ai (a GitHub user) maintains it in linkfox-ai/linkfox-skills, which has 107 GitHub stars. The repository holds 177 skills in this directory. The repository was last updated on September 14, 2026.
Source: linkfox-ai/linkfox-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.