Agent skill

Linkfox Multimodal Recognize Image

by linkfox-ai in linkfox-ai/linkfox-skills

基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual…

MITAuto-check passedAI & LLM Engineering

Install Linkfox Multimodal Recognize Image

skills CLI
$ npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install linkfox-ai/linkfox-skills linkfox-multimodal-recognize-image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/linkfox-ai/linkfox-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/linkfox-multimodal-recognize-image .claude/skills/linkfox-multimodal-recognize-image && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
linkfox-multimodal-recognize-image
GitHub stars
107
Token cost
~1.7k tokens
SKILL.md length
826 words
Files
6 (incl. scripts, references)
Skills in repo
177
Repo updated
First seen
Licence
MIT

At a glance

基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual…

  • Works in 3 steps: Be specific: Instead of "analyze this… → State the goal: If you need text… → Provide context when helpful: For…
  • Tasks that involve Computer vision
  • SKILL.md covers Core Concepts, Parameter Guide, Local Image Upload and Usage Examples, plus 5 more sections
  • Runs Python scripts from its folder; calls python; needs LINKFOX_AGENT_API_KEY and LINKFOXAGENT_API_KEY

What it does

Linkfox Multimodal Recognize Image is an agent skill from linkfox-ai/linkfox-skills. 基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual Q&A时触发此技能。当用户提到图片识别、图片分析、图片描述、识别图片内容、分析产品图、从图片中读取文字、描述图片、提取视觉内容或理解照片内容时触发。当用户提供图片URL并就其视觉内容提问时,即使未明确说"图片识别",也应触发此技能。

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/api.md`, `references/onboarding.md` and `scripts/multimodal_recognize_image.py`).

It sits in AI & LLM Engineering, covering Computer vision. The licence is MIT.

When your agent uses it

  • Tasks that involve Computer vision

Example prompts

  • “/linkfox-multimodal-recognize-image”

Requirements

  • Python 3
  • A credential in LINKFOX_AGENT_API_KEY
  • A credential in LINKFOXAGENT_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Be specific: Instead of "analyze this image", say "List all products visible on the shelf and estimate their category."
  2. State the goal: If you need text extraction, say "Extract all visible text from the image." If you need object identification, say…
  3. Provide context when helpful: For product images, mention "This is an e-commerce product listing image" so the model can tailor its…

What it can do on your machine

Read from SKILL.md and the folder at commit 38fef04. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • skill.linkfox.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LINKFOX_AGENT_API_KEY
    • LINKFOXAGENT_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Linkfox Multimodal Recognize Image loads about 1.7k tokens when it runs, and up to ~2.9k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 826 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from linkfox-ai/linkfox-skills at commit 38fef04, republished under its MIT licence (© linkfox-ai). 826 words, ~1,687 tokens.

Download SKILL.mdSave it as .claude/skills/linkfox-multimodal-recognize-image/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
linkfox-multimodal-recognize-image
description
基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual Q&A时触发此技能。当用户提到图片识别、图片分析、图片描述、识别图片内容、分析产品图、从图片中读取文字、描述图片、提取视觉内容或理解照片内容时触发。当用户提供图片URL并就其视觉内容提问时,即使未明确说"图片识别",也应触发此技能。

Image Recognition

This skill guides you on how to use the multimodal image recognition API to analyze images from URLs and extract meaningful information based on user intent.

Core Concepts

The Image Recognition tool accepts an image URL and an optional natural-language requirement describing what the user wants to know about the image. The backend uses a multimodal AI model to interpret the visual content and return a textual description or analysis.

Supported formats: JPG, JPEG, PNG, GIF, WebP, BMP.

How it works: You provide a publicly accessible image URL and a requirement (what you want to learn from the image). The service downloads the image, runs multimodal analysis, and returns a text-based result.

Parameter Guide

ParameterRequiredDescription
imageUrlYesA publicly accessible URL pointing to the image. Must be JPG, JPEG, PNG, GIF, WebP, or BMP. Maximum 1000 characters.
requirementNoA natural-language description of what to identify or analyze in the image. Defaults to "Describe the content of this image" when omitted. Maximum 1000 characters.
Tips for Writing the requirement Parameter
  1. Be specific: Instead of "analyze this image", say "List all products visible on the shelf and estimate their category."
  2. State the goal: If you need text extraction, say "Extract all visible text from the image." If you need object identification, say "Identify the main objects and their colors."
  3. Provide context when helpful: For product images, mention "This is an e-commerce product listing image" so the model can tailor its analysis.

Local Image Upload

This tool requires a publicly accessible image URL. If the user provides a local image file path (e.g., C:\Users\...\photo.png, /home/.../image.jpg), you must upload it first to obtain a public URL.

Run the upload script:

bash
python scripts/upload_image.py /path/to/local/image.png

The script will return a public URL (valid for 24 hours) that can be used as the image URL parameter.

Usage Examples

1. General Image Description

  • User says: "What is in this picture?"
  • Set imageUrl to the provided URL, leave requirement as default.

2. Product Image Analysis

  • User says: "Analyze this Amazon product image and list the key selling points shown."
  • Set requirement to: "This is an Amazon product listing image. Identify the product, key features, and selling points visible in the image."

3. Text Extraction from an Image

  • User says: "Read the text in this screenshot."
  • Set requirement to: "Extract all visible text from this image, preserving layout where possible."

4. A+ Page Image Review

  • User says: "Describe what this A+ content image communicates."
  • Set requirement to: "This is an Amazon A+ product description image. Describe the visual content, key messaging, and branding elements."

5. Comparison / Detail Inspection

  • User says: "What differences can you spot between the product and its packaging?"
  • Set requirement to: "Identify and describe any differences between the product and its packaging shown in the image."

调用方式

  • API 端点:POST /multimodal/recognizeImage(完整参数/响应/错误码见 references/api.md)
  • Python 脚本:python scripts/multimodal_recognize_image.py '<JSON 参数>' [--inline]
  • 成本约束:本工具会消耗算力;同一会话同一参数组合默认只调用一次,脚本带 24h 本地缓存。失败/空结果不得自动换关键词、翻页或改邮编连续试探;需要继续检索时先向用户说明会产生额外消耗。

输出策略(脚本默认行为):

  • 始终将完整响应写入 <cwd>/linkfox/<YYYY-MM-DD>/<session>/data/linkfox-multimodal-recognize-image-<timestamp>.json(<cwd> 为脚本执行时的工作目录,在 Claude Code 里即当前项目目录;<session> 取自环境变量 SESSION_ID,按用户任务自动聚合;禁止写入 /tmp,当前目录不可写则报错)
  • 响应体 ≤ 8 KB:落盘后把完整 JSON 打印到 stdout
  • 响应体 > 8 KB:落盘后 stdout 只输出摘要(顶层字段、常见计数如 total/costToken、最大列表字段的长度 + 前 3 条样本)
  • 加 --inline 强制全量打印到 stdout(同样落盘)

读数据建议:先看摘要判断是否足够;需要具体字段时优先用 jq或ConvertFrom-Json 从保存的 json 文件按需抽取,避免整份 JSON 进入上下文。

Show full SKILL.md (312 more words)Show less

解决认证和算力问题

发生以下异常情况时,采用 references/onboarding.md 引导解决问题:

异常情况
  • 未配置API Key:环境变量未配置 LINKFOX_AGENT_API_KEY,也未配置 LINKFOXAGENT_API_KEY。
  • 响应401或402状态码
  • 响应提示算力或余额不足:消息含"算力余额不足/计费不足/余额不足/quota exceeded/insufficient balance/套餐到期/需充值/请充值",或类似含义的内容。

Display Rules

  1. Show the analysis result clearly: Present the returned text analysis in a readable format. Use bullet points or paragraphs as appropriate for the content.
  2. No fabrication: Only relay information that the API actually returned. Do not add visual details that were not in the response.
  3. Format support: If the image URL is invalid or the format is unsupported, explain the limitation and list the supported formats (JPG, JPEG, PNG, GIF, WebP, BMP).
  4. Error handling: When the API returns an error status, explain the issue based on the response and suggest corrective actions (e.g., check that the URL is publicly accessible, verify the image format).
  5. Token usage: If the user asks about cost, you may mention the costToken value from the response.

User Expression & Scenario Quick Reference

Applicable -- Image analysis tasks:

User SaysScenario
"What's in this image/picture/photo"General image description
"Analyze this product image"Product visual analysis
"Read the text in this image"OCR / text extraction
"Describe the A+ page images"E-commerce content review
"What does this screenshot show"Screenshot interpretation
"Identify objects in this photo"Object detection / listing

Not applicable -- Needs beyond image recognition: ``

  • Generating or editing images
  • Video analysis
  • Analyzing images from local file paths (only URLs are supported)
  • Image search or reverse image lookup

算力消耗规则

按动态规则计费:消耗算力 = 输入消耗的算力 + 输出结果消耗的算力。

重要:本技能的服务按倍数动态计算,可能一次性消耗大量算力,必须提醒用户,由用户决定是否继续。

Feedback:

Auto-detect and report feedback via the Feedback API when any of the following apply:

  1. The functionality or purpose described in this skill does not match actual behavior
  2. The skill's results do not match the user's intent
  3. The user expresses dissatisfaction or praise about this skill
  4. Anything you believe could be improved

Call the feedback API as specified in references/api.md. Do not interrupt the user's flow.


For more high-quality, professional cross-border e-commerce skills, set LinkFox Skills.

© linkfox-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/linkfox-multimodal-recognize-image of linkfox-ai/linkfox-skills.

  • SKILL.md
  • references/api.md
  • references/onboarding.md
  • scripts/multimodal_recognize_image.py
  • scripts/onboarding.py
  • scripts/upload_image.py

Open the folder on GitHubat commit 38fef04

Compare with similar skills

Linkfox Multimodal Recognize Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Linkfox Multimodal Recognize Image compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Linkfox Multimodal Recognize Image this skilllinkfox-ai/linkfox-skills107—~1.7kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Yolo Master AgentTencent/YOLO-Master745—~755Automated safety check: PassAGPL-3.0
Video Understandjjyaoao/HelloAgents3.2k1 repos~6.2kAutomated safety check: PassMIT
LLaVA Vision-Language ModelOrchestra-Research/AI-Research-SKILLs13k6 repos~2kAutomated safety check: PassMIT

Similar skills

  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Yolo Master Agent

    Tencent/YOLO-Master

    A skill your agent uses when the user wants to run a YOLO-Master task (train/val/predict/track/export/benchmark) or use the Agent Skill dispatcher.

    745 GitHub stars~755 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Video Understand

    jjyaoao/HelloAgents

    Implement specialized video understanding capabilities using the z-ai-web-dev-sdk.

    3.2k GitHub starsUsed in 1 repo~6.2k tokens
    AI & LLM EngineeringAuto-check passed
  • LLaVA Vision-Language Model

    Orchestra-Research/AI-Research-SKILLs

    Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.

    13k GitHub starsUsed in 6 repos~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Motioneyes Visual Analysis

    edwardsanchez/MotionEyes

    Pixel-based motion and UI change analysis from frame sequences or screenshots using computer vision and visual comparison.

    229 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from linkfox-ai/linkfox-skills

All 177 skills in this repo
  • Linkfox 1688 Search By Image

    linkfox-ai/linkfox-skills

    1688平台以图搜图,通过商品图片精准检索外观相似或同款的1688货源,返回标题、价格、起批量、月销量、复购率、交易评分等核心数据。当用户提到1688以图搜图、1688找货源、以图找同款、跨境找工厂、1688识图、图片找货源、找相似货源、image search 1688、find supplier by…

    107 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Linkfox Aba Intelligent Query

    linkfox-ai/linkfox-skills

    亚马逊ABA(品牌分析)搜索词数据的查询与分析,涵盖15个站点近3年的周维度数据。当用户提到ABA数据、亚马逊搜索词分析、关键词挖掘、搜索排名趋势、市场机会分析、季节性关键词、高点击低转化分析、蓝海词发现、竞品关键词分析、ABA data, search term report, keyword mining, search ranking trends, blue ocean…

    107 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Linkfox Amazon Alexa Search

    linkfox-ai/linkfox-skills

    通过亚马逊前台的 Alexa 购物助手发起自然语言问答,获取与问题相关的导购回答、推荐商品分组、ASIN 列表,以及可继续追问的问题。每次调用仅支持 1 条 prompt,如需追问须由 agent 总结上下文后拼接新问题发起新请求。可用 url 补充亚马逊页面上下文。当用户提到亚马逊 Alexa、Alexa 购物助手、亚马逊智能助手、AI…

    107 GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • 亚马逊反向选品:基于历史商业洞察报告沉淀的指标数据池,按 30+ 项商业维度(市场规模与增长、价格区间与档位份额、竞争密度与头部集中度、人群画像如年龄/性别/收入、评论卖点与痛点等)反向筛选亚马逊赛道与关键词。当用户提到反向选品、指标筛选、细分市场反查、蓝海赛道挖掘、低竞争赛道、新人友好赛道、品牌分散市场、痛点切入、卖点反查、定价档位机会、人群画像选品、Amazon niche reverse…

    107 GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • Linkfox Amazon Product Detail

    linkfox-ai/linkfox-skills

    通过ASIN获取亚马逊商品详细信息,包括标题、图片、五点描述、规格参数、A+页面、价格、评分评论、变体等;可在取得原始HTML时尝试提取Item Highlights(商品亮点)。当用户提到亚马逊商品详情、ASIN查询、商品页面数据、Listing分析、五点描述提取、Item…

    107 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Linkfox Amazon Reviews List

    linkfox-ai/linkfox-skills

    按ASIN获取并分析亚马逊商品评论,支持15个站点(含美国站),按星级筛选评论。当用户提到亚马逊评论、美国站评论、商品评价、买家投诉、差评、好评、星级评分、评论分析、评论情感、产品改良建议、Vine评论、已验证购买评论、竞品评论研究、Amazon reviews, US reviews, Amazon.com reviews, product feedback, negative review…

    107 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed

Questions about Linkfox Multimodal Recognize Image

What does Linkfox Multimodal Recognize Image do?

基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image analysis, image description, image content understanding, OCR text recognition, visual…. Linkfox Multimodal Recognize Image is an agent skill from linkfox-ai/linkfox-skills.

When should I use Linkfox Multimodal Recognize Image?

Linkfox Multimodal Recognize Image fits situations like: tasks that involve Computer vision.

How do I install Linkfox Multimodal Recognize Image in Claude Code?

Run `npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a claude-code`. Or copy the skill folder (skills/linkfox-multimodal-recognize-image in linkfox-ai/linkfox-skills) into .claude/skills/linkfox-multimodal-recognize-image in your project. Claude Code loads it when a task matches its description.

How do I install Linkfox Multimodal Recognize Image in Codex?

Run `npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a codex`. Or copy the skill folder (skills/linkfox-multimodal-recognize-image in linkfox-ai/linkfox-skills) into .agents/skills/linkfox-multimodal-recognize-image in your project. Codex loads it when a task matches its description.

Can I use Linkfox Multimodal Recognize Image in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add linkfox-ai/linkfox-skills --skill linkfox-multimodal-recognize-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/linkfox-multimodal-recognize-image, .gemini/skills/linkfox-multimodal-recognize-image, .github/skills/linkfox-multimodal-recognize-image and .opencode/skills/linkfox-multimodal-recognize-image in your project.

What does Linkfox Multimodal Recognize Image need to run?

Going by SKILL.md and its folder, Linkfox Multimodal Recognize Image needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named LINKFOX_AGENT_API_KEY and LINKFOXAGENT_API_KEY. Our summary lists: Python 3; A credential in LINKFOX_AGENT_API_KEY; A credential in LINKFOXAGENT_API_KEY.

Does Linkfox Multimodal Recognize Image access the network?

SKILL.md names 1 domain. As links in the text: skill.linkfox.com. This is read from the text; nothing was executed.

Is Linkfox Multimodal Recognize Image safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Linkfox Multimodal Recognize Image use?

Linkfox Multimodal Recognize Image is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Linkfox Multimodal Recognize Image use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Linkfox Multimodal Recognize Image?

Skills that share tags, products or a category with Linkfox Multimodal Recognize Image: Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Yolo Master Agent (Tencent/YOLO-Master, 745 stars) and Video Understand (jjyaoao/HelloAgents, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Linkfox Multimodal Recognize Image?

linkfox-ai (a GitHub user) maintains it in linkfox-ai/linkfox-skills, which has 107 GitHub stars. The repository holds 177 skills in this directory. The repository was last updated on September 14, 2026.

Source: linkfox-ai/linkfox-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.