Agent skill

Iflytek Image Understanding

by iflytek in iflytek/iFly-Skills

A skill your agent uses when user asks to analyze an image, describe image contents, or answer questions about a picture.

Apache-2.0Auto-check passedBackend & APIs

Install Iflytek Image Understanding

skills CLI
$ npx skills add iflytek/iFly-Skills --skill iflytek-image-understanding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install iflytek/iFly-Skills iflytek-image-understanding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/iflytek/iFly-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/iflytek-image-understanding .claude/skills/iflytek-image-understanding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
iflytek-image-understanding
GitHub stars
209
Token cost
~949 tokens
SKILL.md length
200 words
Files
4 (incl. scripts)
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when user asks to analyze an image, describe image contents, or answer questions about a picture.

  • Works in 2 steps: Create an app at 讯飞控制台 with 图片理解 service… → Set environment variables
  • User asks to analyze an image
  • SKILL.md covers Setup, Usage, Notes and 错误码说明 😢, plus 2 more sections
  • Runs Python scripts from its folder; calls python3; needs IFLY_API_KEY and IFLY_API_SECRET

What it does

Iflytek Image Understanding is an agent skill from iflytek/iFly-Skills. Use when user asks to analyze an image, describe image contents, or answer questions about a picture. iFlytek Image Understanding (图片理解) — analyze and answer questions about images using Spark Vision model. WebSocket API, pure Python stdlib, no pip dependencies.

Its SKILL.md is about 950 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `README.md`, `_meta.json` and `scripts/image_understanding.py`).

It sits in Backend & APIs, covering Realtime and WebSockets and Computer vision. It works with Python. The repository describes itself as: Official collection of iFLYTEK skills for speech, OCR, translation, proofreading, and multimodal AI capabilities. The licence is Apache-2.0.

When your agent uses it

  • User asks to analyze an image
  • Describe image contents
  • Answer questions about a picture

Example prompts

  • “/iflytek-image-understanding”

Requirements

  • Python 3
  • A credential in IFLY_API_KEY
  • A credential in IFLY_API_SECRET

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Create an app at 讯飞控制台 with 图片理解 service enabled
  2. Set environment variables

What it can do on your machine

Read from SKILL.md and the folder at commit 58dc114. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • console.xfyun.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • IFLY_API_KEY
    • IFLY_API_SECRET

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Iflytek Image Understanding loads about 949 tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 200 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~949

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from iflytek/iFly-Skills at commit 58dc114, republished under its Apache-2.0 licence (© iflytek). 200 words, ~949 tokens.

Download SKILL.mdSave it as .claude/skills/iflytek-image-understanding/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
iflytek-image-understanding
description
Use when user asks to analyze an image, describe image contents, or answer questions about a picture. iFlytek Image Understanding (图片理解) — analyze and answer questions about images using Spark Vision model. WebSocket API, pure Python stdlib, no pip dependencies.
metadata.homepage
https://www.xfyun.cn/services/image_understanding
metadata.openclaw
{"emoji":"🖼️","dimensions":["图片理解","视觉问答"],"user_instructions":["分析这张图片","描述一下这张图","这张图里有什么"],"requires":{"bins":["python3"],"env":["IFLY_APP_ID","IFLY_API_KE…

ifly-image-understanding

Analyze images and answer questions about their content using iFlytek's Spark Vision model (图片理解).

Setup

  1. Create an app at 讯飞控制台 with 图片理解 service enabled
  2. Set environment variables:
    bash
    export IFLY_APP_ID="your_app_id"
    export IFLY_API_KEY="your_api_key"
    export IFLY_API_SECRET="your_api_secret"

Usage

Describe an image
bash
python3 scripts/image_understanding.py photo.jpg
Ask a question about an image
bash
python3 scripts/image_understanding.py photo.jpg -q "图片里有什么动物?"
Use basic model (lower token cost)
bash
python3 scripts/image_understanding.py photo.jpg --domain general
Options
FlagShortDescription
imageImage file path (.jpg, .jpeg, .png)
--question-qQuestion about the image (default: describe)
--domain-dimagev3 (advanced, default) or general (basic, fixed 273 tokens/image)
--temperature-tSampling temperature (0,1], default 0.5
--max-tokensMax response tokens 1-8192, default 2048
--rawOutput raw WebSocket JSON frames
Examples
bash
# OCR a receipt
python3 scripts/image_understanding.py receipt.png -q "总金额是多少?"

# Identify objects
python3 scripts/image_understanding.py scene.jpg -q "图片中有哪些物体?"

# Low-cost basic model
python3 scripts/image_understanding.py chart.png -q "图表的趋势是什么?" -d general

Notes

  • Image formats: .jpg, .jpeg, .png
  • Max image size: 4MB
  • Max tokens: 8192 (input + output combined)
  • Auth: HMAC-SHA256 signed WebSocket URL
  • Endpoint: wss://spark-api.cn-huabei-1.xf-yun.com/v2.1/image
  • Pure stdlib: No pip dependencies — uses built-in socket + ssl for WebSocket
  • Model versions: imagev3 (advanced, dynamic token cost) vs general (basic, fixed 273 tokens/image)

错误码说明 😢

遇到错误先别慌~看看下面找到对应的解决方法吧!✨

错误码错误信息解决办法
0🎉 成功恭喜你!请求正常完成啦~
10003用户的消息格式有错误检查一下你的请求格式是否正确哦~确保发送的是合法的JSON格式呢!
10004用户数据的schema错误看起来数据结构有点问题~请检查一下字段名称和类型是否正确呀!
10005用户参数值有错误参数值可能不太对呢~仔细核对一下每个参数的有效范围吧!
10006用户并发错误:同一用户不能多处同时连接检测到重复连接啦!请确保只有一个客户端在连接同一个用户ID哦~
10013用户问题涉及敏感信息,审核不通过哎呀,你的问题可能包含了一些不太合适的内容~换个问题试试看吧!
10022模型生产的图片涉及敏感信息,审核不通过生成的图片没有通过审核呢...很抱歉,换张图片再试一下吧!
10029图片任何一边的长度超过12800图片尺寸太大啦!请确保图片宽高都不超过12800像素哦~
10041图片分辨率不符合要求图片尺寸不合适的呢~要求是:50×50 < 图片总像素值 < 6000×6000 哦!
10907Token数量超过上限内容太丰富啦!对话历史+问题的字数太多,需要精简一下输入哦~

💡 小贴士:如果还有其他问题,可以查看官方文档或者联系技术支持哦!


常见问题 🤔

图片理解的主要功能是什么呀?🐱

答:用户输入一张图片和问题,从而识别出图片中的对象、场景等信息,然后回答你的问题~是不是很方便呢!✨

图片理解支持什么应用平台呢?📱

答:目前支持 Web API 应用平台哦!直接在代码里调用就可以啦~

图片理解的文本大小限制多少呀?📝

答:有效内容不能超过 8192 Token 呢~如果超过了就要精简一下输入啦!


更多资源 📚

有更多问题随时来问我哦~祝你使用愉快!🌸

© iflytek, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/iflytek-image-understanding of iflytek/iFly-Skills.

  • SKILL.md
  • README.md
  • _meta.json
  • scripts/image_understanding.py

Open the folder on GitHubat commit 58dc114

Compare with similar skills

Iflytek Image Understanding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Iflytek Image Understanding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Iflytek Image Understanding this skilliflytek/iFly-Skills209—~949Automated safety check: PassApache-2.0
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0
Azure AI Voicelive Pymicrosoft/skills3.1k6 repos~2.9kAutomated safety check: PassMIT
Deepstream DevNVIDIA/skills3.5k—~3.3kAutomated safety check: PassApache-2.0
Gemini Live API DevJetBrains/skills364—~2.6kAutomated safety check: PassNone
Web3 PolymarketPolymarket/agent-skills1912 repos~2kAutomated safety check: PassNone

Similar skills

  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Azure AI Voicelive Py

    microsoft/skills

    Official

    Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive).

    3.1k GitHub starsUsed in 6 repos~2.9k tokens
    Backend & APIsAuto-check passed
  • Deepstream Dev

    NVIDIA/skills

    Official

    NVIDIA DeepStream SDK development with Python pyservicemaker API.

    3.5k GitHub stars~3.3k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Gemini Live API Dev

    JetBrains/skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API.

    364 GitHub stars~2.6k tokensUpdated 3 mo ago
    Backend & APIsAuto-check passed
  • Web3 Polymarket

    Polymarket/agent-skills

    Polymarket integration for prediction market trading on Polygon.

    191 GitHub starsUsed in 2 repos~2k tokens
    Backend & APIsAuto-check passed
  • Spider King

    aoyunyang/spider-king-skill

    Pure-web protocol reverse skill: turn hostile browser clients into browser-free Python collectors.

    507 GitHub stars~7.3k tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed

More from iflytek/iFly-Skills

All 11 skills in this repo
  • Animated Sketch Diagram

    iflytek/iFly-Skills

    生成"黑墨手绘涂鸦"风格的动画架构图/流程图:米色纸面、针管笔墨线、极淡水洗色块、简笔涂鸦图标、序号章、连线上的流动圆点动画、图标微动效。产出单文件自包含动画 HTML(SVG+CSS),可一键导出无缝循环 GIF。当用户想画架构图、流程图、信息图、技术示意图、对比图、pipeline/workflow 可视化,或提到"手绘风""涂鸦风""动图""animated diagram""GIF…

    209 GitHub stars~999 tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when user asks to review contracts, detect contract risks, or perform compliance checks.

    209 GitHub stars~773 tokensUpdated yesterday
    Auto-check passed
  • Iflytek Hyper Tts

    iflytek/iFly-Skills

    A skill your agent uses when user asks to synthesize speech, convert text to audio, or read text aloud.

    209 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Iflytek OCR Invoice

    iflytek/iFly-Skills

    A skill your agent uses when user asks to recognize invoices, extract receipt data, or OCR bills and tickets.

    209 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Iflytek PDF Image OCR

    iflytek/iFly-Skills

    ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition.

    209 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Iflytek Speed Transcription

    iflytek/iFly-Skills

    Ultra-fast speech transcription using iFLYTEK Speed Transcription API.

    209 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Iflytek Image Understanding

What does Iflytek Image Understanding do?

A skill your agent uses when user asks to analyze an image, describe image contents, or answer questions about a picture. Iflytek Image Understanding is an agent skill from iflytek/iFly-Skills. Use when user asks to analyze an image, describe image contents, or answer questions about a picture.

When should I use Iflytek Image Understanding?

Iflytek Image Understanding fits situations like: user asks to analyze an image; describe image contents; answer questions about a picture.

How do I install Iflytek Image Understanding in Claude Code?

Run `npx skills add iflytek/iFly-Skills --skill iflytek-image-understanding -a claude-code`. Or copy the skill folder (skills/iflytek-image-understanding in iflytek/iFly-Skills) into .claude/skills/iflytek-image-understanding in your project. Claude Code loads it when a task matches its description.

How do I install Iflytek Image Understanding in Codex?

Run `npx skills add iflytek/iFly-Skills --skill iflytek-image-understanding -a codex`. Or copy the skill folder (skills/iflytek-image-understanding in iflytek/iFly-Skills) into .agents/skills/iflytek-image-understanding in your project. Codex loads it when a task matches its description.

Can I use Iflytek Image Understanding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add iflytek/iFly-Skills --skill iflytek-image-understanding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/iflytek-image-understanding, .gemini/skills/iflytek-image-understanding, .github/skills/iflytek-image-understanding and .opencode/skills/iflytek-image-understanding in your project.

What does Iflytek Image Understanding need to run?

Going by SKILL.md and its folder, Iflytek Image Understanding needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named IFLY_API_KEY and IFLY_API_SECRET. Our summary lists: Python 3; A credential in IFLY_API_KEY; A credential in IFLY_API_SECRET.

Does Iflytek Image Understanding access the network?

SKILL.md names 1 domain. As links in the text: console.xfyun.cn. This is read from the text; nothing was executed.

Is Iflytek Image Understanding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Iflytek Image Understanding use?

Iflytek Image Understanding is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Iflytek Image Understanding use?

About 949 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Iflytek Image Understanding?

Skills that share tags, products or a category with Iflytek Image Understanding: Gemini Live API Dev (google-gemini/gemini-skills, 4.3k stars), Azure AI Voicelive Py (microsoft/skills, 3.1k stars), Deepstream Dev (NVIDIA/skills, 3.5k stars) and Gemini Live API Dev (JetBrains/skills, 364 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Iflytek Image Understanding?

iflytek (a GitHub organization) maintains it in iflytek/iFly-Skills, which has 209 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 7, 2026.

Source: iflytek/iFly-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.