Agent skill

Iflytek PDF Image OCR

by iflytek in iflytek/iFly-Skills

ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition.

Apache-2.0Auto-check passedDocuments & Office

Install Iflytek PDF Image OCR

skills CLI
$ npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install iflytek/iFly-Skills iflytek-pdf-image-ocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/iflytek/iFly-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/iflytek-pdf-image-ocr .claude/skills/iflytek-pdf-image-ocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
iflytek-pdf-image-ocr
GitHub stars
209
Token cost
~2.3k tokens
SKILL.md length
719 words
Files
5 (incl. scripts)
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition.

  • Works in 6 steps: Generate RFC1123 format date: EEE, dd… → Create signature origin: host:… → Calculate signature:… → …
  • User asks to OCR images
  • SKILL.md covers Quick Start, Setup, Features and API Parameters, plus 6 more sections
  • Runs Python scripts from its folder; calls python3; reaches bjcdn.openstorage.cn; needs API_SECRET and API_KEY

What it does

Iflytek PDF Image OCR is an agent skill from iflytek/iFly-Skills. ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition. Use when user asks to OCR images, extract text from images/PDFs, convert PDF to Word/Markdown, or perform any OCR tasks on images or PDFs. Supports multi-language text extraction, document layout understanding, and various output formats.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `README.md`, `_meta.json` and `scripts/image_ocr.py`).

It sits in Documents & Office, covering PDF. The repository describes itself as: Official collection of iFLYTEK skills for speech, OCR, translation, proofreading, and multimodal AI capabilities. The licence is Apache-2.0.

When your agent uses it

  • User asks to OCR images
  • Extract text from images/PDFs
  • Convert PDF to Word/Markdown
  • Perform any OCR tasks on images

Example prompts

  • “/iflytek-pdf-image-ocr”

Requirements

  • Python 3
  • A credential in IFLY_API_KEY
  • A credential in IFLY_API_SECRET

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Generate RFC1123 format date: EEE, dd MMM yyyy HH:mm:ss GMT
  2. Create signature origin: host: {host}\ndate: {date}\nPOST {path} HTTP/1.1
  3. Calculate signature: HMAC-SHA256(signature_origin, apiSecret)
  4. Build authorization: hmac username="{apiKey}", algorithm="hmac-sha256", headers="host date request-line", signature="{signature}"
  5. Encode authorization in base64
  6. Send as query parameters: ?authorization={auth}&host={host}&date={date}

What it can do on your machine

Read from SKILL.md and the folder at commit 58dc114. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • bjcdn.openstorage.cn

    Also links to:

    • console.xfyun.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • API_SECRET
    • API_KEY
    • IFLY_API_KEY
    • IFLY_API_SECRET

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Iflytek PDF Image OCR loads about 2.3k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 719 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from iflytek/iFly-Skills at commit 58dc114, republished under its Apache-2.0 licence (© iflytek). 719 words, ~2,299 tokens.

Download SKILL.mdSave it as .claude/skills/iflytek-pdf-image-ocr/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
iflytek-pdf-image-ocr
description
ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition. Use when user asks to OCR images, extract text from images/PDFs, convert PDF to Word/Markdown, or perform any OCR tasks on images or PDFs. Supports multi-language text extraction, document layout understanding, and various output formats.
metadata.homepage
https://www.xfyun.cn/services/ocr
metadata.openclaw
{"emoji":"📄","dimensions":["文档OCR","PDF识别"],"user_instructions":["OCR这个PDF","提取图片中的文字","帮我把这个文档转成文字"],"requires":{"bins":["python3"],"env":["IFLY_APP_ID","IFL…

ifly-pdf-image-ocr

AI-powered OCR service for images and PDF documents using iFlytek's advanced recognition APIs.

Quick Start

Image OCR (LLM OCR)
bash
# OCR an image and extract text
python3 scripts/image_ocr.py /path/to/image.jpg

# Save result to file
python3 scripts/image_ocr.py /path/to/image.jpg -o output.txt

# Specify output format
python3 scripts/image_ocr.py /path/to/image.jpg --format json
python3 scripts/image_ocr.py /path/to/image.jpg --format markdown
PDF OCR
bash
# Convert PDF to Word (default)
python3 scripts/pdf_ocr.py document.pdf

# Convert PDF to Markdown
python3 scripts/pdf_ocr.py document.pdf --format markdown

# Convert PDF to JSON
python3 scripts/pdf_ocr.py document.pdf --format json

# From public URL
python3 scripts/pdf_ocr.py --pdf-url "https://example.com/doc.pdf" --format word

Setup

API Credentials

Get credentials from iFlytek Open Platform:

For Image OCR:

  • APP_ID: Application ID
  • API_KEY: API key for authentication
  • API_SECRET: API secret for signing requests

For PDF OCR:

  • APP_ID: Application ID
  • API_SECRET: Application secret (for signature generation)
Environment Variables
bash
# Required for both Image OCR and PDF OCR
export IFLY_APP_ID="your_app_id"

# Required for Image OCR
export IFLY_API_KEY="your_api_key"

# Required for PDF OCR
export IFLY_API_SECRET="your_api_secret"

Features

Image OCR (LLM OCR)
  • AI-powered: Advanced LLM-based OCR for high accuracy
  • Multi-format output: JSON, Markdown, or both
  • Layout understanding: Preserves document structure
  • Multi-language: Supports text extraction in multiple languages
  • Image preprocessing: Automatic rotation correction, noise removal
PDF OCR
  • AI-powered OCR: Advanced AI model for accurate text extraction
  • Multiple output formats:
    • Word (.docx) - Editable Word document
    • Markdown - Plain text with formatting
    • JSON - Structured data
  • Large PDF support: Up to 100 pages per document
  • Page-by-page results: Access individual page results
  • Download URLs: Direct links to processed files

API Parameters

Image OCR Parameters
ParameterTypeRequiredDescription
image_pathstringYesPath to image file
--formatstringNoOutput format: json, markdown, json,markdown (default: json,markdown)
--outputstringNoSave result to file
PDF OCR Parameters
ParameterTypeRequiredDescription
pdf_pathstringYes*Path to PDF file
--pdf-urlstringNo*Public URL of PDF file
--formatstringNoOutput format: word, markdown, json (default: word)
--no-pollflagNoReturn task ID without polling
--poll-intervalintNoPolling interval in seconds (min 5, default: 5)
--max-waitintNoMaximum wait time in seconds (default: 300)

*Either pdf_path or --pdf-url must be provided

Authentication

Image OCR (HMAC-SHA256)

Uses HMAC-SHA256 signature authentication:

  1. Generate RFC1123 format date: EEE, dd MMM yyyy HH:mm:ss GMT
  2. Create signature origin: host: {host}\\ndate: {date}\\nPOST {path} HTTP/1.1
  3. Calculate signature: HMAC-SHA256(signature_origin, apiSecret)
  4. Build authorization: hmac username="{apiKey}", algorithm="hmac-sha256", headers="host date request-line", signature="{signature}"
  5. Encode authorization in base64
  6. Send as query parameters: ?authorization={auth}&host={host}&date={date}
PDF OCR (MD5 + HMAC-SHA1)

Uses MD5 + HMAC-SHA1 signature authentication:

  1. Generate timestamp (Unix epoch in seconds)
  2. Calculate auth = MD5(appId + timestamp)
  3. Calculate signature = Base64(HMAC-SHA1(auth, apiSecret))
  4. Send headers:
    • appId: Application ID
    • timestamp: Timestamp in seconds
    • signature: Generated signature

Important: Timestamp must be within 5 minutes of server time.

Response Format

Image OCR Response
json
{
  "header": {
    "code": 0,
    "message": "success"
  },
  "payload": {
    "result": {
      "text": "Base64-encoded OCR text..."
    }
  }
}
PDF OCR Start Response
json
{
  "flag": true,
  "code": 0,
  "desc": "成功",
  "data": {
    "taskNo": "25082744936879",
    "status": "CREATE",
    "tip": "任务创建成功"
  }
}
PDF OCR Status Response
json
{
  "flag": true,
  "code": 0,
  "desc": "成功",
  "data": {
    "taskNo": "25082759289333",
    "exportFormat": "word",
    "status": "FINISH",
    "downUrl": "http://bjcdn.openstorage.cn/...",
    "tip": "已完成",
    "pageList": [...]
  }
}

Task Status (PDF OCR)

StatusDescription
CREATETask created successfully
WAITINGWaiting in queue
DOINGProcessing
FINISHCompleted
FAILEDFailed
ANY_FAILEDPartially completed (some pages failed)
STOPPaused

Error Codes

(。・ω・。) 嗨遇到错误码了吗?来看看怎么解决吧 ✧⁺⸜(●˙▾˙●)⸝⁺✧

Show full SKILL.md (322 more words)Show less
Platform Common Error Codes
CodeDescriptionHintSolution
10009input invalid data(◎_◎;) 哎呀~数据格式不太对呢检查输入数据是否符合要求
10010service license not enough(╯°□°)╯︵ ┻━┻ 授权数量不足或已过期!提交工单联系客服
10019service read buffer timeout(。-`ω´-) session超时啦~检查是否数据发送完毕但未关闭连接
10043Syscall AudioCodingDecode error(◎_◎;) 音频解码失败惹...检查aue参数,如果为speex,请确保音频是speex音频并分段压缩且与帧大小一致
10114session timeout(。-`ω´-) 会话时间超时啦~检查是否发送数据时间超过了60s
10139invalid param(◎_◎;) 参数好像不太对呢检查参数是否正确
10160parse request json error(◎_◎;) 请求数据格式有误~检查请求数据是否是合法的json
10161parse base64 string error(◎_◎;) Base64解码失败啦检查发送的数据是否使用base64编码了
10163param validate error(◎_◎;) 参数校验没通过呢具体原因见详细的描述
10200read data timeout(。-`ω´-) 读取数据超时了~检查是否累计10s未发送数据并且未关闭连接
10222context deadline exceeded(╯°□°)╯︵ ┻━┻ 出错啦!1.检查上传数据是否超过接口上限;2.SSL证书无效请提交工单
10223RemoteLB: can't find valued addr(◎_◎;) 找不到服务节点呢提交工单联系技术人员
10313invalid appid(◎_◎;) appid和apikey不匹配哦检查appid是否合法
10317invalid version(◎_◎;) 版本号有问题呢请到控制台提交工单联系技术人员
10700not authority(╯°□°)╯︵ ┻━┻ 权限不足!按照报错原因对照开发文档检查,如仍无法解决,请提供sid及错误信息提交工单
11200auth no license(╯°□°)╯︵ ┻━┻ 功能未授权!检查appid是否正确,确认是否添加了相关服务,检查调用量是否超限或授权是否到期
11201auth no enough license(╯°□°)╯︵ ┻━┻ 每日交互次数超限啦!提交应用审核提额或联系商务购买企业级接口
11503server error: atmos return error(。-`ω´-) 服务器返回了错误数据...提交工单
11502server error: too many datas(。-`ω´-) 服务器配置有问题呢提交工单
100001~100010WrapperInitErr(◎_◎;) 引擎调用出错啦!请根据message中的errno查看引擎错误码说明
Additional Resources

Original API Error Codes
CodeDescriptionSolution
10000System errorCheck auth info, request method, parameters
10001Signature authentication failedCheck credentials
10002Business processing errorCheck error message
10003Quota/insufficient balanceCheck account balance

Limitations

Image OCR
  • Format: Common image formats (JPG, PNG, etc.)
  • Size: Reasonable file sizes for web upload
  • Rate limiting: Follow API rate limits
PDF OCR
  • Max pages: 100 pages per PDF
  • Protected PDFs: Not supported (password/encrypted)
  • Rate limiting: Status query limited to once per 5 seconds
  • Time limit: Timestamp must be within ±5 minutes of server time

Tips

Image OCR
  1. High-quality images: Use clear, high-resolution images for best results
  2. Multiple formats: Use json,markdown to get both structured and formatted output
  3. Save results: Use -o flag to save OCR results to file
PDF OCR
  1. Math formulas: Use markdown format for PDFs with mathematical formulas
  2. Large PDFs: Split into sections if > 100 pages
  3. Polling interval: Minimum 5 seconds between status queries
  4. Network URLs: Ensure PDF URLs are publicly accessible
  5. Download URLs: Download files promptly as URLs may expire

© iflytek, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in skills/iflytek-pdf-image-ocr of iflytek/iFly-Skills.

  • SKILL.md
  • README.md
  • _meta.json
  • scripts/image_ocr.py
  • scripts/pdf_ocr.py

Open the folder on GitHubat commit 58dc114

Compare with similar skills

Iflytek PDF Image OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Iflytek PDF Image OCR compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Iflytek PDF Image OCR this skilliflytek/iFly-Skills209—~2.3kAutomated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice8.9k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3141 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    8.9k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    314 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from iflytek/iFly-Skills

All 11 skills in this repo
  • Animated Sketch Diagram

    iflytek/iFly-Skills

    生成"黑墨手绘涂鸦"风格的动画架构图/流程图:米色纸面、针管笔墨线、极淡水洗色块、简笔涂鸦图标、序号章、连线上的流动圆点动画、图标微动效。产出单文件自包含动画 HTML(SVG+CSS),可一键导出无缝循环 GIF。当用户想画架构图、流程图、信息图、技术示意图、对比图、pipeline/workflow 可视化,或提到"手绘风""涂鸦风""动图""animated diagram""GIF…

    209 GitHub stars~999 tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when user asks to review contracts, detect contract risks, or perform compliance checks.

    209 GitHub stars~773 tokensUpdated yesterday
    Auto-check passed
  • Iflytek Hyper Tts

    iflytek/iFly-Skills

    A skill your agent uses when user asks to synthesize speech, convert text to audio, or read text aloud.

    209 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Iflytek Image Understanding

    iflytek/iFly-Skills

    A skill your agent uses when user asks to analyze an image, describe image contents, or answer questions about a picture.

    209 GitHub stars~949 tokensUpdated yesterday
    Auto-check passed
  • Iflytek OCR Invoice

    iflytek/iFly-Skills

    A skill your agent uses when user asks to recognize invoices, extract receipt data, or OCR bills and tickets.

    209 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Iflytek Speed Transcription

    iflytek/iFly-Skills

    Ultra-fast speech transcription using iFLYTEK Speed Transcription API.

    209 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Iflytek PDF Image OCR

What does Iflytek PDF Image OCR do?

ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition. Iflytek PDF Image OCR is an agent skill from iflytek/iFly-Skills. ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition.

When should I use Iflytek PDF Image OCR?

Iflytek PDF Image OCR fits situations like: user asks to OCR images; extract text from images/PDFs; convert PDF to Word/Markdown; perform any OCR tasks on images.

How do I install Iflytek PDF Image OCR in Claude Code?

Run `npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a claude-code`. Or copy the skill folder (skills/iflytek-pdf-image-ocr in iflytek/iFly-Skills) into .claude/skills/iflytek-pdf-image-ocr in your project. Claude Code loads it when a task matches its description.

How do I install Iflytek PDF Image OCR in Codex?

Run `npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a codex`. Or copy the skill folder (skills/iflytek-pdf-image-ocr in iflytek/iFly-Skills) into .agents/skills/iflytek-pdf-image-ocr in your project. Codex loads it when a task matches its description.

Can I use Iflytek PDF Image OCR in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/iflytek-pdf-image-ocr, .gemini/skills/iflytek-pdf-image-ocr, .github/skills/iflytek-pdf-image-ocr and .opencode/skills/iflytek-pdf-image-ocr in your project.

What does Iflytek PDF Image OCR need to run?

Going by SKILL.md and its folder, Iflytek PDF Image OCR needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named API_SECRET, API_KEY, IFLY_API_KEY and IFLY_API_SECRET. Our summary lists: Python 3; A credential in IFLY_API_KEY; A credential in IFLY_API_SECRET.

Does Iflytek PDF Image OCR access the network?

SKILL.md names 2 domains. In commands or code: bjcdn.openstorage.cn; the agent is likely to contact it when it follows the instructions. As links in the text: console.xfyun.cn. This is read from the text; nothing was executed.

Is Iflytek PDF Image OCR safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Iflytek PDF Image OCR use?

Iflytek PDF Image OCR is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Iflytek PDF Image OCR use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Iflytek PDF Image OCR?

Skills that share tags, products or a category with Iflytek PDF Image OCR: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.9k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Iflytek PDF Image OCR?

iflytek (a GitHub organization) maintains it in iflytek/iFly-Skills, which has 209 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 7, 2026.

Source: iflytek/iFly-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.