Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition.
$ npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install iflytek/iFly-Skills iflytek-pdf-image-ocr --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/iflytek/iFly-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/iflytek-pdf-image-ocr .claude/skills/iflytek-pdf-image-ocr && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "iflytek-pdf-image-ocr" agent skill from https://github.com/iflytek/iFly-Skills/tree/main/skills/iflytek-pdf-image-ocr into .claude/skills/iflytek-pdf-image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "iflytek-pdf-image-ocr", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/iflytek/iFly-Skills/tree/main/skills/iflytek-pdf-image-ocrType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install iflytek/iFly-Skills iflytek-pdf-image-ocr --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iflytek/iFly-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/iflytek-pdf-image-ocr .agents/skills/iflytek-pdf-image-ocr && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "iflytek-pdf-image-ocr" agent skill from https://github.com/iflytek/iFly-Skills/tree/main/skills/iflytek-pdf-image-ocr into .agents/skills/iflytek-pdf-image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "iflytek-pdf-image-ocr", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install iflytek/iFly-Skills iflytek-pdf-image-ocr --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iflytek/iFly-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/iflytek-pdf-image-ocr .cursor/skills/iflytek-pdf-image-ocr && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "iflytek-pdf-image-ocr" agent skill from https://github.com/iflytek/iFly-Skills/tree/main/skills/iflytek-pdf-image-ocr into .cursor/skills/iflytek-pdf-image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "iflytek-pdf-image-ocr", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/iflytek/iFly-Skills.git --path skills/iflytek-pdf-image-ocr--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install iflytek/iFly-Skills iflytek-pdf-image-ocr --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iflytek/iFly-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/iflytek-pdf-image-ocr .gemini/skills/iflytek-pdf-image-ocr && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "iflytek-pdf-image-ocr" agent skill from https://github.com/iflytek/iFly-Skills/tree/main/skills/iflytek-pdf-image-ocr into .gemini/skills/iflytek-pdf-image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "iflytek-pdf-image-ocr", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install iflytek/iFly-Skills iflytek-pdf-image-ocrInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/iflytek/iFly-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/iflytek-pdf-image-ocr .github/skills/iflytek-pdf-image-ocr && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "iflytek-pdf-image-ocr" agent skill from https://github.com/iflytek/iFly-Skills/tree/main/skills/iflytek-pdf-image-ocr into .github/skills/iflytek-pdf-image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "iflytek-pdf-image-ocr", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install iflytek/iFly-Skills iflytek-pdf-image-ocr --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/iflytek/iFly-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/iflytek-pdf-image-ocr .opencode/skills/iflytek-pdf-image-ocr && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "iflytek-pdf-image-ocr" agent skill from https://github.com/iflytek/iFly-Skills/tree/main/skills/iflytek-pdf-image-ocr into .opencode/skills/iflytek-pdf-image-ocr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "iflytek-pdf-image-ocr", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
iflytek-pdf-image-ocrifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition.
Iflytek PDF Image OCR is an agent skill from iflytek/iFly-Skills. ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition. Use when user asks to OCR images, extract text from images/PDFs, convert PDF to Word/Markdown, or perform any OCR tasks on images or PDFs. Supports multi-language text extraction, document layout understanding, and various output formats.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `README.md`, `_meta.json` and `scripts/image_ocr.py`).
It sits in Documents & Office, covering PDF. The repository describes itself as: Official collection of iFLYTEK skills for speech, OCR, translation, proofreading, and multimodal AI capabilities. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 58dc114. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
bjcdn.openstorage.cnAlso links to:
console.xfyun.cnFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
API_SECRETAPI_KEYIFLY_API_KEYIFLY_API_SECRETFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Iflytek PDF Image OCR loads about 2.3k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 719 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from iflytek/iFly-Skills at commit 58dc114, republished under its Apache-2.0 licence (© iflytek). 719 words, ~2,299 tokens.
.claude/skills/iflytek-pdf-image-ocr/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.AI-powered OCR service for images and PDF documents using iFlytek's advanced recognition APIs.
# OCR an image and extract text
python3 scripts/image_ocr.py /path/to/image.jpg
# Save result to file
python3 scripts/image_ocr.py /path/to/image.jpg -o output.txt
# Specify output format
python3 scripts/image_ocr.py /path/to/image.jpg --format json
python3 scripts/image_ocr.py /path/to/image.jpg --format markdown# Convert PDF to Word (default)
python3 scripts/pdf_ocr.py document.pdf
# Convert PDF to Markdown
python3 scripts/pdf_ocr.py document.pdf --format markdown
# Convert PDF to JSON
python3 scripts/pdf_ocr.py document.pdf --format json
# From public URL
python3 scripts/pdf_ocr.py --pdf-url "https://example.com/doc.pdf" --format wordGet credentials from iFlytek Open Platform:
For Image OCR:
For PDF OCR:
# Required for both Image OCR and PDF OCR
export IFLY_APP_ID="your_app_id"
# Required for Image OCR
export IFLY_API_KEY="your_api_key"
# Required for PDF OCR
export IFLY_API_SECRET="your_api_secret"| Parameter | Type | Required | Description |
|---|---|---|---|
image_path | string | Yes | Path to image file |
--format | string | No | Output format: json, markdown, json,markdown (default: json,markdown) |
--output | string | No | Save result to file |
| Parameter | Type | Required | Description |
|---|---|---|---|
pdf_path | string | Yes* | Path to PDF file |
--pdf-url | string | No* | Public URL of PDF file |
--format | string | No | Output format: word, markdown, json (default: word) |
--no-poll | flag | No | Return task ID without polling |
--poll-interval | int | No | Polling interval in seconds (min 5, default: 5) |
--max-wait | int | No | Maximum wait time in seconds (default: 300) |
*Either pdf_path or --pdf-url must be provided
Uses HMAC-SHA256 signature authentication:
EEE, dd MMM yyyy HH:mm:ss GMThost: {host}\\ndate: {date}\\nPOST {path} HTTP/1.1HMAC-SHA256(signature_origin, apiSecret)hmac username="{apiKey}", algorithm="hmac-sha256", headers="host date request-line", signature="{signature}"?authorization={auth}&host={host}&date={date}Uses MD5 + HMAC-SHA1 signature authentication:
auth = MD5(appId + timestamp)signature = Base64(HMAC-SHA1(auth, apiSecret))appId: Application IDtimestamp: Timestamp in secondssignature: Generated signatureImportant: Timestamp must be within 5 minutes of server time.
{
"header": {
"code": 0,
"message": "success"
},
"payload": {
"result": {
"text": "Base64-encoded OCR text..."
}
}
}{
"flag": true,
"code": 0,
"desc": "成功",
"data": {
"taskNo": "25082744936879",
"status": "CREATE",
"tip": "任务创建成功"
}
}{
"flag": true,
"code": 0,
"desc": "成功",
"data": {
"taskNo": "25082759289333",
"exportFormat": "word",
"status": "FINISH",
"downUrl": "http://bjcdn.openstorage.cn/...",
"tip": "已完成",
"pageList": [...]
}
}| Status | Description |
|---|---|
CREATE | Task created successfully |
WAITING | Waiting in queue |
DOING | Processing |
FINISH | Completed |
FAILED | Failed |
ANY_FAILED | Partially completed (some pages failed) |
STOP | Paused |
(。・ω・。) 嗨
遇到错误码了吗?来看看怎么解决吧✧⁺⸜(●˙▾˙●)⸝⁺✧
| Code | Description | Hint | Solution |
|---|---|---|---|
| 10009 | input invalid data | (◎_◎;) 哎呀~数据格式不太对呢 | 检查输入数据是否符合要求 |
| 10010 | service license not enough | (╯°□°)╯︵ ┻━┻ 授权数量不足或已过期! | 提交工单联系客服 |
| 10019 | service read buffer timeout | (。-`ω´-) session超时啦~ | 检查是否数据发送完毕但未关闭连接 |
| 10043 | Syscall AudioCodingDecode error | (◎_◎;) 音频解码失败惹... | 检查aue参数,如果为speex,请确保音频是speex音频并分段压缩且与帧大小一致 |
| 10114 | session timeout | (。-`ω´-) 会话时间超时啦~ | 检查是否发送数据时间超过了60s |
| 10139 | invalid param | (◎_◎;) 参数好像不太对呢 | 检查参数是否正确 |
| 10160 | parse request json error | (◎_◎;) 请求数据格式有误~ | 检查请求数据是否是合法的json |
| 10161 | parse base64 string error | (◎_◎;) Base64解码失败啦 | 检查发送的数据是否使用base64编码了 |
| 10163 | param validate error | (◎_◎;) 参数校验没通过呢 | 具体原因见详细的描述 |
| 10200 | read data timeout | (。-`ω´-) 读取数据超时了~ | 检查是否累计10s未发送数据并且未关闭连接 |
| 10222 | context deadline exceeded | (╯°□°)╯︵ ┻━┻ 出错啦! | 1.检查上传数据是否超过接口上限;2.SSL证书无效请提交工单 |
| 10223 | RemoteLB: can't find valued addr | (◎_◎;) 找不到服务节点呢 | 提交工单联系技术人员 |
| 10313 | invalid appid | (◎_◎;) appid和apikey不匹配哦 | 检查appid是否合法 |
| 10317 | invalid version | (◎_◎;) 版本号有问题呢 | 请到控制台提交工单联系技术人员 |
| 10700 | not authority | (╯°□°)╯︵ ┻━┻ 权限不足! | 按照报错原因对照开发文档检查,如仍无法解决,请提供sid及错误信息提交工单 |
| 11200 | auth no license | (╯°□°)╯︵ ┻━┻ 功能未授权! | 检查appid是否正确,确认是否添加了相关服务,检查调用量是否超限或授权是否到期 |
| 11201 | auth no enough license | (╯°□°)╯︵ ┻━┻ 每日交互次数超限啦! | 提交应用审核提额或联系商务购买企业级接口 |
| 11503 | server error: atmos return error | (。-`ω´-) 服务器返回了错误数据... | 提交工单 |
| 11502 | server error: too many datas | (。-`ω´-) 服务器配置有问题呢 | 提交工单 |
| 100001~100010 | WrapperInitErr | (◎_◎;) 引擎调用出错啦! | 请根据message中的errno查看引擎错误码说明 |
| Code | Description | Solution |
|---|---|---|
| 10000 | System error | Check auth info, request method, parameters |
| 10001 | Signature authentication failed | Check credentials |
| 10002 | Business processing error | Check error message |
| 10003 | Quota/insufficient balance | Check account balance |
json,markdown to get both structured and formatted output-o flag to save OCR results to file© iflytek, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts) in skills/iflytek-pdf-image-ocr of iflytek/iFly-Skills.
Open the folder on GitHubat commit 58dc114
Iflytek PDF Image OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Iflytek PDF Image OCR this skilliflytek/iFly-Skills | 209 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Gzh Designisjiamu/gzh-design-skill | 3.9k | 1 repos | ~2.2k | Automated safety check: Pass | AGPL-3.0 | |
| GenOffice Document CLIgenspark-ai/genoffice | 8.9k | — | ~19k | Automated safety check: Pass | Apache-2.0 | |
| Harness Book Best Practicewquguru/harness-books | 3.2k | — | ~4.1k | Automated safety check: Pass | None | |
| Bookforge Korean Ebook PDF Makergongnyang/bookforge | 314 | 1 repos | ~1.7k | Automated safety check: Pass | MIT |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
isjiamu/gzh-design-skill
微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…
genspark-ai/genoffice
Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.
wquguru/harness-books
Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.
gongnyang/bookforge
Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.
aws-samples/amazon-bedrock-agents-healthcare-lifesciences
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.
iflytek/iFly-Skills
生成"黑墨手绘涂鸦"风格的动画架构图/流程图:米色纸面、针管笔墨线、极淡水洗色块、简笔涂鸦图标、序号章、连线上的流动圆点动画、图标微动效。产出单文件自包含动画 HTML(SVG+CSS),可一键导出无缝循环 GIF。当用户想画架构图、流程图、信息图、技术示意图、对比图、pipeline/workflow 可视化,或提到"手绘风""涂鸦风""动图""animated diagram""GIF…
iflytek/iFly-Skills
A skill your agent uses when user asks to review contracts, detect contract risks, or perform compliance checks.
iflytek/iFly-Skills
A skill your agent uses when user asks to synthesize speech, convert text to audio, or read text aloud.
iflytek/iFly-Skills
A skill your agent uses when user asks to analyze an image, describe image contents, or answer questions about a picture.
iflytek/iFly-Skills
A skill your agent uses when user asks to recognize invoices, extract receipt data, or OCR bills and tickets.
iflytek/iFly-Skills
Ultra-fast speech transcription using iFLYTEK Speed Transcription API.
Categories
ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition. Iflytek PDF Image OCR is an agent skill from iflytek/iFly-Skills. ifly-pdf-image-ocr skill supporting both image OCR (AI-powered LLM OCR) and PDF document recognition.
Iflytek PDF Image OCR fits situations like: user asks to OCR images; extract text from images/PDFs; convert PDF to Word/Markdown; perform any OCR tasks on images.
Run `npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a claude-code`. Or copy the skill folder (skills/iflytek-pdf-image-ocr in iflytek/iFly-Skills) into .claude/skills/iflytek-pdf-image-ocr in your project. Claude Code loads it when a task matches its description.
Run `npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a codex`. Or copy the skill folder (skills/iflytek-pdf-image-ocr in iflytek/iFly-Skills) into .agents/skills/iflytek-pdf-image-ocr in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add iflytek/iFly-Skills --skill iflytek-pdf-image-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/iflytek-pdf-image-ocr, .gemini/skills/iflytek-pdf-image-ocr, .github/skills/iflytek-pdf-image-ocr and .opencode/skills/iflytek-pdf-image-ocr in your project.
Going by SKILL.md and its folder, Iflytek PDF Image OCR needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named API_SECRET, API_KEY, IFLY_API_KEY and IFLY_API_SECRET. Our summary lists: Python 3; A credential in IFLY_API_KEY; A credential in IFLY_API_SECRET.
SKILL.md names 2 domains. In commands or code: bjcdn.openstorage.cn; the agent is likely to contact it when it follows the instructions. As links in the text: console.xfyun.cn. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Iflytek PDF Image OCR is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Iflytek PDF Image OCR: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.9k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
iflytek (a GitHub organization) maintains it in iflytek/iFly-Skills, which has 209 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 7, 2026.
Source: iflytek/iFly-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.