Agent skill

PDF Image Text Extractor

by redfox-data in redfox-data/redfox-community

从图片或 PDF 文档中识别并提取文字内容,支持多种图片格式和 PDF 文件,自动判断是否包含文字并保留原始格式输出结构化结果;v2.1 采用零额外依赖方案:扫描版 PDF 自动渲染为图片交由 AI 视觉识别(无需 tesseract/rapidocr)、表格用 pymupdf 内置 findtables 结构化提取(无需…

No licenceAuto-check passedDocuments & Office

Install PDF Image Text Extractor

skills CLI
$ npx skills add redfox-data/redfox-community --skill pdf-image-text-extractor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install redfox-data/redfox-community pdf-image-text-extractor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/redfox-data/redfox-community.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/pdf-image-text-extractor .claude/skills/pdf-image-text-extractor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf-image-text-extractor
GitHub stars
425
Token cost
~2.1k tokens
SKILL.md length
529 words
Files
7 (incl. scripts)
Skills in repo
123
Repo updated
First seen
Licence
None found

At a glance

从图片或 PDF 文档中识别并提取文字内容,支持多种图片格式和 PDF 文件,自动判断是否包含文字并保留原始格式输出结构化结果;v2.1 采用零额外依赖方案:扫描版 PDF 自动渲染为图片交由 AI 视觉识别(无需 tesseract/rapidocr)、表格用 pymupdf 内置 findtables 结构化提取(无需…

  • Works in 2 steps: :鉴权检查(必须首先执行) → 5:版本更新提示(每次执行)
  • Tasks that involve PDF
  • SKILL.md covers 任务目标, 🔑 鉴权, 前置准备 and 操作步骤, plus 3 more sections
  • Runs Python scripts from its folder; calls python3 and pip; reaches redfox.hk; needs REDFOX_API_KEY

What it does

PDF Image Text Extractor is an agent skill from redfox-data/redfox-community. 从图片或 PDF 文档中识别并提取文字内容,支持多种图片格式和 PDF 文件,自动判断是否包含文字并保留原始格式输出结构化结果;v2.1 采用零额外依赖方案:扫描版 PDF 自动渲染为图片交由 AI 视觉识别(无需 tesseract/rapidocr)、表格用 pymupdf 内置 findtables 结构化提取(无需 pdfplumber)、批量处理目录(PDF+图片一次性提取);当用户需要从图片或 PDF 提取文字、进行 OCR 识别、处理含文字的文档、提取 PDF 表格、批量处理文件夹或转换为可编辑文本时使用。该skill能力来自RedFoxHub,官网:https://redfox.hk/skills。

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts (for example `README.en.md`, `README.md` and `scripts/batch_extractor.py`).

It sits in Documents & Office, covering PDF. It works with pypdf. The repository describes itself as: 红狐数据(RedFoxHub) 技能合集:面向 Agent 的可复用 SKILL 集合,覆盖灵感、选题、文案创作、数据复盘等场景,持续更新。

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/pdf-image-text-extractor”

Requirements

  • Python 3
  • A credential in REDFOX_API_KEY

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. :鉴权检查(必须首先执行)
  2. 5:版本更新提示(每次执行)

What it can do on your machine

Read from SKILL.md and the folder at commit 8263643. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • redfox.hk

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • REDFOX_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

PDF Image Text Extractor loads about 2.1k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 529 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 529 words (~2,071 tokens).

“本 Skill 完全免费,但需配置红狐 API Key 请求使用权限,Key 本身不扣积分。”

— opening of SKILL.md by redfox-data
name
pdf-image-text-extractor
slug
pdf-image-text-extractor
version
2.1.0
displayName
PDF和图片文字提取
dependency.python
pymupdf>=1.23.0, requests>=2.28.0

Read the full SKILL.md on GitHub

Files

SKILL.md and 6 other files (scripts) in skills/pdf-image-text-extractor of redfox-data/redfox-community.

  • SKILL.md
  • README.en.md
  • README.md
  • scripts/batch_extractor.py
  • scripts/changelog.py
  • scripts/pdf_text_extractor.py
  • scripts/record.py

Open the folder on GitHubat commit 8263643

Compare with similar skills

PDF Image Text Extractor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

PDF Image Text Extractor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
PDF Image Text Extractor this skillredfox-data/redfox-community425—~2.1kAutomated safety check: PassNone
PDFnuoyimanaituling/manus-x830—~985Automated safety check: PassNone
PDFeinverne/dotfiles12146 repos~1.8kAutomated safety check: PassProprietary
Reportlabjimmc414/Kosmos5941 repos~4.2kAutomated safety check: PassNone
PDFguyi-a/pi-ling106—~3.3kAutomated safety check: PassMIT
PDF ReadingWide-Moat/open-computer-use1261 repos~2.7kAutomated safety check: PassProprietary

Similar skills

  • PDF

    nuoyimanaituling/manus-x

    Process PDF files - extract text, read content, create PDFs, merge or split documents.

    830 GitHub stars~985 tokensUpdated 8 mo ago
    Documents & OfficeAuto-check passed
  • PDF

    einverne/dotfiles

    Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.

    121 GitHub starsUsed in 46 repos~1.8k tokens
    Documents & OfficeAuto-check passed
  • Reportlab

    jimmc414/Kosmos

    PDF generation toolkit. An agent skill from jimmc414/Kosmos.

    594 GitHub starsUsed in 1 repo~4.2k tokens
    Documents & OfficeAuto-check passed
  • PDF

    guyi-a/pi-ling

    PDF 相关的所有操作:从零生成(reportlab / pypdf)、格式转化(md/html → PDF)、修改(合并 / 拆分 / 旋转 / 加水印 / 提图片 / 元数据)、读内容(pdfplumber / extractdocumenttext)、OCR 扫描件、加密解密。触发场景:用户说"生成 PDF" / "做份 PDF 简历" / "合并这几份 PDF" / "给 PDF…

    106 GitHub stars~3.3k tokensUpdated 21 days ago
    Documents & OfficeAuto-check passed
  • PDF Reading

    Wide-Moat/open-computer-use

    A skill your agent uses when you need to read, inspect, or extract content from PDF files — especially when file content is NOT in your context and you need to read it from disk.

    126 GitHub starsUsed in 1 repo~2.7k tokens
    Documents & OfficeAuto-check passed
  • PDF

    LeastBit/Claude_skills_zh-CN

    全面的 PDF 操作工具包,用于提取文本和表格、创建新 PDF、合并/拆分文档以及处理表单。当 Claude 需要填写 PDF 表单或以编程方式大规模处理、生成或分析 PDF 文档时使用。

    588 GitHub stars~1.5k tokensUpdated 8 mo ago
    Documents & OfficeAuto-check passed

More from redfox-data/redfox-community

All 123 skills in this repo
  • Bilibili Comment

    redfox-data/redfox-community

    B站评论分析工具。输入B站视频链接或BV号即可获取一级评论数据,支持分页浏览、评论情感分析(积极/负面/需求/竞品),生成精美 HTML 报告。当用户需要查看B站视频评论、分析评论舆情、了解用户反馈时使用。触发词:B站评论、B站视频评论、评论查询、评论分析、评论舆情、看评论、bilibili评论。

    425 GitHub stars~763 tokensUpdated 8 days ago
    Auto-check passed
  • Douyin Top Account

    redfox-data/redfox-community

    抖音每日最具影响力账号榜单追踪分析工具;日榜每日17:30更新/回溯7天,周榜每周一17:30更新/回溯3周,月榜每月2号9点更新/回溯3月;当用户需要查询抖音账号排名、抖音日榜/周榜/月榜、抖音赛道TOP账号或下载榜单报告时使用

    425 GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed
  • Global AI News Brief

    redfox-data/redfox-community

    全球AI新闻简报 — 一个关键词同时搜索抖音、小红书、公众号、B站、快手、视频号、今日头条、TikTok、Instagram、X(Twitter)、YouTube 共 11 大平台,跨平台聚合后由 AI 生成智能摘要、热点聚类、舆情分析和深度解读报告,输出终端表格 + 交互式 HTML 报告。包含详细信息源列表,适用于各种AI…

    425 GitHub stars~2.1k tokensUpdated 8 days ago
    Auto-check passed
  • Investor Distiller

    redfox-data/redfox-community

    公众号投资博主蒸馏器。通过公众号文章数据,自动提取交易体系、市场判断、表达风格、内容深度、互动特征、热点图谱、 人设基因七维DNA,输出结构化风格画像 profile。

    425 GitHub stars~2.1k tokensUpdated 8 days ago
    Auto-check passed
  • Multi Content Feed

    redfox-data/redfox-community

    全网内容出海信息源 — 每日扫描全平台(公众号/抖音/视频号/小红书/快手/B站)内容出海爆款作品,按点赞量筛选Top50,智能聚类题材方向后生成包含平台标签、封面、互动数据与创作洞察的HTML日报。支持按平台、关键词、时间范围定向查询。⚠️数据每日15:00更新前一天数据,目标日期无数据时必须先告知用户并等待确认后才能调用接口,禁止自动获取。当用户需要内容出海日报、内容出海爆款、内容出海热点、…

    425 GitHub stars~1.7k tokensUpdated 8 days ago
    Auto-check passed
  • Playlet Wechat Feed

    redfox-data/redfox-community

    短剧-公众号信息源 — 每日扫描公众号短剧爆款文章,按阅读量筛选热门内容,智能聚类题材方向后生成包含封面图、互动数据与创作洞察的HTML日报。支持按题材(穿越/霸总/重生等)、公众号、时间范围定向查询。⚠️查询前脚本先做输入校验:关键词需命中短剧题材词库(topickeywords…

    425 GitHub stars~3.6k tokensUpdated 8 days ago
    Auto-check passed

Works with

Questions about PDF Image Text Extractor

What does PDF Image Text Extractor do?

从图片或 PDF 文档中识别并提取文字内容,支持多种图片格式和 PDF 文件,自动判断是否包含文字并保留原始格式输出结构化结果;v2.1 采用零额外依赖方案:扫描版 PDF 自动渲染为图片交由 AI 视觉识别(无需 tesseract/rapidocr)、表格用 pymupdf 内置 findtables 结构化提取(无需…. PDF Image Text Extractor is an agent skill from redfox-data/redfox-community.

When should I use PDF Image Text Extractor?

PDF Image Text Extractor fits situations like: tasks that involve PDF.

How do I install PDF Image Text Extractor in Claude Code?

Run `npx skills add redfox-data/redfox-community --skill pdf-image-text-extractor -a claude-code`. Or copy the skill folder (skills/pdf-image-text-extractor in redfox-data/redfox-community) into .claude/skills/pdf-image-text-extractor in your project. Claude Code loads it when a task matches its description.

How do I install PDF Image Text Extractor in Codex?

Run `npx skills add redfox-data/redfox-community --skill pdf-image-text-extractor -a codex`. Or copy the skill folder (skills/pdf-image-text-extractor in redfox-data/redfox-community) into .agents/skills/pdf-image-text-extractor in your project. Codex loads it when a task matches its description.

Can I use PDF Image Text Extractor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add redfox-data/redfox-community --skill pdf-image-text-extractor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf-image-text-extractor, .gemini/skills/pdf-image-text-extractor, .github/skills/pdf-image-text-extractor and .opencode/skills/pdf-image-text-extractor in your project.

What does PDF Image Text Extractor need to run?

Going by SKILL.md and its folder, PDF Image Text Extractor needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and pip) and credentials named REDFOX_API_KEY. Our summary lists: Python 3; A credential in REDFOX_API_KEY.

Does PDF Image Text Extractor access the network?

SKILL.md names 1 domain. In commands or code: redfox.hk; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is PDF Image Text Extractor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does PDF Image Text Extractor use?

No licence was found for PDF Image Text Extractor or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does PDF Image Text Extractor use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to PDF Image Text Extractor?

Skills that share tags, products or a category with PDF Image Text Extractor: PDF (nuoyimanaituling/manus-x, 830 stars), PDF (einverne/dotfiles, 121 stars), Reportlab (jimmc414/Kosmos, 594 stars) and PDF (guyi-a/pi-ling, 106 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains PDF Image Text Extractor?

redfox-data (a GitHub user) maintains it in redfox-data/redfox-community, which has 425 GitHub stars. The repository holds 123 skills in this directory. The repository was last updated on September 29, 2026.

Source: redfox-data/redfox-community on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.