Agent skill

Tencentcloud OCR

by infometa in infometa/workbuddyskills

腾讯云通用文字识别(高精度版)(GeneralAccurateOCR) 技能包。当用户发送/粘贴图片、提供图片URL、或要求识别图片中的文字时,应自动调用此技能。支持图像整体文字的检测和识别,支持中文、英文、中英文、数字和特殊字符号的识别,并返回文字框位置和文字内容。适用于文字较多、版式复杂、对识别准召率要求较高的场景,如网络图片、街景店招牌、法律卷宗、多语种简历等场景。支持图片Base64和U…

No licenceAuto-check passedDocuments & Office

Install Tencentcloud OCR

skills CLI
$ npx skills add infometa/workbuddyskills --skill tencentcloud-ocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install infometa/workbuddyskills tencentcloud-ocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/infometa/workbuddyskills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tencentcloud-ocr .claude/skills/tencentcloud-ocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tencentcloud-ocr
GitHub stars
344
Token cost
~1.1k tokens
SKILL.md length
158 words
Files
3 (incl. scripts, references)
Skills in repo
212
Repo updated
First seen
Licence
None found

At a glance

腾讯云通用文字识别(高精度版)(GeneralAccurateOCR) 技能包。当用户发送/粘贴图片、提供图片URL、或要求识别图片中的文字时,应自动调用此技能。支持图像整体文字的检测和识别,支持中文、英文、中英文、数字和特殊字符号的识别,并返回文字框位置和文字内容。适用于文字较多、版式复杂、对识别准召率要求较高的场景,如网络图片、街景店招牌、法律卷宗、多语种简历等场景。支持图片Base64和U…

  • Works in 3 steps: 获取 API 密钥 → 获取/购买 OCR 服务 → 设置环境变量
  • Tasks that involve PDF
  • SKILL.md covers 用途, 📚 可用资源, 使用时机 and 环境要求, plus 2 more sections
  • Runs Python scripts from its folder; calls python and pip; reaches xxx.com and xxx.cos.xxx; needs TENCENTCLOUD_SECRET_KEY

What it does

Tencentcloud OCR is an agent skill from infometa/workbuddyskills. 腾讯云通用文字识别(高精度版)(GeneralAccurateOCR) 技能包。当用户发送/粘贴图片、提供图片URL、或要求识别图片中的文字时,应自动调用此技能。支持图像整体文字的检测和识别,支持中文、英文、中英文、数字和特殊字符号的识别,并返回文字框位置和文字内容。适用于文字较多、版式复杂、对识别准召率要求较高的场景,如网络图片、街景店招牌、法律卷宗、多语种简历等场景。支持图片Base64和URL两种输入方式,同时支持PDF文件识别和单字信息返回。对于简历识别场景,提供专门的结构化解析指引(详见 references/resume-parsing.md)。

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/resume-parsing.md` and `scripts/main.py`).

It sits in Documents & Office, covering PDF. It works with Tencent Cloud and Python. The repository describes itself as: WorkBuddy skills / connectors / experts archive for offline study.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “/tencentcloud-ocr”

Requirements

  • Python 3
  • A credential in TENCENTCLOUD_SECRET_KEY

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. 获取 API 密钥
  2. 获取/购买 OCR 服务
  3. 设置环境变量

What it can do on your machine

Read from SKILL.md and the folder at commit 75a5ad9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • xxx.com
    • xxx.cos.xxx

    Also links to:

    • cloud.tencent.com
    • console.cloud.tencent.com
    • buy.cloud.tencent.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • TENCENTCLOUD_SECRET_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tencentcloud OCR loads about 1.1k tokens when it runs, and up to ~3.6k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 158 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 158 words (~1,077 tokens).

name
tencentcloud-ocr
version
1.0.4
display_name
腾讯云通用文字识别(高精度版)
display_name_en
Tencent Cloud General OCR (High Accuracy)
description_zh
调用腾讯云OCR通用文字识别(高精度版)接口,对图片中的文字进行高精度识别。支持中英文、数字和特殊字符识别,适用于文字较多、版式复杂的场景。支持图片URL、Base64、PDF输入,可返回单字信息和文字框位置。内含简历结构化解析指引。
description_en
High-accuracy OCR powered by Tencent Cloud GeneralAccurateOCR API. Recognizes Chinese, English, numbers and special characters from images. Supports image…
visibility
public

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files (scripts, references) in skills/tencentcloud-ocr of infometa/workbuddyskills.

  • SKILL.md
  • references/resume-parsing.md
  • scripts/main.py

Open the folder on GitHubat commit 75a5ad9

Compare with similar skills

Tencentcloud OCR next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tencentcloud OCR compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tencentcloud OCR this skillinfometa/workbuddyskills344—~1.1kAutomated safety check: PassNone
PDF ToolkitXiaomiMiMo/MiMo-Code14k—~1.7kAutomated safety check: PassApache-2.0
PDF Generation, Forms and Extractionpipeshub-ai/pipeshub-ai3.8k—~2.9kAutomated safety check: PassApache-2.0
MineruNebutra/MinerU-Skill122—~1.4kAutomated safety check: PassMIT
Markdown Exporterbowenliang123/markdown-exporter2711 repos~5.3kAutomated safety check: PassApache-2.0
Kimi PDFthvroyal/kimi-skills238—~1.9kAutomated safety check: PassNone

Similar skills

  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 5 days ago
    Documents & OfficeAuto-check passed
  • Picks the right library for generating a new PDF, filling an existing PDF form, or extracting text and tables, defaulting to Node where possible.

    3.8k GitHub stars~2.9k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~1.4k tokensUpdated 14 days ago
    Documents & OfficeAuto-check passed
  • Markdown Exporter

    bowenliang123/markdown-exporter

    Convert Markdown text to DOCX, PPTX, XLSX, PDF, PNG, SVG, HTML, IPYNB, MD, CSV, JSON, JSONL, XML files, and extract code blocks in Markdown to Python, Bash,JS and etc files.

    271 GitHub starsUsed in 1 repo~5.3k tokens
    Documents & OfficeAuto-check passed
  • Kimi PDF

    thvroyal/kimi-skills

    Professional PDF solution. An agent skill from thvroyal/kimi-skills.

    238 GitHub stars~1.9k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • Markdown Converter

    Team-Commonly/commonly

    Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.

    1.4k GitHub stars~557 tokensUpdated yesterday
    Documents & OfficeAuto-check passed

More from infometa/workbuddyskills

All 212 skills in this repo
  • Campus Event Playbook

    infometa/workbuddyskills

    This skill should be used when planning, coordinating, running, or reviewing student-led campus events, including club recruitment, freshman mixers, welcome activities, small talks, competitions…

    344 GitHub stars~861 tokensUpdated yesterday
    Auto-check passed
  • Teachany

    infometa/workbuddyskills

    K-12 interactive courseware creation. An agent skill from infometa/workbuddyskills.

    344 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check: notes
  • Weight Management HTML

    infometa/workbuddyskills

    A skill your agent uses when the adult weight-management MCP needs to be installed/set up in WorkBuddy (connector) or its result must be rendered as a standalone Chinese HTML plan.

    344 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Agent Browser

    infometa/workbuddyskills

    A skill your agent uses when the user needs browser automation, including opening web pages, taking screenshots, extracting page content, clicking elements, filling forms, or testing web flows.

    344 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • AI Research Radar

    infometa/workbuddyskills

    定时任务:每日研究简报。漏掉重要论文和行业报告?每天帮你盯着,关键信息一条不漏 This skill should be used when the user asks about 定时任务:每日研究简报.

    344 GitHub stars~438 tokensUpdated yesterday
    Auto-check passed
  • Check Deck

    infometa/workbuddyskills

    Investment banking presentation quality checker. An agent skill from infometa/workbuddyskills.

    344 GitHub stars~687 tokensUpdated yesterday
    Auto-check passed

Questions about Tencentcloud OCR

What does Tencentcloud OCR do?

腾讯云通用文字识别(高精度版)(GeneralAccurateOCR) 技能包。当用户发送/粘贴图片、提供图片URL、或要求识别图片中的文字时,应自动调用此技能。支持图像整体文字的检测和识别,支持中文、英文、中英文、数字和特殊字符号的识别,并返回文字框位置和文字内容。适用于文字较多、版式复杂、对识别准召率要求较高的场景,如网络图片、街景店招牌、法律卷宗、多语种简历等场景。支持图片Base64和U…. Tencentcloud OCR is an agent skill from infometa/workbuddyskills.

When should I use Tencentcloud OCR?

Tencentcloud OCR fits situations like: tasks that involve PDF.

How do I install Tencentcloud OCR in Claude Code?

Run `npx skills add infometa/workbuddyskills --skill tencentcloud-ocr -a claude-code`. Or copy the skill folder (skills/tencentcloud-ocr in infometa/workbuddyskills) into .claude/skills/tencentcloud-ocr in your project. Claude Code loads it when a task matches its description.

How do I install Tencentcloud OCR in Codex?

Run `npx skills add infometa/workbuddyskills --skill tencentcloud-ocr -a codex`. Or copy the skill folder (skills/tencentcloud-ocr in infometa/workbuddyskills) into .agents/skills/tencentcloud-ocr in your project. Codex loads it when a task matches its description.

Can I use Tencentcloud OCR in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add infometa/workbuddyskills --skill tencentcloud-ocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tencentcloud-ocr, .gemini/skills/tencentcloud-ocr, .github/skills/tencentcloud-ocr and .opencode/skills/tencentcloud-ocr in your project.

What does Tencentcloud OCR need to run?

Going by SKILL.md and its folder, Tencentcloud OCR needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named TENCENTCLOUD_SECRET_KEY. Our summary lists: Python 3; A credential in TENCENTCLOUD_SECRET_KEY.

Does Tencentcloud OCR access the network?

SKILL.md names 5 domains. In commands or code: xxx.com and xxx.cos.xxx; the agent is likely to contact these when it follows the instructions. As links in the text: cloud.tencent.com, console.cloud.tencent.com and buy.cloud.tencent.com. This is read from the text; nothing was executed.

Is Tencentcloud OCR safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Tencentcloud OCR use?

No licence was found for Tencentcloud OCR or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Tencentcloud OCR use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Tencentcloud OCR?

Skills that share tags, products or a category with Tencentcloud OCR: PDF Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), PDF Generation, Forms and Extraction (pipeshub-ai/pipeshub-ai, 3.8k stars), Mineru (Nebutra/MinerU-Skill, 122 stars) and Markdown Exporter (bowenliang123/markdown-exporter, 271 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tencentcloud OCR?

infometa (a GitHub user) maintains it in infometa/workbuddyskills, which has 344 GitHub stars. The repository holds 212 skills in this directory. The repository was last updated on October 7, 2026.

Source: infometa/workbuddyskills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.