Agent skill

Glmocr

by zai-org in zai-org/GLM-skills

Extract text from images using GLM-OCR API. An agent skill from zai-org/GLM-skills.

Apache-2.0Auto-check passedDocuments & Office

Install Glmocr

skills CLI
$ npx skills add zai-org/GLM-skills --skill glmocr -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zai-org/GLM-skills glmocr --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zai-org/GLM-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/glmocr .claude/skills/glmocr && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
glmocr
GitHub stars
476
Token cost
~1.1k tokens
SKILL.md length
332 words
Files
5 (incl. scripts, references)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract text from images using GLM-OCR API. An agent skill from zai-org/GLM-skills.

  • Works in 5 steps: ONLY use GLM-OCR API - Execute the… → NEVER parse documents directly - Do NOT… → NEVER offer alternatives - Do NOT… → …
  • The user wants to extract text from images
  • SKILL.md covers When to Use, Key Features, Resource Links and Prerequisites, plus 7 more sections
  • Runs Python scripts from its folder; calls python; reaches bigmodel.cn; needs ZHIPU_API_KEY

What it does

Glmocr is an agent skill from zai-org/GLM-skills. Extract text from images using GLM-OCR API. Supports images and PDFs with high accuracy OCR, table recognition, formula extraction, and handwriting recognition. Use this skill whenever the user wants to extract text from images, perform OCR on pictures, scan documents, convert images to text, or process any image files to get their textual content.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/output_schema.md`, `scripts/config_setup.py` and `scripts/glm_ocr_cli.py`).

It sits in Documents & Office, covering PDF. It works with Zhipu GLM. The repository describes itself as: Official skills for the GLM family of models. The licence is Apache-2.0.

When your agent uses it

  • The user wants to extract text from images
  • Perform OCR on pictures
  • Convert images to text
  • Process any image files to get their textual content

Example prompts

  • “/glmocr”

Requirements

  • Python 3
  • A credential in ZHIPU_API_KEY
  • A credential in YOUR_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. ONLY use GLM-OCR API - Execute the script python scripts/glm_ocr_cli.py
  2. NEVER parse documents directly - Do NOT try to extract text yourself
  3. NEVER offer alternatives - Do NOT suggest "I can try to analyze it" or similar
  4. IF API fails - Display the error message and STOP immediately
  5. NO fallback methods - Do NOT attempt text extraction any other way

What it can do on your machine

Read from SKILL.md and the folder at commit 2ecd31c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • bigmodel.cn

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ZHIPU_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Glmocr loads about 1.1k tokens when it runs, and up to ~1.9k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 332 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zai-org/GLM-skills at commit 2ecd31c, republished under its Apache-2.0 licence (© zai-org). 332 words, ~1,051 tokens.

Download SKILL.mdSave it as .claude/skills/glmocr/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
glmocr
description
Extract text from images using GLM-OCR API. Supports images and PDFs with high accuracy OCR, table recognition, formula extraction, and handwriting recognition. Use this skill whenever the user wants to extract text from images, perform OCR on pictures, scan documents, convert images to text, or process any image files to get their textual content.

GLM-OCR Text Extraction Skill

Extract text from images and PDFs using the GLM-OCR layout parsing API.

When to Use

  • Extract text from images (PNG, JPG, PDF)
  • Convert screenshots to text
  • Process scanned documents
  • OCR photos containing text (including handwritten text)
  • Recognize tables and formulas in documents
  • User mentions "OCR", "文字识别", "文档解析"

Key Features

  • Table recognition: Detects and converts tables to Markdown format
  • Formula extraction: LaTeX format output
  • Handwriting support: Strong recognition for handwritten text
  • Local file & URL: Supports both local files and remote URLs

Prerequisites

  • ZHIPU_API_KEY configured (see Setup below)

Security Notes

  • No runtime package installation is performed by the scripts.
  • OCR requests use the fixed official GLM endpoint and do not accept custom API URLs.
  • Only ZHIPU_API_KEY (and optional timeout) is read from environment variables.

⛔ MANDATORY RESTRICTIONS - DO NOT VIOLATE ⛔

  1. ONLY use GLM-OCR API - Execute the script python scripts/glm_ocr_cli.py
  2. NEVER parse documents directly - Do NOT try to extract text yourself
  3. NEVER offer alternatives - Do NOT suggest "I can try to analyze it" or similar
  4. IF API fails - Display the error message and STOP immediately
  5. NO fallback methods - Do NOT attempt text extraction any other way

Setup

  1. Get your API key: https://www.bigmodel.cn/usercenter/proj-mgmt/apikeys
  2. Configure:
    bash
    python scripts/config_setup.py setup --api-key YOUR_KEY

How to Use

Extract from URL
bash
python scripts/glm_ocr_cli.py --file-url "URL provided by user"
Extract from Local File
bash
python scripts/glm_ocr_cli.py --file /path/to/image.jpg
bash
python scripts/glm_ocr_cli.py --file-url "URL" --output result.json

CLI Reference

python {baseDir}/scripts/glm_ocr_cli.py (--file-url URL | --file PATH) [--output FILE] [--pretty]
ParameterRequiredDescription
--file-urlOne ofURL to image/PDF
--fileOne ofLocal file path to image/PDF
--output, -oNoSave result JSON to file
--prettyNoPretty-print JSON output

Response Format

json
{
  "ok": true,
  "text": "# Extracted text in Markdown...",
  "layout_details": [[...]],
  "result": { "raw_api_response": "..." },
  "error": null,
  "source": "/path/to/file.jpg",
  "source_type": "file"
}

Key fields:

  • ok — whether extraction succeeded
  • text — extracted text in Markdown (use this for display)
  • layout_details — layout analysis details
  • result — raw API response
  • error — error details on failure

Error Handling

API key not configured:

Error: ZHIPU_API_KEY not configured. Get your API key at: https://www.bigmodel.cn/usercenter/proj-mgmt/apikeys

→ Show exact error to user, guide them to configure

Authentication failed (401/403): API key invalid/expired → reconfigure

Rate limit (429): Quota exhausted → inform user to wait

File not found: Local file missing → check path

Reference

  • references/output_schema.md — detailed output format specification

© zai-org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/glmocr of zai-org/GLM-skills.

  • SKILL.md
  • references/output_schema.md
  • scripts/config_setup.py
  • scripts/glm_ocr_cli.py
  • scripts/requirements.txt

Open the folder on GitHubat commit 2ecd31c

Compare with similar skills

Glmocr next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Glmocr compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Glmocr this skillzai-org/GLM-skills476—~1.1kAutomated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
GenOffice Document CLIgenspark-ai/genoffice8.9k—~19kAutomated safety check: PassApache-2.0
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone
Bookforge Korean Ebook PDF Makergongnyang/bookforge3141 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    8.9k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Produces book-style Korean ebook PDFs from a topic or finished manuscript, with six design styles, real book parts and quality-check gates before output.

    314 GitHub starsUsed in 1 repo~1.7k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed

More from zai-org/GLM-skills

All 16 skills in this repo
  • Glm Image Gen

    zai-org/GLM-skills

    Official skill for generating high-quality images from text prompts using ZhiPu GLM-Image API.

    476 GitHub stars~2.9k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Formula

    zai-org/GLM-skills

    Official skill for recognizing and extracting mathematical formulas from images and PDFs into LaTeX format using ZhiPu GLM-OCR API.

    476 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Handwriting

    zai-org/GLM-skills

    Official skill for recognizing handwritten text from images using ZhiPu GLM-OCR API.

    476 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Table

    zai-org/GLM-skills

    Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API.

    476 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmv Caption

    zai-org/GLM-skills

    Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series.

    476 GitHub stars~2k tokensUpdated 5 mo ago
    Auto-check: notes
  • Glmv Doc Based Writing

    zai-org/GLM-skills

    Write a textual content based on given document(s) and requirements, using ZhiPu GLM-V multimodal model.

    476 GitHub stars~1.6k tokensUpdated 5 mo ago
    Auto-check passed

Works with

Questions about Glmocr

What does Glmocr do?

Extract text from images using GLM-OCR API. An agent skill from zai-org/GLM-skills. Glmocr is an agent skill from zai-org/GLM-skills. Extract text from images using GLM-OCR API.

When should I use Glmocr?

Glmocr fits situations like: the user wants to extract text from images; perform OCR on pictures; convert images to text; process any image files to get their textual content.

How do I install Glmocr in Claude Code?

Run `npx skills add zai-org/GLM-skills --skill glmocr -a claude-code`. Or copy the skill folder (skills/glmocr in zai-org/GLM-skills) into .claude/skills/glmocr in your project. Claude Code loads it when a task matches its description.

How do I install Glmocr in Codex?

Run `npx skills add zai-org/GLM-skills --skill glmocr -a codex`. Or copy the skill folder (skills/glmocr in zai-org/GLM-skills) into .agents/skills/glmocr in your project. Codex loads it when a task matches its description.

Can I use Glmocr in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zai-org/GLM-skills --skill glmocr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/glmocr, .gemini/skills/glmocr, .github/skills/glmocr and .opencode/skills/glmocr in your project.

What does Glmocr need to run?

Going by SKILL.md and its folder, Glmocr needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named ZHIPU_API_KEY. Our summary lists: Python 3; A credential in ZHIPU_API_KEY; A credential in YOUR_KEY.

Does Glmocr access the network?

SKILL.md names 2 domains. In commands or code: bigmodel.cn; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Glmocr safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Glmocr use?

Glmocr is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Glmocr use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 899 tokens, read only when the agent opens those files.

What are the alternatives to Glmocr?

Skills that share tags, products or a category with Glmocr: Markitdown (ImCa0/just-laws, 781 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.9k stars) and Harness Book Best Practice (wquguru/harness-books, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Glmocr?

zai-org (a GitHub organization) maintains it in zai-org/GLM-skills, which has 476 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on April 15, 2026.

Source: zai-org/GLM-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.