Agent skill

Glmocr Table

by zai-org in zai-org/GLM-skills

Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API.

Apache-2.0Auto-check passedDocuments & Office

Install Glmocr Table

skills CLI
$ npx skills add zai-org/GLM-skills --skill glmocr-table -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zai-org/GLM-skills glmocr-table --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zai-org/GLM-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/glmocr-table .claude/skills/glmocr-table && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
glmocr-table
GitHub stars
475
Token cost
~1.7k tokens
SKILL.md length
549 words
Files
2 (incl. scripts)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API.

  • Works in 3 steps: Global config (recommended) / 全局配置(推荐):… → Skill-level config / Skill 级别配置: Set for… → Shell environment variable / Shell 环境变量:…
  • The user wants to extract tables
  • SKILL.md covers When to Use / 使用场景, Key Features / 核心特性, Resource Links / 资源链接 and Prerequisites / 前置条件, plus 5 more sections
  • Runs Python scripts from its folder; calls python; reaches bigmodel.cn and open.bigmodel.cn; needs ZHIPU_API_KEY

What it does

Glmocr Table is an agent skill from zai-org/GLM-skills. Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API. Supports complex tables, merged cells, and multi-page documents. Use this skill when the user wants to extract tables, recognize spreadsheets, or convert table images to editable format.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/glm_ocr_cli.py`).

It sits in Documents & Office, covering Excel spreadsheets and Markdown. It works with Zhipu GLM. The repository describes itself as: Official skills for the GLM family of models. The licence is Apache-2.0.

When your agent uses it

  • The user wants to extract tables
  • Recognize spreadsheets
  • Convert table images to editable format

Example prompts

  • “/glmocr-table”

Requirements

  • Python 3
  • A credential in ZHIPU_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Global config (recommended) / 全局配置(推荐): Set once in openclaw.json under env.vars, all Zhipu skills will share it
  2. Skill-level config / Skill 级别配置: Set for this skill only in openclaw.json
  3. Shell environment variable / Shell 环境变量: Add to ~/.zshrc

What it can do on your machine

Read from SKILL.md and the folder at commit 2ecd31c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • bigmodel.cn
    • open.bigmodel.cn

    Also links to:

    • docs.bigmodel.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ZHIPU_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Glmocr Table loads about 1.7k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 549 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zai-org/GLM-skills at commit 2ecd31c, republished under its Apache-2.0 licence (© zai-org). 549 words, ~1,694 tokens.

Download SKILL.mdSave it as .claude/skills/glmocr-table/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
glmocr-table
description
Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API. Supports complex tables, merged cells, and multi-page documents. Use this skill when the user wants to extract tables, recognize spreadsheets, or convert table images to editable format.

GLM-OCR Table Recognition Skill / GLM-OCR 表格识别技能

Extract tables from images and PDFs and convert them to Markdown format using the ZhiPu GLM-OCR layout parsing API.

When to Use / 使用场景

  • Extract tables from images or scanned documents / 从图片或扫描件中提取表格
  • Convert table images to Markdown or Excel format / 将表格图片转为 Markdown 或可编辑格式
  • Recognize complex tables with merged cells / 识别含合并单元格的复杂表格
  • Parse financial statements, invoices, reports with tables / 解析财务报表、发票、带表格的报告
  • User mentions "extract table", "recognize table", "表格识别", "提取表格", "表格OCR", "表格转文字"

Key Features / 核心特性

  • Complex table support: Handles merged cells, nested tables, multi-row headers
  • Markdown output: Tables are output in clean Markdown format, easy to edit and convert
  • Multi-page PDF: Supports batch extraction from multi-page PDF documents
  • Local file & URL: Supports both local files and remote URLs

Prerequisites / 前置条件

API Key Setup / API Key 配置(Required / 必需)

脚本通过 ZHIPU_API_KEY 环境变量获取密钥,可与其他智谱技能复用同一个 key。 This script reads the key from the ZHIPU_API_KEY environment variable. Reusing the same key across Zhipu skills is optional.

Get Key / 获取 Key: Visit 智谱开放平台 API Keys to create or copy your key.

Setup options / 配置方式(任选一种):

  1. Global config (recommended) / 全局配置(推荐): Set once in openclaw.json under env.vars, all Zhipu skills will share it:

    json
    {
      "env": {
        "vars": {
          "ZHIPU_API_KEY": "你的密钥"
        }
      }
    }
  2. Skill-level config / Skill 级别配置: Set for this skill only in openclaw.json:

    json
    {
      "skills": {
        "entries": {
          "glmocr-table": {
            "env": {
              "ZHIPU_API_KEY": "你的密钥"
            }
          }
        }
      }
    }
  3. Shell environment variable / Shell 环境变量: Add to ~/.zshrc:

    bash
    export ZHIPU_API_KEY="你的密钥"

💡 如果你已为其他智谱 skill(如 glmocr、glmv-caption、glm-image-generation)配置过 key,它们共享同一个 ZHIPU_API_KEY,无需重复配置。

Security & Transparency / 安全与透明度

  • Environment variables used / 使用的环境变量:
    • ZHIPU_API_KEY (required / 必需)
    • GLM_OCR_TIMEOUT (optional timeout seconds / 可选超时秒数)
  • Fixed endpoint / 固定官方端点: https://open.bigmodel.cn/api/paas/v4/layout_parsing
  • No custom API URL override / 不支持自定义 API URL 覆盖: this avoids accidental key exfiltration via redirected endpoints.
  • Raw upstream response is optional / 原始响应默认不返回: use --include-raw only when needed for debugging.

⛔ MANDATORY RESTRICTIONS / 强制限制 ⛔

  1. ONLY use GLM-OCR API — Execute the script python scripts/glm_ocr_cli.py
  2. NEVER parse tables yourself — Do NOT try to extract tables using built-in vision or any other method
  3. NEVER offer alternatives — Do NOT suggest "I can try to recognize it" or similar
  4. IF API fails — Display the error message and STOP immediately
  5. NO fallback methods — Do NOT attempt table extraction any other way
Show full SKILL.md (204 more words)Show less
📋 Output Display Rules / 输出展示规则

After running the script, present the OCR result clearly and safely.

  • Show extracted table Markdown (text) in full
  • Summarization is allowed, but do not hide important extraction failures
  • If layout_details contains table-related entries, you may highlight them
  • If the result file is saved, tell the user the file path
  • Show raw upstream response only when explicitly requested or debugging (--include-raw)

How to Use / 使用方法

Extract from URL / 从 URL 提取
bash
python scripts/glm_ocr_cli.py --file-url "https://example.com/table.png"
Extract from Local File / 从本地文件提取
bash
python scripts/glm_ocr_cli.py --file /path/to/table.png
Save Result to File / 保存结果到文件
bash
python scripts/glm_ocr_cli.py --file table.png --output result.json --pretty
Include Raw Upstream Response (Debug Only) / 包含原始上游响应(仅调试)
bash
python scripts/glm_ocr_cli.py --file table.png --output result.json --include-raw

CLI Reference / CLI 参数

python {baseDir}/scripts/glm_ocr_cli.py (--file-url URL | --file PATH) [--output FILE] [--pretty] [--include-raw]
ParameterRequiredDescription
--file-urlOne ofURL to image/PDF
--fileOne ofLocal file path to image/PDF
--output, -oNoSave result JSON to file
--prettyNoPretty-print JSON output
--include-rawNoInclude raw upstream API response in result field (debug only)

Response Format / 响应格式

json
{
  "ok": true,
  "text": "| Column 1 | Column 2 |\n|----------|----------|\n| Data     | Data     |",
  "layout_details": [...],
  "result": null,
  "error": null,
  "source": "/path/to/file",
  "source_type": "file",
  "raw_result_included": false
}

Key fields:

  • ok — whether extraction succeeded
  • text — extracted text in Markdown (use this for display)
  • layout_details — layout analysis details
  • error — error details on failure

Error Handling / 错误处理

API key not configured:

ZHIPU_API_KEY not configured. Get your API key at: https://bigmodel.cn/usercenter/proj-mgmt/apikeys

→ Show exact error to user, guide them to configure

Authentication failed (401/403): API key invalid/expired → reconfigure

Rate limit (429): Quota exhausted → inform user to wait

File not found: Local file missing → check path

© zai-org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/glmocr-table of zai-org/GLM-skills.

  • SKILL.md
  • scripts/glm_ocr_cli.py

Open the folder on GitHubat commit 2ecd31c

Compare with similar skills

Glmocr Table next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Glmocr Table compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Glmocr Table this skillzai-org/GLM-skills475—~1.7kAutomated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Markitshift-labs-ai/markit1.3k—~299Automated safety check: PassMIT
Doc Cleanernotoriouslab/doc-cleaner309—~712Automated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
Gdoc To Markdowniurykrieger/claude-bedrock1051 repos~3.8kAutomated safety check: NotesMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markit

    shift-labs-ai/markit

    Convert files and URLs to Markdown. An agent skill from shift-labs-ai/markit.

    1.3k GitHub stars~299 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Doc Cleaner

    notoriouslab/doc-cleaner

    Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

    309 GitHub stars~712 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • Gdoc To Markdown

    iurykrieger/claude-bedrock

    Internal fetcher module for Google Docs and Sheets. An agent skill from iurykrieger/claude-bedrock.

    105 GitHub starsUsed in 1 repo~3.8k tokens
    Documents & OfficeAuto-check: notes
  • Skill Doc Delivery

    nyldn/claude-octopus

    Convert markdown to DOCX, PPTX, XLSX, PDF office documents — use when you need exportable deliverables

    4.2k GitHub starsUsed in 1 repo~2.4k tokens
    Documents & OfficeAuto-check passed

More from zai-org/GLM-skills

All 16 skills in this repo
  • Glmocr

    zai-org/GLM-skills

    Extract text from images using GLM-OCR API. An agent skill from zai-org/GLM-skills.

    475 GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Glm Image Gen

    zai-org/GLM-skills

    Official skill for generating high-quality images from text prompts using ZhiPu GLM-Image API.

    475 GitHub stars~2.9k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Formula

    zai-org/GLM-skills

    Official skill for recognizing and extracting mathematical formulas from images and PDFs into LaTeX format using ZhiPu GLM-OCR API.

    475 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Handwriting

    zai-org/GLM-skills

    Official skill for recognizing handwritten text from images using ZhiPu GLM-OCR API.

    475 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmv Caption

    zai-org/GLM-skills

    Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series.

    475 GitHub stars~2k tokensUpdated 5 mo ago
    Auto-check: notes
  • Glmv Doc Based Writing

    zai-org/GLM-skills

    Write a textual content based on given document(s) and requirements, using ZhiPu GLM-V multimodal model.

    475 GitHub stars~1.6k tokensUpdated 5 mo ago
    Auto-check passed

Works with

Questions about Glmocr Table

What does Glmocr Table do?

Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API. Glmocr Table is an agent skill from zai-org/GLM-skills. Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API.

When should I use Glmocr Table?

Glmocr Table fits situations like: the user wants to extract tables; recognize spreadsheets; convert table images to editable format.

How do I install Glmocr Table in Claude Code?

Run `npx skills add zai-org/GLM-skills --skill glmocr-table -a claude-code`. Or copy the skill folder (skills/glmocr-table in zai-org/GLM-skills) into .claude/skills/glmocr-table in your project. Claude Code loads it when a task matches its description.

How do I install Glmocr Table in Codex?

Run `npx skills add zai-org/GLM-skills --skill glmocr-table -a codex`. Or copy the skill folder (skills/glmocr-table in zai-org/GLM-skills) into .agents/skills/glmocr-table in your project. Codex loads it when a task matches its description.

Can I use Glmocr Table in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zai-org/GLM-skills --skill glmocr-table -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/glmocr-table, .gemini/skills/glmocr-table, .github/skills/glmocr-table and .opencode/skills/glmocr-table in your project.

What does Glmocr Table need to run?

Going by SKILL.md and its folder, Glmocr Table needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named ZHIPU_API_KEY. Our summary lists: Python 3; A credential in ZHIPU_API_KEY.

Does Glmocr Table access the network?

SKILL.md names 3 domains. In commands or code: bigmodel.cn and open.bigmodel.cn; the agent is likely to contact these when it follows the instructions. As links in the text: docs.bigmodel.cn. This is read from the text; nothing was executed.

Is Glmocr Table safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Glmocr Table use?

Glmocr Table is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Glmocr Table use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Glmocr Table?

Skills that share tags, products or a category with Glmocr Table: Markitdown (ImCa0/just-laws, 781 stars), Markit (shift-labs-ai/markit, 1.3k stars), Doc Cleaner (notoriouslab/doc-cleaner, 309 stars) and Mineru (Nebutra/MinerU-Skill, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Glmocr Table?

zai-org (a GitHub organization) maintains it in zai-org/GLM-skills, which has 475 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on April 15, 2026.

Source: zai-org/GLM-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.