Agent skill

Glmv Caption

by zai-org in zai-org/GLM-skills

Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series.

Apache-2.0Auto-check: notesDocuments & Office

Install Glmv Caption

skills CLI
$ npx skills add zai-org/GLM-skills --skill glmv-caption -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zai-org/GLM-skills glmv-caption --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zai-org/GLM-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/glmv-caption .claude/skills/glmv-caption && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
glmv-caption
GitHub stars
475
Token cost
~2k tokens
SKILL.md length
566 words
Files
2 (incl. scripts)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series.

  • Works in 3 steps: OpenClaw config (recommended) / OpenClaw… → Shell environment variable / Shell 环境变量:… → .env file / .env 文件: Create .env in this…
  • The user wants to describe
  • SKILL.md covers When to Use, Supported Input Types, Resource Links and Prerequisites, plus 4 more sections
  • Runs Python scripts from its folder; calls python; reaches bigmodel.cn; needs ZHIPU_API_KEY

What it does

Glmv Caption is an agent skill from zai-org/GLM-skills. Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series. Use this skill whenever the user wants to describe, caption, summarize, or interpret the content of images, videos, or files. Supports single/multiple inputs, URLs, local paths, and base64 (images only).

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/glmv_caption.py`).

It sits in Documents & Office. It works with Zhipu GLM. The repository describes itself as: Official skills for the GLM family of models. The licence is Apache-2.0.

When your agent uses it

  • The user wants to describe
  • Interpret the content of images

Example prompts

  • “/glmv-caption”

Requirements

  • Python 3
  • A credential in ZHIPU_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. OpenClaw config (recommended) / OpenClaw 配置(推荐): Set in openclaw.json under skills.entries.glmv-caption.env
  2. Shell environment variable / Shell 环境变量: Add to ~/.zshrc
  3. .env file / .env 文件: Create .env in this skill directory

What it can do on your machine

Read from SKILL.md and the folder at commit 2ecd31c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • bigmodel.cn

    Also links to:

    • docs.bigmodel.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ZHIPU_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Glmv Caption loads about 2k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 566 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:71
    3. **.env file / .env 文件:** Create `.env` in this skill directory:

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zai-org/GLM-skills at commit 2ecd31c, republished under its Apache-2.0 licence (© zai-org). 566 words, ~2,003 tokens.

Download SKILL.mdSave it as .claude/skills/glmv-caption/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
glmv-caption
description
Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series. Use this skill whenever the user wants to describe, caption, summarize, or interpret the content of images, videos, or files. Supports single/multiple inputs, URLs, local paths, and base64 (images only).

GLM-V Caption Skill

Generate captions for images, videos, and documents using the ZhiPu GLM-V multimodal model.

When to Use

  • Describe, caption, summarize, or interpret image/video/document content
  • User mentions "describe this image", "caption", "summarize this video", "图片描述", "视频摘要", "文档解读", "看图说话"
  • Extract visual or textual information from media files
  • Compare multiple images
  • User provides an image/video/file and asks what's in it

Supported Input Types

TypeFormatsMax SizeMax CountBase64
Imagejpg, png, jpeg5MB / 6000×6000px50✅
Videomp4, mkv, mov200MB—❌
Filepdf, docx, txt, xlsx, pptx, jsonl—50❌

⚠️ file_url cannot mix with image_url or video_url in the same request. ⚠️ Videos and files only support URLs — local paths and base64 are NOT supported (images only).

Prerequisites

API Key Setup / API Key 配置(Required / 必需)

This script reads the key from the ZHIPU_API_KEY environment variable and shares it with other Zhipu skills. 脚本通过 ZHIPU_API_KEY 环境变量获取密钥,与其他智谱技能共用同一个 key。

Get Key / 获取 Key: Visit Zhipu Open Platform API Keys / 智谱开放平台 API Keys to create or copy your key.

Setup options / 配置方式(任选一种):

  1. OpenClaw config (recommended) / OpenClaw 配置(推荐): Set in openclaw.json under skills.entries.glmv-caption.env:

    json
    "glmv-caption": { "enabled": true, "env": { "ZHIPU_API_KEY": "你的密钥" } }
  2. Shell environment variable / Shell 环境变量: Add to ~/.zshrc:

    bash
    export ZHIPU_API_KEY="你的密钥"
  3. .env file / .env 文件: Create .env in this skill directory:

    ZHIPU_API_KEY=你的密钥

⛔ MANDATORY RESTRICTIONS - DO NOT VIOLATE ⛔

  1. ONLY use GLM-V API — Execute the script python scripts/glmv_caption.py
  2. NEVER caption media yourself — Do NOT try to describe content using built-in vision or any other method
  3. NEVER offer alternatives — Do NOT suggest "I can try to describe it" or similar
  4. IF API fails — Display the error message and STOP immediately
  5. NO fallback methods — Do NOT attempt captioning any other way
📋 Output Display Rules (MANDATORY)

After running the script, you must show the full raw output to the user exactly as returned. Do not summarize, truncate, or only say "generated". Users need the original model output to evaluate quality.

  • Image captioning: show the full caption text
  • Multiple images: show each image result
  • Video/files: show the full understanding result
  • If token usage is included, you may optionally display it
Show full SKILL.md (219 more words)Show less

How to Use

Caption an Image
bash
python scripts/glmv_caption.py --images "https://example.com/photo.jpg"
python scripts/glmv_caption.py --images /path/to/photo.png
Caption Multiple Images
bash
python scripts/glmv_caption.py --images img1.jpg img2.png "https://example.com/img3.jpg"
Caption a Video
bash
python scripts/glmv_caption.py --videos "https://example.com/clip.mp4"
Caption a Document
bash
python scripts/glmv_caption.py --files "https://example.com/report.pdf"
python scripts/glmv_caption.py --files "https://example.com/doc1.docx" "https://example.com/doc2.txt"
Custom Prompt
bash
python scripts/glmv_caption.py --images photo.jpg --prompt "Describe the architecture style in detail"
Save Result
bash
python scripts/glmv_caption.py --images photo.jpg --output result.json
Thinking Mode
bash
python scripts/glmv_caption.py --images photo.jpg --thinking

CLI Reference

python {baseDir}/scripts/glmv_caption.py (--images IMG [IMG...] | --videos VID [VID...] | --files FILE [FILE...]) [OPTIONS]
ParameterRequiredDescription
--images, -iOne ofImage paths or URLs (supports multiple, base64 OK)
--videos, -vOne ofVideo paths or URLs (supports multiple, mp4/mkv/mov)
--files, -fOne ofDocument paths or URLs (supports multiple, pdf/docx/txt/xlsx/pptx/jsonl)
--prompt, -pNoCustom prompt (default: "请详细描述这张图片的内容" / "Please describe this image in detail")
--model, -mNoModel name (default: glm-4.6v)
--temperature, -tNoSampling temperature 0-1 (default: 0.8)
--top-pNoNucleus sampling 0.01-1.0 (default: 0.6)
--max-tokensNoMax output tokens (default: 1024, max 32768)
--thinkingNoEnable thinking/reasoning mode
--output, -oNoSave result JSON to file
--prettyNoPretty-print JSON output
--streamNoEnable streaming output

Note: --images, --videos, and --files are mutually exclusive per API limits.

Response Format

json
{
  "success": true,
  "caption": "A landscape photo showing a mountain range at sunset...",
  "usage": {
    "prompt_tokens": 128,
    "completion_tokens": 256,
    "total_tokens": 384
  }
}

Key fields:

  • success — whether the request succeeded
  • caption — the generated caption text
  • usage — token usage statistics
  • warning — present when content was blocked by safety review
  • error — error details on failure

Error Handling

API key not configured:

ZHIPU_API_KEY not configured. Get your API key at: https://bigmodel.cn/usercenter/proj-mgmt/apikeys

→ Show exact error to user, guide them to configure

Authentication failed (401/403): API key invalid/expired → reconfigure

Rate limit (429): Quota exhausted → inform user to wait

File not found: Local file missing → check path

Content filtered: warning field present → content blocked by safety review

© zai-org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/glmv-caption of zai-org/GLM-skills.

  • SKILL.md
  • scripts/glmv_caption.py

Open the folder on GitHubat commit 2ecd31c

Compare with similar skills

Glmv Caption next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Glmv Caption compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Glmv Caption this skillzai-org/GLM-skills475—~2kAutomated safety check: NotesApache-2.0
Markdown Article FormatterJimLiu/baoyu-skills26k6 repos~3.5kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Obsidian MarkdownAtmosphere/atmosphere3.8k20 repos~1.3kAutomated safety check: PassApache-2.0
DOCXrvdbreemen/OTGW-firmware20733 repos~4.3kAutomated safety check: PassProprietary
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0

Similar skills

  • Markdown Article Formatter

    JimLiu/baoyu-skills

    Reformats plain text or Markdown articles with frontmatter, a title, a summary, headings, bold, lists and code blocks, and saves a separate formatted copy.

    26k GitHub starsUsed in 6 repos~3.5k tokens
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Obsidian Markdown

    Atmosphere/atmosphere

    Create and edit Obsidian Flavored Markdown with wikilinks, embeds, callouts, properties, and other Obsidian-specific syntax.

    3.8k GitHub starsUsed in 20 repos~1.3k tokens
    Documents & OfficeAuto-check passed
  • DOCX

    rvdbreemen/OTGW-firmware

    A skill your agent uses whenever the user wants to create, read, edit, or manipulate Word documents (.docx files).

    207 GitHub starsUsed in 33 repos~4.3k tokens
    Documents & OfficeAuto-check passed
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

    41k GitHub stars~2.5k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed

More from zai-org/GLM-skills

All 16 skills in this repo
  • Glmocr

    zai-org/GLM-skills

    Extract text from images using GLM-OCR API. An agent skill from zai-org/GLM-skills.

    475 GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Glm Image Gen

    zai-org/GLM-skills

    Official skill for generating high-quality images from text prompts using ZhiPu GLM-Image API.

    475 GitHub stars~2.9k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Formula

    zai-org/GLM-skills

    Official skill for recognizing and extracting mathematical formulas from images and PDFs into LaTeX format using ZhiPu GLM-OCR API.

    475 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Handwriting

    zai-org/GLM-skills

    Official skill for recognizing handwritten text from images using ZhiPu GLM-OCR API.

    475 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Table

    zai-org/GLM-skills

    Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API.

    475 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmv Doc Based Writing

    zai-org/GLM-skills

    Write a textual content based on given document(s) and requirements, using ZhiPu GLM-V multimodal model.

    475 GitHub stars~1.6k tokensUpdated 5 mo ago
    Auto-check passed

Works with

Questions about Glmv Caption

What does Glmv Caption do?

Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series. Glmv Caption is an agent skill from zai-org/GLM-skills. Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series.

When should I use Glmv Caption?

Glmv Caption fits situations like: the user wants to describe; interpret the content of images.

How do I install Glmv Caption in Claude Code?

Run `npx skills add zai-org/GLM-skills --skill glmv-caption -a claude-code`. Or copy the skill folder (skills/glmv-caption in zai-org/GLM-skills) into .claude/skills/glmv-caption in your project. Claude Code loads it when a task matches its description.

How do I install Glmv Caption in Codex?

Run `npx skills add zai-org/GLM-skills --skill glmv-caption -a codex`. Or copy the skill folder (skills/glmv-caption in zai-org/GLM-skills) into .agents/skills/glmv-caption in your project. Codex loads it when a task matches its description.

Can I use Glmv Caption in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zai-org/GLM-skills --skill glmv-caption -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/glmv-caption, .gemini/skills/glmv-caption, .github/skills/glmv-caption and .opencode/skills/glmv-caption in your project.

What does Glmv Caption need to run?

Going by SKILL.md and its folder, Glmv Caption needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named ZHIPU_API_KEY. Our summary lists: Python 3; A credential in ZHIPU_API_KEY.

Does Glmv Caption access the network?

SKILL.md names 2 domains. In commands or code: bigmodel.cn; the agent is likely to contact it when it follows the instructions. As links in the text: docs.bigmodel.cn. This is read from the text; nothing was executed.

Is Glmv Caption safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Glmv Caption use?

Glmv Caption is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Glmv Caption use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Glmv Caption?

Skills that share tags, products or a category with Glmv Caption: Markdown Article Formatter (JimLiu/baoyu-skills, 26k stars), Markitdown (ImCa0/just-laws, 781 stars), Obsidian Markdown (Atmosphere/atmosphere, 3.8k stars) and DOCX (rvdbreemen/OTGW-firmware, 207 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Glmv Caption?

zai-org (a GitHub organization) maintains it in zai-org/GLM-skills, which has 475 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on April 15, 2026.

Source: zai-org/GLM-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.