Agent skill

Glmv Grounding

by zai-org in zai-org/GLM-skills

A skill that uses GLM-V native grounding capabilities for coordinate conversion, bounding-box visualization, and more.

Apache-2.0Auto-check passedFrontend & Design

Install Glmv Grounding

skills CLI
$ npx skills add zai-org/GLM-skills --skill glmv-grounding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zai-org/GLM-skills glmv-grounding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zai-org/GLM-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/glmv-grounding .claude/skills/glmv-grounding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
glmv-grounding
GitHub stars
476
Token cost
~2.6k tokens
SKILL.md length
835 words
Files
10 (incl. scripts)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill that uses GLM-V native grounding capabilities for coordinate conversion, bounding-box visualization, and more.

  • Works in 2 steps: Get your API key:… → Configure it with
  • Tasks that involve Internationalization
  • SKILL.md covers When to use, Setup your API Key, Security & Transparency and Runtime Dependencies, plus 5 more sections
  • Runs Python scripts from its folder; calls python and pip; needs ZHIPU_API_KEY

What it does

Glmv Grounding is an agent skill from zai-org/GLM-skills. A skill that uses GLM-V native grounding capabilities for coordinate conversion, bounding-box visualization, and more. GLM-V native grounding can locate any target specified by the prompt in an image and output relative coordinates normalized to 0-1000 based on image size. Coordinate formats include 2D bounding box (default), 2D points, and 3D bounding box. GLM-V also supports spatiotemporal localization and tracking of multiple prompt-specified targets in videos, outputting 2D bounding boxes per second.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts (for example `scripts/config_setup.py`, `scripts/glm_grounding_cli.py` and `scripts/utils_3d.py`).

It sits in Frontend & Design, covering Internationalization. It works with Zhipu GLM. The repository describes itself as: Official skills for the GLM family of models. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Internationalization

Example prompts

  • “/glmv-grounding”

Requirements

  • Python 3
  • A credential in ZHIPU_API_KEY
  • A credential in YOUR_KEY

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Get your API key: https://www.bigmodel.cn/usercenter/proj-mgmt/apikeys
  2. Configure it with

What it can do on your machine

Read from SKILL.md and the folder at commit 2ecd31c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 8 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • bigmodel.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ZHIPU_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Glmv Grounding loads about 2.6k tokens when it runs. Until then it costs about 131 tokens; SKILL.md has 835 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~131
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zai-org/GLM-skills at commit 2ecd31c, republished under its Apache-2.0 licence (© zai-org). 835 words, ~2,625 tokens.

Download SKILL.mdSave it as .claude/skills/glmv-grounding/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
glmv-grounding
description
A skill that uses GLM-V native grounding capabilities for coordinate conversion, bounding-box visualization, and more. GLM-V native grounding can locate any target specified by the prompt in an image and output relative coordinates normalized to 0-1000 based on image size. Coordinate formats include 2D bounding box (default), 2D points, and 3D bounding box. GLM-V also supports spatiotemporal localization and tracking of multiple prompt-specified targets in videos, outputting 2D bounding boxes per second.

GLMV-Grounding Skill

Extract and visualize grounding results produced by GLM-V. Depending on the user prompt, grounding coordinates in model outputs may appear in different forms, including 2D bounding boxes, Objects Detection JSON, 2D points, 3D bounding boxes, and target-tracking JSON.

Note: GLM-V outputs coordinates where x and y are relative coordinates normalized from pixel coordinates x_pixel and y_pixel using image width W and height H (range 0-1000), i.e., x=round(x_pixel/W1000), y=round(y_pixel/H1000). The origin of the pixel coordinate system is the top-left corner. Note: If the prompt does not explicitly specify a grounding format (for example, "find the location of xxx" or "draw a box around xxx"), treat the request as 2D bounding boxes by default.

When to use

  • Use GLM-V to ground targets in images: obtain grounding results in an image for any prompt-described target, with output formats such as 2D bounding box (default), 2D points, and 3D bounding box.
  • Use GLM-V to track targets in videos: obtain tracking results in a video for any prompt-described target, with output format like {"0": [{"label": ..., "bbox_2d": ...}, ...], ...}.
  • Use utility functions for extraction, conversion, and visualization: extract coordinates, points, and JSON from natural text; normalize and de-normalize coordinates; visualize boxes, points, 3D boxes, and video tracking results.

Setup your API Key

Configure ZHIPU_API_KEY to call the GLM-V API.

  1. Get your API key: https://www.bigmodel.cn/usercenter/proj-mgmt/apikeys
  2. Configure it with:
python scripts/config_setup.py setup --api-key YOUR_KEY

Security & Transparency

  • Primary API key env: ZHIPU_API_KEY (required).
  • Timeout env: GLM_GROUNDING_TIMEOUT (optional, seconds, default 60).
  • API endpoint: fixed to official Zhipu Chat Completions endpoint in CLI implementation.
  • No dynamic key name switching: the skill expects ZHIPU_API_KEY consistently.
  • URL/local file handling: the skill can read local files or fetch user-provided URLs for processing/visualization; URL inputs are restricted to public http/https targets (localhost/private network targets are rejected).

Runtime Dependencies

Install dependencies before use:

bash
pip install -r scripts/requirements.txt

Main packages used by this skill:

  • requests
  • Pillow
  • opencv-python
  • numpy
  • matplotlib
  • decord

System dependency for video visualization:

  • ffmpeg

General workflow

	Input (image or video + Prompt)
		|
		▼
	Run glm_grounding_cli.py to get grounding results (natural language)
		|
		▼
	Return results (grounding results, visualized image or video)

How to Use

Run glm_grounding_cli.py to get grounding results
  • Ground any target in an image
python scripts/glm_grounding_cli.py --image-url "URL provided by user" --prompt "description of target for grounding"
  • Track any target in a video
python scripts/glm_grounding_cli.py --video-url /path/to/image.jpg --prompt "description of target for tracking" --visualize --visualization-dir "./vis"
Reply with grounding results

After receiving a grounding prompt from the user, your direct reply should be natural language that includes grounding coordinates. Coordinates $x$ and $y$ are relative values in [0, 1000], computed as:

$$ x = round(x_{pixel} / W * 1000) \ y = round(y_{pixel}/H*1000) $$

where $x_{pixel}, y_{pixel}$ are pixel coordinates with origin (0, 0) at the top-left corner of the image, and W/H are the image width/height.

Unless otherwise specified, grounding results should use the following Python data formats:

  • 2D bounding boxes: [[x1, y1, x2, y2], ...], extracted grounding result is a list of boxes, each box has 4 coordinate values
  • 2D points: [[x, y], ...], extracted grounding result is a list of points, each point has 2 coordinate values
  • 2D polygon: [[x1, y1], [x2, y2], ...], extracted grounding result is a polygon coordinate list, each vertex has 2 coordinate values
  • 3D bounding boxes: [{"bbox_3d":[x_center, y_center, z_center, x_size, y_size, z_size, roll, pitch, yaw],"label":"category"}, ...], extracted grounding result is a JSON list where each object contains a category label and one 3D box with 8 coordinate values
  • Objects Detection JSON: [{'label': 'category', 'bbox_2d': [x1, y1, x2, y2]}, ...], extracted grounding result is a JSON list where each object contains a category label and one box
  • Video Objects Tracking JSON: {0: [{'label': 'car-1', 'bbox_2d': [1,2,3,4]}, {'label': 'car-2', 'bbox_2d': [2,3,4,5]}], 1: [{'label': 'car-2', 'bbox_2d': [4,5,6,7]}, {'label': 'person-1', 'bbox_2d': [10,20,30,40]}]}, extracted grounding result is a JSON object whose keys are video frame indices and values are lists of JSON objects, each containing a category label and one 2D box
Show full SKILL.md (253 more words)Show less

Python example

shell

# 1. User grounding request and your reply
image=https://example.com/image.jpg
prompt="Please box all people wearing Santa hats in the image and tell me their coordinates. Use red boxes, line thickness 3, and label format 'SantaHat-i'."


# 2. Get grounding results
python scripts/glm_grounding_cli.py --image-url $image --prompt $prompt --visualize --visualization-dir "./vis"

#  {
#         "ok": True,
#         "grounding_result": [[100, 200, 300, 400], [500, 600, 700, 800]],
#         "visualizations_result": (
#             {"visualized_image": "./vis/image_vis.jpg"}
#         ),
#         "raw_result": "1. Person 1: box [100, 200, 300, 400]\n2. Person 2: box [500, 600, 700, 800]. The box format is [x1, y1, x2, y2], where (x1, y1) is the top-left corner and (x2, y2) is the bottom-right corner.",
#         "error": None,
#         "source": source,
#     }

Utility function quick reference

FunctionPurpose
parse_coordinates_from_response(response_str, coords_type='bbox', init_context_window=2000, max_context_window=-1)Parse and extract all coordinate results from model responses (supports 2D bbox, point, polygon)
parse_3d_boxes_from_response(response_str, max_context_window=-1)Parse and extract all 3D boxes and labels from model responses (strict and loose matching)
parse_detection_from_response(response_str, max_context_window=-1)Parse and extract all 2D detection results from model responses (Objects Detection JSON format)
parse_mot_from_response(response_str, max_context_window=-1)Parse and extract all video object tracking results from model responses (Video Objects Tracking JSON format)
visualize_boxes(img_path=None, img_bytes=None, boxes=[], labels=None, renormalize=False, save_path=None, return_b64=False, save_optimized=True, **kwargs)Draw 2D boxes on images with labels, custom colors, and line thickness
visualize_points(img_path=None, img_bytes=None, points=[], labels=None, renormalize=False, diameters=None, save_path=None, return_b64=False, save_optimized=True, distinct_colors=False, colors=None)Draw points on images with labels, custom size, and colors
visualize_3d_boxes_glmv_simple(image_path, cam_params, bbox_3d_list, image_bytes=None, coord_format='xyzwhlpyr', save_path=None, save_optimized=False, return_b64=False, **kwargs)Draw projected 3D boxes on images using camera intrinsics (supports rotation and multiple coordinate formats)
visualize_mot(video_path=None, video_bytes=None, mot_js=None, renormalize=False, save_path=None, return_b64=False, distinct_colors=True, **kwargs)Draw Video Objects Tracking boxes on each video frame with labels

Common errors

  • Coordinate values exceed 1000: if extracted coordinate values are greater than 1000, the model may have produced unnormalized coordinates due to prompt effects. Extract the target phrase from the user request (for example, "people wearing Santa hats"), then query the model again and explicitly require output coordinates to be relative values normalized to 0-1000 based on image size (for example, "Please box all people wearing Santa hats in the image and tell me their coordinates. Ensure the output coordinates are relative values normalized to 0-1000 based on image size.").

© zai-org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts) in skills/glmv-grounding of zai-org/GLM-skills.

  • SKILL.md
  • .DS_Store
  • scripts/.DS_Store
  • scripts/config_setup.py
  • scripts/glm_grounding_cli.py
  • scripts/requirements.txt
  • scripts/utils_3d.py
  • scripts/utils_boxes.py
  • scripts/utils_detection.py
  • scripts/utils_video.py

Open the folder on GitHubat commit 2ecd31c

Compare with similar skills

Glmv Grounding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Glmv Grounding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Glmv Grounding this skillzai-org/GLM-skills476—~2.6kAutomated safety check: PassApache-2.0
Impeccablebestofjs/bestofjs3.1k27 repos~2.6kAutomated safety check: PassMIT
Chatbox i18n Translatorchatboxai/chatbox42k—~508Automated safety check: PassGPL-3.0
Internationalization Workflow with i18niOfficeAI/AionUi33k1 repos~1.9kAutomated safety check: PassApache-2.0
Enforce Rules For I18nmoeru-ai/airi50k—~1.5kAutomated safety check: PassMIT
Claude Desktop Chinese Localizationjavaht/claude-desktop-zh-cn7.5k—~1.6kAutomated safety check: PassMIT

Similar skills

  • Impeccable

    bestofjs/bestofjs

    A skill your agent uses when the user wants to design, redesign, shape, critique, audit, polish, clarify, distill, harden, optimize, adapt, animate, colorize, extract, or otherwise improve a…

    3.1k GitHub starsUsed in 27 repos~2.6k tokens
    Frontend & DesignAuto-check passed
  • Chatbox i18n Translator

    chatboxai/chatbox

    Translates new or changed i18n keys from a Chatbox Pro diff, staged changes or a commit range, writing the locale JSON files directly with a built-in glossary.

    42k GitHub stars~508 tokensUpdated 14 days ago
    Frontend & DesignAuto-check passed
  • Standards for keeping all user-facing text translatable: read the i18n config first, use namespaced keys, reuse shared strings and follow the key naming rules.

    33k GitHub starsUsed in 1 repo~1.9k tokens
    Frontend & DesignAuto-check passed
  • Review pending AIRI translations on Crowdin in a batch, then sync them into the repository.

    50k GitHub stars~1.5k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Claude Desktop Chinese Localization

    javaht/claude-desktop-zh-cn

    Adds missing Simplified and Traditional Chinese translations to the Claude Desktop Chinese patch across three layers, then checks how many mappings actually hit.

    7.5k GitHub stars~1.6k tokensUpdated 3 days ago
    Frontend & DesignAuto-check passed
  • Taro UI Guide

    jd-opensource/taro-ui

    Guides installing, configuring, styling and using taro-ui (At* components) in Taro apps for WeChat, Alipay, H5 and React Native.

    4.7k GitHub stars~1.3k tokensUpdated 13 days ago
    Frontend & DesignAuto-check passed

More from zai-org/GLM-skills

All 16 skills in this repo
  • Glmocr

    zai-org/GLM-skills

    Extract text from images using GLM-OCR API. An agent skill from zai-org/GLM-skills.

    476 GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Glm Image Gen

    zai-org/GLM-skills

    Official skill for generating high-quality images from text prompts using ZhiPu GLM-Image API.

    476 GitHub stars~2.9k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Formula

    zai-org/GLM-skills

    Official skill for recognizing and extracting mathematical formulas from images and PDFs into LaTeX format using ZhiPu GLM-OCR API.

    476 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Handwriting

    zai-org/GLM-skills

    Official skill for recognizing handwritten text from images using ZhiPu GLM-OCR API.

    476 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmocr Table

    zai-org/GLM-skills

    Official skill for recognizing and extracting tables from images and PDFs into Markdown format using ZhiPu GLM-OCR API.

    476 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Glmv Caption

    zai-org/GLM-skills

    Generate captions (descriptions) for images, videos, and documents using ZhiPu GLM-V multimodal model series.

    476 GitHub stars~2k tokensUpdated 5 mo ago
    Auto-check: notes

Works with

Questions about Glmv Grounding

What does Glmv Grounding do?

A skill that uses GLM-V native grounding capabilities for coordinate conversion, bounding-box visualization, and more. Glmv Grounding is an agent skill from zai-org/GLM-skills. A skill that uses GLM-V native grounding capabilities for coordinate conversion, bounding-box visualization, and more.

When should I use Glmv Grounding?

Glmv Grounding fits situations like: tasks that involve Internationalization.

How do I install Glmv Grounding in Claude Code?

Run `npx skills add zai-org/GLM-skills --skill glmv-grounding -a claude-code`. Or copy the skill folder (skills/glmv-grounding in zai-org/GLM-skills) into .claude/skills/glmv-grounding in your project. Claude Code loads it when a task matches its description.

How do I install Glmv Grounding in Codex?

Run `npx skills add zai-org/GLM-skills --skill glmv-grounding -a codex`. Or copy the skill folder (skills/glmv-grounding in zai-org/GLM-skills) into .agents/skills/glmv-grounding in your project. Codex loads it when a task matches its description.

Can I use Glmv Grounding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zai-org/GLM-skills --skill glmv-grounding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/glmv-grounding, .gemini/skills/glmv-grounding, .github/skills/glmv-grounding and .opencode/skills/glmv-grounding in your project.

What does Glmv Grounding need to run?

Going by SKILL.md and its folder, Glmv Grounding needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named ZHIPU_API_KEY. Our summary lists: Python 3; A credential in ZHIPU_API_KEY; A credential in YOUR_KEY.

Does Glmv Grounding access the network?

SKILL.md names 1 domain. As links in the text: bigmodel.cn. This is read from the text; nothing was executed.

Is Glmv Grounding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Glmv Grounding use?

Glmv Grounding is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Glmv Grounding use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Glmv Grounding?

Skills that share tags, products or a category with Glmv Grounding: Impeccable (bestofjs/bestofjs, 3.1k stars), Chatbox i18n Translator (chatboxai/chatbox, 42k stars), Internationalization Workflow with i18n (iOfficeAI/AionUi, 33k stars) and Enforce Rules For I18n (moeru-ai/airi, 50k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Glmv Grounding?

zai-org (a GitHub organization) maintains it in zai-org/GLM-skills, which has 476 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on April 15, 2026.

Source: zai-org/GLM-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.