Agent skill

Image Understanding

by jeffstric in jeffstric/ZJT

图片理解智能体,负责识别图片内容、反推提示词、分析图片风格、对比图片差异以及文字识别(OCR). An agent skill from jeffstric/ZJT.

Custom licenceAuto-check passedMedia & Creative

Install Image Understanding

skills CLI
$ npx skills add jeffstric/ZJT --skill image-understanding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeffstric/ZJT image-understanding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeffstric/ZJT.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agents/skills/image-understanding .claude/skills/image-understanding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-understanding
GitHub stars
227
Token cost
~701 tokens
SKILL.md length
156 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Custom licence

At a glance

图片理解智能体,负责识别图片内容、反推提示词、分析图片风格、对比图片差异以及文字识别(OCR). An agent skill from jeffstric/ZJT.

  • Works in 2 steps: fetch_image_as_base64(image_url) —… → ask_user(question, options) — 向用户提问
  • Media & Creative work in your project
  • SKILL.md covers 角色定位, 核心工具, ⚠️ 重要:图片来源与加载 and 工作流程, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Image Understanding is an agent skill from jeffstric/ZJT. 图片理解智能体,负责识别图片内容、反推提示词、分析图片风格、对比图片差异以及文字识别(OCR)。

Its SKILL.md is about 700 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative. The repository describes itself as: ZhiJuTong (ZJT) is an AI-powered, open-source platform specifically designed for creating professional short dramas. It automates the entire production pipeline, from script and….

When your agent uses it

  • Media & Creative work in your project

Example prompts

  • “/image-understanding”

Requirements

  • Pre-approved tools (allowed-tools): fetch_image_as_base64, ask_user

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. fetch_image_as_base64(image_url) — 获取图片数据(分析任何图片前的必经步骤)
  2. ask_user(question, options) — 向用户提问

What it can do on your machine

Read from SKILL.md and the folder at commit 5cc1f0c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • fetch_image_as_base64
    • ask_user

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image Understanding loads about 701 tokens when it runs. Until then it costs about 17 tokens; SKILL.md has 156 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~17
When it runs · the whole SKILL.md, loaded when a task matches
~701

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 156 words (~701 tokens).

name
image-understanding
allowed-tools
fetch_image_as_base64, ask_user

Read the full SKILL.md on GitHub

Files

Just SKILL.md in agents/skills/image-understanding of jeffstric/ZJT.

Open the folder on GitHubat commit 5cc1f0c

Compare with similar skills

Image Understanding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image Understanding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image Understanding this skilljeffstric/ZJT227—~701Automated safety check: PassCustom licence
Guizang Social Cardsop7418/guizang-social-card-skill7.4k1 repos~7.8kAutomated safety check: PassAGPL-3.0
Weekly Changelog Videoheygen-com/hyperframes59k—~3.3kAutomated safety check: PassApache-2.0
Anthropic Brand Stylinganthropics/skills180k30 repos~559Automated safety check: PassApache-2.0
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo129k—~2.1kAutomated safety check: WarnMIT
Brag Slim Launch Video Makerlatent-spaces/brag14k1 repos~1.9kAutomated safety check: PassMIT

Similar skills

  • Guizang Social Cards

    op7418/guizang-social-card-skill

    Produces social card sets for Xiaohongshu and WeChat: carousels, Live Photo motion cards and puzzle layouts, and WeChat cover pairs, rendered from single-file HTML.

    7.4k GitHub starsUsed in 1 repo~7.8k tokens
    Media & CreativeAuto-check passed
  • Weekly Changelog Video

    heygen-com/hyperframes

    Turns a weekly changelog markdown file into a branded HyperFrames video with voiceover, animated mock-UI scenes and captions, using fonts, background and scripts bundled in the skill.

    59k GitHub stars~3.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Anthropic Brand Styling

    anthropics/skills

    Official

    Applies Anthropic's brand colors and fonts to artifacts such as PowerPoint slides, using fixed hex values for text and accents, Poppins headings and Lora body text.

    180k GitHub starsUsed in 30 repos~559 tokens
    Media & CreativeAuto-check passed
  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • Builds a short, shareable launch video with music and motion from a project directory or a website URL, using only tools already on the machine.

    14k GitHub starsUsed in 1 repo~1.9k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    59k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed

More from jeffstric/ZJT

All 19 skills in this repo
  • A skill your agent uses when Claude Code needs to operate this project's storyboard automation API or CLI as an agent: exchange agent tokens, discover worlds/scripts/characters/locations/props, call…

    227 GitHub stars~2.8k tokensUpdated 17 days ago
    Auto-check: notes
  • A skill your agent uses when Codex needs to operate this project's storyboard automation API or CLI as an agent: exchange an agent token for authtoken, discover…

    227 GitHub stars~6.6k tokensUpdated 17 days ago
    Auto-check passed
  • Add LLM Model

    jeffstric/ZJT

    新增 LLM 模型专家,指导如何在本项目中添加新的 LLM 供应商和模型。当需要接入新的大语言模型(如 OpenAI、Codex、DeepSeek 等)时使用。

    227 GitHub stars~2.3k tokensUpdated 17 days ago
    Auto-check passed
  • 资产就绪检查专家,负责检查角色(参考图和音色)、场景(参考图)、道具(参考图)的完备性和参考图内容质量(宫格切分污染检测),以及世界画风和构图设定的合理性与精简性。当需要确认所有资产是否准备就绪、是否可以进入制作工坊时使用。

    227 GitHub stars~2k tokensUpdated 17 days ago
    Auto-check passed
  • Character Creator

    jeffstric/ZJT

    角色创建师,基于剧情生成详细的人物角色卡,包括性格、习惯、人物关系网。当需要创建或完善角色信息时使用. An agent skill from jeffstric/ZJT.

    227 GitHub stars~1.5k tokensUpdated 17 days ago
    Auto-check passed
  • 剧本检查师,审核剧本内容是否符合内容规范,检查真实国家名称、黄赌毒、极端封建迷信等违规内容,同时检查角色一致性、大纲一致性、以及每集末尾是否有吸引人的钩子。当需要审核剧本合规性和质量时使用。仅审核任务上下文中的当前剧本,审核完成并记录结果后即结束。

    227 GitHub stars~4k tokensUpdated 17 days ago
    Auto-check passed

Questions about Image Understanding

What does Image Understanding do?

图片理解智能体,负责识别图片内容、反推提示词、分析图片风格、对比图片差异以及文字识别(OCR). An agent skill from jeffstric/ZJT. Image Understanding is an agent skill from jeffstric/ZJT.

When should I use Image Understanding?

Image Understanding fits situations like: media & Creative work in your project.

How do I install Image Understanding in Claude Code?

Run `npx skills add jeffstric/ZJT --skill image-understanding -a claude-code`. Or copy the skill folder (agents/skills/image-understanding in jeffstric/ZJT) into .claude/skills/image-understanding in your project. Claude Code loads it when a task matches its description.

How do I install Image Understanding in Codex?

Run `npx skills add jeffstric/ZJT --skill image-understanding -a codex`. Or copy the skill folder (agents/skills/image-understanding in jeffstric/ZJT) into .agents/skills/image-understanding in your project. Codex loads it when a task matches its description.

Can I use Image Understanding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeffstric/ZJT --skill image-understanding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-understanding, .gemini/skills/image-understanding, .github/skills/image-understanding and .opencode/skills/image-understanding in your project.

What does Image Understanding need to run?

SKILL.md names no scripts, command-line tools or credentials: Image Understanding is instructions for the agent only. Its frontmatter pre-approves these tools: fetch_image_as_base64, ask_user.

Does Image Understanding access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image Understanding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Image Understanding use?

Image Understanding has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Image Understanding use?

About 701 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image Understanding?

Skills that share tags, products or a category with Image Understanding: Guizang Social Cards (op7418/guizang-social-card-skill, 7.4k stars), Weekly Changelog Video (heygen-com/hyperframes, 59k stars), Anthropic Brand Styling (anthropics/skills, 180k stars) and MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 129k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image Understanding?

jeffstric (a GitHub user) maintains it in jeffstric/ZJT, which has 227 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 22, 2026.

Source: jeffstric/ZJT on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.