Agent skill

Ds Vision Skill

by Sorwcyra in Sorwcyra/ds-vision-skill

为纯文本推理模型补充视觉能力。用户提供图片、截图、照片、图表、UI 截图、代码截图、数学题图片、 扫描件、PDF 或文档,并要求描述、理解、推理、阅读、OCR、提取文字、解析图表或分析内容时使用。

MITAuto-check passedDocuments & Office

Install Ds Vision Skill

skills CLI
$ npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Sorwcyra/ds-vision-skill ds-vision-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ds-vision-skill
GitHub stars
163
Token cost
~805 tokens
SKILL.md length
166 words
Files
37 (incl. scripts, references, assets)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

为纯文本推理模型补充视觉能力。用户提供图片、截图、照片、图表、UI 截图、代码截图、数学题图片、 扫描件、PDF 或文档,并要求描述、理解、推理、阅读、OCR、提取文字、解析图表或分析内容时使用。

  • Works in 4 steps: PDF、论文、报告、长文档、多页扫描件:使用… → 图片且需要理解/推理:使用… → 图片默认进入视觉理解免费竞速池;需要纯文字识别时显式使用 -Intent… → …
  • Documents & Office work in your project
  • SKILL.md covers 首选入口, 路由规则, 降级链 and 输出规范, plus 3 more sections
  • Needs MINERU_TOKEN

What it does

Ds Vision Skill is an agent skill from Sorwcyra/ds-vision-skill. 为纯文本推理模型补充视觉能力。用户提供图片、截图、照片、图表、UI 截图、代码截图、数学题图片、 扫描件、PDF 或文档,并要求描述、理解、推理、阅读、OCR、提取文字、解析图表或分析内容时使用。 默认调用 scripts/vision-router.ps1 做自动路由:图片理解先走免费竞速池 GLM/Agnes,再走 custom-1/custom-2/custom-3/local,文档解析走 MinerU, 纯文字识别走 Baidu OCR 或 Windows OCR。所有工具输出标准 JSON,再交给主模型推理和总结。

Its SKILL.md is about 810 tokens, which your agent loads only when the skill is triggered. The skill folder holds 41 other files, including scripts, reference files and assets (for example `.github/workflows/star-history.yml`, `README.md` and `README.zh-CN.md`).

It sits in Documents & Office. It works with PowerShell. The repository describes itself as: ds-vision-skill: Vision skill suite to add multimodal image capabilities for DeepSeek, and other text-only models as well. The licence is MIT.

When your agent uses it

  • Documents & Office work in your project

Example prompts

  • “/ds-vision-skill”

Requirements

  • A credential in YOUR_KEY
  • A credential in MINERU_TOKEN

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. PDF、论文、报告、长文档、多页扫描件:使用 scripts/mineru-extract.ps1 -FilePath -Mode flash -Json。如果配置了 MINERU_TOKEN 且 flash 失败,再尝试 -Mode extract。
  2. 图片且需要理解/推理:使用 scripts/vision-router.ps1,让 agnes-2.5-flash、agnes-2.0-flash、glm、glm-thinking 四个模型同时开始竞速;谁先成功返回就采用谁的结果。如果全部失败,再降级到…
  3. 图片默认进入视觉理解免费竞速池;需要纯文字识别时显式使用 -Intent ocr,优先 scripts/baidu-ocr.ps1 -ImagePath -Json;未配置或失败时用 scripts/windows-ocr.ps1 -ImagePath -Json。
  4. 无法判断时:使用 vision-router.ps1 -Intent auto -Complex -Json。

What it can do on your machine

Read from SKILL.md and the folder at commit e546b60. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MINERU_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ds Vision Skill loads about 805 tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 166 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~805
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Sorwcyra/ds-vision-skill at commit e546b60, republished under its MIT licence (© Sorwcyra). 166 words, ~805 tokens.

Download SKILL.mdSave it as .claude/skills/ds-vision-skill/SKILL.md (or your agent's skills folder). This skill also uses 36 other files; get the full folder from GitHub.
name
ds-vision-skill
description
为纯文本推理模型补充视觉能力。用户提供图片、截图、照片、图表、UI 截图、代码截图、数学题图片、 扫描件、PDF 或文档,并要求描述、理解、推理、阅读、OCR、提取文字、解析图表或分析内容时使用。 默认调用 scripts/vision-router.ps1 做自动路由:图片理解先走免费竞速池 GLM/Agnes,再走 custom-1/custom-2/custom-3/local,文档解析走 MinerU, 纯文字识别走 Baidu OCR 或 Windows OCR。所有工具输出标准 JSON,再交给主模型推理和总结。
metadata.version
0.5.1
metadata.repository
https://github.com/Sorwcyra/ds-vision-skill

DS Vision Skill

这个 skill 负责把视觉输入转换成文本或结构化 JSON。它不替代主模型,只负责识别任务、选择工具、执行视觉/OCR/文档解析,并把结果交给主模型继续推理。

首选入口

优先使用统一路由脚本:

powershell
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts/vision-router.ps1 -Path "path/to/file.png" -Prompt "user request" -Intent auto -Json

如果 harness 默认使用 cmd.exe(例如 Zcode 或某些 Codex/Hermes 包装器),优先调用 scripts/setup.cmd 和 scripts/vision-router.cmd。不要把 PowerShell 专用语法或 <KEY> 占位符粘到 cmd.exe 中;cmd.exe 会把尖括号当成重定向符号,配置 key 时请使用 "YOUR_KEY" 这样的引号占位或真实引号值。

常用参数:

  • -Intent auto|reason|ocr|document:默认 auto。图片默认走视觉理解免费竞速池;纯 OCR 请显式使用 -Intent ocr。
  • -Complex:图表、数学、复杂 UI、代码截图、多步骤视觉推理时启用。
  • -AccurateOcr:票据、扫描件、低清晰度文字识别时启用百度高精度 OCR。
  • -MaxTokens:限制视觉模型输出长度,默认 1024;-Complex 未显式设置时使用 2048。
  • -TimeoutSec:整场视觉竞速的最长等待时间,默认 90 秒。
  • -NoCache:跳过缓存读取,也不写入本次结果。

只有在需要调试单个通道时,才直接调用底层脚本。

路由规则

  1. PDF、论文、报告、长文档、多页扫描件:使用 scripts/mineru-extract.ps1 -FilePath <file> -Mode flash -Json。如果配置了 MINERU_TOKEN 且 flash 失败,再尝试 -Mode extract。
  2. 图片且需要理解/推理:使用 scripts/vision-router.ps1,让 agnes-2.5-flash、agnes-2.0-flash、glm、glm-thinking 四个模型同时开始竞速;谁先成功返回就采用谁的结果。如果全部失败,再降级到 custom-1、custom-2、custom-3 和 local。
  3. 图片默认进入视觉理解免费竞速池;需要纯文字识别时显式使用 -Intent ocr,优先 scripts/baidu-ocr.ps1 -ImagePath <file> -Json;未配置或失败时用 scripts/windows-ocr.ps1 -ImagePath <file> -Json。
  4. 无法判断时:使用 vision-router.ps1 -Intent auto -Complex -Json。

降级链

  • 视觉理解:race(agnes-2.5-flash, agnes-2.0-flash, glm, glm-thinking) -> custom-1 -> custom-2 -> custom-3 -> local。
  • 文档解析:mineru flash -> mineru extract。
  • OCR:baidu-ocr -> windows-ocr -> vision reasoning。
  • 同一通道遇到 401、403、429、网络错误或空结果时,不要反复重试;直接切换下一通道。

输出规范

所有脚本在 -Json 模式下输出:

json
{
  "task_type": "image_reasoning | document_parsing | ocr",
  "tool_used": "actual tool or model",
  "confidence": "high | medium | low",
  "result": "recognized or parsed content",
  "metadata": {}
}

主模型继续推理时,优先使用 result 字段。向用户报告时可简要说明 tool_used 和必要的降级过程。

预检

只在首次配置、诊断问题或所有通道失败时运行;正常执行不要在每次分析前运行:

powershell
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts/preflight.ps1
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts/preflight.ps1 -Json

-Json 用于自动化读取通道、工具和本地运行时状态。

隐私

云端通道会把文件内容发送给对应服务商。用户明确关注隐私、合同、证件、医疗、财务等敏感内容时,优先使用 Windows OCR、本地模型或先征求确认。

维护约定

  • PowerShell 脚本源码保持 ASCII-only,中文通过参数传入。
  • 面向用户的 Markdown 文档使用 UTF-8。
  • 面向 harness 或新用户的可复制 Windows 命令必须显式选择 shell:配置用 scripts/setup.cmd,路由用 scripts/vision-router.cmd,或完整写出 powershell.exe -NoProfile -ExecutionPolicy Bypass -File ...。不要在 cmd 示例中使用 <KEY> 形式的占位符。
  • 新增通道时优先接入 vision-router.ps1,再补充 README 和 references/channels.md。
  • 评估性能或发布版本时先阅读跨版本基准方法与数据,并用 scripts/benchmark-race.ps1 复现;不要仅凭单次 Live 请求下结论。

© Sorwcyra, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 36 other files (scripts, references, assets) in the repository root of Sorwcyra/ds-vision-skill.

  • SKILL.md
  • .gitattributes
  • .github/workflows/star-history.yml
  • .gitignore
  • LICENSE
  • README.md
  • README.zh-CN.md
  • VERSION
  • agents/openai.yaml
  • assets/.star-history.sig
  • assets/benchmark-speed.svg
  • assets/star-history-dark.svg
  • assets/star-history-light.svg
  • assets/star-history.png
  • benchmarks/cross-version-live-2026-08-11.json
  • benchmarks/cross-version-mock-2026-08-11.json
  • … and 21 more

Open the folder on GitHubat commit e546b60

Compare with similar skills

Ds Vision Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ds Vision Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ds Vision Skill this skillSorwcyra/ds-vision-skill163—~805Automated safety check: PassMIT
Repository WorkflowAliceJump/ok-gf2274—~1.4kAutomated safety check: PassNone
Implementing Data Loss Prevention With Microsoft Purviewmukul975/Anthropic-Cybersecurity-Skills34k—~7.9kAutomated safety check: PassApache-2.0
CLI Microsoft365 Scriptpnp/cli-microsoft365-mcp-server132—~3.3kAutomated safety check: PassMIT
Entra Agent Usergithub/awesome-copilot40k1 repos~2.3kAutomated safety check: PassMIT
Ms365 Tenant Managerborghei/Claude-Skills891—~1.8kAutomated safety check: PassMIT

Similar skills

  • Repository Workflow

    AliceJump/ok-gf2

    Apply ok-gf2 repository-wide engineering rules. An agent skill from AliceJump/ok-gf2.

    274 GitHub stars~1.4k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Implementing Data Loss Prevention With Microsoft Purview

    mukul975/Anthropic-Cybersecurity-Skills

    Implements DLP policies using Microsoft Purview PowerShell cmdlets and the Graph API to protect data across Exchange Online, SharePoint, OneDrive, Teams, endpoints, and Power BI, including…

    34k GitHub stars~7.9k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • CLI Microsoft365 Script

    pnp/cli-microsoft365-mcp-server

    Write PowerShell scripts using CLI for Microsoft 365 commands to automate Microsoft 365 management tasks.

    132 GitHub stars~3.3k tokensUpdated 5 days ago
    Documents & OfficeAuto-check passed
  • Entra Agent User

    github/awesome-copilot

    Official

    Create Agent Users in Microsoft Entra ID from Agent Identities, enabling AI agents to act as digital workers with user identity capabilities in Microsoft 365 and Azure environments.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Documents & OfficeAuto-check passed
  • Ms365 Tenant Manager

    borghei/Claude-Skills

    Microsoft 365 tenant administration for Global Administrators.

    891 GitHub stars~1.8k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • Playwright

    6BNBN/FlowPilot

    A skill your agent uses when the task requires automating a real browser from the terminal in Cursor, including navigation, form filling, snapshots, screenshots, data extraction, and UI-flow…

    134 GitHub stars~1k tokensUpdated 7 mo ago
    Testing & QAAuto-check passed

Works with

Questions about Ds Vision Skill

What does Ds Vision Skill do?

为纯文本推理模型补充视觉能力。用户提供图片、截图、照片、图表、UI 截图、代码截图、数学题图片、 扫描件、PDF 或文档,并要求描述、理解、推理、阅读、OCR、提取文字、解析图表或分析内容时使用。. Ds Vision Skill is an agent skill from Sorwcyra/ds-vision-skill.

When should I use Ds Vision Skill?

Ds Vision Skill fits situations like: documents & Office work in your project.

How do I install Ds Vision Skill in Claude Code?

Run `npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a claude-code`. Or copy the skill folder (the Sorwcyra/ds-vision-skill repository) into .claude/skills/ds-vision-skill in your project. Claude Code loads it when a task matches its description.

How do I install Ds Vision Skill in Codex?

Run `npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a codex`. Or copy the skill folder (the Sorwcyra/ds-vision-skill repository) into .agents/skills/ds-vision-skill in your project. Codex loads it when a task matches its description.

Can I use Ds Vision Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ds-vision-skill, .gemini/skills/ds-vision-skill, .github/skills/ds-vision-skill and .opencode/skills/ds-vision-skill in your project.

What does Ds Vision Skill need to run?

Going by SKILL.md and its folder, Ds Vision Skill needs credentials named MINERU_TOKEN. Our summary lists: A credential in YOUR_KEY; A credential in MINERU_TOKEN.

Does Ds Vision Skill access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ds Vision Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ds Vision Skill use?

Ds Vision Skill is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ds Vision Skill use?

About 805 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.4k tokens, read only when the agent opens those files.

What are the alternatives to Ds Vision Skill?

Skills that share tags, products or a category with Ds Vision Skill: Repository Workflow (AliceJump/ok-gf2, 274 stars), Implementing Data Loss Prevention With Microsoft Purview (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), CLI Microsoft365 Script (pnp/cli-microsoft365-mcp-server, 132 stars) and Entra Agent User (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ds Vision Skill?

Sorwcyra (a GitHub user) maintains it in Sorwcyra/ds-vision-skill, which has 163 GitHub stars. The repository was last updated on October 10, 2026.

Source: Sorwcyra/ds-vision-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.