Repository Workflow
AliceJump/ok-gf2
Apply ok-gf2 repository-wide engineering rules. An agent skill from AliceJump/ok-gf2.
为纯文本推理模型补充视觉能力。用户提供图片、截图、照片、图表、UI 截图、代码截图、数学题图片、 扫描件、PDF 或文档,并要求描述、理解、推理、阅读、OCR、提取文字、解析图表或分析内容时使用。
$ npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Sorwcyra/ds-vision-skill ds-vision-skill --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ds-vision-skill" agent skill from https://github.com/Sorwcyra/ds-vision-skill/tree/main into .claude/skills/ds-vision-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ds-vision-skill", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Sorwcyra/ds-vision-skill ds-vision-skill --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ds-vision-skill" agent skill from https://github.com/Sorwcyra/ds-vision-skill/tree/main into .agents/skills/ds-vision-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ds-vision-skill", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Sorwcyra/ds-vision-skill ds-vision-skill --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ds-vision-skill" agent skill from https://github.com/Sorwcyra/ds-vision-skill/tree/main into .cursor/skills/ds-vision-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ds-vision-skill", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Sorwcyra/ds-vision-skill ds-vision-skill --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ds-vision-skill" agent skill from https://github.com/Sorwcyra/ds-vision-skill/tree/main into .gemini/skills/ds-vision-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ds-vision-skill", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Sorwcyra/ds-vision-skill ds-vision-skillInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ds-vision-skill" agent skill from https://github.com/Sorwcyra/ds-vision-skill/tree/main into .github/skills/ds-vision-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ds-vision-skill", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Sorwcyra/ds-vision-skill ds-vision-skill --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ds-vision-skill" agent skill from https://github.com/Sorwcyra/ds-vision-skill/tree/main into .opencode/skills/ds-vision-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ds-vision-skill", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ds-vision-skill为纯文本推理模型补充视觉能力。用户提供图片、截图、照片、图表、UI 截图、代码截图、数学题图片、 扫描件、PDF 或文档,并要求描述、理解、推理、阅读、OCR、提取文字、解析图表或分析内容时使用。
Ds Vision Skill is an agent skill from Sorwcyra/ds-vision-skill. 为纯文本推理模型补充视觉能力。用户提供图片、截图、照片、图表、UI 截图、代码截图、数学题图片、 扫描件、PDF 或文档,并要求描述、理解、推理、阅读、OCR、提取文字、解析图表或分析内容时使用。 默认调用 scripts/vision-router.ps1 做自动路由:图片理解先走免费竞速池 GLM/Agnes,再走 custom-1/custom-2/custom-3/local,文档解析走 MinerU, 纯文字识别走 Baidu OCR 或 Windows OCR。所有工具输出标准 JSON,再交给主模型推理和总结。
Its SKILL.md is about 810 tokens, which your agent loads only when the skill is triggered. The skill folder holds 41 other files, including scripts, reference files and assets (for example `.github/workflows/star-history.yml`, `README.md` and `README.zh-CN.md`).
It sits in Documents & Office. It works with PowerShell. The repository describes itself as: ds-vision-skill: Vision skill suite to add multimodal image capabilities for DeepSeek, and other text-only models as well. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e546b60. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
MINERU_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ds Vision Skill loads about 805 tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 166 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from Sorwcyra/ds-vision-skill at commit e546b60, republished under its MIT licence (© Sorwcyra). 166 words, ~805 tokens.
.claude/skills/ds-vision-skill/SKILL.md (or your agent's skills folder). This skill also uses 36 other files; get the full folder from GitHub.这个 skill 负责把视觉输入转换成文本或结构化 JSON。它不替代主模型,只负责识别任务、选择工具、执行视觉/OCR/文档解析,并把结果交给主模型继续推理。
优先使用统一路由脚本:
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts/vision-router.ps1 -Path "path/to/file.png" -Prompt "user request" -Intent auto -Json如果 harness 默认使用 cmd.exe(例如 Zcode 或某些 Codex/Hermes 包装器),优先调用 scripts/setup.cmd 和 scripts/vision-router.cmd。不要把 PowerShell 专用语法或 <KEY> 占位符粘到 cmd.exe 中;cmd.exe 会把尖括号当成重定向符号,配置 key 时请使用 "YOUR_KEY" 这样的引号占位或真实引号值。
常用参数:
-Intent auto|reason|ocr|document:默认 auto。图片默认走视觉理解免费竞速池;纯 OCR 请显式使用 -Intent ocr。-Complex:图表、数学、复杂 UI、代码截图、多步骤视觉推理时启用。-AccurateOcr:票据、扫描件、低清晰度文字识别时启用百度高精度 OCR。-MaxTokens:限制视觉模型输出长度,默认 1024;-Complex 未显式设置时使用 2048。-TimeoutSec:整场视觉竞速的最长等待时间,默认 90 秒。-NoCache:跳过缓存读取,也不写入本次结果。只有在需要调试单个通道时,才直接调用底层脚本。
scripts/mineru-extract.ps1 -FilePath <file> -Mode flash -Json。如果配置了 MINERU_TOKEN 且 flash 失败,再尝试 -Mode extract。scripts/vision-router.ps1,让 agnes-2.5-flash、agnes-2.0-flash、glm、glm-thinking 四个模型同时开始竞速;谁先成功返回就采用谁的结果。如果全部失败,再降级到 custom-1、custom-2、custom-3 和 local。-Intent ocr,优先 scripts/baidu-ocr.ps1 -ImagePath <file> -Json;未配置或失败时用 scripts/windows-ocr.ps1 -ImagePath <file> -Json。vision-router.ps1 -Intent auto -Complex -Json。race(agnes-2.5-flash, agnes-2.0-flash, glm, glm-thinking) -> custom-1 -> custom-2 -> custom-3 -> local。mineru flash -> mineru extract。baidu-ocr -> windows-ocr -> vision reasoning。所有脚本在 -Json 模式下输出:
{
"task_type": "image_reasoning | document_parsing | ocr",
"tool_used": "actual tool or model",
"confidence": "high | medium | low",
"result": "recognized or parsed content",
"metadata": {}
}主模型继续推理时,优先使用 result 字段。向用户报告时可简要说明 tool_used 和必要的降级过程。
只在首次配置、诊断问题或所有通道失败时运行;正常执行不要在每次分析前运行:
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts/preflight.ps1
powershell.exe -NoProfile -ExecutionPolicy Bypass -File scripts/preflight.ps1 -Json-Json 用于自动化读取通道、工具和本地运行时状态。
云端通道会把文件内容发送给对应服务商。用户明确关注隐私、合同、证件、医疗、财务等敏感内容时,优先使用 Windows OCR、本地模型或先征求确认。
scripts/setup.cmd,路由用 scripts/vision-router.cmd,或完整写出 powershell.exe -NoProfile -ExecutionPolicy Bypass -File ...。不要在 cmd 示例中使用 <KEY> 形式的占位符。vision-router.ps1,再补充 README 和 references/channels.md。scripts/benchmark-race.ps1 复现;不要仅凭单次 Live 请求下结论。© Sorwcyra, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 36 other files (scripts, references, assets) in the repository root of Sorwcyra/ds-vision-skill.
Open the folder on GitHubat commit e546b60
Ds Vision Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ds Vision Skill this skillSorwcyra/ds-vision-skill | 163 | — | ~805 | Automated safety check: Pass | MIT | |
| Repository WorkflowAliceJump/ok-gf2 | 274 | — | ~1.4k | Automated safety check: Pass | None | |
| Implementing Data Loss Prevention With Microsoft Purviewmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~7.9k | Automated safety check: Pass | Apache-2.0 | |
| CLI Microsoft365 Scriptpnp/cli-microsoft365-mcp-server | 132 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Entra Agent Usergithub/awesome-copilot | 40k | 1 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Ms365 Tenant Managerborghei/Claude-Skills | 891 | — | ~1.8k | Automated safety check: Pass | MIT |
AliceJump/ok-gf2
Apply ok-gf2 repository-wide engineering rules. An agent skill from AliceJump/ok-gf2.
mukul975/Anthropic-Cybersecurity-Skills
Implements DLP policies using Microsoft Purview PowerShell cmdlets and the Graph API to protect data across Exchange Online, SharePoint, OneDrive, Teams, endpoints, and Power BI, including…
pnp/cli-microsoft365-mcp-server
Write PowerShell scripts using CLI for Microsoft 365 commands to automate Microsoft 365 management tasks.
github/awesome-copilot
Create Agent Users in Microsoft Entra ID from Agent Identities, enabling AI agents to act as digital workers with user identity capabilities in Microsoft 365 and Azure environments.
borghei/Claude-Skills
Microsoft 365 tenant administration for Global Administrators.
6BNBN/FlowPilot
A skill your agent uses when the task requires automating a real browser from the terminal in Cursor, including navigation, form filling, snapshots, screenshots, data extraction, and UI-flow…
Works with
Categories
为纯文本推理模型补充视觉能力。用户提供图片、截图、照片、图表、UI 截图、代码截图、数学题图片、 扫描件、PDF 或文档,并要求描述、理解、推理、阅读、OCR、提取文字、解析图表或分析内容时使用。. Ds Vision Skill is an agent skill from Sorwcyra/ds-vision-skill.
Ds Vision Skill fits situations like: documents & Office work in your project.
Run `npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a claude-code`. Or copy the skill folder (the Sorwcyra/ds-vision-skill repository) into .claude/skills/ds-vision-skill in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a codex`. Or copy the skill folder (the Sorwcyra/ds-vision-skill repository) into .agents/skills/ds-vision-skill in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Sorwcyra/ds-vision-skill --skill ds-vision-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ds-vision-skill, .gemini/skills/ds-vision-skill, .github/skills/ds-vision-skill and .opencode/skills/ds-vision-skill in your project.
Going by SKILL.md and its folder, Ds Vision Skill needs credentials named MINERU_TOKEN. Our summary lists: A credential in YOUR_KEY; A credential in MINERU_TOKEN.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Ds Vision Skill is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 805 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ds Vision Skill: Repository Workflow (AliceJump/ok-gf2, 274 stars), Implementing Data Loss Prevention With Microsoft Purview (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), CLI Microsoft365 Script (pnp/cli-microsoft365-mcp-server, 132 stars) and Entra Agent User (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Sorwcyra (a GitHub user) maintains it in Sorwcyra/ds-vision-skill, which has 163 GitHub stars. The repository was last updated on October 10, 2026.
Source: Sorwcyra/ds-vision-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.