Agent skill

Deepseek Vision Bridge

by aiskillstore in aiskillstore/marketplace

Deploy and maintain image understanding (OCR + local VLM + cloud VL) and image generation for Codex connected to text-only models like DeepSeek.

MITAuto-check: notesMedia & Creative

Install Deepseek Vision Bridge

skills CLI
$ npx skills add aiskillstore/marketplace --skill deepseek-vision-bridge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aiskillstore/marketplace deepseek-vision-bridge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aiskillstore/marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/lightlw/deepseek-vision-bridge .claude/skills/deepseek-vision-bridge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deepseek-vision-bridge
GitHub stars
433
Token cost
~814 tokens
SKILL.md length
149 words
Files
20 (incl. scripts, references, assets)
Skills in repo
1,044
Repo updated
First seen
Licence
MIT

At a glance

Deploy and maintain image understanding (OCR + local VLM + cloud VL) and image generation for Codex connected to text-only models like DeepSeek.

  • Works in 7 steps: 环境自检 → 确定部署位置与路线 → 复制文件并填配置 → …
  • The user wants to see/read images in the conversation bar with a text-only model
  • SKILL.md covers 部署流程, 故障排查, 引擎选型与架构 and 重要原则
  • Runs JavaScript, PowerShell and Shell scripts from its folder; calls python, ollama and npm; reaches ollama.com; needs DEEPSEEK_API_KEY and SILICONFLOW_API_KEY

What it does

Deepseek Vision Bridge is an agent skill from aiskillstore/marketplace. Deploy and maintain image understanding (OCR + local VLM + cloud VL) and image generation for Codex connected to text-only models like DeepSeek. Use when the user wants to see/read images in the conversation bar with a text-only model, set up automatic OCR or vision-model processing for pasted images, configure image generation via SiliconFlow or local Stable Diffusion/Flux, fix a broken vision proxy, or asks "看图/生图/贴图识别/图片理解/OCR代理" on a Codex+DeepSeek setup. Deploys a local proxy between Codex and the model that…

Its SKILL.md is about 810 tokens, which your agent loads only when the skill is triggered. The skill folder holds 22 other files, including scripts, reference files and assets (for example `README.md`, `agents/openai.yaml` and `assets/image-gen.js`).

It sits in Media & Creative, covering Image generation, Diffusion and image models and Computer vision. It works with DeepSeek, Stable Diffusion and Ollama. The repository describes itself as: Security-audited skills for Claude, Codex & Claude Code. One-click install, quality verified. The licence is MIT.

When your agent uses it

  • The user wants to see/read images in the conversation bar with a text-only model
  • Set up automatic OCR
  • Vision-model processing for pasted images
  • Configure image generation via SiliconFlow

Example prompts

  • “看图/生图/贴图识别/图片理解/OCR代理”
  • “/deepseek-vision-bridge”

Requirements

  • Python 3
  • Node.js
  • A Bash shell
  • PowerShell
  • A credential in DEEPSEEK_API_KEY

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 环境自检
  2. 确定部署位置与路线
  3. 复制文件并填配置
  4. 安装依赖并启动代理
  5. 配置 Codex 指向代理
  6. 设置开机自启(Windows 可选)
  7. 验证

What it can do on your machine

Read from SKILL.md and the folder at commit 44923f3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript, PowerShell and Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • ollama
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ollama.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DEEPSEEK_API_KEY
    • SILICONFLOW_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deepseek Vision Bridge loads about 814 tokens when it runs, and up to ~2.9k if it reads all its reference files. Until then it costs about 146 tokens; SKILL.md has 149 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~814
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:44
    ├── .env                  ← 从 assets/.env.template 复制并填写
  • NoteMentions a .env fileSKILL.md:57
    复制 assets 下文件到代理目录,然后编辑 `.env`:
  • NoteMentions a .env fileSKILL.md:131
    - 每台机器的 key、端口、模型名不同,以用户机器的 .env / config.toml 为准

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aiskillstore/marketplace at commit 44923f3, republished under its MIT licence (© aiskillstore). 149 words, ~814 tokens.

Download SKILL.mdSave it as .claude/skills/deepseek-vision-bridge/SKILL.md (or your agent's skills folder). This skill also uses 19 other files; get the full folder from GitHub.
name
deepseek-vision-bridge
description
Deploy and maintain image understanding (OCR + local VLM + cloud VL) and image generation for Codex connected to text-only models like DeepSeek. Use when the user wants to see/read images in the conversation bar with a text-only model, set up automatic OCR or vision-model processing for pasted images, configure image generation via SiliconFlow or local Stable Diffusion/Flux, fix a broken vision proxy, or asks "看图/生图/贴图识别/图片理解/OCR代理" on a Codex+DeepSeek setup. Deploys a local proxy between Codex and the model that converts images to text before forwarding.

DeepSeek Vision Bridge

为 Codex + 纯文本模型(DeepSeek)部署"看图 + 生图"能力。

核心机制:本地代理拦截 Codex 请求中的图片,用三级引擎(OCR / 本地VLM / 云端VL) 转成文字后再转发给 DeepSeek;生图由 image-gen.js 完成后返回文件路径。

部署流程

Step 1: 环境自检

运行 scripts/check_env.py(需要 Python 3),得到 JSON 报告:

bash
python scripts/check_env.py

关注字段:node、ollama_installed、ollama_models、gpu、 proxy_port_open、codex_config、deepseek_key_configured。

Step 2: 确定部署位置与路线

选定代理目录(例如 ~/codex-vision-bridge/ 或用户偏好位置),创建:

text
<代理目录>/
├── ocr-proxy.js          ← 从 assets/ 复制
├── image-gen.js          ← 从 assets/ 复制
├── config.json           ← 从 assets/config.json.template 复制
├── package.json          ← 从 assets/ 复制
├── .env                  ← 从 assets/.env.template 复制并填写
├── start-proxy.ps1 / .sh ← 从 assets/ 复制(按平台)
└── node_modules/         ← npm install 生成

按 references/model-guide.md 与用户确认路线:

  • 注册了硅基流动 key → 云端 VL 看图 + fast 生图
  • 有 GPU 且愿意装模型 → Ollama 本地 VLM + 本地 SDXL/FLUX 生图
  • 都不愿意 → 仅 OCR(纯文字图可用)
Step 3: 复制文件并填配置

复制 assets 下文件到代理目录,然后编辑 .env:

bash
DEEPSEEK_API_KEY=sk-...          # 必填,DeepSeek 平台获取
SILICONFLOW_API_KEY=sk-...       # 可选,硅基流动获取
LOCAL_VL_MODEL=minicpm-v:8b      # 可选,本地 VLM 模型名
OCR_PROXY_PORT=57323             # 默认

若用户选本地 VLM,先安装 Ollama 并拉模型:

bash
# Windows/macOS: https://ollama.com/download
ollama pull minicpm-v:8b
Step 4: 安装依赖并启动代理
bash
cd <代理目录>
npm install        # 安装 tesseract.js
./start-proxy.sh   # macOS/Linux
# 或 powershell -ExecutionPolicy Bypass -File start-proxy.ps1  # Windows

验证端口监听:57323(或自定义端口)。启动日志在 <代理目录>/outputs/proxy.log。

Step 5: 配置 Codex 指向代理

修改 ~/.codex/config.toml(Windows 为 %USERPROFILE%\.codex\config.toml):

toml
model = "deepseek-v4-flash"
model_provider = "custom"

[model_providers.custom]
name = "custom"
wire_api = "responses"
requires_openai_auth = true
base_url = "http://127.0.0.1:57323/v1"
approvals_reviewer = "user"

关键:model_provider = "custom" + base_url 指向代理端口。 改配置前先备份原文件。注意 config.toml 用 UTF-8 无 BOM。

Step 6: 设置开机自启(Windows 可选)
powershell
powershell -ExecutionPolicy Bypass -File install-autostart.ps1
Step 7: 验证
  1. 重启 Codex 会话(让 config.toml 生效)
  2. 对话栏贴一张图 → 应看到模型基于图片内容回答
  3. 查看代理日志确认路由(OCR / 云端 VL / 本地 VLM)
  4. 要求生图 → 应看到 image-gen.js 生成的图片
  5. 若纯文字截图 → 提示词带"提取文字",验证走 OCR

故障排查

见 references/troubleshooting.md,按症状查: 端口未监听 / 400 / 401 / 超时 / OCR 乱码 / Ollama 连接失败 / 生图失败。

引擎选型与架构

  • 引擎对比与组合策略:references/model-guide.md
  • 请求链路与缓存设计:references/architecture.md

重要原则

  • 图片二进制不进上下文:代理转文字、生图返回路径,禁止 Base64 内联
  • 纯文字图优先 OCR;图文/纯图走 VLM;云端失败自动降级本地
  • 每台机器的 key、端口、模型名不同,以用户机器的 .env / config.toml 为准

© aiskillstore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 19 other files (scripts, references, assets) in skills/lightlw/deepseek-vision-bridge of aiskillstore/marketplace.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.md
  • agents/openai.yaml
  • assets/.env.template
  • assets/config.json.template
  • assets/image-gen.js
  • assets/install-autostart.ps1
  • assets/ocr-proxy.js
  • assets/package.json
  • assets/start-proxy.ps1
  • assets/start-proxy.sh
  • references/architecture.md
  • references/model-guide.md
  • references/multi-agent.md
  • references/troubleshooting.md
  • scripts
  • … and 2 more

Open the folder on GitHubat commit 44923f3

Compare with similar skills

Deepseek Vision Bridge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deepseek Vision Bridge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deepseek Vision Bridge this skillaiskillstore/marketplace433—~814Automated safety check: NotesMIT
Iibzanllp/infinite-image-browsing1.4k—~3.3kAutomated safety check: PassMIT
Stable Diffusion with DiffusersOrchestra-Research/AI-Research-SKILLs13k5 repos~3.2kAutomated safety check: PassMIT
Anima Baseartokun/comfyui-mcp803—~4kAutomated safety check: PassMIT
Imageguaardvark/guaardvark257—~1.8kAutomated safety check: PassMIT
AI Image Prompts SkillLeoYeAI/openclaw-master-skills2.2k—~4.3kAutomated safety check: PassMIT

Similar skills

  • Iib

    zanllp/infinite-image-browsing

    Interact with IIB (Infinite Image Browsing) service for searching, browsing, tagging, and organizing AI-generated images.

    1.4k GitHub stars~3.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Stable Diffusion with Diffusers

    Orchestra-Research/AI-Research-SKILLs

    Generates and edits images with Stable Diffusion through Hugging Face Diffusers, covering text-to-image, image-to-image, inpainting, SDXL and custom pipelines.

    13k GitHub starsUsed in 5 repos~3.2k tokens
    Media & CreativeAuto-check passed
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    803 GitHub stars~4k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Image

    guaardvark/guaardvark

    Generate or edit images on the user's own GPU through Guaardvark: single images, instruction edits, background cut-outs, inpaint and outpaint, consistent characters from the Cast Library, and batch…

    257 GitHub stars~1.8k tokensUpdated today
    Media & CreativeAuto-check passed
  • AI Image Prompts Skill

    LeoYeAI/openclaw-master-skills

    Recommend curated prompts from a 10,000+ real-world image generation prompt library.

    2.2k GitHub stars~4.3k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Comfyui Skill Openclaw

    HuangYuChuh/ComfyUI_Skills_OpenClaw

    Run registered ComfyUI workflows through the fast comfyui-skill CLI, and use the official local Comfy MCP for live template, node, model, validation, and orchestration capabilities.

    413 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from aiskillstore/marketplace

All 1,044 skills in this repo
  • Code Stats

    aiskillstore/marketplace

    Analyze codebase with tokei (fast line counts by language) and difft (semantic AST-aware diffs).

    433 GitHub starsUsed in 1 repo~697 tokens
    Auto-check: notes
  • Data Processing

    aiskillstore/marketplace

    Process JSON with jq and YAML/TOML with yq. An agent skill from aiskillstore/marketplace.

    433 GitHub starsUsed in 1 repo~720 tokens
    Auto-check: notes
  • Doc Scanner

    aiskillstore/marketplace

    Scans for project documentation files (AGENTS.md, CLAUDE.md, GEMINI.md, COPILOT.md, CURSOR.md, WARP.md, and 15+ other formats) and synthesizes guidance.

    433 GitHub starsUsed in 1 repo~644 tokens
    Auto-check: notes
  • File Search

    aiskillstore/marketplace

    Modern file and content search using fd, ripgrep (rg), and fzf.

    433 GitHub starsUsed in 1 repo~598 tokens
    Auto-check: notes
  • Find Replace

    aiskillstore/marketplace

    Modern find-and-replace using sd (simpler than sed) and batch replacement patterns.

    433 GitHub starsUsed in 1 repo~527 tokens
    Auto-check: notes
  • Project Planner

    aiskillstore/marketplace

    Detects stale project plans and suggests session commands. An agent skill from aiskillstore/marketplace.

    433 GitHub starsUsed in 1 repo~504 tokens
    Auto-check passed

Questions about Deepseek Vision Bridge

What does Deepseek Vision Bridge do?

Deploy and maintain image understanding (OCR + local VLM + cloud VL) and image generation for Codex connected to text-only models like DeepSeek. Deepseek Vision Bridge is an agent skill from aiskillstore/marketplace. Deploy and maintain image understanding (OCR + local VLM + cloud VL) and image generation for Codex connected to text-only models like DeepSeek.

When should I use Deepseek Vision Bridge?

Deepseek Vision Bridge fits situations like: the user wants to see/read images in the conversation bar with a text-only model; set up automatic OCR; vision-model processing for pasted images; configure image generation via SiliconFlow.

How do I install Deepseek Vision Bridge in Claude Code?

Run `npx skills add aiskillstore/marketplace --skill deepseek-vision-bridge -a claude-code`. Or copy the skill folder (skills/lightlw/deepseek-vision-bridge in aiskillstore/marketplace) into .claude/skills/deepseek-vision-bridge in your project. Claude Code loads it when a task matches its description.

How do I install Deepseek Vision Bridge in Codex?

Run `npx skills add aiskillstore/marketplace --skill deepseek-vision-bridge -a codex`. Or copy the skill folder (skills/lightlw/deepseek-vision-bridge in aiskillstore/marketplace) into .agents/skills/deepseek-vision-bridge in your project. Codex loads it when a task matches its description.

Can I use Deepseek Vision Bridge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aiskillstore/marketplace --skill deepseek-vision-bridge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepseek-vision-bridge, .gemini/skills/deepseek-vision-bridge, .github/skills/deepseek-vision-bridge and .opencode/skills/deepseek-vision-bridge in your project.

What does Deepseek Vision Bridge need to run?

Going by SKILL.md and its folder, Deepseek Vision Bridge needs JavaScript, PowerShell and a shell for the scripts in its folder, the command-line tools its instructions call (python, ollama and npm) and credentials named DEEPSEEK_API_KEY and SILICONFLOW_API_KEY. Our summary lists: Python 3; Node.js; A Bash shell; PowerShell; A credential in DEEPSEEK_API_KEY.

Does Deepseek Vision Bridge access the network?

SKILL.md names 1 domain. In commands or code: ollama.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Deepseek Vision Bridge safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Deepseek Vision Bridge use?

Deepseek Vision Bridge is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deepseek Vision Bridge use?

About 814 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Deepseek Vision Bridge?

Skills that share tags, products or a category with Deepseek Vision Bridge: Iib (zanllp/infinite-image-browsing, 1.4k stars), Stable Diffusion with Diffusers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Anima Base (artokun/comfyui-mcp, 803 stars) and Image (guaardvark/guaardvark, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deepseek Vision Bridge?

aiskillstore (a GitHub organization) maintains it in aiskillstore/marketplace, which has 433 GitHub stars. The repository holds 1,044 skills in this directory. The repository was last updated on October 10, 2026.

Source: aiskillstore/marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.