Agent skill

Qwen Vision

by davepoon in davepoon/buildwithclaude

A skill your agent uses when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare…

MITAuto-check passed

Install Qwen Vision

skills CLI
$ npx skills add davepoon/buildwithclaude --skill qwen-vision -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davepoon/buildwithclaude qwen-vision --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davepoon/buildwithclaude.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/give-claude-eyes/skills/qwen-vision .claude/skills/qwen-vision && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qwen-vision
GitHub stars
3.6k
Token cost
~1.2k tokens
SKILL.md length
464 words
Files
4 (incl. scripts, references)
Skills in repo
247
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare…

  • Works in 2 steps: DASHSCOPE_API_KEY environment variable… → Python 3.9+ with dashscope package…
  • The user asks to analyze video
  • SKILL.md covers How it works, Prerequisites, Basic usage and Parameters, plus 5 more sections
  • Runs Python scripts from its folder; calls python3 and pip; needs DASHSCOPE_API_KEY

What it does

Qwen Vision is an agent skill from davepoon/buildwithclaude. Use when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare videos", "analyze this image", "what's in this screenshot", or when the user provides a video/image file path and expects visual understanding. Also trigger on: "qwen", "video bridge", "multimodal analysis", "motion analysis", "video reference", "video breakdown", "batch classify", or any task requiring understanding of video content that…

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/batch-pattern.md`, `references/prompt-tips.md` and `scripts/qwen_bridge.py`).

It works with Qwen and Python. The repository describes itself as: A single hub to find Claude Skills, Agents, Commands, Hooks, Plugins, and Marketplace collections to extend Claude Code, Claude Desktop, Agent SDK and OpenClaw. The licence is MIT.

When your agent uses it

  • The user asks to analyze video
  • Watch this video
  • What happens in this video
  • Describe this clip

Example prompts

  • “analyze video”
  • “watch this video”
  • “what happens in this video”
  • “/qwen-vision”

Requirements

  • Python 3
  • A credential in DASHSCOPE_API_KEY

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. DASHSCOPE_API_KEY environment variable set (get one at https://dashscope.console.aliyun.com/ or…
  2. Python 3.9+ with dashscope package installed

What it can do on your machine

Read from SKILL.md and the folder at commit 616deb5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • dashscope.console.aliyun.com
    • modelstudio.console.alibabacloud.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DASHSCOPE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Qwen Vision loads about 1.2k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 138 tokens; SKILL.md has 464 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from davepoon/buildwithclaude at commit 616deb5, republished under its MIT licence (© davepoon). 464 words, ~1,203 tokens.

Download SKILL.mdSave it as .claude/skills/qwen-vision/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
qwen-vision
description
Use when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare videos", "analyze this image", "what's in this screenshot", or when the user provides a video/image file path and expects visual understanding. Also trigger on: "qwen", "video bridge", "multimodal analysis", "motion analysis", "video reference", "video breakdown", "batch classify", or any task requiring understanding of video content that Claude cannot do natively.
version
0.1.0

Qwen Vision Bridge

Claude cannot natively understand video. This skill bridges that gap by calling Qwen Omni — a natively multimodal model that processes video with temporal attention (it sees motion, not just individual frames).

The bridge also handles images, useful when you want Qwen's analysis on screenshots, diagrams, or photos.

How it works

A Python script at ${CLAUDE_PLUGIN_ROOT}/skills/qwen-vision/scripts/qwen_bridge.py sends media files to the Qwen API and returns the analysis as text. Call it via Bash.

Prerequisites

The user must have:

  1. DASHSCOPE_API_KEY environment variable set (get one at https://dashscope.console.aliyun.com/ or https://modelstudio.console.alibabacloud.com/)
  2. Python 3.9+ with dashscope package installed

If the user hasn't set up yet, suggest running /qwen-setup first.

Basic usage

bash
python3 "${CLAUDE_PLUGIN_ROOT}/skills/qwen-vision/scripts/qwen_bridge.py" "/path/to/video.mp4" "Describe what happens in this video"

Parameters

FlagDefaultDescription
(positional 1)requiredPath to video or image file
(positional 2)generic promptAnalysis prompt
--fps2.0Frames per second to sample from video. Lower = cheaper, higher = more detail
--modelqwen-omni-plus-latestQwen model to use
--jsonoffOutput as JSON (for parsing)
--contextnonePath to JSON file with previous conversation (multi-turn)
--save-contextnoneSave conversation context for follow-up questions
--system-promptnoneCustom system prompt for Qwen
--prompt-filenoneRead prompt from a file instead of argument

Supported formats

Video: .mp4, .mov, .avi, .mkv, .webm, .flv, .wmv Image: .png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff

Patterns

Single video analysis
bash
python3 "${CLAUDE_PLUGIN_ROOT}/skills/qwen-vision/scripts/qwen_bridge.py" "/path/to/video.mp4" "Describe the character's body movement, poses, and transitions" --fps 2

Parse the text response and use it in your answer to the user.

Batch analysis

When the user has multiple videos to analyze, write a Python script that loops through files and calls the bridge for each one. Use --json flag for machine-readable output. See references/batch-pattern.md for a template.

Multi-turn (follow-up questions)
bash
# First question
python3 "${CLAUDE_PLUGIN_ROOT}/skills/qwen-vision/scripts/qwen_bridge.py" video.mp4 "General analysis" --save-context /tmp/ctx.json

# Follow-up
python3 "${CLAUDE_PLUGIN_ROOT}/skills/qwen-vision/scripts/qwen_bridge.py" video.mp4 "Tell me more about the lighting" --context /tmp/ctx.json
Image analysis

Same script, just pass an image path instead of video:

bash
python3 "${CLAUDE_PLUGIN_ROOT}/skills/qwen-vision/scripts/qwen_bridge.py" "/path/to/screenshot.png" "What UI elements are visible in this screenshot?"
Show full SKILL.md (186 more words)Show less
Cost-saving tips
  • Use --fps 1 for long videos or when fine detail isn't needed
  • Use --fps 0.5 for very long videos (minutes+)
  • For batch jobs, start with --fps 1 and increase only if results are too vague

Error handling

  • If DASHSCOPE_API_KEY is not set, the script exits with a clear error message. Guide the user to set it up.
  • If dashscope is not installed, suggest pip install dashscope.
  • If the API returns an error, the script prints the error code and message. Common issues: invalid key, quota exceeded, unsupported file format.
  • If a video file is too large for the API, suggest lowering --fps or trimming the video first.

What Qwen sees vs what Claude sees

This is important context for the user: Qwen processes video frames with temporal attention — it understands motion, direction, rhythm, and transitions between frames. Claude analyzing individual screenshots cannot do this. When the user needs to understand what happens in a video (not just what a single frame looks like), this bridge is the right tool.

Additional resources

  • references/batch-pattern.md — template for batch video classification
  • references/prompt-tips.md — effective prompts for different analysis types

© davepoon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in plugins/give-claude-eyes/skills/qwen-vision of davepoon/buildwithclaude.

  • SKILL.md
  • references/batch-pattern.md
  • references/prompt-tips.md
  • scripts/qwen_bridge.py

Open the folder on GitHubat commit 616deb5

Compare with similar skills

Qwen Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Qwen Vision compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Qwen Vision this skilldavepoon/buildwithclaude3.6k—~1.2kAutomated safety check: PassMIT
AI Search Hubminsight-ai-info/AI-Search-Hub1.3k—~1.3kAutomated safety check: PassNone
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Pocketmen With Yousix-nut/PocketMen-with-you310—~2.6kAutomated safety check: PassMIT
Codex Panel Reproducelittlepeachs/NaturePanelForge223—~1.2kAutomated safety check: PassMIT
Crossreviewcoddy-project/coddy-agent168—~3.3kAutomated safety check: PassMIT

Similar skills

  • AI Search Hub

    minsight-ai-info/AI-Search-Hub

    Run the AI Search Hub browser automation scripts for Yuanbao, LongCat, Doubao, Qwen, Gemini, Grok, and MiniMax.

    1.3k GitHub stars~1.3k tokensUpdated 5 mo ago
    Productivity & AutomationAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Pocketmen With You

    six-nut/PocketMen-with-you

    Turn 2+ user reference images into a high-fidelity animated Codex companion using PocketMen's own local stack.

    310 GitHub stars~2.6k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Codex Panel Reproduce

    littlepeachs/NaturePanelForge

    A skill your agent uses when asked to reproduce or refine a single scientific figure panel as editable Python/matplotlib code in NaturePanelForge.

    223 GitHub stars~1.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Crossreview

    coddy-project/coddy-agent

    Run when the user invokes /crossreview (or /crossreview:setup to choose the reviewers) or asks for a cross-review, a quorum review or a second opinion from other agents or models: fan a code review…

    168 GitHub stars~3.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Vision

    xiincs/claude-code-vision-skill

    Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.

    170 GitHub stars~1.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from davepoon/buildwithclaude

All 247 skills in this repo
  • iOS Hig Design Guide

    davepoon/buildwithclaude

    Build, update, and apply iOS design specifications using Apple Human Interface Guidelines (HIG) source data.

    3.6k GitHub stars~735 tokensUpdated yesterday
    Auto-check passed
  • Video Downloader

    davepoon/buildwithclaude

    Download YouTube videos with customizable quality and format options.

    3.6k GitHub starsUsed in 1 repo~871 tokens
    Auto-check passed
  • Atlas Cloud Media

    davepoon/buildwithclaude

    Discover Atlas Cloud image and video models, inspect their live schemas, and submit one confirmed media generation request with bounded GET polling.

    3.6k GitHub stars~852 tokensUpdated yesterday
    Auto-check passed
  • Browser Extension Launch

    davepoon/buildwithclaude

    面向没有编程经验的用户,把想法做成可试用的浏览器插件,并完成检查、商店材料、审核提交和上线验证;也用于继续已有插件、排错和发布新版。用户说“帮我做个插件”“把插件上架”“继续我的插件”时使用。普通网站开发、仅查询插件知识不触发。

    3.6k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Slack Gif Creator

    davepoon/buildwithclaude

    Toolkit for creating animated GIFs optimized for Slack, with validators for size constraints and composable animation primitives.

    3.6k GitHub starsUsed in 12 repos~4.3k tokens
    Auto-check passed
  • Hard Predict Future

    davepoon/buildwithclaude

    Activate this agent for any future-oriented question that requires deep quantitative analysis, historical precedents, and structured scenario planning.

    3.6k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Qwen Vision

What does Qwen Vision do?

A skill your agent uses when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare…. Qwen Vision is an agent skill from davepoon/buildwithclaude. Use when the user asks to "analyze video", "watch this video", "what happens in this video", "describe this clip", "review this footage", "classify these videos", "compare videos", "analyze this image", "what's in this screenshot", or when the user provides a video/image file path and expects visual understanding.

When should I use Qwen Vision?

Qwen Vision fits situations like: the user asks to analyze video; watch this video; what happens in this video; describe this clip.

How do I install Qwen Vision in Claude Code?

Run `npx skills add davepoon/buildwithclaude --skill qwen-vision -a claude-code`. Or copy the skill folder (plugins/give-claude-eyes/skills/qwen-vision in davepoon/buildwithclaude) into .claude/skills/qwen-vision in your project. Claude Code loads it when a task matches its description.

How do I install Qwen Vision in Codex?

Run `npx skills add davepoon/buildwithclaude --skill qwen-vision -a codex`. Or copy the skill folder (plugins/give-claude-eyes/skills/qwen-vision in davepoon/buildwithclaude) into .agents/skills/qwen-vision in your project. Codex loads it when a task matches its description.

Can I use Qwen Vision in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davepoon/buildwithclaude --skill qwen-vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qwen-vision, .gemini/skills/qwen-vision, .github/skills/qwen-vision and .opencode/skills/qwen-vision in your project.

What does Qwen Vision need to run?

Going by SKILL.md and its folder, Qwen Vision needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and pip) and credentials named DASHSCOPE_API_KEY. Our summary lists: Python 3; A credential in DASHSCOPE_API_KEY.

Does Qwen Vision access the network?

SKILL.md names 2 domains. As links in the text: dashscope.console.aliyun.com and modelstudio.console.alibabacloud.com. This is read from the text; nothing was executed.

Is Qwen Vision safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Qwen Vision use?

Qwen Vision is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Qwen Vision use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Qwen Vision?

Skills that share tags, products or a category with Qwen Vision: AI Search Hub (minsight-ai-info/AI-Search-Hub, 1.3k stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Pocketmen With You (six-nut/PocketMen-with-you, 310 stars) and Codex Panel Reproduce (littlepeachs/NaturePanelForge, 223 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Qwen Vision?

davepoon (a GitHub user) maintains it in davepoon/buildwithclaude, which has 3,610 GitHub stars. The repository holds 247 skills in this directory. The repository was last updated on October 9, 2026.

Source: davepoon/buildwithclaude on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.