Topic · AI & LLM Engineering

Best computer vision skills for Claude Code, Codex and other agents.

Skills that build and apply computer-vision and vision-language models.
skills
206
official
35

Computer vision skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Computer vision skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

Orchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT3 mo ago
2

Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

Orchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT3 mo ago
3

Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

liustack/modlens4.1k—~1.3kAutomated safety check: NotesMIT3 days ago
4

A skill your agent uses when the user wants to run a YOLO-Master task (train/val/predict/track/export/benchmark) or use the Agent Skill dispatcher.

Tencent/YOLO-Master742—~755Automated safety check: PassAGPL-3.09 days ago
5

Implement specialized video understanding capabilities using the z-ai-web-dev-sdk.

jjyaoao/HelloAgents3.2k1 repo~6.2kAutomated safety check: PassMITyesterday
6

Composites several moments from real drone footage into one still with ghost trails, then lays out paper figures and an editable PowerPoint file.

XXLiu-HNU/visualize_uav_trajectory242—~535Automated safety check: PassGPL-3.011 days ago
7

Trains and evaluates several WiFi-signal-based pose and sensing models, from unsupervised pose estimation to domain adaptation and publishing.

ruvnet/RuView97k—~1.3kAutomated safety check: NotesMITtoday
8

Adds a new local feature matching model from a GitHub repository to the image-matching-webui project as a working WebUI option.

Vincentqyw/image-matching-webui1.3k—~4.8kAutomated safety check: PassApache-2.0yesterday
9

Pixel-based motion and UI change analysis from frame sequences or screenshots using computer vision and visual comparison.

edwardsanchez/MotionEyes229—~2kAutomated safety check: PassNo licence6 mo ago
10

Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.

Orchestra-Research/AI-Research-SKILLs13k7 repos~2kAutomated safety check: PassMIT3 mo ago
11

Generate a specialized domain-expert research agent modeled on PaperClaw architecture.

guhaohao0991/PaperClaw250—~2kAutomated safety check: PassNo licence7 mo ago
12

Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

huggingface/skills11k1 repo~7.5kAutomated safety check: PassApache-2.06 days ago
13

Interactive click-to-segment using Segment Anything 2 — AI-assisted labeling for Annotation Studio

SharpAI/DeepCamera3.1k—~594Automated safety check: PassMIT20 days ago
14

This skill should be used when user asks to "improve my mAP", "why is my model overfitting", "my training is diverging", "read my results.csv", "interpret my training curves", "my AP50 is good but…

fcakyon/claude-codex-settings1.2k1 repo~1.4kAutomated safety check: PassApache-2.0yesterday
15

Policy for AlbumentationsX transforms that combine multiple images or objects.

albumentations-team/AlbumentationsX566—~1.3kAutomated safety check: PassAGPL-3.0today
16

Command-line access to Z.AI vision analysis, web search, page reading and GitHub repo exploration through npx zai-cli, using an API key.

numman-ali/zai-cli110—~528Automated safety check: PassMIT9 mo ago
17

Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.

xiincs/claude-code-vision-skill170—~1.2kAutomated safety check: PassMIT1 mo ago
18

Query images with a local Ollama vision model without loading the image into the main agent context.

gridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0yesterday
19

Track, retrieve, screen, and synthesize research frontiers for geospatial AI, remote sensing big data, and transferable computer vision methods.

limi124/remote-sensing-research-radar141—~1.3kAutomated safety check: PassNo licence5 mo ago
20

YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.

SharpAI/DeepCamera3.1k—~1.5kAutomated safety check: PassMIT20 days ago
21

Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

waybarrios/opencode-power-pack533—~2.7kAutomated safety check: PassApache-2.0yesterday
22

Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.

Aseiel/VideoHighlighter157—~839Automated safety check: PassAGPL-3.0today
23

Regression test the OpenClaw Codex App Server plugin against a live local OpenClaw instance in Telegram or Discord.

pwrdrvr/openclaw-codex-app-server265—~4kAutomated safety check: PassMIT1 mo ago
24

Convert PDF/EPUB textbooks to searchable markdown files for an AI agent's own reference.

drpwchen/textbook-to-note104—~3.5kAutomated safety check: PassMIT1 mo ago
25

Calls an external vision model through vision.js to analyze an image when the user explicitly invokes /skill luma-vision, for agents whose own model cannot see images.

JochenYang/luma-mcp115—~295Automated safety check: PassMIT1 mo ago
26

Google Coral Edge TPU — real-time object detection natively (macOS / Linux)

SharpAI/DeepCamera3.1k—~1.2kAutomated safety check: PassMIT20 days ago
27

Low-level SenseNova tools for image generation, image editing, image recognition with a VLM and text optimization with an LLM, meant to be called by higher-level skills rather than directly.

OpenSenseNova/SenseNova-Skills5.7k—~3.2kAutomated safety check: PassMIT19 days ago
28

Systematic performance audit for AlbumentationsX runtime code.

albumentations-team/AlbumentationsX566—~1.7kAutomated safety check: PassAGPL-3.0today
29

Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison.

einverne/dotfiles1211 repo~1.6kAutomated safety check: NotesMIT28 days ago
30

Google Coral Edge TPU — real-time object detection natively via Windows WSL

SharpAI/DeepCamera3.1k—~1.1kAutomated safety check: PassMIT20 days ago
31
31.Add Vlm ModelOfficial

Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

intel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0yesterday
32

Vision-language pre-training framework bridging frozen image encoders and LLMs.

Orchestra-Research/AI-Research-SKILLs13k3 repos~4.3kAutomated safety check: PassMIT3 mo ago
33

Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.

TyrealQ/q-skills108—~2kAutomated safety check: NotesMIT13 days ago
34

Agent-only decision procedure for ask-user findings. An agent skill from kunchenguid/firstmate.

kunchenguid/firstmate7.6k—~1.1kAutomated safety check: PassMITtoday
35

Core ML, Create ML, Vision framework, Natural Language framework, on-device ML integration.

gustavscirulis/snapgrid1172 repos~3.8kAutomated safety check: NotesUnknown5 mo ago
36

Processes digital pathology whole slide images with histolab: tissue detection, mask creation, tile extraction and dataset preparation for deep learning.

davila7/claude-code-templates32k12 repos~5.1kAutomated safety check: PassMITtoday
37

Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

davila7/claude-code-templates32k12 repos~1.2kAutomated safety check: PassMITtoday
38

OpenVINO — real-time object detection via Docker (NCS2, Intel GPU, CPU)

SharpAI/DeepCamera3.1k—~1.3kAutomated safety check: PassMIT20 days ago
39

把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.

zenstory-ai/video-recap-skills553—~1.1kAutomated safety check: PassMIT3 days ago
40

Generate text, have conversations, write code, reason, and call functions with Qwen models.

QianWen-AI/qianwen-ai102—~4.7kAutomated safety check: NotesApache-2.016 days ago
41

Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

NVIDIA/skills3.5k—~5kAutomated safety check: NotesApache-2.0today
42

Full checklist for adding a new transform to AlbumentationsX.

albumentations-team/AlbumentationsX566—~1.7kAutomated safety check: PassAGPL-3.0today
43

Review screenshots or other images with OpenRouter vision models via bundled Deno scripts.

mizchi/skills356—~1.3kAutomated safety check: PassNo licence5 days ago
44

Extracts body and hand keypoint trajectories from an authorized reference video, with skeleton previews and confidence data, for pose reference or motion control input.

Pluviobyte/rnskill1.6k—~1kAutomated safety check: PassUnknown16 days ago
45

Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

NVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0today
46

Review public AlbumentationsX docstrings for useful descriptions, runnable examples, parameter semantics, and related transforms.

albumentations-team/AlbumentationsX566—~534Automated safety check: PassAGPL-3.0today
47

CVPR / ICCV / ECCV paper formatting — activate when the user wants to submit to CVPR, ICCV, or ECCV, follow their template, or fix format issues for these computer vision conferences.

nanoAgentTeam/research-claw293—~3.2kAutomated safety check: PassMIT4 mo ago
48

Comprehensive best practices for robot perception systems covering cameras, LiDARs, depth sensors, IMUs, and multi-sensor setups.

arpitg1304/robotics-agent-skills368—~15kAutomated safety check: PassApache-2.01 mo ago

Questions, answered from the data.

What is the best computer vision skill?

Segment Anything Model Guide from Orchestra-Research/AI-Research-SKILLs ranks first of the 206 computer vision skills listed here, with the highest score: its repository has 13k GitHub stars, 9 other GitHub owners carry a copy, its SKILL.md loads about 3.3k tokens and it passes the automated safety check with no findings. Next come CLIP Image-Text Matching and ModLens Image Vision Bridge.

Which computer vision skills are official?

35 of the 206 computer vision skills are official, published by the vendor's own GitHub organization: Hugging Face Vision Trainer, Add Vlm Model, Defect Image Generation with Cosmos AnomalyGen, Physical AI Video Augmentation on OSMO, TAO Detection KPI Analysis and 30 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.