Topic · AI & LLM Engineering
Best computer vision skills for Claude Code, Codex and other agents.
- skills
- 206
- official
- 35
Computer vision skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation. | Orchestra-Research/ | 13k | 9 repos | ~3.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns. | Orchestra-Research/ | 13k | 8 repos | ~1.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 3 | Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics. | liustack/ | 4.1k | — | ~1.3k | Automated safety check: Notes | MIT | 3 days ago |
| 4 | A skill your agent uses when the user wants to run a YOLO-Master task (train/val/predict/track/export/benchmark) or use the Agent Skill dispatcher. | Tencent/ | 742 | — | ~755 | Automated safety check: Pass | AGPL-3.0 | 9 days ago |
| 5 | Implement specialized video understanding capabilities using the z-ai-web-dev-sdk. | jjyaoao/ | 3.2k | 1 repo | ~6.2k | Automated safety check: Pass | MIT | yesterday |
| 6 | Composites several moments from real drone footage into one still with ghost trails, then lays out paper figures and an editable PowerPoint file. | XXLiu-HNU/ | 242 | — | ~535 | Automated safety check: Pass | GPL-3.0 | 11 days ago |
| 7 | Trains and evaluates several WiFi-signal-based pose and sensing models, from unsupervised pose estimation to domain adaptation and publishing. | ruvnet/ | 97k | — | ~1.3k | Automated safety check: Notes | MIT | today |
| 8 | Adds a new local feature matching model from a GitHub repository to the image-matching-webui project as a working WebUI option. | Vincentqyw/ | 1.3k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | Pixel-based motion and UI change analysis from frame sequences or screenshots using computer vision and visual comparison. | edwardsanchez/ | 229 | — | ~2k | Automated safety check: Pass | No licence | 6 mo ago |
| 10 | Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code. | Orchestra-Research/ | 13k | 7 repos | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 11 | Generate a specialized domain-expert research agent modeled on PaperClaw architecture. | guhaohao0991/ | 250 | — | ~2k | Automated safety check: Pass | No licence | 7 mo ago |
| 12 | Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub. | huggingface/ | 11k | 1 repo | ~7.5k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 13 | Interactive click-to-segment using Segment Anything 2 — AI-assisted labeling for Annotation Studio | SharpAI/ | 3.1k | — | ~594 | Automated safety check: Pass | MIT | 20 days ago |
| 14 | This skill should be used when user asks to "improve my mAP", "why is my model overfitting", "my training is diverging", "read my results.csv", "interpret my training curves", "my AP50 is good but… | fcakyon/ | 1.2k | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 15 | Policy for AlbumentationsX transforms that combine multiple images or objects. | albumentations-team/ | 566 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | today |
| 16 | 16.Z.AI CLI Command-line access to Z.AI vision analysis, web search, page reading and GitHub repo exploration through npx zai-cli, using an API key. | numman-ali/ | 110 | — | ~528 | Automated safety check: Pass | MIT | 9 mo ago |
| 17 | 17.Vision Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images. | xiincs/ | 170 | — | ~1.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 18 | 18.Vision Query images with a local Ollama vision model without loading the image into the main agent context. | gridaco/ | 2.7k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 19 | Track, retrieve, screen, and synthesize research frontiers for geospatial AI, remote sensing big data, and transferable computer vision methods. | limi124/ | 141 | — | ~1.3k | Automated safety check: Pass | No licence | 5 mo ago |
| 20 | YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera. | SharpAI/ | 3.1k | — | ~1.5k | Automated safety check: Pass | MIT | 20 days ago |
| 21 | Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs. | waybarrios/ | 533 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 22 | Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better. | Aseiel/ | 157 | — | ~839 | Automated safety check: Pass | AGPL-3.0 | today |
| 23 | Regression test the OpenClaw Codex App Server plugin against a live local OpenClaw instance in Telegram or Discord. | pwrdrvr/ | 265 | — | ~4k | Automated safety check: Pass | MIT | 1 mo ago |
| 24 | Convert PDF/EPUB textbooks to searchable markdown files for an AI agent's own reference. | drpwchen/ | 104 | — | ~3.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 25 | Calls an external vision model through vision.js to analyze an image when the user explicitly invokes /skill luma-vision, for agents whose own model cannot see images. | JochenYang/ | 115 | — | ~295 | Automated safety check: Pass | MIT | 1 mo ago |
| 26 | Google Coral Edge TPU — real-time object detection natively (macOS / Linux) | SharpAI/ | 3.1k | — | ~1.2k | Automated safety check: Pass | MIT | 20 days ago |
| 27 | Low-level SenseNova tools for image generation, image editing, image recognition with a VLM and text optimization with an LLM, meant to be called by higher-level skills rather than directly. | OpenSenseNova/ | 5.7k | — | ~3.2k | Automated safety check: Pass | MIT | 19 days ago |
| 28 | Systematic performance audit for AlbumentationsX runtime code. | albumentations-team/ | 566 | — | ~1.7k | Automated safety check: Pass | AGPL-3.0 | today |
| 29 | Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison. | einverne/ | 121 | 1 repo | ~1.6k | Automated safety check: Notes | MIT | 28 days ago |
| 30 | Google Coral Edge TPU — real-time object detection natively via Windows WSL | SharpAI/ | 3.1k | — | ~1.1k | Automated safety check: Pass | MIT | 20 days ago |
| 31 | Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling. | intel/ | 1.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 32 | Vision-language pre-training framework bridging frozen image encoders and LLMs. | Orchestra-Research/ | 13k | 3 repos | ~4.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 33 | Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API. | TyrealQ/ | 108 | — | ~2k | Automated safety check: Notes | MIT | 13 days ago |
| 34 | Agent-only decision procedure for ask-user findings. An agent skill from kunchenguid/firstmate. | kunchenguid/ | 7.6k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 35 | 35.Core ML Core ML, Create ML, Vision framework, Natural Language framework, on-device ML integration. | gustavscirulis/ | 117 | 2 repos | ~3.8k | Automated safety check: Notes | Unknown | 5 mo ago |
| 36 | Processes digital pathology whole slide images with histolab: tissue detection, mask creation, tile extraction and dataset preparation for deep learning. | davila7/ | 32k | 12 repos | ~5.1k | Automated safety check: Pass | MIT | today |
| 37 | Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets. | davila7/ | 32k | 12 repos | ~1.2k | Automated safety check: Pass | MIT | today |
| 38 | OpenVINO — real-time object detection via Docker (NCS2, Intel GPU, CPU) | SharpAI/ | 3.1k | — | ~1.3k | Automated safety check: Pass | MIT | 20 days ago |
| 39 | 把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills. | zenstory-ai/ | 553 | — | ~1.1k | Automated safety check: Pass | MIT | 3 days ago |
| 40 | 40.Qianwen Text Generate text, have conversations, write code, reason, and call functions with Qwen models. | QianWen-AI/ | 102 | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | 16 days ago |
| 41 | Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Notes | Apache-2.0 | today |
| 42 | Full checklist for adding a new transform to AlbumentationsX. | albumentations-team/ | 566 | — | ~1.7k | Automated safety check: Pass | AGPL-3.0 | today |
| 43 | 43.Review Image Review screenshots or other images with OpenRouter vision models via bundled Deno scripts. | mizchi/ | 356 | — | ~1.3k | Automated safety check: Pass | No licence | 5 days ago |
| 44 | Extracts body and hand keypoint trajectories from an authorized reference video, with skeleton previews and confidence data, for pose reference or motion control input. | Pluviobyte/ | 1.6k | — | ~1k | Automated safety check: Pass | Unknown | 16 days ago |
| 45 | Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download. | NVIDIA/ | 3.5k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | today |
| 46 | Review public AlbumentationsX docstrings for useful descriptions, runnable examples, parameter semantics, and related transforms. | albumentations-team/ | 566 | — | ~534 | Automated safety check: Pass | AGPL-3.0 | today |
| 47 | CVPR / ICCV / ECCV paper formatting — activate when the user wants to submit to CVPR, ICCV, or ECCV, follow their template, or fix format issues for these computer vision conferences. | nanoAgentTeam/ | 293 | — | ~3.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 48 | Comprehensive best practices for robot perception systems covering cameras, LiDARs, depth sensors, IMUs, and multi-sensor setups. | arpitg1304/ | 368 | — | ~15k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
Questions, answered from the data.
What is the best computer vision skill?
Segment Anything Model Guide from Orchestra-Research/AI-Research-SKILLs ranks first of the 206 computer vision skills listed here, with the highest score: its repository has 13k GitHub stars, 9 other GitHub owners carry a copy, its SKILL.md loads about 3.3k tokens and it passes the automated safety check with no findings. Next come CLIP Image-Text Matching and ModLens Image Vision Bridge.
Which computer vision skills are official?
35 of the 206 computer vision skills are official, published by the vendor's own GitHub organization: Hugging Face Vision Trainer, Add Vlm Model, Defect Image Generation with Cosmos AnomalyGen, Physical AI Video Augmentation on OSMO, TAO Detection KPI Analysis and 30 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- LLM inference and serving364
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67
- AI interpretability23