Search
Computer vision
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation. | Orchestra-Research/ | 13k | 9 repos | ~3.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics. | liustack/ | 4.2k | 1 repo | ~1.3k | Automated safety check: Notes | MIT | 4 days ago |
| 3 | Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns. | Orchestra-Research/ | 13k | 8 repos | ~1.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 4 | A skill your agent uses when the user wants to run a YOLO-Master task (train/val/predict/track/export/benchmark) or use the Agent Skill dispatcher. | Tencent/ | 742 | — | ~755 | Automated safety check: Pass | AGPL-3.0 | today |
| 5 | Implement specialized video understanding capabilities using the z-ai-web-dev-sdk. | jjyaoao/ | 3.2k | 1 repo | ~6.2k | Automated safety check: Pass | MIT | 2 days ago |
| 6 | Composites several moments from real drone footage into one still with ghost trails, then lays out paper figures and an editable PowerPoint file. | XXLiu-HNU/ | 242 | — | ~535 | Automated safety check: Pass | GPL-3.0 | 12 days ago |
| 7 | Trains and evaluates several WiFi-signal-based pose and sensing models, from unsupervised pose estimation to domain adaptation and publishing. | ruvnet/ | 97k | — | ~1.3k | Automated safety check: Notes | MIT | today |
| 8 | Adds a new local feature matching model from a GitHub repository to the image-matching-webui project as a working WebUI option. | Vincentqyw/ | 1.3k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 9 | Pixel-based motion and UI change analysis from frame sequences or screenshots using computer vision and visual comparison. | edwardsanchez/ | 229 | — | ~2k | Automated safety check: Pass | No licence | 6 mo ago |
| 10 | Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code. | Orchestra-Research/ | 13k | 7 repos | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 11 | Generate a specialized domain-expert research agent modeled on PaperClaw architecture. | guhaohao0991/ | 250 | — | ~2k | Automated safety check: Pass | No licence | 7 mo ago |
| 12 | Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub. | huggingface/ | 11k | 1 repo | ~7.5k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 13 | Interactive click-to-segment using Segment Anything 2 — AI-assisted labeling for Annotation Studio | SharpAI/ | 3.1k | — | ~594 | Automated safety check: Pass | MIT | 21 days ago |
| 14 | This skill should be used when user asks to "improve my mAP", "why is my model overfitting", "my training is diverging", "read my results.csv", "interpret my training curves", "my AP50 is good but… | fcakyon/ | 1.2k | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 15 | Policy for AlbumentationsX transforms that combine multiple images or objects. | albumentations-team/ | 567 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 16 | 16.Z.AI CLI Command-line access to Z.AI vision analysis, web search, page reading and GitHub repo exploration through npx zai-cli, using an API key. | numman-ali/ | 110 | — | ~528 | Automated safety check: Pass | MIT | 9 mo ago |
| 17 | 17.Vision Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images. | xiincs/ | 170 | — | ~1.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 18 | 18.Vision Query images with a local Ollama vision model without loading the image into the main agent context. | gridaco/ | 2.7k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 19 | Track, retrieve, screen, and synthesize research frontiers for geospatial AI, remote sensing big data, and transferable computer vision methods. | limi124/ | 142 | — | ~1.3k | Automated safety check: Pass | No licence | 5 mo ago |
| 20 | YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera. | SharpAI/ | 3.1k | — | ~1.5k | Automated safety check: Pass | MIT | 21 days ago |
| 21 | Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs. | waybarrios/ | 533 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 22 | Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better. | Aseiel/ | 160 | — | ~839 | Automated safety check: Pass | AGPL-3.0 | today |
| 23 | Regression test the OpenClaw Codex App Server plugin against a live local OpenClaw instance in Telegram or Discord. | pwrdrvr/ | 265 | — | ~4k | Automated safety check: Pass | MIT | 1 mo ago |
| 24 | Convert PDF/EPUB textbooks to searchable markdown files for an AI agent's own reference. | drpwchen/ | 105 | — | ~3.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 25 | Calls an external vision model through vision.js to analyze an image when the user explicitly invokes /skill luma-vision, for agents whose own model cannot see images. | JochenYang/ | 115 | — | ~295 | Automated safety check: Pass | MIT | 2 mo ago |
| 26 | Google Coral Edge TPU — real-time object detection natively (macOS / Linux) | SharpAI/ | 3.1k | — | ~1.2k | Automated safety check: Pass | MIT | 21 days ago |
| 27 | Low-level SenseNova tools for image generation, image editing, image recognition with a VLM and text optimization with an LLM, meant to be called by higher-level skills rather than directly. | OpenSenseNova/ | 5.7k | — | ~3.2k | Automated safety check: Pass | MIT | 20 days ago |
| 28 | Systematic performance audit for AlbumentationsX runtime code. | albumentations-team/ | 567 | — | ~1.7k | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 29 | Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison. | einverne/ | 121 | 1 repo | ~1.6k | Automated safety check: Notes | MIT | 29 days ago |
| 30 | Google Coral Edge TPU — real-time object detection natively via Windows WSL | SharpAI/ | 3.1k | — | ~1.1k | Automated safety check: Pass | MIT | 21 days ago |
| 31 | Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling. | intel/ | 1.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 32 | Vision-language pre-training framework bridging frozen image encoders and LLMs. | Orchestra-Research/ | 13k | 3 repos | ~4.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 33 | Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API. | TyrealQ/ | 108 | — | ~2k | Automated safety check: Notes | MIT | 14 days ago |
| 34 | Agent-only decision procedure for ask-user findings. An agent skill from kunchenguid/firstmate. | kunchenguid/ | 7.7k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 35 | 35.Core ML Core ML, Create ML, Vision framework, Natural Language framework, on-device ML integration. | gustavscirulis/ | 117 | 2 repos | ~3.8k | Automated safety check: Notes | Unknown | 5 mo ago |
| 36 | Processes digital pathology whole slide images with histolab: tissue detection, mask creation, tile extraction and dataset preparation for deep learning. | davila7/ | 32k | 12 repos | ~5.1k | Automated safety check: Pass | MIT | today |
| 37 | Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets. | davila7/ | 32k | 12 repos | ~1.2k | Automated safety check: Pass | MIT | today |
| 38 | OpenVINO — real-time object detection via Docker (NCS2, Intel GPU, CPU) | SharpAI/ | 3.1k | — | ~1.3k | Automated safety check: Pass | MIT | 21 days ago |
| 39 | 39.Qianwen Text Generate text, have conversations, write code, reason, and call functions with Qwen models. | QianWen-AI/ | 104 | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | 17 days ago |
| 40 | Full checklist for adding a new transform to AlbumentationsX. | albumentations-team/ | 567 | — | ~1.7k | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 41 | Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Notes | Apache-2.0 | today |
| 42 | 42.Review Image Review screenshots or other images with OpenRouter vision models via bundled Deno scripts. | mizchi/ | 356 | — | ~1.3k | Automated safety check: Pass | No licence | 6 days ago |
| 43 | Extracts body and hand keypoint trajectories from an authorized reference video, with skeleton previews and confidence data, for pose reference or motion control input. | Pluviobyte/ | 1.6k | — | ~1k | Automated safety check: Pass | Unknown | 17 days ago |
| 44 | 把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills. | zenstory-ai/ | 555 | — | ~1.1k | Automated safety check: Pass | MIT | 4 days ago |
| 45 | Review public AlbumentationsX docstrings for useful descriptions, runnable examples, parameter semantics, and related transforms. | albumentations-team/ | 567 | — | ~534 | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 46 | Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download. | NVIDIA/ | 3.5k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | today |
| 47 | CVPR / ICCV / ECCV paper formatting — activate when the user wants to submit to CVPR, ICCV, or ECCV, follow their template, or fix format issues for these computer vision conferences. | nanoAgentTeam/ | 293 | — | ~3.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 48 | Comprehensive best practices for robot perception systems covering cameras, LiDARs, depth sensors, IMUs, and multi-sensor setups. | arpitg1304/ | 368 | — | ~15k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |