Search

Computer vision

203 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

Orchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT3 mo ago
2

Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

liustack/modlens4.2k1 repo~1.3kAutomated safety check: NotesMIT4 days ago
3

Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

Orchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT3 mo ago
4

A skill your agent uses when the user wants to run a YOLO-Master task (train/val/predict/track/export/benchmark) or use the Agent Skill dispatcher.

Tencent/YOLO-Master742—~755Automated safety check: PassAGPL-3.0today
5

Implement specialized video understanding capabilities using the z-ai-web-dev-sdk.

jjyaoao/HelloAgents3.2k1 repo~6.2kAutomated safety check: PassMIT2 days ago
6

Composites several moments from real drone footage into one still with ghost trails, then lays out paper figures and an editable PowerPoint file.

XXLiu-HNU/visualize_uav_trajectory242—~535Automated safety check: PassGPL-3.012 days ago
7

Trains and evaluates several WiFi-signal-based pose and sensing models, from unsupervised pose estimation to domain adaptation and publishing.

ruvnet/RuView97k—~1.3kAutomated safety check: NotesMITtoday
8

Adds a new local feature matching model from a GitHub repository to the image-matching-webui project as a working WebUI option.

Vincentqyw/image-matching-webui1.3k—~4.8kAutomated safety check: PassApache-2.02 days ago
9

Pixel-based motion and UI change analysis from frame sequences or screenshots using computer vision and visual comparison.

edwardsanchez/MotionEyes229—~2kAutomated safety check: PassNo licence6 mo ago
10

Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.

Orchestra-Research/AI-Research-SKILLs13k7 repos~2kAutomated safety check: PassMIT3 mo ago
11

Generate a specialized domain-expert research agent modeled on PaperClaw architecture.

guhaohao0991/PaperClaw250—~2kAutomated safety check: PassNo licence7 mo ago
12

Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

huggingface/skills11k1 repo~7.5kAutomated safety check: PassApache-2.07 days ago
13

Interactive click-to-segment using Segment Anything 2 — AI-assisted labeling for Annotation Studio

SharpAI/DeepCamera3.1k—~594Automated safety check: PassMIT21 days ago
14

This skill should be used when user asks to "improve my mAP", "why is my model overfitting", "my training is diverging", "read my results.csv", "interpret my training curves", "my AP50 is good but…

fcakyon/claude-codex-settings1.2k1 repo~1.4kAutomated safety check: PassApache-2.0yesterday
15

Policy for AlbumentationsX transforms that combine multiple images or objects.

albumentations-team/AlbumentationsX567—~1.3kAutomated safety check: PassAGPL-3.0yesterday
16

Command-line access to Z.AI vision analysis, web search, page reading and GitHub repo exploration through npx zai-cli, using an API key.

numman-ali/zai-cli110—~528Automated safety check: PassMIT9 mo ago
17

Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images.

xiincs/claude-code-vision-skill170—~1.2kAutomated safety check: PassMIT1 mo ago
18

Query images with a local Ollama vision model without loading the image into the main agent context.

gridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0yesterday
19

Track, retrieve, screen, and synthesize research frontiers for geospatial AI, remote sensing big data, and transferable computer vision methods.

limi124/remote-sensing-research-radar142—~1.3kAutomated safety check: PassNo licence5 mo ago
20

YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.

SharpAI/DeepCamera3.1k—~1.5kAutomated safety check: PassMIT21 days ago
21

Train object-detection, image-classification, or SAM segmentation models on Hugging Face Jobs.

waybarrios/opencode-power-pack533—~2.7kAutomated safety check: PassApache-2.02 days ago
22

Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.

Aseiel/VideoHighlighter160—~839Automated safety check: PassAGPL-3.0today
23

Regression test the OpenClaw Codex App Server plugin against a live local OpenClaw instance in Telegram or Discord.

pwrdrvr/openclaw-codex-app-server265—~4kAutomated safety check: PassMIT1 mo ago
24

Convert PDF/EPUB textbooks to searchable markdown files for an AI agent's own reference.

drpwchen/textbook-to-note105—~3.5kAutomated safety check: PassMIT1 mo ago
25

Calls an external vision model through vision.js to analyze an image when the user explicitly invokes /skill luma-vision, for agents whose own model cannot see images.

JochenYang/luma-mcp115—~295Automated safety check: PassMIT2 mo ago
26

Google Coral Edge TPU — real-time object detection natively (macOS / Linux)

SharpAI/DeepCamera3.1k—~1.2kAutomated safety check: PassMIT21 days ago
27

Low-level SenseNova tools for image generation, image editing, image recognition with a VLM and text optimization with an LLM, meant to be called by higher-level skills rather than directly.

OpenSenseNova/SenseNova-Skills5.7k—~3.2kAutomated safety check: PassMIT20 days ago
28

Systematic performance audit for AlbumentationsX runtime code.

albumentations-team/AlbumentationsX567—~1.7kAutomated safety check: PassAGPL-3.0yesterday
29

Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison.

einverne/dotfiles1211 repo~1.6kAutomated safety check: NotesMIT29 days ago
30

Google Coral Edge TPU — real-time object detection natively via Windows WSL

SharpAI/DeepCamera3.1k—~1.1kAutomated safety check: PassMIT21 days ago
31
31.Add Vlm ModelOfficial

Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

intel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0today
32

Vision-language pre-training framework bridging frozen image encoders and LLMs.

Orchestra-Research/AI-Research-SKILLs13k3 repos~4.3kAutomated safety check: PassMIT3 mo ago
33

Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.

TyrealQ/q-skills108—~2kAutomated safety check: NotesMIT14 days ago
34

Agent-only decision procedure for ask-user findings. An agent skill from kunchenguid/firstmate.

kunchenguid/firstmate7.7k—~1.1kAutomated safety check: PassMITtoday
35

Core ML, Create ML, Vision framework, Natural Language framework, on-device ML integration.

gustavscirulis/snapgrid1172 repos~3.8kAutomated safety check: NotesUnknown5 mo ago
36

Processes digital pathology whole slide images with histolab: tissue detection, mask creation, tile extraction and dataset preparation for deep learning.

davila7/claude-code-templates32k12 repos~5.1kAutomated safety check: PassMITtoday
37

Loads pre-trained Hugging Face Transformers models for text, vision and audio tasks, runs inference with pipelines and fine-tunes on custom datasets.

davila7/claude-code-templates32k12 repos~1.2kAutomated safety check: PassMITtoday
38

OpenVINO — real-time object detection via Docker (NCS2, Intel GPU, CPU)

SharpAI/DeepCamera3.1k—~1.3kAutomated safety check: PassMIT21 days ago
39

Generate text, have conversations, write code, reason, and call functions with Qwen models.

QianWen-AI/qianwen-ai104—~4.7kAutomated safety check: NotesApache-2.017 days ago
40

Full checklist for adding a new transform to AlbumentationsX.

albumentations-team/AlbumentationsX567—~1.7kAutomated safety check: PassAGPL-3.0yesterday
41

Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

NVIDIA/skills3.5k—~5kAutomated safety check: NotesApache-2.0today
42

Review screenshots or other images with OpenRouter vision models via bundled Deno scripts.

mizchi/skills356—~1.3kAutomated safety check: PassNo licence6 days ago
43

Extracts body and hand keypoint trajectories from an authorized reference video, with skeleton previews and confidence data, for pose reference or motion control input.

Pluviobyte/rnskill1.6k—~1kAutomated safety check: PassUnknown17 days ago
44

把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.

zenstory-ai/video-recap-skills555—~1.1kAutomated safety check: PassMIT4 days ago
45

Review public AlbumentationsX docstrings for useful descriptions, runnable examples, parameter semantics, and related transforms.

albumentations-team/AlbumentationsX567—~534Automated safety check: PassAGPL-3.0yesterday
46

Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

NVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0today
47

CVPR / ICCV / ECCV paper formatting — activate when the user wants to submit to CVPR, ICCV, or ECCV, follow their template, or fix format issues for these computer vision conferences.

nanoAgentTeam/research-claw293—~3.2kAutomated safety check: PassMIT4 mo ago
48

Comprehensive best practices for robot perception systems covering cameras, LiDARs, depth sensors, IMUs, and multi-sensor setups.

arpitg1304/robotics-agent-skills368—~15kAutomated safety check: PassApache-2.01 mo ago