Topic · AI & LLM Engineering
Best computer vision skills, page 2
Computer vision skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Converts PDFs, scans, and Word documents into text or markdown with the gaik toolkit's parsers, choosing the parser that will not silently destroy the structure the downstream task depends on. | GAIK-project/ | 100 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 50 | Build production computer vision pipelines for object detection, tracking, and video analysis. | curiositech/ | 243 | 1 repo | ~4k | Automated safety check: Pass | MIT | 1 mo ago |
| 51 | 51.Ops Yolo OPS on-demand: This skill should be used when the user asks to "yolo mode", "run the business today"… | Lifecycle-Innovations-Limited/ | 540 | — | ~4k | Automated safety check: Notes | MIT | 2 days ago |
| 52 | Understand images and videos with Qwen vision models. An agent skill from QianWen-AI/qianwen-ai. | QianWen-AI/ | 102 | — | ~4.9k | Automated safety check: Notes | Apache-2.0 | 16 days ago |
| 53 | Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV. | NVIDIA/ | 3.5k | — | ~2.7k | Automated safety check: Notes | Apache-2.0 | today |
| 54 | Maintain AlbumentationsX license, CLA, provenance notices, and packaged legal artifacts consistently. | albumentations-team/ | 566 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | today |
| 55 | Runs pre-trained Hugging Face models in JavaScript or TypeScript with Transformers.js, in browsers or Node.js, Bun and Deno, for text, vision, audio and multimodal tasks. | huggingface/ | 11k | 1 repo | ~6.2k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 56 | A skill your agent uses when user asks to analyze an image, describe image contents, or answer questions about a picture. | iflytek/ | 209 | — | ~949 | Automated safety check: Pass | Apache-2.0 | today |
| 57 | World-class computer vision skill for image/video processing, object detection, segmentation, and visual AI systems. | davila7/ | 32k | 3 repos | ~1.4k | Automated safety check: Pass | MIT | today |
| 58 | Deep expertise in ML/CV model selection, training pipelines, and inference architecture. | alirezarezvani/ | 117 | — | ~3.1k | Automated safety check: Pass | MIT | 9 mo ago |
| 59 | Runs TAO Data Services gap analysis that compares ground-truth and predicted boxes to find weak images by per-class recall, precision and AP50. | NVIDIA/ | 3.5k | — | ~1.8k | Automated safety check: Notes | Apache-2.0 | today |
| 60 | After completing code changes, runs tests and pre-commit, then iteratively fixes failures until all pass. | albumentations-team/ | 566 | — | ~847 | Automated safety check: Pass | AGPL-3.0 | today |
| 61 | Build an end-to-end UAV object detection and telemetry overlay application on Intel hardware using DL Streamer Pipeline Server with MAVLink telemetry. | open-edge-platform/ | 140 | — | ~2.6k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 62 | Build image analysis applications with Azure AI Vision SDK for Java. | microsoft/ | 3.1k | 6 repos | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 63 | Azure AI Vision Image Analysis SDK for captions, tags, objects, OCR, people detection, and smart cropping. | microsoft/ | 3.1k | 6 repos | ~2.5k | Automated safety check: Pass | MIT | yesterday |
| 64 | Computer vision engineering skill for object detection, image segmentation, and visual AI systems. | alirezarezvani/ | 28k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 65 | Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report. | NVIDIA/ | 3.5k | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | today |
| 66 | 66.Fei Fei Li Applies the reasoning, frameworks, and mental models of Fei-Fei Li, computer vision pioneer, ImageNet creator, and co-director of Stanford HAI. | K-Dense-AI/ | 282 | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 67 | Process and generate multimedia content using Google Gemini API. | Microck/ | 401 | 1 repo | ~2.7k | Automated safety check: Notes | MIT | 1 mo ago |
| 68 | Use the repo internal/ directory for anything that must not be committed — scratch files, temporary outputs, local demos, Codex artifacts, or one-off scripts. | albumentations-team/ | 566 | — | ~338 | Automated safety check: Pass | AGPL-3.0 | today |
| 69 | Agent-driven YOLO fine-tuning — annotate, train, export, deploy | SharpAI/ | 3.1k | — | ~985 | Automated safety check: Pass | MIT | 20 days ago |
| 70 | Looks up Microsoft Learn guidance for Azure AI Vision: Image Analysis, Read OCR containers, smart-crop thumbnails, background removal and video frame analysis, plus limits and deployment. | MicrosoftDocs/ | 775 | — | ~1.6k | Automated safety check: Pass | CC-BY-4.0 | yesterday |
| 71 | Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining. | NVIDIA/ | 3.5k | — | ~2k | Automated safety check: Notes | Apache-2.0 | today |
| 72 | 72.Pi Agent Builds with and operates Pi, the minimal terminal coding harness. | K-Dense-AI/ | 48k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | 2 days ago |
| 73 | A skill your agent uses when running a Docker Agent with docker agent run, choosing a safety/approval mode, using the --sandbox isolation flag, setting up aliases, or troubleshooting a run (missing… | docker/ | 539 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 74 | Best practices for image classification tasks. An agent skill from aiming-lab/AutoResearchClaw. | aiming-lab/ | 15k | — | ~304 | Automated safety check: Pass | MIT | 1 mo ago |
| 75 | 75.Benchmark Measure AlbumentationsX runtime changes with paired baseline and candidate benchmarks on the affected routes. | albumentations-team/ | 566 | — | ~900 | Automated safety check: Pass | AGPL-3.0 | today |
| 76 | Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs. | einverne/ | 121 | 1 repo | ~2.6k | Automated safety check: Notes | MIT | 28 days ago |
| 77 | 77.Visual QA Use vision models to self-review screenshots against design intent. | dylanfeltus/ | 179 | — | ~2.4k | Automated safety check: Pass | MIT | 20 days ago |
| 78 | Run TAO Data Services TMM unique-neighbor matching mining from embedding parquet files for object detection workflows. | NVIDIA/ | 3.5k | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | today |
| 79 | Review an AlbumentationsX transform for correctness, public API coherence, performance, documentation, and test coverage. | albumentations-team/ | 566 | — | ~973 | Automated safety check: Pass | AGPL-3.0 | today |
| 80 | 80.Cv Detection Best practices for object detection tasks. An agent skill from aiming-lab/AutoResearchClaw. | aiming-lab/ | 15k | — | ~257 | Automated safety check: Pass | MIT | 1 mo ago |
| 81 | Train object detection, image classification, and SAM or SAM2 segmentation models locally or on Hugging Face Jobs, with dataset validation and results saved to the Hub. | sickn33/ | 47k | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 82 | 82.On Device AI Build on-device AI features in React Native and Expo apps with React Native ExecuTorch. | software-mansion-labs/ | 291 | — | ~2.3k | Automated safety check: Pass | No licence | 9 days ago |
| 83 | 83.Vision Sft Fine-tune vision-language models (VLMs) with supervised learning on image+text data. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 2 days ago |
| 84 | Implement computer vision features including text recognition (OCR), face detection, barcode scanning, image segmentation, object tracking, and document scanning in iOS apps. | dpearson2699/ | 1.2k | — | ~4.7k | Automated safety check: Pass | Unknown | 2 mo ago |
| 85 | 85.Fal Redesign Upgrade a coded website to award-tier, editorially-crafted design using fal.ai. | fal-ai-community/ | 249 | — | ~1.4k | Automated safety check: Pass | MIT | 8 days ago |
| 86 | 86.Transformers Work with state-of-the-art machine learning models for NLP, computer vision, audio, and multimodal tasks using HuggingFace Transformers. | ynulihao/ | 617 | — | ~2.9k | Automated safety check: Pass | No licence | 7 mo ago |
| 87 | Meta Quest Passthrough Camera Access (PCA) for Unity — access the forward-facing RGB cameras on Quest 3 / Quest 3S to feed Computer Vision and Machine Learning pipelines. | meta-quest/ | 213 | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 88 | 火山视频理解 - 使用火山方舟视频理解 API 分析视频内容。通过 Files API 上传视频(推荐),支持大文件(最大512MB),可用于视频内容分析、物体识别、动作理解等。当用户需要分析视频、理解视频内容、提取视频信息时激活此技能。 | freestylefly/ | 461 | — | ~1k | Automated safety check: Notes | No licence | 4 mo ago |
| 89 | Run the full DEFT smart-data-augmentation loop for NVIDIA TAO Grounding DINO object detection: zero-shot baseline inference, KPI analysis, per-class gap analysis, SigLIP embedding of weak images… | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Notes | Apache-2.0 | today |
| 90 | Port a published computer vision paper's official code and training recipe onto a customer's own dataset, or diagnose why such a transfer produced bad numbers. | NVIDIA/ | 3.5k | — | ~4.2k | Automated safety check: Notes | Apache-2.0 | today |
| 91 | Grade or filter workflow HDF5 episodes with an OpenAI-compatible vision model. | NVIDIA/ | 3.5k | 1 repo | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 92 | 92.Spec Lite Lightweight, Chorus-native local specs for Chorus PM workflows in Hermes — a durable local spec .chorus/specs/<slug/spec.md (one per capability/feature) edited in place and NEVER synced (git history… | Chorus-AIDLC/ | 1.2k | — | ~2.2k | Automated safety check: Pass | AGPL-3.0 | today |
| 93 | 93.Spec Lite Lightweight, Chorus-native local specs for Chorus PM workflows in Pi — a durable local spec .chorus/specs/<slug/spec.md (one per capability/feature) edited in place and NEVER synced (git history is… | Chorus-AIDLC/ | 1.2k | — | ~2.1k | Automated safety check: Pass | AGPL-3.0 | today |
| 94 | 94.Spec Lite Lightweight, Chorus-native local specs for Chorus PM workflows on OpenClaw — a durable local spec .chorus/specs/<slug/spec.md (one per capability/feature) edited in place and NEVER synced (git… | Chorus-AIDLC/ | 1.2k | — | ~2.4k | Automated safety check: Pass | AGPL-3.0 | today |
| 95 | 95.Spec Lite Lightweight, Chorus-native local specs for Chorus PM workflows — a durable local spec .chorus/specs/<slug/spec.md (one per capability/feature) edited in place and NEVER synced (git history is its… | Chorus-AIDLC/ | 1.2k | — | ~2.1k | Automated safety check: Pass | AGPL-3.0 | today |
| 96 | Lightweight, Chorus-native local specs for Chorus PM workflows on dsh — a durable local spec .chorus/specs/<slug/spec.md (one per capability/feature) edited in place and NEVER synced (git history is… | Chorus-AIDLC/ | 1.2k | — | ~2.4k | Automated safety check: Pass | AGPL-3.0 | today |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- LLM inference and serving364
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67
- AI interpretability23