Agent skill

Senior Computer Vision

by borghei in borghei/Claude-Skills

Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment.

MITAuto-check passedAI & LLM Engineering

Install Senior Computer Vision

skills CLI
$ npx skills add borghei/Claude-Skills --skill senior-computer-vision -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills senior-computer-vision --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/senior-computer-vision .claude/skills/senior-computer-vision && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
senior-computer-vision
GitHub stars
874
Token cost
~1.8k tokens
SKILL.md length
650 words
Files
10 (incl. scripts, references)
Skills in repo
364
Repo updated
First seen
Licence
MIT

At a glance

Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment.

  • Building detection pipelines
  • SKILL.md covers Core Capabilities, When to Use, Clarify First and Tools, plus 3 more sections
  • Runs Python scripts from its folder; calls python
  • Training models

What it does

Senior Computer Vision is an agent skill from borghei/Claude-Skills. Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment. Use when building detection pipelines, training models, or optimizing inference.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `references/commands-targets-and-troubleshooting.md`, `references/computer_vision_architectures.md` and `references/detection-workflows.md`).

It sits in AI & LLM Engineering, covering Computer vision. It works with NVIDIA AI Platform and ONNX. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Building detection pipelines
  • Training models
  • Optimizing inference

Example prompts

  • “/senior-computer-vision”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Senior Computer Vision loads about 1.8k tokens when it runs, and up to ~26k if it reads all its reference files. Until then it costs about 66 tokens; SKILL.md has 650 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~26k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 650 words, ~1,754 tokens.

Download SKILL.mdSave it as .claude/skills/senior-computer-vision/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
senior-computer-vision
description
Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment. Use when building detection pipelines, training models, or optimizing inference.
license
MIT + Commons Clause
metadata.version
1.1.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
computer-vision
metadata.updated
2026-06-17
metadata.tags
object-detection, image-segmentation, computer-vision, model-training

Senior Computer Vision Engineer

Design end-to-end computer vision pipelines for object detection, instance/semantic segmentation, and production deployment. Generates training configurations for YOLO/Detectron2/MMDetection, optimizes models for ONNX/TensorRT/OpenVINO runtimes, and builds dataset preparation workflows with format conversion and augmentation.

Core Capabilities

  • Detection pipeline design — requirements analysis, architecture selection (YOLO/RT-DETR/Faster R-CNN/DINO), dataset prep, training config, and metric evaluation.
  • Model optimization & deployment — baseline benchmarking, ONNX export, INT8/FP16 quantization, and conversion to TensorRT/OpenVINO/CoreML/TFLite per target platform.
  • Dataset engineering — audit, cleaning, format conversion (COCO/YOLO/VOC/CVAT/LabelMe), augmentation config, and stratified train/val/test splits.
  • Architecture guidance — detection and segmentation architecture trade-offs plus CNN vs Vision Transformer selection.
  • Production targets — FPS, mAP, latency P99, memory, and model-size budgets for real-time, high-accuracy, and edge deployments.

When to Use

  • Building an object detection or segmentation system from scratch.
  • Optimizing and deploying a trained model to GPU, edge, or mobile.
  • Preparing, converting, or auditing a computer vision dataset.
  • Choosing an architecture for a speed/accuracy/deployment trade-off.

Clarify First

Before generating training configs or pipelines, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Task — detection / instance or semantic segmentation / classification (selects the architecture and --task)
  • Dataset — location and format (COCO / YOLO / VOC) to analyze or convert (the input to dataset_pipeline_builder.py)
  • Deployment target — GPU / edge / mobile (drives architecture choice and inference_optimizer --target)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Tools

ToolPurposeCommand
vision_model_trainer.pyGenerate training configs for YOLO / Detectron2 / MMDetectionpython scripts/vision_model_trainer.py data/coco/ --task detection --arch yolov8m -o configs/train.yaml
inference_optimizer.pyAnalyze, benchmark, and recommend optimizations for a modelpython scripts/inference_optimizer.py model.pt --analyze --benchmark --recommend --target edge
dataset_pipeline_builder.pyAnalyze/convert/split/augment/validate CV datasets (subcommands)python scripts/dataset_pipeline_builder.py analyze --input data/coco/

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

  • references/detection-workflows.md — quick-start commands and the three end-to-end workflows (detection pipeline, model optimization/deployment, dataset prep) plus the architecture selection guide. Read when executing a pipeline.
  • references/commands-targets-and-troubleshooting.md — framework command catalogs (YOLO/Detectron2/MMDetection/optimization), performance targets, anti-patterns, troubleshooting table, and success criteria. Read while running training or deployment.
  • references/tool-reference.md — full parameter, example, and output-format reference for the three scripts. Read when scripting the tools.
  • references/computer_vision_architectures.md — CNN backbones (ResNet, EfficientNet, ConvNeXt), ViT variants (ViT, DeiT, Swin), detection heads, and FPN/BiFPN/PANet necks. Read when choosing or tuning architectures.
  • references/object_detection_optimization.md — NMS variants, anchor optimization, loss design (focal, GIoU/CIoU/DIoU), training strategies, and detection augmentation. Read when improving detection accuracy.
  • references/production_vision_systems.md — ONNX/TensorRT export, batch inference, edge deployment (Jetson, Intel NCS), Triton serving, and video pipelines. Read when deploying to production.
Show full SKILL.md (226 more words)Show less

Scope & Limitations

This skill covers:

  • End-to-end object detection and segmentation pipeline design (data preparation through production deployment)
  • Training configuration generation for Ultralytics YOLO, Detectron2, and MMDetection frameworks
  • Model optimization and export to ONNX, TensorRT, OpenVINO, and CoreML runtimes
  • Dataset format conversion (COCO, YOLO, Pascal VOC, CVAT), splitting, validation, and augmentation configuration

This skill does NOT cover:

  • Generative vision tasks (image generation, style transfer, super-resolution) -- see dedicated generative AI skills
  • 3D reconstruction, SLAM, or point cloud processing beyond basic depth estimation
  • Medical imaging regulatory compliance (DICOM, FDA 510(k)) -- see ra-qm-team/ compliance skills
  • Real-time video streaming infrastructure (RTSP, WebRTC, GStreamer pipeline design) -- see senior-devops for infrastructure

Integration Points

SkillIntegrationData Flow
senior-ml-engineerModel serving and MLOps pipeline setupTrained model artifacts (.pt, .onnx) flow into model_deployment_pipeline.py for containerized serving and monitoring
senior-data-engineerDataset ETL and storage pipelinesRaw image data ingested via pipeline_orchestrator.py; cleaned datasets flow into dataset_pipeline_builder.py for CV formatting
senior-data-scientistExperiment design and statistical analysisExperiment parameters from experiment_designer.py guide hyperparameter search; model metrics feed back for significance testing
senior-devopsCI/CD and GPU infrastructure provisioningOptimized model artifacts deployed via CI/CD pipelines; GPU node scaling managed through infrastructure-as-code
senior-prompt-engineerMultimodal RAG and vision-language integrationVision model embeddings and detections feed into rag_system_builder.py for multimodal retrieval pipelines
senior-cloud-architectCloud GPU resource planning and cost optimizationBenchmark results from inference_optimizer.py inform instance type selection and auto-scaling policies

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references) in engineering/senior-computer-vision of borghei/Claude-Skills.

  • SKILL.md
  • references/commands-targets-and-troubleshooting.md
  • references/computer_vision_architectures.md
  • references/detection-workflows.md
  • references/object_detection_optimization.md
  • references/production_vision_systems.md
  • references/tool-reference.md
  • scripts/dataset_pipeline_builder.py
  • scripts/inference_optimizer.py
  • scripts/vision_model_trainer.py

Open the folder on GitHubat commit c9a1487

Compare with similar skills

Senior Computer Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Senior Computer Vision compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Senior Computer Vision this skillborghei/Claude-Skills874—~1.8kAutomated safety check: PassMIT
Senior Computer Visionalirezarezvani/claude-skills28k2 repos~3.2kAutomated safety check: PassMIT
Deepstream Import Vision ModelNVIDIA/skills3.5k—~3.6kAutomated safety check: PassApache-2.0
Tao Finetune ClipNVIDIA/skills3.5k—~4kAutomated safety check: NotesApache-2.0
Tao Port Huggingface ModelNVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0
Computer Vision Engineertheneoai/awesome-skills183—~2.2kAutomated safety check: PassMIT

Similar skills

  • Senior Computer Vision

    alirezarezvani/claude-skills

    Computer vision engineering skill for object detection, image segmentation, and visual AI systems.

    28k GitHub starsUsed in 2 repos~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    A skill your agent uses to bring a supported object-detection vision model from HuggingFace or NVIDIA NGC into an NVIDIA DeepStream pipeline with end-to-end automation: ONNX download, SafeTensors…

    3.5k GitHub stars~3.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Tao Finetune Clip

    NVIDIA/skills

    Official

    CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

    3.5k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Official

    Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

    3.5k GitHub stars~4.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Computer Vision Engineer

    theneoai/awesome-skills

    Elite Computer Vision Engineer skill with expertise in deep learning for images and video (CNNs, Transformers), object detection (YOLO, DETR), segmentation, OCR, and production CV deployment…

    183 GitHub stars~2.2k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Robot Perception Engineer

    theneoai/awesome-skills

    Expert robot perception engineer specializing in 3D point cloud processing, multi-modal sensor fusion (camera+LiDAR+IMU), real-time SLAM, and edge-optimized deep learning inference via TensorRT/ONNX…

    183 GitHub stars~1.5k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed

More from borghei/Claude-Skills

All 364 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    874 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    874 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    874 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    874 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Senior Computer Vision

What does Senior Computer Vision do?

Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment. Senior Computer Vision is an agent skill from borghei/Claude-Skills. Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment.

When should I use Senior Computer Vision?

Senior Computer Vision fits situations like: building detection pipelines; training models; optimizing inference.

How do I install Senior Computer Vision in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill senior-computer-vision -a claude-code`. Or copy the skill folder (engineering/senior-computer-vision in borghei/Claude-Skills) into .claude/skills/senior-computer-vision in your project. Claude Code loads it when a task matches its description.

How do I install Senior Computer Vision in Codex?

Run `npx skills add borghei/Claude-Skills --skill senior-computer-vision -a codex`. Or copy the skill folder (engineering/senior-computer-vision in borghei/Claude-Skills) into .agents/skills/senior-computer-vision in your project. Codex loads it when a task matches its description.

Can I use Senior Computer Vision in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill senior-computer-vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/senior-computer-vision, .gemini/skills/senior-computer-vision, .github/skills/senior-computer-vision and .opencode/skills/senior-computer-vision in your project.

What does Senior Computer Vision need to run?

Going by SKILL.md and its folder, Senior Computer Vision needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Senior Computer Vision access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Senior Computer Vision safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Senior Computer Vision use?

Senior Computer Vision is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Senior Computer Vision use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 24k tokens, read only when the agent opens those files.

What are the alternatives to Senior Computer Vision?

Skills that share tags, products or a category with Senior Computer Vision: Senior Computer Vision (alirezarezvani/claude-skills, 28k stars), Deepstream Import Vision Model (NVIDIA/skills, 3.5k stars), Tao Finetune Clip (NVIDIA/skills, 3.5k stars) and Tao Port Huggingface Model (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Senior Computer Vision?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.