Agent skill

Computer Vision

by FerroxLabs in FerroxLabs/wayland

Computer vision implementation covering image classification, object detection (YOLO), image segmentation, OCR (Tesseract), face detection, image preprocessing, data augmentation, transfer learning…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Computer Vision

skills CLI
$ npx skills add FerroxLabs/wayland --skill computer-vision -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FerroxLabs/wayland computer-vision --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/ai-machine-learning/computer-vision .claude/skills/computer-vision && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
computer-vision
GitHub stars
608
Token cost
~3.9k tokens
SKILL.md length
518 words
Files
1
Skills in repo
1,194
Repo updated
First seen
Licence
Apache-2.0

At a glance

Computer vision implementation covering image classification, object detection (YOLO), image segmentation, OCR (Tesseract), face detection, image preprocessing, data augmentation, transfer learning…

  • The user asks about computer vision
  • SKILL.md covers Overview, Image Preprocessing, Data Augmentation and Image Classification, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Computer vision best practices

What it does

Computer Vision is an agent skill from FerroxLabs/wayland. Computer vision implementation covering image classification, object detection (YOLO), image segmentation, OCR (Tesseract), face detection, image preprocessing, data augmentation, transfer learning, model deployment, and edge inference. Use when the user asks about computer vision, computer vision best practices, or needs guidance on computer vision implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Computer vision. It works with PyTorch and OpenCV. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.

When your agent uses it

  • The user asks about computer vision
  • Computer vision best practices
  • Needs guidance on computer vision implementation
  • The user needs a different specialized skill

Example prompts

  • “/computer-vision”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, yaml and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Computer Vision loads about 3.9k tokens when it runs. Until then it costs about 124 tokens; SKILL.md has 518 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 518 words, ~3,899 tokens.

Download SKILL.mdSave it as .claude/skills/computer-vision/SKILL.md (or your agent's skills folder).
name
computer-vision
description
Computer vision implementation covering image classification, object detection (YOLO), image segmentation, OCR (Tesseract), face detection, image preprocessing, data augmentation, transfer learning, model deployment, and edge inference. Use when the user asks about computer vision, computer vision best practices, or needs guidance on computer vision implementation. Do NOT use when the user needs a different specialized skill or is asking about an unrelated technology domain.
license
Apache-2.0
metadata.author
foundry-skills
metadata.version
1.0.0
metadata.tags
ai-ml deep-learning guide
metadata.category
ai-machine-learning
metadata.subcategory
applied-ai
metadata.disclaimer
none
metadata.difficulty
intermediate

Computer Vision

Overview

Computer vision enables machines to interpret and act on visual data. This skill covers practical implementation of core CV tasks -- classification, detection, segmentation, and OCR -- using modern frameworks and pre-trained models, with guidance on data augmentation, transfer learning, and deployment.

Image Preprocessing

Standard Pipeline
python
import cv2
import numpy as np
from PIL import Image
import torchvision.transforms as T

class ImagePreprocessor:
    """Standard image preprocessing for CV models."""

    def __init__(self, target_size: tuple = (224, 224), normalize: bool = True):
        self.target_size = target_size
        self.normalize = normalize

    def preprocess_opencv(self, image_path: str) -> np.ndarray:
        """Preprocess with OpenCV."""
        img = cv2.imread(image_path)
        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)

        # Resize
        # ... (condensed) ...

        transforms.extend([
            T.ToTensor(),
            T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
        ])

        return T.Compose(transforms)
Common Preprocessing Operations
OperationWhen to UseLibrary
ResizeAlways (model input size)cv2, PIL, torchvision
NormalizeAlways (ImageNet stats for pretrained)torchvision
GrayscaleOCR, edge detectioncv2
Histogram equalizationLow contrast imagescv2 (CLAHE)
DenoisingNoisy imagescv2.fastNlMeansDenoising
CropFocus on region of interestcv2, PIL

Data Augmentation

python
import albumentations as A
from albumentations.pytorch import ToTensorV2

def get_training_augmentations(image_size: int = 224) -> A.Compose:
    """Production training augmentation pipeline."""
    return A.Compose([
        A.RandomResizedCrop(height=image_size, width=image_size, scale=(0.8, 1.0)),
        A.HorizontalFlip(p=0.5),
        A.VerticalFlip(p=0.1),
        A.ShiftScaleRotate(
            shift_limit=0.1, scale_limit=0.15, rotate_limit=15, p=0.5
        ),
        A.OneOf([
            A.GaussNoise(var_limit=(10.0, 50.0)),
            A.GaussianBlur(blur_limit=(3, 7)),
            A.MotionBlur(blur_limit=5),
        ], p=0.3),
        A.OneOf([
            # ... (condensed) ...
        A.RandomResizedCrop(height=image_size, width=image_size, scale=(0.5, 1.0)),
        A.HorizontalFlip(p=0.5),
        A.RandomBrightnessContrast(p=0.3),
        A.HueSaturationValue(p=0.3),
        A.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),
        ToTensorV2(),
    ], bbox_params=A.BboxParams(format='pascal_voc', label_fields=['labels']))
Augmentation Strategy by Dataset Size
Dataset size < 500 images:
  -> Heavy augmentation + transfer learning essential
  -> Use MixUp, CutMix, Mosaic
  -> Consider synthetic data generation

Dataset size 500-5000:
  -> Moderate augmentation
  -> Standard flips, crops, color jitter
  -> Transfer learning highly recommended

Dataset size > 5000:
  -> Light augmentation
  -> Basic flips and normalization
  -> Can train from scratch for simple tasks

Image Classification

Transfer Learning with PyTorch
python
import torch
import torch.nn as nn
from torchvision import models

def create_classifier(
    num_classes: int,
    backbone: str = "resnet50",
    pretrained: bool = True,
    freeze_backbone: bool = True,
) -> nn.Module:
    """Create a classifier with transfer learning."""

    if backbone == "resnet50":
        model = models.resnet50(weights="IMAGENET1K_V2" if pretrained else None)
        num_features = model.fc.in_features
        model.fc = nn.Sequential(
            nn.Dropout(0.3),
            nn.Linear(num_features, num_classes),
        # ... (condensed) ...

    # Freeze backbone layers
    if freeze_backbone:
        for param in list(model.parameters())[:-2]:
            param.requires_grad = False

    return model
Hugging Face Image Classification
python
from transformers import pipeline, AutoModelForImageClassification, AutoImageProcessor

# Quick inference with pipeline
classifier = pipeline("image-classification", model="google/vit-base-patch16-224")
result = classifier("photo.jpg")
# [{"label": "golden retriever", "score": 0.95}, ...]

# Fine-tuning
from transformers import TrainingArguments, Trainer
from datasets import load_dataset

dataset = load_dataset("imagefolder", data_dir="./data")

model = AutoModelForImageClassification.from_pretrained(
    "google/vit-base-patch16-224",
    num_labels=len(dataset["train"].features["label"].names),
    ignore_mismatched_sizes=True,
)
# ... (condensed) ...
        load_best_model_at_end=True,
    ),
    train_dataset=dataset["train"],
    eval_dataset=dataset["test"],
)

trainer.train()
Model Selection Guide
ModelParamsAccuracy (ImageNet)SpeedBest For
MobileNetV3-Small2.5M67.4%Very fastMobile/edge
EfficientNet-B05.3M77.1%FastBalanced
ResNet-5025M80.4%MediumGeneral purpose
EfficientNet-V2-S21M84.2%MediumHigh accuracy
ViT-B/1686M84.5%SlowerMax accuracy
ConvNeXt-Base88M85.8%SlowerSOTA accuracy

Object Detection

YOLOv8 (Ultralytics)
python
from ultralytics import YOLO

# Load pretrained model
model = YOLO("yolov8n.pt")  # nano (fastest)
# model = YOLO("yolov8s.pt")  # small
# model = YOLO("yolov8m.pt")  # medium
# model = YOLO("yolov8l.pt")  # large

# Inference
results = model("image.jpg")

for result in results:
    boxes = result.boxes
    for box in boxes:
        cls = int(box.cls[0])
        conf = float(box.conf[0])
        xyxy = box.xyxy[0].tolist()  # [x1, y1, x2, y2]
        label = model.names[cls]
        # ... (condensed) ...
    patience=20,
    device="0",  # GPU
)

# Export for deployment
model.export(format="onnx")
model.export(format="tflite")
YOLO Dataset Format
yaml
# dataset.yaml
path: /path/to/dataset
train: images/train
val: images/val
test: images/test

names:
  0: person
  1: car
  2: bicycle
# Label format (one .txt per image)
# class_id center_x center_y width height (all normalized 0-1)
0 0.5 0.6 0.3 0.4
1 0.2 0.3 0.15 0.2
Detection Model Comparison
ModelmAP@50 (COCO)Speed (ms)ParamsBest For
YOLOv8n37.31.23.2MReal-time, edge
YOLOv8s44.92.011.2MBalanced
YOLOv8m50.24.225.9MHigh accuracy
YOLOv8x53.98.768.2MMax accuracy
RT-DETR-L53.09.332MTransformer-based

Image Segmentation

Semantic Segmentation
python
from transformers import pipeline

segmenter = pipeline("image-segmentation", model="nvidia/segformer-b0-finetuned-ade-512-512")

results = segmenter("street_scene.jpg")
for segment in results:
    print(f"{segment['label']}: score={segment['score']:.2f}")
Instance Segmentation with SAM (Segment Anything)
python
from segment_anything import sam_model_registry, SamPredictor, SamAutomaticMaskGenerator

# Load SAM model
sam = sam_model_registry["vit_h"](checkpoint="sam_vit_h.pth")
sam.to("cuda")

# Automatic mask generation
mask_generator = SamAutomaticMaskGenerator(sam)
masks = mask_generator.generate(image)

# Point-prompted segmentation
predictor = SamPredictor(sam)
predictor.set_image(image)

# Click a point to segment the object there
masks, scores, logits = predictor.predict(
    point_coords=np.array([[500, 375]]),
    point_labels=np.array([1]),  # 1 = foreground
    multimask_output=True,
)

# Use the highest-scoring mask
best_mask = masks[np.argmax(scores)]
YOLO Segmentation
python
from ultralytics import YOLO

model = YOLO("yolov8n-seg.pt")  # Instance segmentation

results = model("image.jpg")

for result in results:
    if result.masks:
        for mask, box in zip(result.masks.data, result.boxes):
            cls = int(box.cls[0])
            label = model.names[cls]
            binary_mask = mask.cpu().numpy()  # H x W binary mask

OCR (Optical Character Recognition)

Tesseract OCR
python
import pytesseract
from PIL import Image
import cv2

def extract_text(image_path: str, preprocess: bool = True) -> str:
    """Extract text from image using Tesseract."""
    img = cv2.imread(image_path)

    if preprocess:
        # Convert to grayscale
        gray = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)

        # Adaptive thresholding
        binary = cv2.adaptiveThreshold(
            gray, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C,
            cv2.THRESH_BINARY, 11, 2
        )

        # ... (condensed) ...
                    'x': data['left'][i],
                    'y': data['top'][i],
                    'width': data['width'][i],
                    'height': data['height'][i],
                },
            })
    return results
EasyOCR (Multi-language)
python
import easyocr

reader = easyocr.Reader(['en', 'fr', 'de'])

results = reader.readtext('document.jpg')
for (bbox, text, confidence) in results:
    print(f"'{text}' (confidence: {confidence:.2f})")

Face Detection

MediaPipe (Fastest)
python
import mediapipe as mp

mp_face = mp.solutions.face_detection
mp_drawing = mp.solutions.drawing_utils

def detect_faces(image_path: str) -> list[dict]:
    """Detect faces using MediaPipe."""
    with mp_face.FaceDetection(model_selection=1, min_detection_confidence=0.5) as detector:
        image = cv2.imread(image_path)
        rgb = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)

        results = detector.process(rgb)

        faces = []
        if results.detections:
            h, w = image.shape[:2]
            for detection in results.detections:
                bbox = detection.location_data.relative_bounding_box
                faces.append({
                    "confidence": detection.score[0],
                    "bbox": {
                        "x": int(bbox.xmin * w),
                        "y": int(bbox.ymin * h),
                        "width": int(bbox.width * w),
                        "height": int(bbox.height * h),
                    },
                })
        return faces

Model Deployment

ONNX Export
python
import torch

def export_to_onnx(model, sample_input, output_path: str = "model.onnx"):
    """Export PyTorch model to ONNX format."""
    model.cpu()
    torch.onnx.export(
        model,
        sample_input,
        output_path,
        export_params=True,
        opset_version=17,
        input_names=['input'],
        output_names=['output'],
        dynamic_axes={
            'input': {0: 'batch_size'},
            'output': {0: 'batch_size'},
        },
    )
# ... (condensed) ...
    session = ort.InferenceSession(
        model_path,
        providers=['CUDAExecutionProvider', 'CPUExecutionProvider']
    )
    input_name = session.get_inputs()[0].name
    output = session.run(None, {input_name: input_data})
    return output[0]
Edge Inference with TFLite
python
import tensorflow as tf

def convert_to_tflite(saved_model_path: str, output_path: str, quantize: bool = True):
    """Convert model to TFLite for edge deployment."""
    converter = tf.lite.TFLiteConverter.from_saved_model(saved_model_path)

    if quantize:
        converter.optimizations = [tf.lite.Optimize.DEFAULT]
        converter.target_spec.supported_types = [tf.float16]

    tflite_model = converter.convert()
    with open(output_path, 'wb') as f:
        f.write(tflite_model)
Deployment Decision Tree
Target platform?
  CLOUD (GPU available):
    -> PyTorch/TensorRT (best performance)
    -> ONNX Runtime (portable)

  EDGE (mobile, IoT):
    -> TFLite (Android, Raspberry Pi)
    -> CoreML (iOS, macOS)
    -> ONNX Runtime Mobile

  BROWSER:
    -> TensorFlow.js
    -> ONNX Runtime Web

Latency requirement?
  < 10ms: TensorRT or hardware-specific optimization
  10-100ms: ONNX Runtime with GPU
  > 100ms: Any framework works

Checklist

  • Set up image preprocessing pipeline with proper normalization
  • Implement data augmentation appropriate to dataset size
  • Choose model architecture based on accuracy vs speed tradeoffs
  • Use transfer learning for datasets under 10K images
  • Set up proper train/val/test splits
  • Track training metrics with experiment tracker
  • Export model to deployment format (ONNX, TFLite, TensorRT)
  • Benchmark inference latency on target hardware
  • Implement preprocessing in the deployment pipeline (not just training)
  • Test with edge cases (poor lighting, occlusion, unusual angles)
  • Monitor model performance in production
Show full SKILL.md (198 more words)Show less

When to Use

Use this skill when:

  • Designing or implementing computer vision solutions
  • Reviewing or improving existing computer vision approaches
  • Making architectural or implementation decisions about computer vision
  • Learning computer vision patterns and best practices
  • Troubleshooting computer vision-related issues

Do NOT use this skill when:

  • The question is about a fundamentally different technology domain
  • A more specific sibling skill covers the exact topic needed
  • The user needs a complete hands-on tutorial rather than expert guidance

Output Format

markdown
# Computer Vision Analysis

## Context Assessment
[Situation summary and constraints]

## Recommended Approach
[Primary recommendation with rationale]

## Implementation Steps
1. [Step with specific details]
2. [Step with specific details]
3. [Step with specific details]

## Trade-offs and Considerations
- [Key trade-off 1]
- [Key trade-off 2]

## Next Steps
- [Immediate action item]
- [Follow-up action item]

Example

Input: "Help me implement computer vision for a medium-scale production application"

Output: A structured analysis covering current state assessment, recommended computer vision approach with specific patterns, implementation roadmap with milestones, and risk mitigation strategies tailored to the application scale and constraints.

Edge Cases

  • Legacy system integration: When computer vision must coexist with legacy approaches, provide a gradual migration path rather than a complete rewrite
  • Scale mismatch: When the solution complexity exceeds the project scale, recommend a simpler approach and note when to revisit
  • Team skill gaps: When the team lacks experience with the recommended approach, include learning resources and simpler alternatives
  • Conflicting requirements: When constraints conflict (e.g., performance vs. maintainability), explicitly state the trade-off and recommend based on stated priorities

© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/process/resources/skills-library/bodies/skills/ai-machine-learning/computer-vision of FerroxLabs/wayland.

Open the folder on GitHubat commit 4c030c7

Compare with similar skills

Computer Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Computer Vision compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Computer Vision this skillFerroxLabs/wayland608—~3.9kAutomated safety check: PassApache-2.0
Senior Computer Visiondavila7/claude-code-templates32k3 repos~1.4kAutomated safety check: PassMIT
Computer Vision Pipelinecuriositech/some_claude_skills2431 repos~4kAutomated safety check: PassMIT
Torch Performance Optimizationalbumentations-team/albucore123—~895Automated safety check: PassMIT
Senior Computer Visionalirezarezvani/claude-skills28k2 repos~3.2kAutomated safety check: PassMIT
Tao Finetune ClipNVIDIA/skills3.5k—~4kAutomated safety check: NotesApache-2.0

Similar skills

  • Senior Computer Vision

    davila7/claude-code-templates

    World-class computer vision skill for image/video processing, object detection, segmentation, and visual AI systems.

    32k GitHub starsUsed in 3 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Computer Vision Pipeline

    curiositech/some_claude_skills

    Build production computer vision pipelines for object detection, tracking, and video analysis.

    243 GitHub starsUsed in 1 repo~4k tokens
    AI & LLM EngineeringAuto-check passed
  • Torch Performance Optimization

    albumentations-team/albucore

    Optimize or review eager CPU-only Albucore PyTorch runtime paths with benchmark-backed decisions.

    123 GitHub stars~895 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Senior Computer Vision

    alirezarezvani/claude-skills

    Computer vision engineering skill for object detection, image segmentation, and visual AI systems.

    28k GitHub starsUsed in 2 repos~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Tao Finetune Clip

    NVIDIA/skills

    Official

    CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

    3.5k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Official

    Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches.

    3.5k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from FerroxLabs/wayland

All 1,194 skills in this repo
  • Star Office Helper

    FerroxLabs/wayland

    Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.

    608 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Openclaw Setup

    FerroxLabs/wayland

    OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.

    608 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Tvcontrol Setup

    FerroxLabs/wayland

    Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.

    608 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Ab Testing Specialist

    FerroxLabs/wayland

    End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

    608 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Academic Writer

    FerroxLabs/wayland

    Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…

    608 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Accessibility Auditor

    FerroxLabs/wayland

    Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…

    608 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Computer Vision

What does Computer Vision do?

Computer vision implementation covering image classification, object detection (YOLO), image segmentation, OCR (Tesseract), face detection, image preprocessing, data augmentation, transfer learning…. Computer Vision is an agent skill from FerroxLabs/wayland. Computer vision implementation covering image classification, object detection (YOLO), image segmentation, OCR (Tesseract), face detection, image preprocessing, data augmentation, transfer learning, model deployment, and edge inference.

When should I use Computer Vision?

Computer Vision fits situations like: the user asks about computer vision; computer vision best practices; needs guidance on computer vision implementation; the user needs a different specialized skill.

How do I install Computer Vision in Claude Code?

Run `npx skills add FerroxLabs/wayland --skill computer-vision -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/ai-machine-learning/computer-vision in FerroxLabs/wayland) into .claude/skills/computer-vision in your project. Claude Code loads it when a task matches its description.

How do I install Computer Vision in Codex?

Run `npx skills add FerroxLabs/wayland --skill computer-vision -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/ai-machine-learning/computer-vision in FerroxLabs/wayland) into .agents/skills/computer-vision in your project. Codex loads it when a task matches its description.

Can I use Computer Vision in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill computer-vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/computer-vision, .gemini/skills/computer-vision, .github/skills/computer-vision and .opencode/skills/computer-vision in your project.

What does Computer Vision need to run?

SKILL.md names no scripts, command-line tools or credentials: Computer Vision is instructions for the agent only. Our summary lists: Python 3.

Does Computer Vision access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Computer Vision safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Computer Vision use?

Computer Vision is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Computer Vision use?

About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Computer Vision?

Skills that share tags, products or a category with Computer Vision: Senior Computer Vision (davila7/claude-code-templates, 32k stars), Computer Vision Pipeline (curiositech/some_claude_skills, 243 stars), Torch Performance Optimization (albumentations-team/albucore, 123 stars) and Senior Computer Vision (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Computer Vision?

FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 1,194 skills in this directory. The repository was last updated on October 6, 2026.

Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.