Agent skill

Detection

by VectorSpaceLab in VectorSpaceLab/AREX-Skill

A skill your agent uses for PaddleViT object detection workflows with DETR, Swin, or PVTv2: validate COCO data, select configs, build/train/evaluate models, reason about transforms, losses…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Detection

skills CLI
$ npx skills add VectorSpaceLab/AREX-Skill --skill detection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install VectorSpaceLab/AREX-Skill detection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/VectorSpaceLab/AREX-Skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/repositories/repo-skills/paddlevit/sub-skills/detection .claude/skills/detection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
detection
GitHub stars
328
Token cost
~2.4k tokens
SKILL.md length
1,046 words
Files
7 (incl. scripts, references)
Skills in repo
157
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses for PaddleViT object detection workflows with DETR, Swin, or PVTv2: validate COCO data, select configs, build/train/evaluate models, reason about transforms, losses…

  • Works in 3 steps: Identify the family before changing a… → If the request says only “PaddleViT… → Keep segmentation separate, even if a…
  • PaddleViT object detection workflows with DETR
  • SKILL.md covers Route the request, Preflight before a source run, Select a source root and config and CLI flag contract, plus 3 more sections
  • Runs Python scripts from its folder; calls python

What it does

Detection is an agent skill from VectorSpaceLab/AREX-Skill. Use for PaddleViT object detection workflows with DETR, Swin, or PVTv2: validate COCO data, select configs, build/train/evaluate models, reason about transforms, losses, post-processing, and run safe utility smokes. Excludes segmentation and generic export; cross-link deployment-and-operations for shared runtime concerns.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/data-formats.md`, `references/model-overview.md` and `references/troubleshooting.md`).

It sits in AI & LLM Engineering, covering Computer vision and Deployment. The repository describes itself as: A Skill Library for Automated Machine Learning. The licence is Apache-2.0.

When your agent uses it

  • PaddleViT object detection workflows with DETR
  • PVTv2: validate COCO data
  • Build/train/evaluate models
  • Reason about transforms

Example prompts

  • “/detection”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Identify the family before changing a config or launching a command
  2. If the request says only “PaddleViT detection,” ask which family and
  3. Keep segmentation separate, even if a request mentions masks or a DETR

What it can do on your machine

Read from SKILL.md and the folder at commit ac3fe1a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Detection loads about 2.4k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 1,046 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from VectorSpaceLab/AREX-Skill at commit ac3fe1a, republished under its Apache-2.0 licence (© VectorSpaceLab). 1,046 words, ~2,411 tokens.

Download SKILL.mdSave it as .claude/skills/detection/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
detection
description
Use for PaddleViT object detection workflows with DETR, Swin, or PVTv2: validate COCO data, select configs, build/train/evaluate models, reason about transforms, losses, post-processing, and run safe utility smokes. Excludes segmentation and generic export; cross-link deployment-and-operations for shared runtime concerns.
disable-model-invocation
true
metadata.disco-role
operating
license
Apache 2.0

PaddleViT object detection

Use this route for standalone PaddleViT object detection under object_detection/: DETR, Swin, and PVTv2 with COCO boxes. It covers family/config selection, data preflight, transforms and targets, safe model-contract checks, checkpoint triage, training/evaluation boundaries, and backend-aware diagnosis. It does not own semantic segmentation, generic export/inference, quantization, or weight porting; route those requests to deployment-and-operations.

The bundled helpers are self-contained and do not import the original source checkout, download data or weights, or modify a user-owned dataset. A source checkout is optional and is needed only for a source-model build or native candidate test. Details that are useful during a focused investigation live in model overview, data formats, workflows, and troubleshooting.

Route the request

  1. Identify the family before changing a config or launching a command:
    • DETR for query-based, end-to-end detection with transformer decoder queries, Hungarian matching, and direct box post-processing.
    • Swin or PVTv2 for hierarchical backbones feeding FPN, RPN, and RoI-style heads with anchors, proposals, and NMS.
  2. If the request says only “PaddleViT detection,” ask which family and whether the goal is a utility check, source build, evaluation, or training. Do not mix family directories in one Python process: common modules such as config, coco, box_ops, and utils can resolve incorrectly.
  3. Keep segmentation separate, even if a request mentions masks or a DETR import. For export, predictor, AMP, distributed, or shared runtime issues, use deployment-and-operations.

Preflight before a source run

Validate the COCO root, not a split directory. The expected root contains annotations/instances_{train,val}2017.json and the matching train2017/ or val2017/ image directory. Run the read-only helper from the skill root (replace placeholders with caller-owned values):

bash
python <skill-root>/scripts/check_coco_layout.py <coco-root> --split val

Add --check-images to decode referenced images and compare dimensions, or --check-api to parse the annotation through pycocotools. Use --json for a machine-readable report and --demo for a temporary one-image validator fixture. A failed check is a stop condition: repair the dataset outside this skill and rerun; do not download or silently substitute another dataset.

Before importing a source model, run the bounded synthetic contract smoke:

bash
python <skill-root>/scripts/detection_model_smoke.py --model all --device cpu

This checks tiny DETR and four-level anchor-family shapes and finite values. It is not a source-model build, checkpoint test, COCO test, mAP result, or benchmark reproduction. Use --device gpu:0 only when a GPU claim is required and the backend has already passed its environment probe.

Select a source root and config

The three projects are standalone. For a source-backed run, enter exactly one family directory and expose only that directory (and its intended parent) to imports:

bash
cd <source-root>/object_detection/DETR   # or Swin or PVTv2
export PYTHONPATH="$PWD:$PWD/..:${PYTHONPATH:-}"

Resolve configuration in this order: Python defaults, recursive YAML BASE files, then CLI overrides. Print and inspect the effective config before model construction. Check backbone output channels against FPN inputs, pyramid strides and anchor levels, class/category conventions, image divisibility, and DETR embedding-dimension/head divisibility. Swin/PVTv2 configs preserve the historical ROI.NUM_ClASSES spelling; do not replace it with a guessed key. The family-specific config and output/target details are in model overview and data formats.

CLI flag contract

The family main_single_gpu.py and main_multi_gpu.py scripts use these short, single-dash options:

text
-cfg PATH             YAML config (BASE files are recursively merged)
-dataset coco         dataset selector
-data_path PATH       COCO root, not train2017/val2017
-batch_size N         per-process/per-GPU batch size
-eval                 evaluation-only mode
-pretrained PREFIX    source appends .pdparams
-resume PREFIX        source expects .pdparams and .pdopt
-last_epoch N         resume epoch metadata
-ngpus N              configured multi-GPU worker count

-cfg=... and -cfg ... are both acceptable. Pass checkpoint prefixes without .pdparams when following the source loader. Inspect launcher shell files rather than executing them blindly: paths, visible devices, output prefixes, and training duration are caller-owned. A representative command shape is:

bash
cd <source-root>/object_detection/DETR
CUDA_VISIBLE_DEVICES=0 python main_single_gpu.py \
  -cfg=./configs/detr_resnet50.yaml -dataset=coco \
  -batch_size=1 -data_path=<coco-root> -eval \
  -pretrained=<checkpoint-prefix>

Treat real training/evaluation as expensive and data/checkpoint/GPU dependent; use the workflow reference for gated sequencing.

Family contracts to preserve

  • DETR: expect logits shaped [B,Q,C+1], normalized center-size boxes [B,Q,4], optional auxiliary decoder outputs, and losses including classification, L1, and GIoU. Post-processing needs target sizes in [height,width] order and emits absolute xyxy boxes, scores, and labels.
  • Swin/PVTv2: expect hierarchical features, FPN/RPN/RoI losses during training, and post-NMS rows shaped [label, score, xmin, ymin, xmax, ymax] during evaluation. Their target path uses absolute boxes and contiguous classes; the two families share the broad head/neck contract but not every backbone channel/config value.
  • All families: COCO boxes begin as [x,y,width,height]; transformations must update geometry and area. A valid synthetic training fixture needs at least one non-empty target. Preserve original image IDs for COCO results. Full transform and output schemas are in the linked references.
Show full SKILL.md (374 more words)Show less

Verification tiers

Choose the cheapest tier that answers the question and record command, device, Paddle version, config, and result:

  1. check_coco_layout.py for root layout, JSON arrays, IDs, boxes, files, and optional Pillow/COCO-API checks.
  2. detection_model_smoke.py for standalone tiny shape and finite-value contracts; a non-32-divisible size intentionally reports the padding caveat.
  3. Native DETR box/transform/model tests when a source checkout and dependencies are available. They are evidence candidates, not runtime prerequisites; some are skipped or require an untracked fixture.
  4. A one-batch source build/forward using one family root and a tiny local fixture, only after config/import checks pass.
  5. Real COCO evaluation or training only with explicit data, checkpoint, GPU, and time approval. Never call a smoke or random-weight forward mAP evidence.

CPU can validate parsing, layout, box utilities, and tiny diagnostics. It does not establish CUDA, AMP, distributed, or full detector claims. For those boundaries use deployment-and-operations.

Stop and recover

  • Missing files/API: -data_path must be the root; verify exact split filenames. If pycocotools is absent, omit --check-api but do not claim COCO evaluation is verified.
  • Import/YAML errors: start a fresh process, isolate one family root, resolve BASE relative to its YAML, compare keys to that family's config, and inspect the final config. Do not mix module paths or claim a build that never constructed a model.
  • Shape/checkpoint errors: verify family, backbone/FPN channels, attention divisibility, query/class counts, head spelling, and checkpoint provenance. A prefix normally maps to .pdparams; resume also needs .pdopt. Do not reshape incompatible state dictionaries.
  • Bad boxes/empty batches: inspect crowd flags, category mapping, positive area, coordinate format, resize/flip/crop order, and image IDs. Keep one valid target in a fixture; an all-empty batch is not a valid smoke.
  • NaN/Inf or no detections: run CPU geometry checks, verify normalized DETR boxes or absolute Swin/PVTv2 boxes, then inspect thresholds, NMS, and loaded keys. Random weights cannot support an accuracy claim.
  • CUDA OOM/unavailable: lower per-GPU batch size or input scale for OOM. If CUDA is unavailable, report CPU-only partial evidence; never substitute it for a required GPU result. For multi-GPU, check visible devices, -ngpus, per-process batch semantics, NCCL, and rank gathering before retrying.

See troubleshooting for expanded recovery branches. Do not run network, full-native, long-training, or multi-GPU workflows as part of this route unless explicitly authorized.

© VectorSpaceLab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/repositories/repo-skills/paddlevit/sub-skills/detection of VectorSpaceLab/AREX-Skill.

  • SKILL.md
  • references/data-formats.md
  • references/model-overview.md
  • references/troubleshooting.md
  • references/workflows.md
  • scripts/check_coco_layout.py
  • scripts/detection_model_smoke.py

Open the folder on GitHubat commit ac3fe1a

Compare with similar skills

Detection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Detection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Detection this skillVectorSpaceLab/AREX-Skill328—~2.4kAutomated safety check: PassApache-2.0
Cvatmajiayu000/claude-skill-registry6661 repos~909Automated safety check: PassMIT
Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit1.1k—~3.1kAutomated safety check: PassCustom licence
Azure AI Vision ReferenceMicrosoftDocs/Agent-Skills777—~1.6kAutomated safety check: PassCC-BY-4.0
Uav Vision Analyticsopen-edge-platform/edge-ai-suites140—~2.6kAutomated safety check: NotesApache-2.0
Azure Custom VisionMicrosoftDocs/Agent-Skills777—~1.6kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Cvat

    majiayu000/claude-skill-registry

    Operate CVAT for computer-vision annotation, dataset workflows, SDK/CLI automation, auto-annotation, and self-hosted deployment.

    666 GitHub starsUsed in 1 repo~909 tokens
    AI & LLM EngineeringAuto-check passed
  • Matlab Use Visual Inspection

    matlab/matlab-agentic-toolkit

    Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.

    1.1k GitHub stars~3.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Azure AI Vision Reference

    MicrosoftDocs/Agent-Skills

    Official

    Looks up Microsoft Learn guidance for Azure AI Vision: Image Analysis, Read OCR containers, smart-crop thumbnails, background removal and video frame analysis, plus limits and deployment.

    777 GitHub stars~1.6k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Uav Vision Analytics

    open-edge-platform/edge-ai-suites

    Build an end-to-end UAV object detection and telemetry overlay application on Intel hardware using DL Streamer Pipeline Server with MAVLink telemetry.

    140 GitHub stars~2.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Azure Custom Vision

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure AI Custom Vision development including best practices, decision making, limits & quotas, security, integrations & coding patterns, and deployment.

    777 GitHub stars~1.6k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed

More from VectorSpaceLab/AREX-Skill

All 157 skills in this repo
  • Agent Lightning

    VectorSpaceLab/AREX-Skill

    Use this repo skill for Agent Lightning package tasks: authoring trainable agents, tracing rewards and spans, running LightningStore/Trainer loops, using agl CLI services, choosing examples, and…

    328 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Tools

    VectorSpaceLab/AREX-Skill

    A skill your agent uses when configuring LiteLLM for MCP tools, A2A agents, Claude Code/Cursor agent gateway traffic, MCP auth/OAuth, tool permissions, semantic filtering, or agent-specific proxy…

    328 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents And Awel

    VectorSpaceLab/AREX-Skill

    Build and debug DB-GPT agents, tools, skills, teams, and AWEL workflows, including deterministic local DAG runs and HTTP-trigger topology without assuming an LLM, credential, or external service.

    328 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents And Middleware

    VectorSpaceLab/AREX-Skill

    Work on the actively maintained LangChain v1 agent package: initchatmodel, createagent, structured output, tools, middleware, embeddings initialization, provider routing, and agent runtime…

    328 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents Workflows

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for giskard.agents async chat workflows, tools, prompt templates, structured outputs, retries, rate limiting, embeddings, and optional LiteLLM backend.

    328 GitHub stars~500 tokensUpdated 1 mo ago
    Auto-check passed
  • Alphafold3

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for AlphaFold 3 input preparation, prediction command planning, output interpretation, and Python API inspection.

    328 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Detection

What does Detection do?

A skill your agent uses for PaddleViT object detection workflows with DETR, Swin, or PVTv2: validate COCO data, select configs, build/train/evaluate models, reason about transforms, losses…. Detection is an agent skill from VectorSpaceLab/AREX-Skill. Use for PaddleViT object detection workflows with DETR, Swin, or PVTv2: validate COCO data, select configs, build/train/evaluate models, reason about transforms, losses, post-processing, and run safe utility smokes.

When should I use Detection?

Detection fits situations like: paddleViT object detection workflows with DETR; PVTv2: validate COCO data; build/train/evaluate models; reason about transforms.

How do I install Detection in Claude Code?

Run `npx skills add VectorSpaceLab/AREX-Skill --skill detection -a claude-code`. Or copy the skill folder (skills/repositories/repo-skills/paddlevit/sub-skills/detection in VectorSpaceLab/AREX-Skill) into .claude/skills/detection in your project. Claude Code loads it when a task matches its description.

How do I install Detection in Codex?

Run `npx skills add VectorSpaceLab/AREX-Skill --skill detection -a codex`. Or copy the skill folder (skills/repositories/repo-skills/paddlevit/sub-skills/detection in VectorSpaceLab/AREX-Skill) into .agents/skills/detection in your project. Codex loads it when a task matches its description.

Can I use Detection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add VectorSpaceLab/AREX-Skill --skill detection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/detection, .gemini/skills/detection, .github/skills/detection and .opencode/skills/detection in your project.

What does Detection need to run?

Going by SKILL.md and its folder, Detection needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Detection access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Detection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Detection use?

Detection is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Detection use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.6k tokens, read only when the agent opens those files.

What are the alternatives to Detection?

Skills that share tags, products or a category with Detection: Cvat (majiayu000/claude-skill-registry, 666 stars), Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars), Azure AI Vision Reference (MicrosoftDocs/Agent-Skills, 777 stars) and Uav Vision Analytics (open-edge-platform/edge-ai-suites, 140 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Detection?

VectorSpaceLab (a GitHub organization) maintains it in VectorSpaceLab/AREX-Skill, which has 328 GitHub stars. The repository holds 157 skills in this directory. The repository was last updated on September 3, 2026.

Source: VectorSpaceLab/AREX-Skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.