Agent skill

Training Evaluation

by VectorSpaceLab in VectorSpaceLab/AREX-Skill

Use Torch Points3D Hydra training, evaluation, Trainer, checkpoints, forward inference, visualization, and experiment-output helpers.

Custom licenceAuto-check passed

Install Training Evaluation

skills CLI
$ npx skills add VectorSpaceLab/AREX-Skill --skill training-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install VectorSpaceLab/AREX-Skill training-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/VectorSpaceLab/AREX-Skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/repositories/repo-skills/torch-points3d/sub-skills/training-evaluation .claude/skills/training-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
training-evaluation
GitHub stars
330
Token cost
~1.2k tokens
SKILL.md length
410 words
Files
9 (incl. scripts, references)
Skills in repo
159
Repo updated
First seen
Licence
Custom licence

At a glance

Use Torch Points3D Hydra training, evaluation, Trainer, checkpoints, forward inference, visualization, and experiment-output helpers.

  • SKILL.md covers Read First, Main Workflows, Boundary Rules and Safety Checklist
  • Runs Python scripts from its folder; calls python

What it does

Training Evaluation is an agent skill from VectorSpaceLab/AREX-Skill. Use Torch Points3D Hydra training, evaluation, Trainer, checkpoints, forward inference, visualization, and experiment-output helpers.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `references/configuration-and-checkpoints.md`, `references/forward-inference.md` and `references/training-evaluation-workflows.md`).

The repository describes itself as: A Skill Library for Automated Machine Learning.

Example prompts

  • “/training-evaluation”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit ac3fe1a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Training Evaluation loads about 1.2k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 38 tokens; SKILL.md has 410 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 410 words (~1,195 tokens).

“Use this sub-skill when the user asks to compose Torch Points3D train.py or eval.py commands, debug Trainer, resume or evaluate checkpoints, inspect outputs// runs, configure W&B/TensorBoard/visualization, or run forward inference on unlabeled data from a trained checkpoint.”

— opening of SKILL.md by VectorSpaceLab, Custom licence
name
training-evaluation
disable-model-invocation
true
metadata.disco-role
operating
license
NOASSERTION

Read the full SKILL.md on GitHub

Files

SKILL.md and 8 other files (scripts, references) in skills/repositories/repo-skills/torch-points3d/sub-skills/training-evaluation of VectorSpaceLab/AREX-Skill.

  • SKILL.md
  • references/configuration-and-checkpoints.md
  • references/forward-inference.md
  • references/training-evaluation-workflows.md
  • references/troubleshooting.md
  • scripts/compose_config_smoke.py
  • scripts/convert_checkpoint_omegaconf.py
  • scripts/forward_preflight.py
  • scripts/summarize_runs.py

Open the folder on GitHubat commit ac3fe1a

Compare with similar skills

Training Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Training Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Training Evaluation this skillVectorSpaceLab/AREX-Skill330—~1.2kAutomated safety check: PassCustom licence
Arize Evaluatorgithub/awesome-copilot40k1 repos~8.1kAutomated safety check: NotesMIT
Ray Train Distributed TrainingOrchestra-Research/AI-Research-SKILLs13k2 repos~2.7kAutomated safety check: PassMIT
Ito Trainingaffaan-m/ECC276k1 repos~1.5kAutomated safety check: PassMIT
LLM Evaluationdavila7/claude-code-templates32k12 repos~3.5kAutomated safety check: PassMIT
Train Poseruvnet/RuView97k—~504Automated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • Ray Train Distributed Training

    Orchestra-Research/AI-Research-SKILLs

    Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

    13k GitHub starsUsed in 2 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Ito Training

    affaan-m/ECC

    Inspect the availability of ML training on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed training manifest.

    276k GitHub starsUsed in 1 repo~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Pose

    ruvnet/RuView

    Train/evaluate WiFi pose models honestly — camera-supervised (MediaPipe + CSI) and camera-free (WiFlow), always checked against the mean-pose baseline before any PCK is quoted.

    97k GitHub stars~504 tokensUpdated today
    DevelopmentAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed

More from VectorSpaceLab/AREX-Skill

All 159 skills in this repo
  • Agent Lightning

    VectorSpaceLab/AREX-Skill

    Use this repo skill for Agent Lightning package tasks: authoring trainable agents, tracing rewards and spans, running LightningStore/Trainer loops, using agl CLI services, choosing examples, and…

    330 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Tools

    VectorSpaceLab/AREX-Skill

    A skill your agent uses when configuring LiteLLM for MCP tools, A2A agents, Claude Code/Cursor agent gateway traffic, MCP auth/OAuth, tool permissions, semantic filtering, or agent-specific proxy…

    330 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents And Awel

    VectorSpaceLab/AREX-Skill

    Build and debug DB-GPT agents, tools, skills, teams, and AWEL workflows, including deterministic local DAG runs and HTTP-trigger topology without assuming an LLM, credential, or external service.

    330 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents And Middleware

    VectorSpaceLab/AREX-Skill

    Work on the actively maintained LangChain v1 agent package: initchatmodel, createagent, structured output, tools, middleware, embeddings initialization, provider routing, and agent runtime…

    330 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents Workflows

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for giskard.agents async chat workflows, tools, prompt templates, structured outputs, retries, rate limiting, embeddings, and optional LiteLLM backend.

    330 GitHub stars~500 tokensUpdated 1 mo ago
    Auto-check passed
  • Alphafold3

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for AlphaFold 3 input preparation, prediction command planning, output interpretation, and Python API inspection.

    330 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Training Evaluation

What does Training Evaluation do?

Use Torch Points3D Hydra training, evaluation, Trainer, checkpoints, forward inference, visualization, and experiment-output helpers. Training Evaluation is an agent skill from VectorSpaceLab/AREX-Skill. Use Torch Points3D Hydra training, evaluation, Trainer, checkpoints, forward inference, visualization, and experiment-output helpers.

How do I install Training Evaluation in Claude Code?

Run `npx skills add VectorSpaceLab/AREX-Skill --skill training-evaluation -a claude-code`. Or copy the skill folder (skills/repositories/repo-skills/torch-points3d/sub-skills/training-evaluation in VectorSpaceLab/AREX-Skill) into .claude/skills/training-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Training Evaluation in Codex?

Run `npx skills add VectorSpaceLab/AREX-Skill --skill training-evaluation -a codex`. Or copy the skill folder (skills/repositories/repo-skills/torch-points3d/sub-skills/training-evaluation in VectorSpaceLab/AREX-Skill) into .agents/skills/training-evaluation in your project. Codex loads it when a task matches its description.

Can I use Training Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add VectorSpaceLab/AREX-Skill --skill training-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/training-evaluation, .gemini/skills/training-evaluation, .github/skills/training-evaluation and .opencode/skills/training-evaluation in your project.

What does Training Evaluation need to run?

Going by SKILL.md and its folder, Training Evaluation needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Training Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Training Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Training Evaluation use?

Training Evaluation has a licence file (declared in SKILL.md) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Training Evaluation use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Training Evaluation?

Skills that share tags, products or a category with Training Evaluation: Arize Evaluator (github/awesome-copilot, 40k stars), Ray Train Distributed Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), Ito Training (affaan-m/ECC, 276k stars) and LLM Evaluation (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Training Evaluation?

VectorSpaceLab (a GitHub organization) maintains it in VectorSpaceLab/AREX-Skill, which has 330 GitHub stars. The repository holds 159 skills in this directory. The repository was last updated on September 3, 2026.

Source: VectorSpaceLab/AREX-Skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.