Agent skill

Training And Evaluation

by VectorSpaceLab in VectorSpaceLab/AREX-Skill

A skill your agent uses for ModelScope trainer construction, TrainingArgs conversion, fine-tuning and evaluation preflight, checkpoint hooks, and safe train/eval command planning.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Training And Evaluation

skills CLI
$ npx skills add VectorSpaceLab/AREX-Skill --skill training-and-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install VectorSpaceLab/AREX-Skill training-and-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/VectorSpaceLab/AREX-Skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/repositories/repo-skills/modelscope/sub-skills/training-and-evaluation .claude/skills/training-and-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
training-and-evaluation
GitHub stars
331
Token cost
~1.4k tokens
SKILL.md length
596 words
Files
5 (incl. scripts, references)
Skills in repo
157
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses for ModelScope trainer construction, TrainingArgs conversion, fine-tuning and evaluation preflight, checkpoint hooks, and safe train/eval command planning.

  • Works in 5 steps: Identify whether the task is… → For preview/config planning, run the… → For real training/evaluation, require… → …
  • ModelScope trainer construction
  • SKILL.md covers Read this when, Route elsewhere, Primary references and helper and Safe default workflow, plus 3 more sections
  • Runs Python scripts from its folder; calls python

What it does

Training And Evaluation is an agent skill from VectorSpaceLab/AREX-Skill. Use for ModelScope trainer construction, TrainingArgs conversion, fine-tuning and evaluation preflight, checkpoint hooks, and safe train/eval command planning.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/training-args-reference.md`, `references/troubleshooting.md` and `references/workflows.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: A Skill Library for Automated Machine Learning. The licence is Apache-2.0.

When your agent uses it

  • ModelScope trainer construction
  • TrainingArgs conversion
  • Fine-tuning and evaluation preflight
  • Checkpoint hooks

Example prompts

  • “/training-and-evaluation”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Identify whether the task is preview/config planning, real training, real evaluation, or checkpoint/inference handoff.
  2. For preview/config planning, run the bundled helper with the proposed flags. Prefer previewing before writing a config file or launching…
  3. For real training/evaluation, require the user or environment to supply all external resources first: local model cache or trusted model…
  4. Build the trainer only after checking the config, dataset columns, metrics, checkpoint strategy, and backend requirements.
  5. Treat any CUDA, DeepSpeed, Megatron, vLLM, TensorFlow, audio/CV/NLP domain-extra execution, or large-model fine-tuning as optional and…

What it can do on your machine

Read from SKILL.md and the folder at commit ac3fe1a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Training And Evaluation loads about 1.4k tokens when it runs, and up to ~9.4k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 596 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from VectorSpaceLab/AREX-Skill at commit ac3fe1a, republished under its Apache-2.0 licence (© VectorSpaceLab). 596 words, ~1,415 tokens.

Download SKILL.mdSave it as .claude/skills/training-and-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
training-and-evaluation
description
Use for ModelScope trainer construction, TrainingArgs conversion, fine-tuning and evaluation preflight, checkpoint hooks, and safe train/eval command planning.
disable-model-invocation
true
metadata.disco-role
operating
license
Apache 2.0

ModelScope Training and Evaluation

Use this sub-skill when a task asks to train, fine-tune, evaluate, configure, or preflight a ModelScope training job. It focuses on package-level trainer APIs and safe planning. It does not perform long-running training, download models or datasets, or verify broad CUDA/domain recipes by itself.

Read this when

  • The user needs build_trainer(name, default_args) or an EpochBasedTrainer-style workflow.
  • The user has CLI-like fine-tuning flags and wants the effective ModelScope config before launch.
  • The user needs a config-file shape for train, evaluation, optimizer, LR scheduler, hooks, checkpointing, or distributed launch.
  • The user is adapting a ModelScope example recipe but still needs model/data/cache/GPU preflight.
  • The user asks about modelscope.tools.train or modelscope.tools.eval command shape.

Route elsewhere

  • For inference from a trained checkpoint or exported model, read ../pipelines-and-models/SKILL.md.
  • For MsDataset.load, local dataset recipes, config file parsing, or file IO details, read ../datasets-config/SKILL.md first, then return here for trainer wiring.
  • For Hub authentication, cache layout, model or dataset snapshot downloads, and offline/local-files-only planning, use ../hub-and-cli/SKILL.md.
  • For serving, export, checkpoint conversion utilities, or vLLM/server launch, use ../serving-export-and-tools/SKILL.md.
  • For implementing new ModelScope trainer classes or repository contribution tests, use ../customization-and-development/SKILL.md.

Primary references and helper

  • Read references/workflows.md for safe end-to-end training/evaluation recipes, preflight checklists, API and real-job CLI command shapes.
  • Read references/training-args-reference.md for TrainingArgs, parse_cli, to_config, config-node mapping, flattened optimizer/LR scheduler values, and dataset column mapping details.
  • Read references/troubleshooting.md for symptoms, likely causes, and recovery steps for trainer construction, data/config errors, checkpoints, distributed launch, and optional GPU/domain failures.
  • Run scripts/build_training_args_preview.py --help to inspect the bundled safe preview tool. The helper parses TrainingArgs-style flags and prints the effective config summary without importing ModelScope, downloading models, reading datasets, writing files, or launching training.

Safe default workflow

  1. Identify whether the task is preview/config planning, real training, real evaluation, or checkpoint/inference handoff.
  2. For preview/config planning, run the bundled helper with the proposed flags. Prefer previewing before writing a config file or launching any train/eval command.
  3. For real training/evaluation, require the user or environment to supply all external resources first: local model cache or trusted model id, dataset access or local files, optional extras, credentials, GPU/VRAM if needed, and a writable work directory.
  4. Build the trainer only after checking the config, dataset columns, metrics, checkpoint strategy, and backend requirements.
  5. Treat any CUDA, DeepSpeed, Megatron, vLLM, TensorFlow, audio/CV/NLP domain-extra execution, or large-model fine-tuning as optional and unverified for this production scope unless a later task explicitly verifies it in the target environment.
Show full SKILL.md (191 more words)Show less

Minimal API pattern

python
from modelscope.trainers import build_trainer

kwargs = dict(
    model="local-model-dir-or-trusted-model-id",
    train_dataset=train_dataset,
    eval_dataset=eval_dataset,
    max_epochs=1,
    work_dir="./work_dir",
)
trainer = build_trainer(name="trainer", default_args=kwargs)
trainer.train()
metrics = trainer.evaluate()

Important safety notes:

  • Passing a remote model id can trigger model/config access and optional remote-code/plugin checks. Use a local model directory or a verified cache when operating offline or in restricted environments.
  • build_trainer accepts name='trainer' and default_args=None by default; task-specific trainers use registered names such as NLP, CV, audio, or multi-modal trainer ids.
  • trainer.train() and trainer.evaluate() are real execution calls. Do not run them as a harmless smoke test.

Real-job command shapes

These commands are documented here only so an agent can recognize or prepare them. They launch real jobs and can download models, allocate GPUs, read datasets, and write checkpoints/logs.

bash
python -m modelscope.tools.train CONFIG_PATH TRAINER_NAME
python -m modelscope.tools.eval CONFIG_PATH --trainer_name TRAINER_NAME --checkpoint_path CHECKPOINT_PATH

Before using either command, complete the preflight in references/workflows.md and preview TrainingArgs-derived config with the bundled helper when the job is being constructed from flags.

Evidence basis

This sub-skill distills public behavior from the README training example, trainer builder and trainer implementation, TrainingArgs and CLI argument parser implementation, hook/checkpoint/distributed trainer modules, train/eval tool modules, representative PyTorch finetuning examples, the TrainingArgs unit test, and repository developer test-level guidance. Source paths are evidence only; future agents should use the bundled references and helper instead of reopening the original checkout.

© VectorSpaceLab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/repositories/repo-skills/modelscope/sub-skills/training-and-evaluation of VectorSpaceLab/AREX-Skill.

  • SKILL.md
  • references/training-args-reference.md
  • references/troubleshooting.md
  • references/workflows.md
  • scripts/build_training_args_preview.py

Open the folder on GitHubat commit ac3fe1a

Compare with similar skills

Training And Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Training And Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Training And Evaluation this skillVectorSpaceLab/AREX-Skill331—~1.4kAutomated safety check: PassApache-2.0
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0
Qwopus27b Rl TrainingR6410418/Jackrong-llm-finetuning-guide1.7k—~830Automated safety check: PassApache-2.0
Dataset Evaluationawslabs/agent-plugins9161 repos~1.3kAutomated safety check: PassApache-2.0
Train SftOpenPipe/ART11k—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    916 GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Sft

    OpenPipe/ART

    SFT training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Fine Tuning With Trl

    Orchestra-Research/AI-Research-SKILLs

    Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training.

    13k GitHub starsUsed in 6 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from VectorSpaceLab/AREX-Skill

All 157 skills in this repo
  • Agent Lightning

    VectorSpaceLab/AREX-Skill

    Use this repo skill for Agent Lightning package tasks: authoring trainable agents, tracing rewards and spans, running LightningStore/Trainer loops, using agl CLI services, choosing examples, and…

    331 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Tools

    VectorSpaceLab/AREX-Skill

    A skill your agent uses when configuring LiteLLM for MCP tools, A2A agents, Claude Code/Cursor agent gateway traffic, MCP auth/OAuth, tool permissions, semantic filtering, or agent-specific proxy…

    331 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents And Awel

    VectorSpaceLab/AREX-Skill

    Build and debug DB-GPT agents, tools, skills, teams, and AWEL workflows, including deterministic local DAG runs and HTTP-trigger topology without assuming an LLM, credential, or external service.

    331 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents And Middleware

    VectorSpaceLab/AREX-Skill

    Work on the actively maintained LangChain v1 agent package: initchatmodel, createagent, structured output, tools, middleware, embeddings initialization, provider routing, and agent runtime…

    331 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents Workflows

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for giskard.agents async chat workflows, tools, prompt templates, structured outputs, retries, rate limiting, embeddings, and optional LiteLLM backend.

    331 GitHub stars~500 tokensUpdated 1 mo ago
    Auto-check passed
  • Alphafold3

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for AlphaFold 3 input preparation, prediction command planning, output interpretation, and Python API inspection.

    331 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Training And Evaluation

What does Training And Evaluation do?

A skill your agent uses for ModelScope trainer construction, TrainingArgs conversion, fine-tuning and evaluation preflight, checkpoint hooks, and safe train/eval command planning. Training And Evaluation is an agent skill from VectorSpaceLab/AREX-Skill. Use for ModelScope trainer construction, TrainingArgs conversion, fine-tuning and evaluation preflight, checkpoint hooks, and safe train/eval command planning.

When should I use Training And Evaluation?

Training And Evaluation fits situations like: modelScope trainer construction; trainingArgs conversion; fine-tuning and evaluation preflight; checkpoint hooks.

How do I install Training And Evaluation in Claude Code?

Run `npx skills add VectorSpaceLab/AREX-Skill --skill training-and-evaluation -a claude-code`. Or copy the skill folder (skills/repositories/repo-skills/modelscope/sub-skills/training-and-evaluation in VectorSpaceLab/AREX-Skill) into .claude/skills/training-and-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Training And Evaluation in Codex?

Run `npx skills add VectorSpaceLab/AREX-Skill --skill training-and-evaluation -a codex`. Or copy the skill folder (skills/repositories/repo-skills/modelscope/sub-skills/training-and-evaluation in VectorSpaceLab/AREX-Skill) into .agents/skills/training-and-evaluation in your project. Codex loads it when a task matches its description.

Can I use Training And Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add VectorSpaceLab/AREX-Skill --skill training-and-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/training-and-evaluation, .gemini/skills/training-and-evaluation, .github/skills/training-and-evaluation and .opencode/skills/training-and-evaluation in your project.

What does Training And Evaluation need to run?

Going by SKILL.md and its folder, Training And Evaluation needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Training And Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Training And Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Training And Evaluation use?

Training And Evaluation is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Training And Evaluation use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8k tokens, read only when the agent opens those files.

What are the alternatives to Training And Evaluation?

Skills that share tags, products or a category with Training And Evaluation: Sentence-Transformers Training Router (huggingface/skills, 11k stars), Train Rl (OpenPipe/ART, 11k stars), Qwopus27b Rl Training (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars) and Dataset Evaluation (awslabs/agent-plugins, 916 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Training And Evaluation?

VectorSpaceLab (a GitHub organization) maintains it in VectorSpaceLab/AREX-Skill, which has 331 GitHub stars. The repository holds 157 skills in this directory. The repository was last updated on September 3, 2026.

Source: VectorSpaceLab/AREX-Skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.