Agent skill

Testlib Cpp Judging

by VectorSpaceLab in VectorSpaceLab/AREX-Skill

Guide C++ competitive-programming judging with the bundled testlib.h.

Apache-2.0Auto-check passed

Install Testlib Cpp Judging

skills CLI
$ npx skills add VectorSpaceLab/AREX-Skill --skill testlib-cpp-judging -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install VectorSpaceLab/AREX-Skill testlib-cpp-judging --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/VectorSpaceLab/AREX-Skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/task-oriented/FrontierCS/algorithmic-problem-solving/sub-skills/testlib-cpp-judging .claude/skills/testlib-cpp-judging && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testlib-cpp-judging
GitHub stars
328
Token cost
~1.9k tokens
SKILL.md length
849 words
Files
4 (incl. scripts, references)
Skills in repo
159
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide C++ competitive-programming judging with the bundled testlib.h.

  • Works in 5 steps: Read the testlib usage guide before… → Copy the bundled testlib.h into the… → Put #include "testlib.h" before other… → …
  • The user asks to author special
  • SKILL.md covers Scope, Start Here, Choose the Executable Role and Generate Data as Reproducible…, plus 5 more sections
  • Scored checkers

What it does

Testlib Cpp Judging is an agent skill from VectorSpaceLab/AREX-Skill. Guide C++ competitive-programming judging with the bundled testlib.h. Use when the user asks to author special or scored checkers, strict validators, deterministic generators, basic interactors, or a minimal local solution-checker workflow.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/testlib-usage.md` and `references/troubleshooting.md`).

It works with C++. The repository describes itself as: A Skill Library for Automated Machine Learning. The licence is Apache-2.0.

When your agent uses it

  • The user asks to author special
  • Scored checkers
  • Strict validators
  • Deterministic generators

Example prompts

  • “/testlib-cpp-judging”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Read the testlib usage guide before writing
  2. Copy the bundled testlib.h into the working directory
  3. Put #include "testlib.h" before other includes.
  4. Select exactly one registration function for each executable.
  5. For ordinary offline judging, use the minimal commands below rather than

What it can do on your machine

Read from SKILL.md and the folder at commit ac3fe1a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testlib Cpp Judging loads about 1.9k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 849 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from VectorSpaceLab/AREX-Skill at commit ac3fe1a, republished under its Apache-2.0 licence (© VectorSpaceLab). 849 words, ~1,882 tokens.

Download SKILL.mdSave it as .claude/skills/testlib-cpp-judging/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
testlib-cpp-judging
description
Guide C++ competitive-programming judging with the bundled testlib.h. Use when the user asks to author special or scored checkers, strict validators, deterministic generators, basic interactors, or a minimal local solution-checker workflow.

Testlib C++ Judging

Scope

Use this skill when a task involves testlib.h, a special judge, a custom checker or scorer, an input validator, a deterministic test generator, or a testlib interactor. It is intentionally C++-only and header-only. Do not create a Python harness for the ordinary local judging workflow.

This skill owns concrete Testlib APIs, streams, verdicts, command lines, and templates. When evaluator-facing evidence has triggered independent contract or oracle work, use checker and local evaluation to choose checker/scorer/interactor architecture, reconstruct legality and objectives, and validate the evaluator itself. Do not enter that route merely because a previously specified generator or validator is being implemented.

Start Here

  1. Read the testlib usage guide before writing checker, validator, generator, or interactor code.
  2. Copy the bundled testlib.h into the working directory beside the C++ sources, or add its directory to the compiler include path.
  3. Put #include "testlib.h" before other includes.
  4. Select exactly one registration function for each executable.
  5. For ordinary offline judging, use the minimal commands below rather than building a separate evaluation framework.

Choose the Executable Role

RoleRegistrationMain testlib interfacesPurpose
CheckerregisterTestlibCmd(argc, argv)inf, ouf, ans, quitfJudge contestant output against the input and reference answer
ValidatorregisterValidation(argc, argv)strict inf, ensurefReject malformed or out-of-constraint input
GeneratorregisterGen(argc, argv, 1)rnd, opt, printlnProduce deterministic tests from command-line parameters
InteractorregisterInteraction(argc, argv)inf, ouf, tout, stdoutExchange a protocol with an interactive solution

Generate Data as Reproducible Code

Testlib includes input-generation support; it is not limited to checkers and validators. When a local problem-finding campaign needs randomized, batch, or maximum-scale data, implement a problem-specific generator.cpp instead of hand-authoring large inputs or copying many variants. Parameterize the relevant case family, size, density/bias, structure, and seed tag with opt; use rnd for all randomness. The full command line determines Testlib's deterministic seed, so preserve that command exactly.

Define the coverage families and their expected oracle, invariant, or failure target before generating volume. Use one deterministic generator invocation per case, or a small coded batch driver that records every invocation in a manifest. Validate every intended-valid generated case with an independent validator before running the solution, and retain the generator command, input hash, validator result, and any failing seed and artifacts. Follow Checker and Local Evaluation for the complete coverage-to-retention workflow.

Critical Checker Contract

Always invoke a checker in this exact order:

checker <input-file> <contestant-output> <standard-answer>

After registration, the mapping is:

  • inf: input file
  • ouf: contestant or participant output
  • ans: standard or jury answer

Never swap ouf and ans. Even a checker that does not use the standard answer still needs the third file argument; pass an empty placeholder file if necessary.

Minimal Local Judging Flow

Run these commands from a directory containing testlib.h, solution.cpp, checker.cpp, case.in, and case.ans:

bash
g++ -std=c++17 -O2 solution.cpp -o solution
g++ -std=c++17 -O2 -I. checker.cpp -o checker

timeout 2s ./solution < case.in > case.out
solver_status=$?
if [ "$solver_status" -ne 0 ]; then
  echo "solution failed with status $solver_status" >&2
  exit "$solver_status"
fi
./checker case.in case.out case.ans
checker_status=$?
echo "$checker_status"
exit "$checker_status"

If the problem has validator.cpp, compile it and validate the input before running the solution:

bash
g++ -std=c++17 -O2 -I. validator.cpp -o validator
./validator < case.in || exit 1

Capture solver_status before running the checker; a timeout, signal, or other nonzero solution exit is a solver failure and must not be relabeled as an output verdict. Capture checker_status immediately after the checker. For a verdict-only checker under the default local Testlib configuration, status 0 means accepted and nonzero means rejection or judge failure. A points checker using quitp instead returns a partial-points status (default 7) that a points-aware runner must parse separately. Preserve checker diagnostics and any reported points with the status.

Show full SKILL.md (292 more words)Show less

Authoring Rules

  • Use testlib readers instead of raw parsing for judged files.
  • Return _wa for a semantically wrong contestant answer, _pe for malformed contestant output when you detect it explicitly, _fail for a broken jury answer or checker invariant, and _ok only after all required checks pass.
  • In validators, describe the exact grammar with readSpace, readEoln, and readEof; use bounded reads and ensuref for semantic constraints.
  • In generators, use rnd rather than rand, srand, or random_shuffle; identical command lines should reproduce identical data.
  • Implement batch and large-case generation in generator code; do not maintain hand-edited large input files as the source of truth.
  • Run every intended-valid generated case through the independent validator and preserve its full generator command and seed tag.
  • Flush every interactor query with std::endl or an explicit flush.
  • Keep solution.cpp, checker logic, validator logic, and answer generation as separate concerns.

Boundaries

  • This skill explains the public testlib.h workflow, not development of the testlib repository itself.
  • For evaluator roles, this skill implements a previously derived contract. Use checker and local evaluation only when contract, reconstruction, score-transform, or fidelity uncertainty is evidence-triggered and blocks the current decision, or when official and local behavior disagree.
  • This skill owns concrete testlib implementation. Use interactive problem solving for protocol modeling, hidden hypotheses, query design, and adversarial strategy.
  • It does not bundle a Python evaluator, a build system, repository tests, or CI configuration.
  • The simple three-file flow is for non-interactive judging. Interactive solutions require a bidirectional process runner supplied by the judge.
  • Partial scoring needs a runner that understands testlib points verdicts; do not interpret every nonzero checker status as ordinary wrong answer in that mode.

Troubleshooting

  • Read troubleshooting when compilation, checker arguments, strict whitespace, exit status, generator reproducibility, or interactor flushing causes a failure.

© VectorSpaceLab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/task-oriented/FrontierCS/algorithmic-problem-solving/sub-skills/testlib-cpp-judging of VectorSpaceLab/AREX-Skill.

  • SKILL.md
  • references/testlib-usage.md
  • references/troubleshooting.md
  • scripts/testlib.h

Open the folder on GitHubat commit ac3fe1a

Compare with similar skills

Testlib Cpp Judging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testlib Cpp Judging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testlib Cpp Judging this skillVectorSpaceLab/AREX-Skill328—~1.9kAutomated safety check: PassApache-2.0
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0
Fory Releaseapache/fory4.6k—~2.9kAutomated safety check: PassApache-2.0
ONNX Runtime Shape Inference Safety Auditmicrosoft/onnxruntime22k—~3.3kAutomated safety check: PassMIT
Code Audit3stoneBrother/code-audit8931 repos~2.7kAutomated safety check: PassNone
Qt C++ Code Reviewx-tools-author/x-tools1.1k2 repos~4.3kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Fory Release

    apache/fory

    Prepare an Apache Fory release candidate from a clean release branch, including the version bump, RC tag, JVM staging, ASF source artifacts, SVN upload, and vote email.

    4.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Finds and fixes out-of-range output writes in ONNX Runtime operator shape-inference functions where a getNumOutputs guard admits too few outputs.

    22k GitHub stars~3.3k tokensUpdated yesterday
    SecurityAuto-check passed
  • Code Audit

    3stoneBrother/code-audit

    Professional code security audit skill covering 55+ vulnerability types.

    893 GitHub starsUsed in 1 repo~2.7k tokens
    SecurityAuto-check passed
  • Qt C++ Code Review

    x-tools-author/x-tools

    Read-only review of Qt6 C++ code that combines a deterministic lint script with six parallel analysis agents and reports only high-confidence issues.

    1.1k GitHub starsUsed in 2 repos~4.3k tokens
    DevelopmentAuto-check passed
  • Translation

    doxygen/doxygen

    Keeps all Doxygen and Doxywizard translations up to date across three mechanisms: translator C++ classes (src/translatorxx.h), Qt .ts locale files for the Doxywizard GUI (addon/doxywizard/i18n/)…

    6.6k GitHub stars~5.2k tokensUpdated 7 days ago
    Writing & ContentAuto-check passed

More from VectorSpaceLab/AREX-Skill

All 159 skills in this repo
  • Agent Lightning

    VectorSpaceLab/AREX-Skill

    Use this repo skill for Agent Lightning package tasks: authoring trainable agents, tracing rewards and spans, running LightningStore/Trainer loops, using agl CLI services, choosing examples, and…

    328 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Tools

    VectorSpaceLab/AREX-Skill

    A skill your agent uses when configuring LiteLLM for MCP tools, A2A agents, Claude Code/Cursor agent gateway traffic, MCP auth/OAuth, tool permissions, semantic filtering, or agent-specific proxy…

    328 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents And Awel

    VectorSpaceLab/AREX-Skill

    Build and debug DB-GPT agents, tools, skills, teams, and AWEL workflows, including deterministic local DAG runs and HTTP-trigger topology without assuming an LLM, credential, or external service.

    328 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents And Middleware

    VectorSpaceLab/AREX-Skill

    Work on the actively maintained LangChain v1 agent package: initchatmodel, createagent, structured output, tools, middleware, embeddings initialization, provider routing, and agent runtime…

    328 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Agents Workflows

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for giskard.agents async chat workflows, tools, prompt templates, structured outputs, retries, rate limiting, embeddings, and optional LiteLLM backend.

    328 GitHub stars~500 tokensUpdated 1 mo ago
    Auto-check passed
  • Alphafold3

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for AlphaFold 3 input preparation, prediction command planning, output interpretation, and Python API inspection.

    328 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Testlib Cpp Judging

What does Testlib Cpp Judging do?

Guide C++ competitive-programming judging with the bundled testlib.h. Testlib Cpp Judging is an agent skill from VectorSpaceLab/AREX-Skill.h.

When should I use Testlib Cpp Judging?

Testlib Cpp Judging fits situations like: the user asks to author special; scored checkers; strict validators; deterministic generators.

How do I install Testlib Cpp Judging in Claude Code?

Run `npx skills add VectorSpaceLab/AREX-Skill --skill testlib-cpp-judging -a claude-code`. Or copy the skill folder (skills/task-oriented/FrontierCS/algorithmic-problem-solving/sub-skills/testlib-cpp-judging in VectorSpaceLab/AREX-Skill) into .claude/skills/testlib-cpp-judging in your project. Claude Code loads it when a task matches its description.

How do I install Testlib Cpp Judging in Codex?

Run `npx skills add VectorSpaceLab/AREX-Skill --skill testlib-cpp-judging -a codex`. Or copy the skill folder (skills/task-oriented/FrontierCS/algorithmic-problem-solving/sub-skills/testlib-cpp-judging in VectorSpaceLab/AREX-Skill) into .agents/skills/testlib-cpp-judging in your project. Codex loads it when a task matches its description.

Can I use Testlib Cpp Judging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add VectorSpaceLab/AREX-Skill --skill testlib-cpp-judging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testlib-cpp-judging, .gemini/skills/testlib-cpp-judging, .github/skills/testlib-cpp-judging and .opencode/skills/testlib-cpp-judging in your project.

What does Testlib Cpp Judging need to run?

SKILL.md names no scripts, command-line tools or credentials: Testlib Cpp Judging is instructions for the agent only. Our summary lists: Python 3.

Does Testlib Cpp Judging access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Testlib Cpp Judging safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Testlib Cpp Judging use?

Testlib Cpp Judging is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testlib Cpp Judging use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.

What are the alternatives to Testlib Cpp Judging?

Skills that share tags, products or a category with Testlib Cpp Judging: Paddle Build (PaddlePaddle/Paddle, 24k stars), Fory Release (apache/fory, 4.6k stars), ONNX Runtime Shape Inference Safety Audit (microsoft/onnxruntime, 22k stars) and Code Audit (3stoneBrother/code-audit, 893 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testlib Cpp Judging?

VectorSpaceLab (a GitHub organization) maintains it in VectorSpaceLab/AREX-Skill, which has 328 GitHub stars. The repository holds 159 skills in this directory. The repository was last updated on September 3, 2026.

Source: VectorSpaceLab/AREX-Skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.