Corvus Bowtie Testing
corvus-dotnet/Corvus.JsonSchema
Test Corvus.JsonSchema against the JSON Schema Test Suite using Bowtie, the cross-implementation meta-validator.
Creates new Margin Eval test suites from scratch. An agent skill from Margin-Lab/evals.
$ npx skills add Margin-Lab/evals --skill suite-creator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Margin-Lab/evals suite-creator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/suite-creator .claude/skills/suite-creator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "suite-creator" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-creator into .claude/skills/suite-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-creator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-creatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Margin-Lab/evals --skill suite-creator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Margin-Lab/evals suite-creator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/suite-creator .agents/skills/suite-creator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "suite-creator" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-creator into .agents/skills/suite-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-creator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Margin-Lab/evals --skill suite-creator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Margin-Lab/evals suite-creator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/suite-creator .cursor/skills/suite-creator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "suite-creator" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-creator into .cursor/skills/suite-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-creator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Margin-Lab/evals.git --path .agents/skills/suite-creator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Margin-Lab/evals --skill suite-creator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Margin-Lab/evals suite-creator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/suite-creator .gemini/skills/suite-creator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "suite-creator" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-creator into .gemini/skills/suite-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-creator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Margin-Lab/evals suite-creatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Margin-Lab/evals --skill suite-creator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/suite-creator .github/skills/suite-creator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "suite-creator" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-creator into .github/skills/suite-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-creator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Margin-Lab/evals --skill suite-creator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Margin-Lab/evals suite-creator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/suite-creator .opencode/skills/suite-creator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "suite-creator" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-creator into .opencode/skills/suite-creator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-creator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
suite-creatorCreates new Margin Eval test suites from scratch. An agent skill from Margin-Lab/evals.
Suite Creator is an agent skill from Margin-Lab/evals. Creates new Margin Eval test suites from scratch. Use this skill whenever the user wants to author, build, or scaffold a new eval test suite, define test cases for evaluating coding agents, or create tasks with Dockerfiles, prompts, and grading scripts for the Margin Eval framework.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Test generation and Containers. It works with Docker. The repository describes itself as: Fast, robust, configurable agent evals. The licence is AGPL-3.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b57dfe9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pytestnpmpythonkindFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Suite Creator loads about 2.3k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 978 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Margin-Lab/evals at commit b57dfe9, republished under its AGPL-3.0 licence (© Margin-Lab). 978 words, ~2,259 tokens.
.claude/skills/suite-creator/SKILL.md (or your agent's skills folder).Creates new Margin Eval test suites from scratch, including the directory structure, configuration files, prompts, Dockerfiles, and grading scripts.
<suite-name>/
├── suite.toml
└── cases/
└── <case-name>/
├── case.toml
├── prompt.md
├── env/
│ └── Dockerfile
├── tests/
│ ├── test.sh
│ └── <supporting test files>
└── oracle/ # optional
└── solve.shkind = "test_suite"
name = "<suite-name>"
description = "<what this suite evaluates>"
cases = [
"<case-1>",
"<case-2>",
]The cases array lists directory names under cases/, in the order they should run.
kind = "test_case"
name = "<case-name>"
description = "<one-line description>"
test_cwd = "/"
test_timeout_seconds = 1800
[metadata]
difficulty = "easy"
category = "programming"
tags = ["tag1", "tag2"]Key rules:
name must match its directory name exactlykind is always "test_case"test_timeout_seconds is an integer (seconds)test_cwd is the working directory where test.sh runs inside the containerImage handling — exactly one of:
image = "registry/repo@sha256:<64hex>" for a pre-built, digest-pinned imageimage and place a Dockerfile at env/Dockerfile to build at compile timeThe task description sent to the agent as its initial prompt. Must not be empty.
The grading script — this is the evaluator. It must terminate with exactly one of:
0 = pass1 = fail2 = infraDo not write reward.txt. The verifier process exit code is the authoritative result. The script must be executable. All files in tests/ are packaged together and staged at {test_cwd}/tests/ in the container.
Reference solution. Not executed during normal eval runs — useful for dry runs and validation.
Container environment. The entire env/ directory is the build context, so supporting files (setup scripts, seed data, config) can live alongside the Dockerfile.
Clarify with the user:
Create the suite directory, cases/ subdirectory, and a skeleton for each case. It helps to create all case directories first, then fill them in one at a time.
Start with the Dockerfile because it defines the environment everything else runs in. The Dockerfile should produce a container that:
WORKDIR (this becomes test_cwd in case.toml)Keep images minimal — install only what the task requires. Use specific version tags, not latest.
FROM python:3.12-slim
WORKDIR /app
# Pre-install project dependencies
COPY requirements.txt .
RUN pip install -r requirements.txt
# Seed the workspace with starter code
COPY src/ ./src/The prompt is the only input the agent receives. Write it as if briefing a developer who has just been dropped into the container. It should include:
Avoid leaking test implementation details. The agent should not know how it will be graded — only what the correct behavior is.
A good prompt is specific enough that a competent developer could complete the task without asking clarifying questions, but does not prescribe a particular implementation approach.
tests/test.sh determines pass/fail. The grading approach depends on what's being tested:
Use a simple explicit verdict API in every new harness:
#!/bin/bash
set -euo pipefail
pass() { printf 'VERDICT: PASS\n'; exit 0; }
fail() { printf 'VERDICT: FAIL\n'; exit 1; }
infra() { printf 'VERDICT: INFRA\n' >&2; exit 2; }Policy:
pass or fail only when the harness reached a trustworthy verdict about the candidateinfra when the harness cannot reach a trustworthy verdict for reasons not attributable to the candidatefailinfraFile/output verification — check that the agent produced the right files with the right content:
#!/bin/bash
set -euo pipefail
pass() { exit 0; }
fail() { exit 1; }
infra() { exit 2; }
if [ ! -f /app/output.json ]; then
fail
fi
set +e
pytest tests/test_outputs.py -rA
exit_code=$?
set -e
case "$exit_code" in
0) pass ;;
1) fail ;;
*) infra ;;
esacTest suite pass-through — run the project's own test suite against the agent's changes:
#!/bin/bash
set -euo pipefail
pass() { exit 0; }
fail() { exit 1; }
infra() { exit 2; }
cd /app
set +e
npm test
exit_code=$?
set -e
case "$exit_code" in
0) pass ;;
1) fail ;;
*) infra ;;
esacCustom validation — for tasks where correctness is more nuanced:
#!/bin/bash
set -euo pipefail
python tests/validate.pyGuidelines for grading scripts:
infra unless you can directly attribute the failure to the candidateoracle/solve.sh applies the known-correct fix. Useful for validating that the grading script works — run the solution, then run test.sh, and confirm it passes.
#!/bin/bash
cd /app
# Apply the fix
sed -i 's/old_pattern/new_pattern/' src/module.pyFill in the case config:
test_cwd to match the Dockerfile's WORKDIRtest_timeout_seconds generously — allow 2-3x the expected solve timeThen generate suite.toml listing all cases.
Check every case:
case.toml exists, name matches directory name, kind = "test_case"prompt.md exists and is non-emptytests/test.sh exists and is executableenv/Dockerfile exists (or image is set in case.toml)suite.toml lists all case directory nameschmod +x on all .sh filesIf possible, build the Docker image and verify:
oracle/solve.sh followed by tests/test.sh exits 0tests/test.sh without the solution exits 1Ambiguous prompts make it hard to distinguish agent capability from prompt interpretation. If an agent fails, you want to be confident it's because the agent couldn't do the task, not because the instructions were unclear.
The grading script should accept any correct solution, not just the reference solution. Avoid:
Categorize cases by difficulty to make results more informative:
Each case should be self-contained. Cases should not depend on each other or share state. Every case runs in a fresh container.
© Margin-Lab, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/suite-creator of Margin-Lab/evals.
Open the folder on GitHubat commit b57dfe9
Suite Creator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Suite Creator this skillMargin-Lab/evals | 161 | — | ~2.3k | Automated safety check: Pass | AGPL-3.0 | |
| Corvus Bowtie Testingcorvus-dotnet/Corvus.JsonSchema | 199 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Model Download Devopen-edge-platform/edge-ai-libraries | 169 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Liveblog Devliveblog/liveblog | 119 | — | ~1.9k | Automated safety check: Pass | AGPL-3.0 | |
| Docker Testrojopolis/spellcheck-github-actions | 151 | — | ~148 | Automated safety check: Pass | MIT | |
| Run SDK Testsrestatedev/sdk-typescript | 125 | — | ~745 | Automated safety check: Pass | MIT |
corvus-dotnet/Corvus.JsonSchema
Test Corvus.JsonSchema against the JSON Schema Test Suite using Bowtie, the cross-implementation meta-validator.
open-edge-platform/edge-ai-libraries
Extend, test, debug, or integrate the Model Download microservice codebase.
liveblog/liveblog
Run a local Liveblog development environment. An agent skill from liveblog/liveblog.
rojopolis/spellcheck-github-actions
Build the Docker image as :local and run the full Bats test suite
restatedev/sdk-typescript
Run the Restate SDK conformance test suite locally against this SDK's Docker image.
rhesis-ai/rhesis
Run the backend test suite correctly — working directory, Docker requirement, single-test vs full-suite commands.
Margin-Lab/evals
Converts test suites from external eval frameworks into the Margin Eval suite format.
Margin-Lab/evals
Creates or updates Margin Eval agent definitions for new CLI coding agents.
Works with
Categories
Creates new Margin Eval test suites from scratch. An agent skill from Margin-Lab/evals. Suite Creator is an agent skill from Margin-Lab/evals. Creates new Margin Eval test suites from scratch.
Suite Creator fits situations like: the user wants to author; scaffold a new eval test suite; define test cases for evaluating coding agents; create tasks with Dockerfiles.
Run `npx skills add Margin-Lab/evals --skill suite-creator -a claude-code`. Or copy the skill folder (.agents/skills/suite-creator in Margin-Lab/evals) into .claude/skills/suite-creator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Margin-Lab/evals --skill suite-creator -a codex`. Or copy the skill folder (.agents/skills/suite-creator in Margin-Lab/evals) into .agents/skills/suite-creator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Margin-Lab/evals --skill suite-creator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/suite-creator, .gemini/skills/suite-creator, .github/skills/suite-creator and .opencode/skills/suite-creator in your project.
Going by SKILL.md and its folder, Suite Creator needs the command-line tools its instructions call (pytest, npm, python and kind). Our summary lists: Python 3; Docker.
SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Suite Creator is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Suite Creator: Corvus Bowtie Testing (corvus-dotnet/Corvus.JsonSchema, 199 stars), Model Download Dev (open-edge-platform/edge-ai-libraries, 169 stars), Liveblog Dev (liveblog/liveblog, 119 stars) and Docker Test (rojopolis/spellcheck-github-actions, 151 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Margin-Lab (a GitHub organization) maintains it in Margin-Lab/evals, which has 161 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on July 31, 2026.
Source: Margin-Lab/evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.