Agent skill

Suite Creator

by Margin-Lab in Margin-Lab/evals

Creates new Margin Eval test suites from scratch. An agent skill from Margin-Lab/evals.

AGPL-3.0Auto-check passedDevOps & Cloud

Install Suite Creator

skills CLI
$ npx skills add Margin-Lab/evals --skill suite-creator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Margin-Lab/evals suite-creator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/suite-creator .claude/skills/suite-creator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
suite-creator
GitHub stars
161
Token cost
~2.3k tokens
SKILL.md length
978 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Creates new Margin Eval test suites from scratch. An agent skill from Margin-Lab/evals.

  • Works in 9 steps: Define the suite scope → Scaffold the directory structure → Write the Dockerfile → …
  • The user wants to author
  • SKILL.md covers Suite Structure, Workflow and Writing Good Test Cases
  • Calls pytest, npm and python

What it does

Suite Creator is an agent skill from Margin-Lab/evals. Creates new Margin Eval test suites from scratch. Use this skill whenever the user wants to author, build, or scaffold a new eval test suite, define test cases for evaluating coding agents, or create tasks with Dockerfiles, prompts, and grading scripts for the Margin Eval framework.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Test generation and Containers. It works with Docker. The repository describes itself as: Fast, robust, configurable agent evals. The licence is AGPL-3.0.

When your agent uses it

  • The user wants to author
  • Scaffold a new eval test suite
  • Define test cases for evaluating coding agents
  • Create tasks with Dockerfiles

Example prompts

  • “Use the suite-creator skill to create new Margin Eval test suites from scratch. An agent skill from Margin-Lab/evals”
  • “/suite-creator”

Requirements

  • Python 3
  • Docker

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Define the suite scope
  2. Scaffold the directory structure
  3. Write the Dockerfile
  4. Write the prompt
  5. Write the grading script
  6. Write the reference solution (optional)
  7. Write case.toml and suite.toml
  8. Validate
  9. Smoke test

What it can do on your machine

Read from SKILL.md and the folder at commit b57dfe9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest
    • npm
    • python
    • kind

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Suite Creator loads about 2.3k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 978 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Margin-Lab/evals at commit b57dfe9, republished under its AGPL-3.0 licence (© Margin-Lab). 978 words, ~2,259 tokens.

Download SKILL.mdSave it as .claude/skills/suite-creator/SKILL.md (or your agent's skills folder).
name
suite-creator
description
Creates new Margin Eval test suites from scratch. Use this skill whenever the user wants to author, build, or scaffold a new eval test suite, define test cases for evaluating coding agents, or create tasks with Dockerfiles, prompts, and grading scripts for the Margin Eval framework.

Suite Creator

Creates new Margin Eval test suites from scratch, including the directory structure, configuration files, prompts, Dockerfiles, and grading scripts.

Suite Structure

<suite-name>/
├── suite.toml
└── cases/
    └── <case-name>/
        ├── case.toml
        ├── prompt.md
        ├── env/
        │   └── Dockerfile
        ├── tests/
        │   ├── test.sh
        │   └── <supporting test files>
        └── oracle/              # optional
            └── solve.sh
suite.toml
toml
kind = "test_suite"
name = "<suite-name>"
description = "<what this suite evaluates>"
cases = [
  "<case-1>",
  "<case-2>",
]

The cases array lists directory names under cases/, in the order they should run.

case.toml
toml
kind = "test_case"
name = "<case-name>"
description = "<one-line description>"
test_cwd = "/"
test_timeout_seconds = 1800

[metadata]
difficulty = "easy"
category = "programming"
tags = ["tag1", "tag2"]

Key rules:

  • name must match its directory name exactly
  • kind is always "test_case"
  • test_timeout_seconds is an integer (seconds)
  • test_cwd is the working directory where test.sh runs inside the container

Image handling — exactly one of:

  • image = "registry/repo@sha256:<64hex>" for a pre-built, digest-pinned image
  • Omit image and place a Dockerfile at env/Dockerfile to build at compile time
prompt.md

The task description sent to the agent as its initial prompt. Must not be empty.

tests/test.sh

The grading script — this is the evaluator. It must terminate with exactly one of:

  • 0 = pass
  • 1 = fail
  • 2 = infra

Do not write reward.txt. The verifier process exit code is the authoritative result. The script must be executable. All files in tests/ are packaged together and staged at {test_cwd}/tests/ in the container.

oracle/solve.sh (optional)

Reference solution. Not executed during normal eval runs — useful for dry runs and validation.

env/Dockerfile

Container environment. The entire env/ directory is the build context, so supporting files (setup scripts, seed data, config) can live alongside the Dockerfile.


Workflow

1. Define the suite scope

Clarify with the user:

  • What capability or behavior is being evaluated?
  • How many cases? What range of difficulty?
  • What does the agent's environment look like? (language, frameworks, tools available)
2. Scaffold the directory structure

Create the suite directory, cases/ subdirectory, and a skeleton for each case. It helps to create all case directories first, then fill them in one at a time.

3. Write the Dockerfile

Start with the Dockerfile because it defines the environment everything else runs in. The Dockerfile should produce a container that:

  • Has all dependencies the agent might need pre-installed
  • Sets a WORKDIR (this becomes test_cwd in case.toml)
  • Contains any seed data, starter code, or project scaffolding the task requires
  • Does not contain the solution or hints toward it

Keep images minimal — install only what the task requires. Use specific version tags, not latest.

dockerfile
FROM python:3.12-slim

WORKDIR /app

# Pre-install project dependencies
COPY requirements.txt .
RUN pip install -r requirements.txt

# Seed the workspace with starter code
COPY src/ ./src/
4. Write the prompt

The prompt is the only input the agent receives. Write it as if briefing a developer who has just been dropped into the container. It should include:

  • What to do — the task objective, stated clearly
  • Where things are — relevant file paths, project structure
  • Constraints — specific requirements, output format, things to avoid
  • Success criteria — what "done" looks like, in concrete terms

Avoid leaking test implementation details. The agent should not know how it will be graded — only what the correct behavior is.

A good prompt is specific enough that a competent developer could complete the task without asking clarifying questions, but does not prescribe a particular implementation approach.

5. Write the grading script

tests/test.sh determines pass/fail. The grading approach depends on what's being tested:

Use a simple explicit verdict API in every new harness:

bash
#!/bin/bash
set -euo pipefail

pass()  { printf 'VERDICT: PASS\n'; exit 0; }
fail()  { printf 'VERDICT: FAIL\n'; exit 1; }
infra() { printf 'VERDICT: INFRA\n' >&2; exit 2; }

Policy:

  • return pass or fail only when the harness reached a trustworthy verdict about the candidate
  • return infra when the harness cannot reach a trustworthy verdict for reasons not attributable to the candidate
  • missing candidate artifact, candidate compile/import/runtime failure, candidate timeout after candidate logic starts, and wrong output are all fail
  • verifier/bootstrap/parser/config failures are infra

File/output verification — check that the agent produced the right files with the right content:

bash
#!/bin/bash
set -euo pipefail

pass()  { exit 0; }
fail()  { exit 1; }
infra() { exit 2; }

if [ ! -f /app/output.json ]; then
  fail
fi

set +e
pytest tests/test_outputs.py -rA
exit_code=$?
set -e

case "$exit_code" in
  0) pass ;;
  1) fail ;;
  *) infra ;;
esac

Test suite pass-through — run the project's own test suite against the agent's changes:

bash
#!/bin/bash
set -euo pipefail

pass()  { exit 0; }
fail()  { exit 1; }
infra() { exit 2; }

cd /app
set +e
npm test
exit_code=$?
set -e

case "$exit_code" in
  0) pass ;;
  1) fail ;;
  *) infra ;;
esac

Custom validation — for tasks where correctness is more nuanced:

bash
#!/bin/bash
set -euo pipefail
python tests/validate.py

Guidelines for grading scripts:

  • Install test dependencies inside the script (they run in a separate context from the agent)
  • Test observable outcomes, not implementation details — the agent should be free to solve the problem however it sees fit
  • Make sure the script is deterministic — same agent output should always produce the same verdict
  • Keep terminal verdict paths explicit and easy to audit
  • Prefer verifier-level changes over Dockerfile changes when clarifying verdict semantics
  • If a parser or verifier cannot interpret the result, default to infra unless you can directly attribute the failure to the candidate
Show full SKILL.md (310 more words)Show less
6. Write the reference solution (optional)

oracle/solve.sh applies the known-correct fix. Useful for validating that the grading script works — run the solution, then run test.sh, and confirm it passes.

bash
#!/bin/bash
cd /app
# Apply the fix
sed -i 's/old_pattern/new_pattern/' src/module.py
7. Write case.toml and suite.toml

Fill in the case config:

  • Set test_cwd to match the Dockerfile's WORKDIR
  • Set test_timeout_seconds generously — allow 2-3x the expected solve time
  • Add meaningful metadata (difficulty, category, tags)

Then generate suite.toml listing all cases.

8. Validate

Check every case:

  • case.toml exists, name matches directory name, kind = "test_case"
  • prompt.md exists and is non-empty
  • tests/test.sh exists and is executable
  • env/Dockerfile exists (or image is set in case.toml)
  • suite.toml lists all case directory names
  • chmod +x on all .sh files
9. Smoke test

If possible, build the Docker image and verify:

  1. The container starts and the workspace is set up correctly
  2. Running oracle/solve.sh followed by tests/test.sh exits 0
  3. Running tests/test.sh without the solution exits 1
  4. An induced verifier/setup failure exits 2

Writing Good Test Cases

Prompt clarity

Ambiguous prompts make it hard to distinguish agent capability from prompt interpretation. If an agent fails, you want to be confident it's because the agent couldn't do the task, not because the instructions were unclear.

Grading robustness

The grading script should accept any correct solution, not just the reference solution. Avoid:

  • Checking for specific variable names or function signatures (unless the prompt requires them)
  • Exact string matching on output when approximate matching would do
  • Depending on file modification timestamps or execution order
Difficulty calibration

Categorize cases by difficulty to make results more informative:

  • easy — a competent developer solves it in under 15 minutes
  • medium — requires understanding the codebase or domain (15 min to 1 hour)
  • hard — requires significant reasoning, multi-file changes, or domain expertise (1+ hours)
Independence

Each case should be self-contained. Cases should not depend on each other or share state. Every case runs in a fresh container.

© Margin-Lab, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/suite-creator of Margin-Lab/evals.

Open the folder on GitHubat commit b57dfe9

Compare with similar skills

Suite Creator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Suite Creator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Suite Creator this skillMargin-Lab/evals161—~2.3kAutomated safety check: PassAGPL-3.0
Corvus Bowtie Testingcorvus-dotnet/Corvus.JsonSchema199—~1.4kAutomated safety check: PassApache-2.0
Model Download Devopen-edge-platform/edge-ai-libraries169—~2.8kAutomated safety check: PassApache-2.0
Liveblog Devliveblog/liveblog119—~1.9kAutomated safety check: PassAGPL-3.0
Docker Testrojopolis/spellcheck-github-actions151—~148Automated safety check: PassMIT
Run SDK Testsrestatedev/sdk-typescript125—~745Automated safety check: PassMIT

Similar skills

  • Corvus Bowtie Testing

    corvus-dotnet/Corvus.JsonSchema

    Test Corvus.JsonSchema against the JSON Schema Test Suite using Bowtie, the cross-implementation meta-validator.

    199 GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Model Download Dev

    open-edge-platform/edge-ai-libraries

    Extend, test, debug, or integrate the Model Download microservice codebase.

    169 GitHub stars~2.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Liveblog Dev

    liveblog/liveblog

    Run a local Liveblog development environment. An agent skill from liveblog/liveblog.

    119 GitHub stars~1.9k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Docker Test

    rojopolis/spellcheck-github-actions

    Build the Docker image as :local and run the full Bats test suite

    151 GitHub stars~148 tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • Run SDK Tests

    restatedev/sdk-typescript

    Run the Restate SDK conformance test suite locally against this SDK's Docker image.

    125 GitHub stars~745 tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Backend Testing

    rhesis-ai/rhesis

    Run the backend test suite correctly — working directory, Docker requirement, single-test vs full-suite commands.

    397 GitHub stars~400 tokensUpdated yesterday
    Testing & QAAuto-check passed

More from Margin-Lab/evals

  • Suite Converter

    Margin-Lab/evals

    Converts test suites from external eval frameworks into the Margin Eval suite format.

    161 GitHub stars~3.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Agent Definition Creator

    Margin-Lab/evals

    Creates or updates Margin Eval agent definitions for new CLI coding agents.

    161 GitHub stars~2.4k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Suite Creator

What does Suite Creator do?

Creates new Margin Eval test suites from scratch. An agent skill from Margin-Lab/evals. Suite Creator is an agent skill from Margin-Lab/evals. Creates new Margin Eval test suites from scratch.

When should I use Suite Creator?

Suite Creator fits situations like: the user wants to author; scaffold a new eval test suite; define test cases for evaluating coding agents; create tasks with Dockerfiles.

How do I install Suite Creator in Claude Code?

Run `npx skills add Margin-Lab/evals --skill suite-creator -a claude-code`. Or copy the skill folder (.agents/skills/suite-creator in Margin-Lab/evals) into .claude/skills/suite-creator in your project. Claude Code loads it when a task matches its description.

How do I install Suite Creator in Codex?

Run `npx skills add Margin-Lab/evals --skill suite-creator -a codex`. Or copy the skill folder (.agents/skills/suite-creator in Margin-Lab/evals) into .agents/skills/suite-creator in your project. Codex loads it when a task matches its description.

Can I use Suite Creator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Margin-Lab/evals --skill suite-creator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/suite-creator, .gemini/skills/suite-creator, .github/skills/suite-creator and .opencode/skills/suite-creator in your project.

What does Suite Creator need to run?

Going by SKILL.md and its folder, Suite Creator needs the command-line tools its instructions call (pytest, npm, python and kind). Our summary lists: Python 3; Docker.

Does Suite Creator access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Suite Creator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Suite Creator use?

Suite Creator is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Suite Creator use?

About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Suite Creator?

Skills that share tags, products or a category with Suite Creator: Corvus Bowtie Testing (corvus-dotnet/Corvus.JsonSchema, 199 stars), Model Download Dev (open-edge-platform/edge-ai-libraries, 169 stars), Liveblog Dev (liveblog/liveblog, 119 stars) and Docker Test (rojopolis/spellcheck-github-actions, 151 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Suite Creator?

Margin-Lab (a GitHub organization) maintains it in Margin-Lab/evals, which has 161 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on July 31, 2026.

Source: Margin-Lab/evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.