Agent skill

Suite Converter

by Margin-Lab in Margin-Lab/evals

Converts test suites from external eval frameworks into the Margin Eval suite format.

AGPL-3.0Auto-check passedTesting & QA

Install Suite Converter

skills CLI
$ npx skills add Margin-Lab/evals --skill suite-converter -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Margin-Lab/evals suite-converter --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/suite-converter .claude/skills/suite-converter && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
suite-converter
GitHub stars
161
Token cost
~3.7k tokens
SKILL.md length
1,833 words
Files
2 (incl. references)
Skills in repo
3
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Converts test suites from external eval frameworks into the Margin Eval suite format.

  • Works in 12 steps: Identify the source format by inspecting… → Choose a suite name — derive from the… → Create the output structure: /cases/ → …
  • The user wants to import
  • SKILL.md covers Purpose, Target Margin Contract, Conversion Invariants and Decision Rules, plus 4 more sections
  • Calls kind

What it does

Suite Converter is an agent skill from Margin-Lab/evals. Converts test suites from external eval frameworks into the Margin Eval suite format. Use this skill whenever the user wants to import, convert, translate, or migrate an eval dataset or test suite into Margin Eval format, or when they mention converting tasks from other benchmarking frameworks into Margin's structure.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/harbor-mapping.md`).

It sits in Testing & QA, covering Test generation, Containers and LLM evaluation. It works with Docker. The repository describes itself as: Fast, robust, configurable agent evals. The licence is AGPL-3.0.

When your agent uses it

  • The user wants to import
  • Migrate an eval dataset
  • Test suite into Margin Eval format
  • They mention converting tasks from other benchmarking frameworks into Margins structure

Example prompts

  • “Use the suite-converter skill to convert test suites from external eval frameworks into the Margin Eval suite format”
  • “/suite-converter”

Requirements

  • Docker

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Identify the source format by inspecting the input directory (look for characteristic files like task.toml, case.toml, etc.)
  2. Choose a suite name — derive from the dataset name or ask the user
  3. Create the output structure: /cases/
  4. For each task/case in the source, create a case directory and convert
  5. Generate suite.toml listing all case names (sorted alphabetically)
  6. Validate: every case must have case.toml, prompt.md, tests/test.sh, and either image or env/Dockerfile
  7. Set permissions: chmod +x on tests/test.sh and oracle/solve.sh
  8. Ensure unique case names: if sanitization causes collisions, append a stable suffix (for example part of the source task ID/UUID)
  9. Smoke test a sample first: convert a small sample of 2-5 representative cases before committing to the full dataset
  10. Run Margin on that sample: execute the sample with margin run and carefully inspect the results, logs, and verifier behavior to identify…
  11. Validate verifier behavior: confirm the converted tests/test.sh locates its test assets under {test_cwd}/tests, returns a non-zero exit…
  12. Validate agent starting context: confirm agent_cwd points at the directory where the agent should actually begin work for the task

What it can do on your machine

Read from SKILL.md and the folder at commit b57dfe9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kind

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Suite Converter loads about 3.7k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 1,833 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Margin-Lab/evals at commit b57dfe9, republished under its AGPL-3.0 licence (© Margin-Lab). 1,833 words, ~3,739 tokens.

Download SKILL.mdSave it as .claude/skills/suite-converter/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
suite-converter
description
Converts test suites from external eval frameworks into the Margin Eval suite format. Use this skill whenever the user wants to import, convert, translate, or migrate an eval dataset or test suite into Margin Eval format, or when they mention converting tasks from other benchmarking frameworks into Margin's structure.

Suite Converter

Converts downloaded eval datasets into valid Margin Eval test suites.

Purpose

Use this skill when converting an external eval dataset into the Margin Eval suite format.

The job has three parts:

  • map the source format into Margin's filesystem contract
  • preserve the source verifier's intended behavior without breaking Margin's execution model
  • validate the conversion on a small sample with margin run before scaling up

Target Margin Contract

<suite-name>/
├── suite.toml
└── cases/
    └── <case-name>/
        ├── case.toml
        ├── prompt.md
        ├── env/
        │   └── Dockerfile
        ├── tests/
        │   ├── test.sh
        │   └── <other test files>
        └── oracle/              # optional
            └── solve.sh
suite.toml
toml
kind = "test_suite"
name = "<suite-name>"
description = "<description>"
cases = [
  "<case-1>",
  "<case-2>",
]

The cases array lists directory names under cases/, in the order they should run.

case.toml
toml
kind = "test_case"
name = "<case-name>"
description = "<one-line description>"
agent_cwd = "/"
test_cwd = "/"
test_timeout_seconds = 1800

[metadata]
difficulty = "easy"
category = "programming"
tags = ["tag1", "tag2"]

[metadata] can be preserved for human context and future tooling, but the current Margin compiler/runtime primarily cares about the execution fields such as kind, name, description, image, agent_cwd, test_cwd, and test_timeout_seconds.

Key rules:

  • name must match its directory name exactly
  • kind is always "test_case"
  • test_timeout_seconds is an integer (seconds)
  • agent_cwd is the directory where the agent is expected to start and do its work inside the container
  • test_cwd is the working directory where test.sh runs inside the container

Image handling, exactly one of:

  • image = "registry/repo@sha256:<64hex>" for a pre-built, digest-pinned image
  • Omit image and place a Dockerfile at env/Dockerfile to build at compile time
prompt.md

The full task description sent to the agent as its initial prompt. Must not be empty. Copy the source's instruction/prompt file as-is — don't summarize or reformat it.

tests/test.sh

The grading script. This is the evaluator. There is no separate grader abstraction.

env/Dockerfile

Container environment. Can include supporting files alongside the Dockerfile. The entire env/ directory is the build context.

oracle/solve.sh

Optional reference solution. Not executed during normal eval runs. In the current Margin implementation, oracle/ is informational only and is not part of the active compile/runtime path.

Conversion Invariants

These rules are cross-source and should hold for every conversion.

Verifier rules
  • Must be executable (chmod +x)
  • Exit 0 = pass, 1 = fail, 2 = infra
  • Do not write reward.txt. The verifier process exit code is the authoritative result.
  • All files in tests/ are packaged together and staged at {test_cwd}/tests/ in the container
  • If the source suite uses absolute test-asset paths such as /tests/..., rewrite them to Margin-compatible paths such as tests/... or "${PWD}/tests/...".
  • If the source verifier derives test-asset paths from environment variables, normalize or override those variables to Margin-compatible values. Do not rely only on fallback expressions when the source environment may already set an incompatible path such as /tests.
  • Ensure the verifier creates any directories it expects before writing output artifacts.
  • Use exactly one authoritative test location. If the verifier copies tests from tests/ into the workspace, its test command must avoid rediscovering the mounted tests/ tree. If it runs tests directly from tests/, do not also copy them into the workspace.
  • Never leave a wrapper that always exits 0. Every verifier must terminate through explicit pass, fail, or infra paths.
  • If the source suite expects status or report artifacts in addition to the exit code, preserve that behavior, but do not treat any single artifact format as universal.
Verdict policy

Use this attribution rule across conversions:

  • pass: the harness reached a trustworthy verdict and the candidate satisfied the task
  • fail: the harness reached a trustworthy verdict and the candidate did not satisfy the task
  • infra: the harness could not reach a trustworthy verdict for reasons not attributable to the candidate

Classify these as fail:

  • missing candidate-owned artifact
  • candidate compile/import/build/runtime failure
  • candidate timeout after candidate logic begins
  • wrong output, failed assertions, malformed candidate-generated output
  • candidate-caused repo state that prevents intended test execution

Classify these as infra:

  • verifier/bootstrap/parser dependency failure
  • suite-owned config or hidden test asset is missing or malformed
  • verifier cannot interpret logs or required verifier artifacts are missing
  • harness-side timeout before candidate evaluation meaningfully begins

If a parser or verifier layer cannot produce a trustworthy verdict, default to infra unless the harness already has direct evidence of a candidate-caused failure.

Case identity rules
  • case.toml.name must match the case directory name exactly
  • If source-name sanitization causes collisions, append a stable suffix such as part of the source task ID or UUID
Working directory rules
  • agent_cwd and test_cwd are separate concepts:
    • agent_cwd is where the agent starts
    • test_cwd is where the verifier runs
  • Set agent_cwd to the directory where the source task expects the agent to operate on the codebase or files
  • Set test_cwd to the directory where the verifier should execute
  • Do not assume they are the same. Only use the same value for both when the source task clearly indicates that the agent and verifier operate from the same directory
Validation rules
  • Do not assume the conversion is correct until a sample run succeeds without harness-level issues
  • Distinguish real task failures from conversion failures
  • If the sample exposes harness issues, fix them and rerun the sample before scaling up
  • Validate all three terminal states when practical:
    • a known-good sample yields 0
    • a known-bad or unsolved sample yields 1
    • an induced verifier/setup failure yields 2

Decision Rules

Inferring working directories

Parse the Dockerfile for WORKDIR directives. Use the last WORKDIR value found as the default working directory inside the container.

Use that information carefully:

  • infer test_cwd from the verifier and Dockerfile execution context
  • infer agent_cwd from where the source task expects the agent to work on the repo or files
  • do not default agent_cwd to test_cwd unless the source task clearly uses the same directory for both

If no better signal exists:

  • default test_cwd to the last WORKDIR, or "/" if none exists
  • choose agent_cwd from the most plausible agent workspace for the task, rather than assuming it matches test_cwd
Clarifying Docker behavior

If the source suite provides more than one viable environment path and the user has not explicitly said which to use, ask before converting.

Common ambiguous cases:

  • The source provides both a Dockerfile-like environment and a prebuilt image reference
  • The source provides environment files that could be rebuilt, but the user may prefer to keep the original image
  • The user may want you to rebuild the image, push it to a registry, and reference the new image in case.toml

When Docker behavior is ambiguous, ask which of these they want:

  • Preserve the original image reference in image
  • Convert using env/Dockerfile
  • Rebuild and publish a new image, then use that in image

Do not modify Dockerfiles just to express pass/fail/infra policy unless the user explicitly asks for environment changes. Prefer verifier-level changes for verdict semantics.

Show full SKILL.md (789 more words)Show less

Conversion Workflow

  1. Identify the source format by inspecting the input directory (look for characteristic files like task.toml, case.toml, etc.)
  2. Choose a suite name — derive from the dataset name or ask the user
  3. Create the output structure: <suite-name>/cases/
  4. For each task/case in the source, create a case directory and convert:
    • Config file → case.toml
    • Prompt/instruction → prompt.md
    • Test scripts → tests/
    • Dockerfile/environment → env/
    • Solution (if any) → oracle/
  5. Generate suite.toml listing all case names (sorted alphabetically)
  6. Validate: every case must have case.toml, prompt.md, tests/test.sh, and either image or env/Dockerfile
  7. Set permissions: chmod +x on tests/test.sh and oracle/solve.sh
  8. Ensure unique case names: if sanitization causes collisions, append a stable suffix (for example part of the source task ID/UUID)
  9. Smoke test a sample first: convert a small sample of 2-5 representative cases before committing to the full dataset
  10. Run Margin on that sample: execute the sample with margin run and carefully inspect the results, logs, and verifier behavior to identify conversion issues before scaling up
  11. Validate verifier behavior: confirm the converted tests/test.sh locates its test assets under {test_cwd}/tests, returns a non-zero exit code when the underlying tests fail, and does not discover the same tests from both the workspace and tests/
  12. Validate agent starting context: confirm agent_cwd points at the directory where the agent should actually begin work for the task
  13. Fix and rerun before scaling up: if the sample exposes harness issues, adjust the conversion so the verifier has a single authoritative test location, then rerun the sample before converting the full dataset

Generated Wrapper Guidance

If the source has test scripts (e.g., pytest files) but no test.sh, generate a wrapper:

bash
#!/bin/bash
set -euo pipefail

pass()  { printf 'VERDICT: PASS\n'; exit 0; }
fail()  { printf 'VERDICT: FAIL\n'; exit 1; }
infra() { printf 'VERDICT: INFRA\n' >&2; exit 2; }

<dependency installation commands or bootstrap checks>

set +e
<test runner command, e.g., pytest tests/test_outputs.py -rA>
exit_code=$?
set -e

case "$exit_code" in
  0) pass ;;
  1) fail ;;
  *) infra ;;
esac

Common failure signatures to look for during sample validation:

  • Path mismatch: the verifier still points at /tests/... instead of {test_cwd}/tests/...
  • Duplicate discovery: the same tests are collected from both the workspace and tests/
  • Masked failures: the wrapper reaches an ambiguous state but still exits 0 or 1

Source-Specific Adapters

Harbor

See references/harbor-mapping.md for the complete field-by-field mapping.

Harbor datasets (downloaded via harbor datasets download) have this layout:

<output-dir>/
├── <uuid-1>/<task-name>/
│   ├── task.toml
│   ├── instruction.md
│   ├── environment/
│   │   ├── Dockerfile
│   │   └── <setup scripts...>
│   ├── tests/
│   │   ├── test.sh
│   │   └── test_*.py
│   └── solution/
│       └── solve.sh
├── <uuid-2>/<task-name>/
│   └── ...

Each UUID directory wraps exactly one named task subdirectory.

Harbor-specific steps
  1. Scan the input directory for UUID subdirectories (each contains one task folder)
  2. For each task: a. Read task.toml b. Use the inner directory name, not the UUID, as the case name c. Sanitize the case name to be filesystem-safe d. If sanitization collides with an existing case name, append a stable suffix derived from the Harbor UUID e. Create cases/<case-name>/ f. Generate case.toml per references/harbor-mapping.md g. Copy instruction.md to prompt.md h. If Harbor provides both a reusable image path and rebuildable environment files, and the user has not chosen a Docker strategy, stop and ask which behavior they want i. If task.toml declares [environment].docker_image, map it to Margin image when the user wants to preserve or reuse the source image j. If the user wants a Dockerfile-backed case, copy environment/ to env/ k. If the user wants a rebuilt and published image, build from environment/, publish to the selected registry, and write the published image reference to Margin image l. Copy tests/ to tests/ m. Apply the general verifier rules above, especially path normalization, environment-variable overrides, single test location, and exit-code propagation n. If solution/ exists, copy it to oracle/ as reference material only
  3. Parse WORKDIR from env/Dockerfile to help infer working directories when using a Dockerfile-backed environment or rebuilding/publishing a new image. Use it directly for test_cwd when it matches the verifier context, and infer agent_cwd separately from the source task layout or expected repo workspace. If using a Harbor docker_image without a Dockerfile, infer both directories from the source task and verifier rather than assuming they match.
  4. Generate suite.toml with all case names
  5. chmod +x all .sh files in tests/ and oracle/
What to drop

These Harbor fields have no Margin Eval equivalent and are safely dropped:

  • version (replaced by kind = "test_case")
  • [environment].build_timeout_sec
  • [environment].allow_internet
  • [environment].mcp_servers
  • [verifier.env], [solution.env]
What to preserve as metadata

Resource constraints are useful context even though Margin doesn't enforce them at the case level. Store them in [metadata] for documentation only:

toml
[metadata]
# ... standard fields ...
harbor_cpus = 1
harbor_memory_mb = 2048
harbor_gpus = 0
harbor_agent_timeout_sec = 120

Validation Checklist

Before declaring a conversion complete, verify:

  • Every case has case.toml, prompt.md, tests/test.sh, and either image or env/Dockerfile
  • case.toml.name matches the directory name
  • agent_cwd points at the directory where the agent should actually start
  • test_cwd resolves to a real working directory assumption for the case
  • Required shell scripts are executable
  • The sample suite runs under margin run
  • The verifier reads test assets from the correct location
  • The verifier does not discover the same tests from both the workspace and tests/
  • The verifier propagates real failures with a non-zero exit code
  • Any remaining sample failures are real task failures, not harness failures

© Margin-Lab, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .agents/skills/suite-converter of Margin-Lab/evals.

  • SKILL.md
  • references/harbor-mapping.md

Open the folder on GitHubat commit b57dfe9

Compare with similar skills

Suite Converter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Suite Converter compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Suite Converter this skillMargin-Lab/evals161—~3.7kAutomated safety check: PassAGPL-3.0
Benchflow Experiment Reviewbenchflow-ai/benchflow353—~4kAutomated safety check: PassApache-2.0
Run SDK Testsrestatedev/sdk-typescript125—~745Automated safety check: PassMIT
Backend Testingrhesis-ai/rhesis397—~400Automated safety check: PassCustom licence
Corvus Bowtie Testingcorvus-dotnet/Corvus.JsonSchema199—~1.4kAutomated safety check: PassApache-2.0
PR TestElite588/AUTOGPT103—~9.4kAutomated safety check: NotesCustom licence

Similar skills

  • Benchflow Experiment Review

    benchflow-ai/benchflow

    Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.

    353 GitHub stars~4k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Run SDK Tests

    restatedev/sdk-typescript

    Run the Restate SDK conformance test suite locally against this SDK's Docker image.

    125 GitHub stars~745 tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Backend Testing

    rhesis-ai/rhesis

    Run the backend test suite correctly — working directory, Docker requirement, single-test vs full-suite commands.

    397 GitHub stars~400 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Corvus Bowtie Testing

    corvus-dotnet/Corvus.JsonSchema

    Test Corvus.JsonSchema against the JSON Schema Test Suite using Bowtie, the cross-implementation meta-validator.

    199 GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • PR Test

    Elite588/AUTOGPT

    E2E manual testing of PRs/branches using docker compose, agent-browser, and API calls.

    103 GitHub stars~9.4k tokensUpdated 5 mo ago
    Testing & QAAuto-check: notes
  • Dokan Automation

    getdokan/dokan

    Build, scaffold, and run the Dokan Lite/Pro Playwright suite.

    288 GitHub stars~6.1k tokensUpdated yesterday
    Testing & QAAuto-check: notes

More from Margin-Lab/evals

  • Agent Definition Creator

    Margin-Lab/evals

    Creates or updates Margin Eval agent definitions for new CLI coding agents.

    161 GitHub stars~2.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Suite Creator

    Margin-Lab/evals

    Creates new Margin Eval test suites from scratch. An agent skill from Margin-Lab/evals.

    161 GitHub stars~2.3k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Suite Converter

What does Suite Converter do?

Converts test suites from external eval frameworks into the Margin Eval suite format. Suite Converter is an agent skill from Margin-Lab/evals. Converts test suites from external eval frameworks into the Margin Eval suite format.

When should I use Suite Converter?

Suite Converter fits situations like: the user wants to import; migrate an eval dataset; test suite into Margin Eval format; they mention converting tasks from other benchmarking frameworks into Margins structure.

How do I install Suite Converter in Claude Code?

Run `npx skills add Margin-Lab/evals --skill suite-converter -a claude-code`. Or copy the skill folder (.agents/skills/suite-converter in Margin-Lab/evals) into .claude/skills/suite-converter in your project. Claude Code loads it when a task matches its description.

How do I install Suite Converter in Codex?

Run `npx skills add Margin-Lab/evals --skill suite-converter -a codex`. Or copy the skill folder (.agents/skills/suite-converter in Margin-Lab/evals) into .agents/skills/suite-converter in your project. Codex loads it when a task matches its description.

Can I use Suite Converter in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Margin-Lab/evals --skill suite-converter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/suite-converter, .gemini/skills/suite-converter, .github/skills/suite-converter and .opencode/skills/suite-converter in your project.

What does Suite Converter need to run?

Going by SKILL.md and its folder, Suite Converter needs the command-line tools its instructions call (kind). Our summary lists: Docker.

Does Suite Converter access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Suite Converter safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Suite Converter use?

Suite Converter is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Suite Converter use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to Suite Converter?

Skills that share tags, products or a category with Suite Converter: Benchflow Experiment Review (benchflow-ai/benchflow, 353 stars), Run SDK Tests (restatedev/sdk-typescript, 125 stars), Backend Testing (rhesis-ai/rhesis, 397 stars) and Corvus Bowtie Testing (corvus-dotnet/Corvus.JsonSchema, 199 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Suite Converter?

Margin-Lab (a GitHub organization) maintains it in Margin-Lab/evals, which has 161 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on July 31, 2026.

Source: Margin-Lab/evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.