Benchflow Experiment Review
benchflow-ai/benchflow
Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.
Converts test suites from external eval frameworks into the Margin Eval suite format.
$ npx skills add Margin-Lab/evals --skill suite-converter -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Margin-Lab/evals suite-converter --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/suite-converter .claude/skills/suite-converter && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "suite-converter" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-converter into .claude/skills/suite-converter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-converter", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-converterType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Margin-Lab/evals --skill suite-converter -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Margin-Lab/evals suite-converter --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/suite-converter .agents/skills/suite-converter && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "suite-converter" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-converter into .agents/skills/suite-converter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-converter", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Margin-Lab/evals --skill suite-converter -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Margin-Lab/evals suite-converter --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/suite-converter .cursor/skills/suite-converter && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "suite-converter" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-converter into .cursor/skills/suite-converter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-converter", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Margin-Lab/evals.git --path .agents/skills/suite-converter--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Margin-Lab/evals --skill suite-converter -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Margin-Lab/evals suite-converter --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/suite-converter .gemini/skills/suite-converter && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "suite-converter" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-converter into .gemini/skills/suite-converter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-converter", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Margin-Lab/evals suite-converterInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Margin-Lab/evals --skill suite-converter -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/suite-converter .github/skills/suite-converter && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "suite-converter" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-converter into .github/skills/suite-converter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-converter", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Margin-Lab/evals --skill suite-converter -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Margin-Lab/evals suite-converter --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Margin-Lab/evals.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/suite-converter .opencode/skills/suite-converter && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "suite-converter" agent skill from https://github.com/Margin-Lab/evals/tree/main/.agents/skills/suite-converter into .opencode/skills/suite-converter/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "suite-converter", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
suite-converterConverts test suites from external eval frameworks into the Margin Eval suite format.
Suite Converter is an agent skill from Margin-Lab/evals. Converts test suites from external eval frameworks into the Margin Eval suite format. Use this skill whenever the user wants to import, convert, translate, or migrate an eval dataset or test suite into Margin Eval format, or when they mention converting tasks from other benchmarking frameworks into Margin's structure.
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/harbor-mapping.md`).
It sits in Testing & QA, covering Test generation, Containers and LLM evaluation. It works with Docker. The repository describes itself as: Fast, robust, configurable agent evals. The licence is AGPL-3.0.
12 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit b57dfe9. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kindFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Suite Converter loads about 3.7k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 84 tokens; SKILL.md has 1,833 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Margin-Lab/evals at commit b57dfe9, republished under its AGPL-3.0 licence (© Margin-Lab). 1,833 words, ~3,739 tokens.
.claude/skills/suite-converter/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Converts downloaded eval datasets into valid Margin Eval test suites.
Use this skill when converting an external eval dataset into the Margin Eval suite format.
The job has three parts:
margin run before scaling up<suite-name>/
├── suite.toml
└── cases/
└── <case-name>/
├── case.toml
├── prompt.md
├── env/
│ └── Dockerfile
├── tests/
│ ├── test.sh
│ └── <other test files>
└── oracle/ # optional
└── solve.shkind = "test_suite"
name = "<suite-name>"
description = "<description>"
cases = [
"<case-1>",
"<case-2>",
]The cases array lists directory names under cases/, in the order they should run.
kind = "test_case"
name = "<case-name>"
description = "<one-line description>"
agent_cwd = "/"
test_cwd = "/"
test_timeout_seconds = 1800
[metadata]
difficulty = "easy"
category = "programming"
tags = ["tag1", "tag2"][metadata] can be preserved for human context and future tooling, but the current Margin compiler/runtime primarily cares about the execution fields such as kind, name, description, image, agent_cwd, test_cwd, and test_timeout_seconds.
Key rules:
name must match its directory name exactlykind is always "test_case"test_timeout_seconds is an integer (seconds)agent_cwd is the directory where the agent is expected to start and do its work inside the containertest_cwd is the working directory where test.sh runs inside the containerImage handling, exactly one of:
image = "registry/repo@sha256:<64hex>" for a pre-built, digest-pinned imageimage and place a Dockerfile at env/Dockerfile to build at compile timeThe full task description sent to the agent as its initial prompt. Must not be empty. Copy the source's instruction/prompt file as-is — don't summarize or reformat it.
The grading script. This is the evaluator. There is no separate grader abstraction.
Container environment. Can include supporting files alongside the Dockerfile. The entire env/ directory is the build context.
Optional reference solution. Not executed during normal eval runs. In the current Margin implementation, oracle/ is informational only and is not part of the active compile/runtime path.
These rules are cross-source and should hold for every conversion.
chmod +x)0 = pass, 1 = fail, 2 = infrareward.txt. The verifier process exit code is the authoritative result.tests/ are packaged together and staged at {test_cwd}/tests/ in the container/tests/..., rewrite them to Margin-compatible paths such as tests/... or "${PWD}/tests/..."./tests.tests/ into the workspace, its test command must avoid rediscovering the mounted tests/ tree. If it runs tests directly from tests/, do not also copy them into the workspace.0. Every verifier must terminate through explicit pass, fail, or infra paths.Use this attribution rule across conversions:
pass: the harness reached a trustworthy verdict and the candidate satisfied the taskfail: the harness reached a trustworthy verdict and the candidate did not satisfy the taskinfra: the harness could not reach a trustworthy verdict for reasons not attributable to the candidateClassify these as fail:
Classify these as infra:
If a parser or verifier layer cannot produce a trustworthy verdict, default to infra unless the harness already has direct evidence of a candidate-caused failure.
case.toml.name must match the case directory name exactlyagent_cwd and test_cwd are separate concepts:agent_cwd is where the agent startstest_cwd is where the verifier runsagent_cwd to the directory where the source task expects the agent to operate on the codebase or filestest_cwd to the directory where the verifier should execute012Parse the Dockerfile for WORKDIR directives. Use the last WORKDIR value found as the default working directory inside the container.
Use that information carefully:
test_cwd from the verifier and Dockerfile execution contextagent_cwd from where the source task expects the agent to work on the repo or filesagent_cwd to test_cwd unless the source task clearly uses the same directory for bothIf no better signal exists:
test_cwd to the last WORKDIR, or "/" if none existsagent_cwd from the most plausible agent workspace for the task, rather than assuming it matches test_cwdIf the source suite provides more than one viable environment path and the user has not explicitly said which to use, ask before converting.
Common ambiguous cases:
case.tomlWhen Docker behavior is ambiguous, ask which of these they want:
imageenv/DockerfileimageDo not modify Dockerfiles just to express pass/fail/infra policy unless the user explicitly asks for environment changes. Prefer verifier-level changes for verdict semantics.
task.toml, case.toml, etc.)<suite-name>/cases/case.tomlprompt.mdtests/env/oracle/suite.toml listing all case names (sorted alphabetically)case.toml, prompt.md, tests/test.sh, and either image or env/Dockerfilechmod +x on tests/test.sh and oracle/solve.shmargin run and carefully inspect the results, logs, and verifier behavior to identify conversion issues before scaling uptests/test.sh locates its test assets under {test_cwd}/tests, returns a non-zero exit code when the underlying tests fail, and does not discover the same tests from both the workspace and tests/agent_cwd points at the directory where the agent should actually begin work for the taskIf the source has test scripts (e.g., pytest files) but no test.sh, generate a wrapper:
#!/bin/bash
set -euo pipefail
pass() { printf 'VERDICT: PASS\n'; exit 0; }
fail() { printf 'VERDICT: FAIL\n'; exit 1; }
infra() { printf 'VERDICT: INFRA\n' >&2; exit 2; }
<dependency installation commands or bootstrap checks>
set +e
<test runner command, e.g., pytest tests/test_outputs.py -rA>
exit_code=$?
set -e
case "$exit_code" in
0) pass ;;
1) fail ;;
*) infra ;;
esacCommon failure signatures to look for during sample validation:
/tests/... instead of {test_cwd}/tests/...tests/0 or 1See references/harbor-mapping.md for the complete field-by-field mapping.
Harbor datasets (downloaded via harbor datasets download) have this layout:
<output-dir>/
├── <uuid-1>/<task-name>/
│ ├── task.toml
│ ├── instruction.md
│ ├── environment/
│ │ ├── Dockerfile
│ │ └── <setup scripts...>
│ ├── tests/
│ │ ├── test.sh
│ │ └── test_*.py
│ └── solution/
│ └── solve.sh
├── <uuid-2>/<task-name>/
│ └── ...Each UUID directory wraps exactly one named task subdirectory.
task.toml
b. Use the inner directory name, not the UUID, as the case name
c. Sanitize the case name to be filesystem-safe
d. If sanitization collides with an existing case name, append a stable suffix derived from the Harbor UUID
e. Create cases/<case-name>/
f. Generate case.toml per references/harbor-mapping.md
g. Copy instruction.md to prompt.md
h. If Harbor provides both a reusable image path and rebuildable environment files, and the user has not chosen a Docker strategy, stop and ask which behavior they want
i. If task.toml declares [environment].docker_image, map it to Margin image when the user wants to preserve or reuse the source image
j. If the user wants a Dockerfile-backed case, copy environment/ to env/
k. If the user wants a rebuilt and published image, build from environment/, publish to the selected registry, and write the published image reference to Margin image
l. Copy tests/ to tests/
m. Apply the general verifier rules above, especially path normalization, environment-variable overrides, single test location, and exit-code propagation
n. If solution/ exists, copy it to oracle/ as reference material onlyWORKDIR from env/Dockerfile to help infer working directories when using a Dockerfile-backed environment or rebuilding/publishing a new image. Use it directly for test_cwd when it matches the verifier context, and infer agent_cwd separately from the source task layout or expected repo workspace. If using a Harbor docker_image without a Dockerfile, infer both directories from the source task and verifier rather than assuming they match.suite.toml with all case nameschmod +x all .sh files in tests/ and oracle/These Harbor fields have no Margin Eval equivalent and are safely dropped:
version (replaced by kind = "test_case")[environment].build_timeout_sec[environment].allow_internet[environment].mcp_servers[verifier.env], [solution.env]Resource constraints are useful context even though Margin doesn't enforce them at the case level. Store them in [metadata] for documentation only:
[metadata]
# ... standard fields ...
harbor_cpus = 1
harbor_memory_mb = 2048
harbor_gpus = 0
harbor_agent_timeout_sec = 120Before declaring a conversion complete, verify:
case.toml, prompt.md, tests/test.sh, and either image or env/Dockerfilecase.toml.name matches the directory nameagent_cwd points at the directory where the agent should actually starttest_cwd resolves to a real working directory assumption for the casemargin runtests/© Margin-Lab, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in .agents/skills/suite-converter of Margin-Lab/evals.
Open the folder on GitHubat commit b57dfe9
Suite Converter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Suite Converter this skillMargin-Lab/evals | 161 | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | |
| Benchflow Experiment Reviewbenchflow-ai/benchflow | 353 | — | ~4k | Automated safety check: Pass | Apache-2.0 | |
| Run SDK Testsrestatedev/sdk-typescript | 125 | — | ~745 | Automated safety check: Pass | MIT | |
| Backend Testingrhesis-ai/rhesis | 397 | — | ~400 | Automated safety check: Pass | Custom licence | |
| Corvus Bowtie Testingcorvus-dotnet/Corvus.JsonSchema | 199 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| PR TestElite588/AUTOGPT | 103 | — | ~9.4k | Automated safety check: Notes | Custom licence |
benchflow-ai/benchflow
Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.
restatedev/sdk-typescript
Run the Restate SDK conformance test suite locally against this SDK's Docker image.
rhesis-ai/rhesis
Run the backend test suite correctly — working directory, Docker requirement, single-test vs full-suite commands.
corvus-dotnet/Corvus.JsonSchema
Test Corvus.JsonSchema against the JSON Schema Test Suite using Bowtie, the cross-implementation meta-validator.
Elite588/AUTOGPT
E2E manual testing of PRs/branches using docker compose, agent-browser, and API calls.
getdokan/dokan
Build, scaffold, and run the Dokan Lite/Pro Playwright suite.
Margin-Lab/evals
Creates or updates Margin Eval agent definitions for new CLI coding agents.
Margin-Lab/evals
Creates new Margin Eval test suites from scratch. An agent skill from Margin-Lab/evals.
Works with
Categories
Converts test suites from external eval frameworks into the Margin Eval suite format. Suite Converter is an agent skill from Margin-Lab/evals. Converts test suites from external eval frameworks into the Margin Eval suite format.
Suite Converter fits situations like: the user wants to import; migrate an eval dataset; test suite into Margin Eval format; they mention converting tasks from other benchmarking frameworks into Margins structure.
Run `npx skills add Margin-Lab/evals --skill suite-converter -a claude-code`. Or copy the skill folder (.agents/skills/suite-converter in Margin-Lab/evals) into .claude/skills/suite-converter in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Margin-Lab/evals --skill suite-converter -a codex`. Or copy the skill folder (.agents/skills/suite-converter in Margin-Lab/evals) into .agents/skills/suite-converter in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Margin-Lab/evals --skill suite-converter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/suite-converter, .gemini/skills/suite-converter, .github/skills/suite-converter and .opencode/skills/suite-converter in your project.
Going by SKILL.md and its folder, Suite Converter needs the command-line tools its instructions call (kind). Our summary lists: Docker.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Suite Converter is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Suite Converter: Benchflow Experiment Review (benchflow-ai/benchflow, 353 stars), Run SDK Tests (restatedev/sdk-typescript, 125 stars), Backend Testing (rhesis-ai/rhesis, 397 stars) and Corvus Bowtie Testing (corvus-dotnet/Corvus.JsonSchema, 199 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Margin-Lab (a GitHub organization) maintains it in Margin-Lab/evals, which has 161 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on July 31, 2026.
Source: Margin-Lab/evals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.