Benchflow
benchflow-ai/benchflow
Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.
Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.
$ npx skills add tokencanopy/e2a --skill email-evals -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install tokencanopy/e2a email-evals --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/tokencanopy/e2a.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/e2a/skills/email-evals .claude/skills/email-evals && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "email-evals" agent skill from https://github.com/tokencanopy/e2a/tree/main/plugins/e2a/skills/email-evals into .claude/skills/email-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "email-evals", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/tokencanopy/e2a/tree/main/plugins/e2a/skills/email-evalsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add tokencanopy/e2a --skill email-evals -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install tokencanopy/e2a email-evals --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokencanopy/e2a.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/e2a/skills/email-evals .agents/skills/email-evals && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "email-evals" agent skill from https://github.com/tokencanopy/e2a/tree/main/plugins/e2a/skills/email-evals into .agents/skills/email-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "email-evals", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tokencanopy/e2a --skill email-evals -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install tokencanopy/e2a email-evals --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokencanopy/e2a.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/e2a/skills/email-evals .cursor/skills/email-evals && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "email-evals" agent skill from https://github.com/tokencanopy/e2a/tree/main/plugins/e2a/skills/email-evals into .cursor/skills/email-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "email-evals", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/tokencanopy/e2a.git --path plugins/e2a/skills/email-evals--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add tokencanopy/e2a --skill email-evals -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install tokencanopy/e2a email-evals --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokencanopy/e2a.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/e2a/skills/email-evals .gemini/skills/email-evals && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "email-evals" agent skill from https://github.com/tokencanopy/e2a/tree/main/plugins/e2a/skills/email-evals into .gemini/skills/email-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "email-evals", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install tokencanopy/e2a email-evalsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add tokencanopy/e2a --skill email-evals -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/tokencanopy/e2a.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/e2a/skills/email-evals .github/skills/email-evals && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "email-evals" agent skill from https://github.com/tokencanopy/e2a/tree/main/plugins/e2a/skills/email-evals into .github/skills/email-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "email-evals", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tokencanopy/e2a --skill email-evals -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install tokencanopy/e2a email-evals --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tokencanopy/e2a.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/e2a/skills/email-evals .opencode/skills/email-evals && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "email-evals" agent skill from https://github.com/tokencanopy/e2a/tree/main/plugins/e2a/skills/email-evals into .opencode/skills/email-evals/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "email-evals", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
email-evalsAuthor and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.
Email Evals is an agent skill from tokencanopy/e2a. Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents. Use when a user wants to define synthetic email cases, validate a suite, inspect its dry-run plan, run it after approval, or regrade changed assertions.
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 104 other files (for example `email-evals.sh`, `runtime/THIRD_PARTY_NOTICES.md` and `runtime/package-lock.json`).
It sits in AI & LLM Engineering, covering LLM evaluation, Agent evaluation and testing and Transactional email. The repository describes itself as: Open-source email API for applications and AI agents. Managed hosting at e2a.dev, or self-host with Docker. The licence is Apache-2.0.
12 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 776fe2c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (JavaScript and Shell, from the files we listed), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.e2a.devFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
E2A_EVAL_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Email Evals loads about 2.1k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 953 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from tokencanopy/e2a at commit 776fe2c, republished under its Apache-2.0 licence (© tokencanopy). 953 words, ~2,095 tokens.
.claude/skills/email-evals/SKILL.md (or your agent's skills folder). This skill also uses 102 other files; get the full folder from GitHub.Ask one logical question at a time; do not dump a questionnaire. Ask exactly one prompt, wait for the answer, then advance. Conditionally omit a later prompt only when an earlier answer proves its field irrelevant; never bundle prompts.
<!-- email-evals:authoring-prompts:start -->
<!-- email-evals:field=use-case --> **use case** — Ask: “What is the synthetic use case?”
<!-- email-evals:field=existing-target-runtime --> **existing target runtime** — Ask: “Does the existing target runtime already run?”
<!-- email-evals:field=dedicated-actor-environment-name --> **dedicated actor environment name** — Ask: “What is the dedicated actor environment name?”
<!-- email-evals:field=dedicated-target-environment-name --> **dedicated target environment name** — Ask: “What is the dedicated target environment name?”
<!-- email-evals:field=expected-action --> **expected action** — Ask: “What is the expected action?”
<!-- email-evals:field=exact-allowed-recipients --> **exact allowed recipients** — Ask: “Which exact allowed recipients are required?”
<!-- email-evals:field=sender --> **sender** — Ask: “What is the sender?”
<!-- email-evals:field=reply-to --> **Reply-To** — Ask: “What is the Reply-To expectation?”
<!-- email-evals:field=thread --> **thread** — Ask: “What is the thread expectation?”
<!-- email-evals:field=subject --> **subject** — Ask: “What is the subject expectation?”
<!-- email-evals:field=required-facts --> **required facts** — Ask: “What required facts apply?”
<!-- email-evals:field=forbidden-patterns --> **forbidden patterns** — Ask: “What forbidden patterns apply?”
<!-- email-evals:field=attachments --> **attachments** — Ask: “What attachments are expected?”
<!-- email-evals:field=timeout --> **timeout** — Ask: “What timeout applies?”
<!-- email-evals:field=lifecycle --> **lifecycle** — Ask: “What lifecycle outcome is required?”
<!-- email-evals:authoring-prompts:end -->
This skill does not build or start the target agent runtime.
Refuse real customer messages, identifiers, customer domains, or production-derived fixtures. Immediately propose a synthetic replacement, such as a fictional order identifier and a .test mailbox. Keep every case and attachment synthetic.
Do not scaffold until these answers are sufficient to make deterministic assertions.
Never change e2a protection. Tell the user to configure containment separately, in the same dedicated account for e2a, with an account-scoped API key:
allowlist/block [target]allowlist/block [actor, probes...]Use dedicated actor and target agents only. Do not infer, broaden, or repair protection settings; surface protection failures during validation instead.
After gathering sufficient answers, scaffold the suite:
Resolve the absolute directory containing this loaded SKILL.md resource.
For the tool call, set EMAIL_EVALS_LAUNCHER to the absolute path formed by
joining that resource directory with email-evals.sh. This is an agent-local
value, not a persisted environment variable. Do not derive it from the user's
current directory or a client-specific plugin-root variable.
"$EMAIL_EVALS_LAUNCHER" scaffold --root <suite-root> --name <suite-name> --target-env <target-env> --actor-env <actor-env>Edit the generated suite.yaml and case YAML files to reflect the answers. The
credential name is fixed outside suite authority: export only
E2A_EVAL_API_KEY. YAML interpolation is limited to the documented actor,
target, and probe mailbox fields; never interpolate names, actions, timing,
subjects, bodies, patterns, or attachment metadata.
Then verify the checked-in, single-file plugin runtime bundle and validate the suite. Setup performs no package installation and never copies or executes JavaScript or dependencies beneath the suite root:
Resolve the absolute directory containing this loaded SKILL.md resource.
For the tool call, set EMAIL_EVALS_LAUNCHER to the absolute path formed by
joining that resource directory with email-evals.sh. This is an agent-local
value, not a persisted environment variable. Do not derive it from the user's
current directory or a client-specific plugin-root variable.
"$EMAIL_EVALS_LAUNCHER" setup --root <suite-root>
"$EMAIL_EVALS_LAUNCHER" validate --suite <suite-root>/suite.yamlShow the complete alias-only dry-run plan, its approvalDigest, and any
protection failures. Validation is the preflight gate; do not run a suite
while it reports a protection or capability failure. Treat the digest as an
opaque value; it binds the resolved identities, origin, containment posture,
stimuli, assertions, recipients, and execution limits without revealing them.
The public https://api.e2a.dev origin is the default. If the suite names a
custom/local origin, the operator must independently authorize the exact
origin on every command with --trusted-origin <origin>; cleartext origins
are restricted to loopback. Never infer this flag from suite content.
Ask for explicit user approval immediately before run, because it sends real email between the dedicated agents. Do not treat earlier authoring answers as approval.
Only after that approval, run:
Resolve the absolute directory containing this loaded SKILL.md resource.
For the tool call, set EMAIL_EVALS_LAUNCHER to the absolute path formed by
joining that resource directory with email-evals.sh. This is an agent-local
value, not a persisted environment variable. Do not derive it from the user's
current directory or a client-specific plugin-root variable.
"$EMAIL_EVALS_LAUNCHER" run --suite <suite-root>/suite.yaml --approval-digest <approvalDigest-from-validate>If the suite, resolved mailbox identities, custom origin, protection posture,
or plan changed since validation, run fails closed. Validate again, show the
new complete plan, and request fresh approval; never substitute or guess a
digest.
Read the generated report.md. Summarize deterministic failures without hiding errors, then propose the smallest case/agent change that addresses each failure.
When only assertions changed, use regrade instead of run; regrade performs no sends:
Resolve the absolute directory containing this loaded SKILL.md resource.
For the tool call, set EMAIL_EVALS_LAUNCHER to the absolute path formed by
joining that resource directory with email-evals.sh. This is an agent-local
value, not a persisted environment variable. Do not derive it from the user's
current directory or a client-specific plugin-root variable.
"$EMAIL_EVALS_LAUNCHER" regrade --suite <suite-root>/suite.yaml --run <run-dir>Regrade accepts a changed full-suite digest only when the execution digest is
unchanged. It always uses the current validated assertions and rejects changes
to sending, correlation, timing, containment, actor, target, or origin inputs.
The output root also contains the private mode-0600
.email-evals-artifact-auth-key, which authenticates redaction-loss metadata
independently of the rotatable API key. Keep that hidden file with its run
directories; never publish it or place it inside a suite. Every ordinary case
record carries authenticated redaction state, including an explicit empty
declaration when no text was erased. Regrade fails closed if this authentication
root or a case's authenticated redaction state is missing, replaced, or altered.
If body evidence required configured-pattern redaction, the forbidden-pattern
set must remain identical (reordering is allowed); changing that set fails
closed because arbitrary new regexes cannot be evaluated against erased text.
Keep claims within the launch slice: no semantic judge, no deep HTML equivalence, no scheduled-send proof, and no full review/bounce/complaint matrix. Do not promise simulator, model, or data-generator features.
© tokencanopy, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 102 other files in plugins/e2a/skills/email-evals of tokencanopy/e2a.
Open the folder on GitHubat commit 776fe2c
Email Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Email Evals this skilltokencanopy/e2a | 192 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Benchflowbenchflow-ai/benchflow | 353 | — | ~1.9k | Automated safety check: Notes | Apache-2.0 | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Autocontextgreyhaven-ai/autocontext | 1.3k | — | ~892 | Automated safety check: Pass | Apache-2.0 | |
| GAIA Agent Benchmarkingamd/gaia | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Bkit Evalsww-w-ai/bkit-claude-code | 600 | — | ~1k | Automated safety check: Notes | Apache-2.0 |
benchflow-ai/benchflow
Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
greyhaven-ai/autocontext
Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.
amd/gaia
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
ww-w-ai/bkit-claude-code
Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.
windmill-labs/windmill
Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.
tokencanopy/e2a
Conversationally configure and operate a policy-first, always-on local e2a email agent.
tokencanopy/e2a
A skill your agent uses when operating an already-connected e2a inbox over MCP: reading, composing, sending, replying, forwarding, handling attachments, managing contacts/outreach, scheduling mail…
tokencanopy/e2a
Beta — Stay in the loop over email during a long-running coding session.
tokencanopy/e2a
Beta — Deploy the autonomous-repo feedback loop into a GitHub repo.
tokencanopy/e2a
A skill your agent uses when an existing e2a MCP connection, inbox, custom domain, protection policy, webhook, or message delivery is failing or unclear.
tokencanopy/e2a
A skill your agent uses when adding e2a email capabilities to an application or codebase: outbound sending, inbound signed webhooks, REST polling, or SDK integration.
Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents. Email Evals is an agent skill from tokencanopy/e2a. Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.
Email Evals fits situations like: A user wants to define synthetic email cases; validate a suite; inspect its dry-run plan; run it after approval.
Run `npx skills add tokencanopy/e2a --skill email-evals -a claude-code`. Or copy the skill folder (plugins/e2a/skills/email-evals in tokencanopy/e2a) into .claude/skills/email-evals in your project. Claude Code loads it when a task matches its description.
Run `npx skills add tokencanopy/e2a --skill email-evals -a codex`. Or copy the skill folder (plugins/e2a/skills/email-evals in tokencanopy/e2a) into .agents/skills/email-evals in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tokencanopy/e2a --skill email-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/email-evals, .gemini/skills/email-evals, .github/skills/email-evals and .opencode/skills/email-evals in your project.
Going by SKILL.md and its folder, Email Evals needs JavaScript and a shell for the scripts in its folder and credentials named E2A_EVAL_API_KEY. Our summary lists: Node.js; A Bash shell; A credential in E2A_EVAL_API_KEY.
SKILL.md names 1 domain. In commands or code: api.e2a.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Email Evals is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Email Evals: Benchflow (benchflow-ai/benchflow, 353 stars), Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Autocontext (greyhaven-ai/autocontext, 1.3k stars) and GAIA Agent Benchmarking (amd/gaia, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
tokencanopy (a GitHub organization) maintains it in tokencanopy/e2a, which has 192 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 6, 2026.
Source: tokencanopy/e2a on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.