Agent skill

Run Evals

by Devin-AXIS in Devin-AXIS/iPolloWork

do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals.

Custom licenceAuto-check passedAI & LLM Engineering

Install Run Evals

skills CLI
$ npx skills add Devin-AXIS/iPolloWork --skill run-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Devin-AXIS/iPolloWork run-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Devin-AXIS/iPolloWork.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.opencode/skills/run-evals .claude/skills/run-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-evals
GitHub stars
6.7k
Token cost
~851 tokens
SKILL.md length
311 words
Files
1
Skills in repo
30
Repo updated
First seen
Licence
Custom licence

At a glance

do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals.

  • Tasks that involve LLM evaluation
  • SKILL.md covers Prerequisites, Preferred path: Daytona sandbox, Run the flows and Recording (motion only), plus 2 more sections
  • Calls pnpm, bash and curl
  • Tasks that involve End-to-end testing

What it does

Run Evals is an agent skill from Devin-AXIS/iPolloWork. do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. Launches iPolloWork on Daytona or local Electron and runs the coded eval flows via CDP. Launch + run mechanics; the proof loop itself is the fraimz skill.

Its SKILL.md is about 850 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and End-to-end testing. It works with pnpm and Bash. The repository describes itself as: Enterprise-grade, local-first Agent Workbench for people and agent teams. A unified multi-engine workspace for Codex Harness, DeepSeek Harness, and OpenCode, with unified plugins…

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve End-to-end testing

Example prompts

  • “/run-evals”

What it can do on your machine

Read from SKILL.md and the folder at commit 71a74ad. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • bash
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Evals loads about 851 tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 311 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~851

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 311 words (~851 tokens).

“Launch a real iPolloWork app and run coded eval flows against it. This skill owns launch + run; the prove/repair/verdict loop and evidence standard live in the fraimz skill — load that too for anything that ends in a verdict.”

— opening of SKILL.md by Devin-AXIS, Custom licence
name
run-evals

Read the full SKILL.md on GitHub

Files

Just SKILL.md in .opencode/skills/run-evals of Devin-AXIS/iPolloWork.

Open the folder on GitHubat commit 71a74ad

Compare with similar skills

Run Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Evals this skillDevin-AXIS/iPolloWork6.7k—~851Automated safety check: PassCustom licence
Paperclip Evalspaperclipai/paperclip98k—~839Automated safety check: PassMIT
Chatbox Session RAG Evalchatboxai/chatbox42k—~758Automated safety check: PassGPL-3.0
Run Evalslycorp-jp/sim-use1.4k—~1.4kAutomated safety check: PassApache-2.0
Write A Specdifferent-ai/openwork24k—~3.3kAutomated safety check: PassCustom licence
Phoenix Pxi PlaywrightArize-ai/phoenix12k—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • Paperclip Evals

    paperclipai/paperclip

    Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

    98k GitHub stars~839 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Chatbox Session RAG Eval

    chatboxai/chatbox

    Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.

    42k GitHub stars~758 tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Run Evals

    lycorp-jp/sim-use

    Prepare the environment and run the LLM-driven agent evals (e2e/agent-evals/) against a chosen sim-use binary.

    1.4k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Write A Spec

    different-ai/openwork

    Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer.

    24k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Phoenix Pxi Playwright

    Arize-ai/phoenix

    Write, extend, and debug PXI Playwright E2E tests for Phoenix.

    12k GitHub stars~2.6k tokensUpdated today
    Testing & QAAuto-check passed
  • E2E

    kortix-ai/suna

    Agentic end-to-end tests with e2e, the e2e runner. An agent skill from kortix-ai/suna.

    20k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed

More from Devin-AXIS/iPolloWork

All 30 skills in this repo
  • A code-change gate for the iPolloWork repository: search and reuse first, keep one source of truth, justify every new file or dependency, and audit the change.

    6.7k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Reproduces iPolloWork enterprise TLS behavior in a Daytona Windows sandbox by installing a fake corporate CA and proving the app and its runtimes use the OS trust store.

    6.7k GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Build iPolloWork Templates

    Devin-AXIS/iPolloWork

    Builds a reusable iPolloWork template for Design, Slides or PPT, or HyperFrames Video through conversation, keeping a manifest, reusable variables and a validated package current.

    6.7k GitHub stars~925 tokensUpdated yesterday
    Auto-check passed
  • OpenCode Plugin Creator

    Devin-AXIS/iPolloWork

    Scaffolds an OpenCode plugin for iPolloWork with the right async factory shape, zod-based tool definitions and hook registration, and explains where to place and register it.

    6.7k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Agent-First Screenshots

    Devin-AXIS/iPolloWork

    Has an agent drive the real app through CDP and capture clean, defect-free product screenshots, gated by DOM, pixel and vision checks in a loop.

    6.7k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Daytona Chrome CDP

    Devin-AXIS/iPolloWork

    Launches a standalone Chrome or Chromium in a Daytona sandbox and drives it over CDP for web sign-in, OAuth and other flows that should bypass the Electron app.

    6.7k GitHub stars~582 tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Run Evals

What does Run Evals do?

do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. Run Evals is an agent skill from Devin-AXIS/iPolloWork. do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals.

When should I use Run Evals?

Run Evals fits situations like: tasks that involve LLM evaluation; tasks that involve End-to-end testing.

How do I install Run Evals in Claude Code?

Run `npx skills add Devin-AXIS/iPolloWork --skill run-evals -a claude-code`. Or copy the skill folder (.opencode/skills/run-evals in Devin-AXIS/iPolloWork) into .claude/skills/run-evals in your project. Claude Code loads it when a task matches its description.

How do I install Run Evals in Codex?

Run `npx skills add Devin-AXIS/iPolloWork --skill run-evals -a codex`. Or copy the skill folder (.opencode/skills/run-evals in Devin-AXIS/iPolloWork) into .agents/skills/run-evals in your project. Codex loads it when a task matches its description.

Can I use Run Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Devin-AXIS/iPolloWork --skill run-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-evals, .gemini/skills/run-evals, .github/skills/run-evals and .opencode/skills/run-evals in your project.

What does Run Evals need to run?

Going by SKILL.md and its folder, Run Evals needs the command-line tools its instructions call (pnpm, bash and curl).

Does Run Evals access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Run Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Evals use?

Run Evals has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Run Evals use?

About 851 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Run Evals?

Skills that share tags, products or a category with Run Evals: Paperclip Evals (paperclipai/paperclip, 98k stars), Chatbox Session RAG Eval (chatboxai/chatbox, 42k stars), Run Evals (lycorp-jp/sim-use, 1.4k stars) and Write A Spec (different-ai/openwork, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Evals?

Devin-AXIS (a GitHub user) maintains it in Devin-AXIS/iPolloWork, which has 6,717 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on October 6, 2026.

Source: Devin-AXIS/iPolloWork on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.