Agent skill

Paperclip Evals

by paperclipai in paperclipai/paperclip

Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

MITAuto-check passedAI & LLM Engineering

Install Paperclip Evals

skills CLI
$ npx skills add paperclipai/paperclip --skill paperclip-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install paperclipai/paperclip paperclip-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/paperclipai/paperclip.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/paperclip-evals .claude/skills/paperclip-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
paperclip-evals
GitHub stars
100k
Token cost
~839 tokens
SKILL.md length
386 words
Files
1
Skills in repo
60
Repo updated
First seen
Licence
MIT

At a glance

Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

  • Tasks that involve LLM evaluation
  • SKILL.md covers Discover the repository, Route the request and Work safely
  • Calls pnpm and git
  • Tasks that involve End-to-end testing

What it does

Paperclip Evals is an agent skill from paperclipai/paperclip. Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

Its SKILL.md is about 840 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and End-to-end testing. It works with pnpm. The repository describes itself as: The open-source app everyone uses to manage agents at work. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve End-to-end testing

Example prompts

  • “/paperclip-evals”

What it can do on your machine

Read from SKILL.md and the folder at commit de9ab8e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paperclip Evals loads about 839 tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 386 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~839

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from paperclipai/paperclip at commit de9ab8e, republished under its MIT licence (© paperclipai). 386 words, ~839 tokens.

Download SKILL.mdSave it as .claude/skills/paperclip-evals/SKILL.md (or your agent's skills folder).
name
paperclip-evals
description
Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

Paperclip evals

Use this skill when a request concerns Paperclip evaluation selection, interpretation, evidence, history, or a live run. Read doc/evals.md in the Paperclip repository first. It defines the two families and their boundaries.

Discover the repository

Do not assume the skill's installed location is inside a checkout. Locate the repo explicitly with git rev-parse --show-toplevel from the current directory, or inspect likely workspace roots and select the checkout containing package.json, tests/runner-e2e, and packages/paperclip-runner. Locate the private sibling paperclip-evals only when a Runner Eval needs its definitions; use an explicit PAPERCLIP_EVALS_ROOT or a discovered sibling checkout. Never invent a relative path from this copied skill into the repository.

Route the request

Choose Runner Evals for real runner/provider protocol behavior against the mock control plane. Authoritative details are in packages/paperclip-runner/docs/runner-protocol-live-evals.md and the sibling paperclip-evals/evals/paperclip-runner definitions.

Choose Product E2E Evals for real browser/server/database/runner/provider workflows, including local and Daytona environments. Read tests/runner-e2e/README.md, then FIXTURES.md, SECURITY.md, or EVERYDAY-WORKFLOWS.md as relevant. Everyday Workflows remain Product E2E even when imported into Evalbook. “Headless” is a browser mode, not a family.

Show full SKILL.md (212 more words)Show less

Work safely

Start with read-only catalog inspection and credential-free validation. For Product E2E use pnpm test:e2e:runner:typecheck, pnpm test:e2e:runner:unit, and pnpm test:e2e:runner -- --list; run one explicit cell only when the user has authorized a live/paid run and the needed credentials and immutable Daytona image are configured. For Runner Evals use the pinned eval revision and the documented workflow/CLI. Never use a partial selector as evidence of full coverage.

Keep source revisions, definition/catalog fingerprints, model/profile, environment, selected cells, retries, timing, usage/cost coverage, and grader version attached to every interpretation. Preserve partial attempts and classify failures as product, model/provider behavior, grading/evidence, or infrastructure from the observed failure and supported cause. A usable completed behavior failure is not infrastructure; missing provider/profile, transport, startup, or evidence requires examining the evidence before choosing the cause.

Use the existing family generator and viewer. Public projections may contain sanitized fixture conversation and allowlisted tool outcomes/evidence; follow the family's projection and publisher checks. Do not expose raw trusted artifacts, credentials, secrets, private data, provider session IDs, or hidden reasoning. A refresh from retained evidence has zero provider calls and remains the original measurement with a new presentation. Link the public histories and hub from doc/evals.md when reporting results.

For adding a case or fixture, use the narrower add-runner-eval or add-product-e2e-eval skill.

© paperclipai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/paperclip-evals of paperclipai/paperclip.

Open the folder on GitHubat commit de9ab8e

Compare with similar skills

Paperclip Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paperclip Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paperclip Evals this skillpaperclipai/paperclip100k—~839Automated safety check: PassMIT
Run EvalsDevin-AXIS/iPolloWork6.8k—~851Automated safety check: PassCustom licence
Chatbox Session RAG Evalchatboxai/chatbox42k—~758Automated safety check: PassGPL-3.0
Run Evalslycorp-jp/sim-use1.4k—~1.4kAutomated safety check: PassApache-2.0
Write A Specdifferent-ai/openwork24k—~3.3kAutomated safety check: PassCustom licence
Phoenix Pxi PlaywrightArize-ai/phoenix12k—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • Run Evals

    Devin-AXIS/iPolloWork

    do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals.

    6.8k GitHub stars~851 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Chatbox Session RAG Eval

    chatboxai/chatbox

    Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.

    42k GitHub stars~758 tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed
  • Run Evals

    lycorp-jp/sim-use

    Prepare the environment and run the LLM-driven agent evals (e2e/agent-evals/) against a chosen sim-use binary.

    1.4k GitHub stars~1.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Write A Spec

    different-ai/openwork

    Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer.

    24k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Phoenix Pxi Playwright

    Arize-ai/phoenix

    Write, extend, and debug PXI Playwright E2E tests for Phoenix.

    12k GitHub stars~2.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Eval Authoring Workflow

    Azure/azure-sdk-tools

    Official

    Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows.

    134 GitHub stars~972 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from paperclipai/paperclip

All 60 skills in this repo
  • Garden Inbox

    paperclipai/paperclip

    Scan a Paperclip user's Mine inbox, classify reversible archive candidates, request checkbox confirmation, and archive only accepted selections.

    100k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Paperclip

    paperclipai/paperclip

    Interact with the Paperclip control plane API for task coordination and governance.

    100k GitHub stars~9.6k tokensUpdated today
    Auto-check passed
  • Paperclip

    paperclipai/paperclip

    A skill your agent uses for Paperclip-managed tasks and heartbeats: reading task context, delivering task documents or files, updating completion or blockers, coordinating or delegating work, and…

    100k GitHub stars~18k tokensUpdated today
    Auto-check passed
  • Design Guide

    paperclipai/paperclip

    Paperclip UI design system guide for building consistent, reusable frontend components.

    100k GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Paperclip Page

    paperclipai/paperclip

    Publish static HTML pages and asset folders to the Paperclip S3/CloudFront page host.

    100k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Paperclip Create Agent

    paperclipai/paperclip

    Create new agents in Paperclip with governance-aware hiring.

    100k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed

Works with

Questions about Paperclip Evals

What does Paperclip Evals do?

Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification. Paperclip Evals is an agent skill from paperclipai/paperclip. Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

When should I use Paperclip Evals?

Paperclip Evals fits situations like: tasks that involve LLM evaluation; tasks that involve End-to-end testing.

How do I install Paperclip Evals in Claude Code?

Run `npx skills add paperclipai/paperclip --skill paperclip-evals -a claude-code`. Or copy the skill folder (.agents/skills/paperclip-evals in paperclipai/paperclip) into .claude/skills/paperclip-evals in your project. Claude Code loads it when a task matches its description.

How do I install Paperclip Evals in Codex?

Run `npx skills add paperclipai/paperclip --skill paperclip-evals -a codex`. Or copy the skill folder (.agents/skills/paperclip-evals in paperclipai/paperclip) into .agents/skills/paperclip-evals in your project. Codex loads it when a task matches its description.

Can I use Paperclip Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add paperclipai/paperclip --skill paperclip-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paperclip-evals, .gemini/skills/paperclip-evals, .github/skills/paperclip-evals and .opencode/skills/paperclip-evals in your project.

What does Paperclip Evals need to run?

Going by SKILL.md and its folder, Paperclip Evals needs the command-line tools its instructions call (pnpm and git).

Does Paperclip Evals access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Paperclip Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Paperclip Evals use?

Paperclip Evals is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Paperclip Evals use?

About 839 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Paperclip Evals?

Skills that share tags, products or a category with Paperclip Evals: Run Evals (Devin-AXIS/iPolloWork, 6.8k stars), Chatbox Session RAG Eval (chatboxai/chatbox, 42k stars), Run Evals (lycorp-jp/sim-use, 1.4k stars) and Write A Spec (different-ai/openwork, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paperclip Evals?

paperclipai (a GitHub organization) maintains it in paperclipai/paperclip, which has 99,905 GitHub stars. The repository holds 60 skills in this directory. The repository was last updated on October 11, 2026.

Source: paperclipai/paperclip on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.