Agent skill

Add Product E2E Eval

by paperclipai in paperclipai/paperclip

Add or extend a Paperclip full-stack runner E2E workflow, fixture, matcher, or report evidence path for local or Daytona execution.

MITAuto-check passedTesting & QA

Install Add Product E2E Eval

skills CLI
$ npx skills add paperclipai/paperclip --skill add-product-e2e-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install paperclipai/paperclip add-product-e2e-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/paperclipai/paperclip.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-product-e2e-eval .claude/skills/add-product-e2e-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-product-e2e-eval
GitHub stars
99k
Token cost
~1.1k tokens
SKILL.md length
522 words
Files
1
Skills in repo
60
Repo updated
First seen
Licence
MIT

At a glance

Add or extend a Paperclip full-stack runner E2E workflow, fixture, matcher, or report evidence path for local or Daytona execution.

  • Tasks that involve End-to-end testing
  • Calls pnpm and git

What it does

Add Product E2E Eval is an agent skill from paperclipai/paperclip. Add or extend a Paperclip full-stack runner E2E workflow, fixture, matcher, or report evidence path for local or Daytona execution.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing. The repository describes itself as: The open-source app everyone uses to manage agents at work. The licence is MIT.

When your agent uses it

  • Tasks that involve End-to-end testing

Example prompts

  • “/add-product-e2e-eval”

What it can do on your machine

Read from SKILL.md and the folder at commit b9750b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Product E2E Eval loads about 1.1k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 522 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from paperclipai/paperclip at commit b9750b1, republished under its MIT licence (© paperclipai). 522 words, ~1,113 tokens.

Download SKILL.mdSave it as .claude/skills/add-product-e2e-eval/SKILL.md (or your agent's skills folder).
name
add-product-e2e-eval
description
Add or extend a Paperclip full-stack runner E2E workflow, fixture, matcher, or report evidence path for local or Daytona execution.

Add a Product E2E Eval

Use this skill for Product E2E Evals: real Chromium, Paperclip server, database, runner, provider, and optionally Daytona. Everyday Workflows are in this family even when their packaged results are imported into Evalbook. Runner protocol cases against the mock control plane belong in add-runner-eval.

Locate the repository using PAPERCLIP_ROOT when supplied, or git rev-parse --show-toplevel from a checkout. From outside Git, inspect workspace roots such as ~/paperclipai/paperclip; verify the selected root contains tests/runner-e2e and packages/paperclip-runner. Run commands from that repository root. The copied skill may live outside the checkout. Read doc/evals.md, then the authoritative tests/runner-e2e/README.md, FIXTURES.md, SECURITY.md, and EVERYDAY-WORKFLOWS.md for the selected area. Inspect the nearest existing catalog entry, case, harness flow, matcher, evidence writer, and report test before changing anything. Keep the user journey on production browser/API surfaces; do not add private runner hooks or direct fixture database writes.

Define one bounded workflow with a clear user outcome, durable state assertions, and independent evidence. Declare profile, environment, expected provider turns, timeout, cleanup, screenshots, artifact checks, and billing scope. Keep credentials and secrets out of catalog data, screenshots, fixture metadata, logs, and tracked files. Daytona images must be immutable digest references. Use the existing result validator, failure classifier, screenshot policy, and Product E2E report pipeline rather than duplicating them.

Register the workflow in tests/runner-e2e/catalog.ts and the relevant fixture-registry.ts paths. Everyday cases and actions live in everyday-cases.ts and everyday-flow.ts; match the existing suite's structure. Wire its profile/environment/case IDs through existing selector and matcher tables, and add report/catalog coverage tests where the surrounding suite does so. Calibrate new grading assertions with a valid outcome and a plausible wrong outcome; missing evidence must not produce a pass. Confirm discovery before running it:

sh
pnpm test:e2e:runner -- --list --suite <suite-name>
pnpm test:e2e:runner -- --list --suite everyday-workflows

--all intentionally excludes the manual everyday-workflows suite and other explicit-only cells. Use the exact suite or execution ID for those. The Product E2E generator is pnpm test:e2e:runner:report; an Everyday Workflows result may also be imported into Evalbook with its canonical importer, but it remains a Product E2E run and should use its packaged dashboard/report first.

Show full SKILL.md (184 more words)Show less

Run credential-free checks first:

sh
pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
pnpm test:e2e:runner -- --list

For an authorized live check, use the smallest explicit local --id selector. Select Daytona only with the configured immutable image and credentials. A full --all campaign is paid and is for the governed workflow. Verify results from the packaged attempt evidence and dashboard; passing model-authored tests cannot override the independent oracle.

Classify a completed wrong workflow as product or model/provider behavior as the assertions warrant. Startup, transport, and timeout symptoms require evidence-based attribution: they may indicate a Paperclip/Runner product bug, provider behavior, or infrastructure. Preserve the observed failure and cause separately, keep the existing machine grade/classifier unchanged, and retain partial attempts, retries, source SHA, catalog/definition digest, model/profile, environment, grader version, timing, tokens, runtime estimates, and cost coverage. Do not claim qualification from a partial/manual selection.

Update README.md, FIXTURES.md, SECURITY.md, or EVERYDAY-WORKFLOWS.md when their authoritative contract changes, and link from doc/evals.md. Do not recreate Evalbook HTML or publish raw trusted traces, videos, archives, databases, workspaces, credentials, SVG, provider session IDs, or hidden reasoning. Public sanitized fixture conversation, marked screenshots, and allowlisted structured evidence are expected when the existing publisher permits them.

© paperclipai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/add-product-e2e-eval of paperclipai/paperclip.

Open the folder on GitHubat commit b9750b1

Compare with similar skills

Add Product E2E Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Product E2E Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Product E2E Eval this skillpaperclipai/paperclip99k—~1.1kAutomated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
Uloop Replay Inputkurotu/VRCQuestTools3733 repos~615Automated safety check: PassMIT
Ui4 Convert Testspayloadcms/payload45k—~3.5kAutomated safety check: PassMIT
E2Estackia/rtp2httpd2.2k—~517Automated safety check: PassGPL-2.0

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • Uloop Replay Input

    kurotu/VRCQuestTools

    Replay recorded PlayMode keyboard and mouse input. An agent skill from kurotu/VRCQuestTools.

    373 GitHub starsUsed in 3 repos~615 tokens
    Testing & QAAuto-check passed
  • Ui4 Convert Tests

    payloadcms/payload

    A skill your agent uses when UI changes are complete and e2e tests need updating.

    45k GitHub stars~3.5k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • E2E

    stackia/rtp2httpd

    Write, run, review, or debug rtp2httpd E2E tests and their harness in e2e/ and scripts/run-e2e.sh.

    2.2k GitHub stars~517 tokensUpdated 7 days ago
    Testing & QAAuto-check passed
  • Moav E2E

    MotherofallVPNs/MoaV

    Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.

    448 GitHub stars~1.9k tokensUpdated 2 days ago
    Testing & QAAuto-check: notes

More from paperclipai/paperclip

All 60 skills in this repo
  • Garden Inbox

    paperclipai/paperclip

    Scan a Paperclip user's Mine inbox, classify reversible archive candidates, request checkbox confirmation, and archive only accepted selections.

    99k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Paperclip

    paperclipai/paperclip

    Interact with the Paperclip control plane API for task coordination and governance.

    99k GitHub stars~9.6k tokensUpdated today
    Auto-check passed
  • Paperclip

    paperclipai/paperclip

    A skill your agent uses for Paperclip-managed tasks and heartbeats: reading task context, delivering task documents or files, updating completion or blockers, coordinating or delegating work, and…

    99k GitHub stars~17k tokensUpdated today
    Auto-check passed
  • Design Guide

    paperclipai/paperclip

    Paperclip UI design system guide for building consistent, reusable frontend components.

    99k GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Paperclip Page

    paperclipai/paperclip

    Publish static HTML pages and asset folders to the Paperclip S3/CloudFront page host.

    99k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Paperclip Create Agent

    paperclipai/paperclip

    Create new agents in Paperclip with governance-aware hiring.

    99k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed

Categories

Questions about Add Product E2E Eval

What does Add Product E2E Eval do?

Add or extend a Paperclip full-stack runner E2E workflow, fixture, matcher, or report evidence path for local or Daytona execution. Add Product E2E Eval is an agent skill from paperclipai/paperclip. Add or extend a Paperclip full-stack runner E2E workflow, fixture, matcher, or report evidence path for local or Daytona execution.

When should I use Add Product E2E Eval?

Add Product E2E Eval fits situations like: tasks that involve End-to-end testing.

How do I install Add Product E2E Eval in Claude Code?

Run `npx skills add paperclipai/paperclip --skill add-product-e2e-eval -a claude-code`. Or copy the skill folder (.agents/skills/add-product-e2e-eval in paperclipai/paperclip) into .claude/skills/add-product-e2e-eval in your project. Claude Code loads it when a task matches its description.

How do I install Add Product E2E Eval in Codex?

Run `npx skills add paperclipai/paperclip --skill add-product-e2e-eval -a codex`. Or copy the skill folder (.agents/skills/add-product-e2e-eval in paperclipai/paperclip) into .agents/skills/add-product-e2e-eval in your project. Codex loads it when a task matches its description.

Can I use Add Product E2E Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add paperclipai/paperclip --skill add-product-e2e-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-product-e2e-eval, .gemini/skills/add-product-e2e-eval, .github/skills/add-product-e2e-eval and .opencode/skills/add-product-e2e-eval in your project.

What does Add Product E2E Eval need to run?

Going by SKILL.md and its folder, Add Product E2E Eval needs the command-line tools its instructions call (pnpm and git).

Does Add Product E2E Eval access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Add Product E2E Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Product E2E Eval use?

Add Product E2E Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Product E2E Eval use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Product E2E Eval?

Skills that share tags, products or a category with Add Product E2E Eval: Web Application Testing (anthropics/skills, 180k stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Uloop Replay Input (kurotu/VRCQuestTools, 373 stars) and Ui4 Convert Tests (payloadcms/payload, 45k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Product E2E Eval?

paperclipai (a GitHub organization) maintains it in paperclipai/paperclip, which has 98,967 GitHub stars. The repository holds 60 skills in this directory. The repository was last updated on October 9, 2026.

Source: paperclipai/paperclip on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.