Agent skill

OpenWork Test Runner

by different-ai in different-ai/openwork

Runs OpenWork's agent-first test specs one at a time, locally or on Daytona, and records exact results, skips and runtime mismatches instead of reporting a loose pass.

Custom licenceAuto-check passedTesting & QA

Install OpenWork Test Runner

skills CLI
$ npx skills add different-ai/openwork --skill run-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install different-ai/openwork run-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/different-ai/openwork.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.opencode/skills/run-tests .claude/skills/run-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-tests
GitHub stars
24k
Token cost
~764 tokens
SKILL.md length
329 words
Files
1
Skills in repo
32
Repo updated
First seen
Licence
Custom licence

At a glance

Runs OpenWork's agent-first test specs one at a time, locally or on Daytona, and records exact results, skips and runtime mismatches instead of reporting a loose pass.

  • Running a single OpenWork spec before a pull request lands
  • SKILL.md covers Run the landable tree, Choose the execution environment, Run the core journey and Prepare local fallback, plus 4 more sections
  • Calls pnpm and node; needs FREESTYLE_API_KEY
  • Running end-to-end tests locally or on Daytona

What it does

The skill sets the rules for running OpenWork testkit specs. Tests run against the exact pull request head that will land, with any earlier verdict thrown out after a rebase or cherry-pick, and one test at a time so each failure has a single owner. The pnpm evals:pr command runs one app-less spec and pnpm evals:e2e runs an app or Den driving E2E test. The CLI picks and prints where it runs, Daytona or local, and that line goes into the report. Local is used only when you ask for it, and a red Daytona run is never rerun in another lane to turn it green.

It also describes the core journey test every PR runs, which needs the evals install and a Freestyle API key, the local fallback that builds workspace packages first and needs MySQL at 127.0.0.1:3306, and a workaround for checkout paths containing spaces. The server runs on Bun in evals while Desktop runs the same code on Electron's Node, so changes touching fetch, streams, signals, GC or timers should run on the shipping runtime too. Results record the command, exit code and passed, failed and skipped counts, with each skip reported as skipped along with what it needs.

When your agent uses it

  • Running a single OpenWork spec before a pull request lands
  • Running end-to-end tests locally or on Daytona
  • Investigating why a spec was skipped
  • Checking a change on the runtime the desktop app actually ships with

Example prompts

  • “Run the core chat E2E spec against the pushed head of this PR and report the exit code and counts.”
  • “Run specs/core-chat.e2e.test.ts locally, since I want it on my machine and not on Daytona.”
  • “One spec shows as skipped in the last run, so find out what it needs before we call it done.”

Requirements

  • pnpm and the OpenWork evals install
  • A Freestyle API key for the core journey test
  • MySQL at 127.0.0.1:3306 for the local fallback
  • Daytona access for the Daytona lane

What it can do on your machine

Read from SKILL.md and the folder at commit e7d03dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FREESTYLE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OpenWork Test Runner loads about 764 tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 329 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~764

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 329 words (~764 tokens).

“discard the old verdict and run again.”

— opening of SKILL.md by different-ai, Custom licence
name
run-tests

Read the full SKILL.md on GitHub

Files

Just SKILL.md in .opencode/skills/run-tests of different-ai/openwork.

Open the folder on GitHubat commit e7d03dd

Compare with similar skills

OpenWork Test Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OpenWork Test Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OpenWork Test Runner this skilldifferent-ai/openwork24k—~764Automated safety check: PassCustom licence
Acceptance Evidence for Deliverieslobehub/lobehub83k—~9.7kAutomated safety check: PassApache-2.0
Airbyte source-mysql E2E Testsairbytehq/airbyte22k—~2.6kAutomated safety check: PassCustom licence
E2Egronxb/hot-updater1.8k—~1.6kAutomated safety check: PassCustom licence
Local CImodule-federation/core2.7k—~914Automated safety check: PassMIT
Verdaccio Change Testing Selectorverdaccio/verdaccio18k—~1.6kAutomated safety check: WarnMIT

Similar skills

  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Stands up a throwaway local MySQL 8.0 backend, applies SQL fixtures and sweeps the Airbyte spec, check, discover and read commands against a source-mysql image.

    22k GitHub stars~2.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • E2E

    gronxb/hot-updater

    Run end-to-end OTA verification for examples/v0.85.0 with agent-device.

    1.8k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Local CI

    module-federation/core

    Run this repository's local CI parity commands and pnpm run ci:local jobs.

    2.7k GitHub stars~914 tokensUpdated today
    Testing & QAAuto-check passed
  • Figures out which rebuild and test suites actually cover a change in the verdaccio monorepo, instead of a scoped run that passes untested.

    18k GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check: warnings
  • E2E Verification

    Chorus-AIDLC/Chorus

    A skill your agent uses when manually verifying a Chorus frontend change in a real browser — finding local login credentials, driving the running dev server with the Playwright MCP, logging in…

    1.2k GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check: notes

More from different-ai/openwork

All 32 skills in this repo
  • OpenWork Desktop CDP Driver

    different-ai/openwork

    Drives a running OpenWork desktop window over CDP from the shell to evaluate JS, take screenshots, start sessions and send prompts for hand checks.

    24k GitHub stars~465 tokensUpdated today
    Auto-check passed
  • Fake Model Provider Faults

    different-ai/openwork

    Makes the desktop app's model provider fail on demand, with refused connections, resets, stalls and HTTP 4xx and 5xx errors, so error and retry states can be reproduced.

    24k GitHub stars~642 tokensUpdated today
    Auto-check passed
  • OpenWork Model Alias Manager

    different-ai/openwork

    Manages OpenWork's inference model aliases, discounts and overlays over the upstream OpenRouter catalog, and triggers the GitHub workflow that refreshes base models.

    24k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Reproduce Chat States

    different-ai/openwork

    Fires known chat states in the running OpenWork desktop app, such as provider errors, retries and tool steps, so you can check how each renders.

    24k GitHub stars~673 tokensUpdated today
    Auto-check passed
  • Attaches OpenCode browser tools to the OpenWork Electron dev app through CDP to explore its UI, send a composer task and debug, not to give test verdicts.

    24k GitHub stars~780 tokensUpdated today
    Auto-check passed
  • Daytona Sandbox Operations

    different-ai/openwork

    Covers Daytona CLI setup, sandbox debugging, keeping a sandbox alive and which credentials the CLI uses, for when Daytona itself is the problem rather than the tests.

    24k GitHub stars~917 tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about OpenWork Test Runner

What does OpenWork Test Runner do?

Runs OpenWork's agent-first test specs one at a time, locally or on Daytona, and records exact results, skips and runtime mismatches instead of reporting a loose pass. The skill sets the rules for running OpenWork testkit specs. Tests run against the exact pull request head that will land, with any earlier verdict thrown out after a rebase or cherry-pick, and one test at a time so each failure has a single owner.

When should I use OpenWork Test Runner?

OpenWork Test Runner fits situations like: running a single OpenWork spec before a pull request lands; running end-to-end tests locally or on Daytona; investigating why a spec was skipped; checking a change on the runtime the desktop app actually ships with.

How do I install OpenWork Test Runner in Claude Code?

Run `npx skills add different-ai/openwork --skill run-tests -a claude-code`. Or copy the skill folder (.opencode/skills/run-tests in different-ai/openwork) into .claude/skills/run-tests in your project. Claude Code loads it when a task matches its description.

How do I install OpenWork Test Runner in Codex?

Run `npx skills add different-ai/openwork --skill run-tests -a codex`. Or copy the skill folder (.opencode/skills/run-tests in different-ai/openwork) into .agents/skills/run-tests in your project. Codex loads it when a task matches its description.

Can I use OpenWork Test Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add different-ai/openwork --skill run-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-tests, .gemini/skills/run-tests, .github/skills/run-tests and .opencode/skills/run-tests in your project.

What does OpenWork Test Runner need to run?

Going by SKILL.md and its folder, OpenWork Test Runner needs the command-line tools its instructions call (pnpm and node) and credentials named FREESTYLE_API_KEY. Our summary lists: pnpm and the OpenWork evals install; A Freestyle API key for the core journey test; MySQL at 127.0.0.1:3306 for the local fallback; Daytona access for the Daytona lane.

Does OpenWork Test Runner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is OpenWork Test Runner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does OpenWork Test Runner use?

OpenWork Test Runner has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does OpenWork Test Runner use?

About 764 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to OpenWork Test Runner?

Skills that share tags, products or a category with OpenWork Test Runner: Acceptance Evidence for Deliveries (lobehub/lobehub, 83k stars), Airbyte source-mysql E2E Tests (airbytehq/airbyte, 22k stars), E2E (gronxb/hot-updater, 1.8k stars) and Local CI (module-federation/core, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OpenWork Test Runner?

different-ai (a GitHub organization) maintains it in different-ai/openwork, which has 23,920 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 7, 2026.

Source: different-ai/openwork on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.