Acceptance Evidence for Deliveries
lobehub/lobehub
Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.
Runs OpenWork's agent-first test specs one at a time, locally or on Daytona, and records exact results, skips and runtime mismatches instead of reporting a loose pass.
$ npx skills add different-ai/openwork --skill run-tests -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install different-ai/openwork run-tests --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/different-ai/openwork.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.opencode/skills/run-tests .claude/skills/run-tests && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "run-tests" agent skill from https://github.com/different-ai/openwork/tree/dev/.opencode/skills/run-tests into .claude/skills/run-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-tests", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/different-ai/openwork/tree/dev/.opencode/skills/run-testsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add different-ai/openwork --skill run-tests -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install different-ai/openwork run-tests --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/different-ai/openwork.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.opencode/skills/run-tests .agents/skills/run-tests && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "run-tests" agent skill from https://github.com/different-ai/openwork/tree/dev/.opencode/skills/run-tests into .agents/skills/run-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-tests", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add different-ai/openwork --skill run-tests -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install different-ai/openwork run-tests --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/different-ai/openwork.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.opencode/skills/run-tests .cursor/skills/run-tests && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "run-tests" agent skill from https://github.com/different-ai/openwork/tree/dev/.opencode/skills/run-tests into .cursor/skills/run-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-tests", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/different-ai/openwork.git --path .opencode/skills/run-tests--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add different-ai/openwork --skill run-tests -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install different-ai/openwork run-tests --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/different-ai/openwork.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.opencode/skills/run-tests .gemini/skills/run-tests && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "run-tests" agent skill from https://github.com/different-ai/openwork/tree/dev/.opencode/skills/run-tests into .gemini/skills/run-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-tests", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install different-ai/openwork run-testsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add different-ai/openwork --skill run-tests -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/different-ai/openwork.git skills-src && mkdir -p .github/skills && cp -r skills-src/.opencode/skills/run-tests .github/skills/run-tests && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "run-tests" agent skill from https://github.com/different-ai/openwork/tree/dev/.opencode/skills/run-tests into .github/skills/run-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-tests", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add different-ai/openwork --skill run-tests -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install different-ai/openwork run-tests --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/different-ai/openwork.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.opencode/skills/run-tests .opencode/skills/run-tests && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "run-tests" agent skill from https://github.com/different-ai/openwork/tree/dev/.opencode/skills/run-tests into .opencode/skills/run-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-tests", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
run-testsRuns OpenWork's agent-first test specs one at a time, locally or on Daytona, and records exact results, skips and runtime mismatches instead of reporting a loose pass.
The skill sets the rules for running OpenWork testkit specs. Tests run against the exact pull request head that will land, with any earlier verdict thrown out after a rebase or cherry-pick, and one test at a time so each failure has a single owner. The pnpm evals:pr command runs one app-less spec and pnpm evals:e2e runs an app or Den driving E2E test. The CLI picks and prints where it runs, Daytona or local, and that line goes into the report. Local is used only when you ask for it, and a red Daytona run is never rerun in another lane to turn it green.
It also describes the core journey test every PR runs, which needs the evals install and a Freestyle API key, the local fallback that builds workspace packages first and needs MySQL at 127.0.0.1:3306, and a workaround for checkout paths containing spaces. The server runs on Bun in evals while Desktop runs the same code on Electron's Node, so changes touching fetch, streams, signals, GC or timers should run on the shipping runtime too. Results record the command, exit code and passed, failed and skipped counts, with each skip reported as skipped along with what it needs.
Read from SKILL.md and the folder at commit e7d03dd. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pnpmnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
FREESTYLE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
OpenWork Test Runner loads about 764 tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 329 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 329 words (~764 tokens).
“discard the old verdict and run again.”
Just SKILL.md in .opencode/skills/run-tests of different-ai/openwork.
Open the folder on GitHubat commit e7d03dd
OpenWork Test Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| OpenWork Test Runner this skilldifferent-ai/openwork | 24k | — | ~764 | Automated safety check: Pass | Custom licence | |
| Acceptance Evidence for Deliverieslobehub/lobehub | 83k | — | ~9.7k | Automated safety check: Pass | Apache-2.0 | |
| Airbyte source-mysql E2E Testsairbytehq/airbyte | 22k | — | ~2.6k | Automated safety check: Pass | Custom licence | |
| E2Egronxb/hot-updater | 1.8k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Local CImodule-federation/core | 2.7k | — | ~914 | Automated safety check: Pass | MIT | |
| Verdaccio Change Testing Selectorverdaccio/verdaccio | 18k | — | ~1.6k | Automated safety check: Warn | MIT |
lobehub/lobehub
Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.
airbytehq/airbyte
Stands up a throwaway local MySQL 8.0 backend, applies SQL fixtures and sweeps the Airbyte spec, check, discover and read commands against a source-mysql image.
gronxb/hot-updater
Run end-to-end OTA verification for examples/v0.85.0 with agent-device.
module-federation/core
Run this repository's local CI parity commands and pnpm run ci:local jobs.
verdaccio/verdaccio
Figures out which rebuild and test suites actually cover a change in the verdaccio monorepo, instead of a scoped run that passes untested.
Chorus-AIDLC/Chorus
A skill your agent uses when manually verifying a Chorus frontend change in a real browser — finding local login credentials, driving the running dev server with the Playwright MCP, logging in…
different-ai/openwork
Drives a running OpenWork desktop window over CDP from the shell to evaluate JS, take screenshots, start sessions and send prompts for hand checks.
different-ai/openwork
Makes the desktop app's model provider fail on demand, with refused connections, resets, stalls and HTTP 4xx and 5xx errors, so error and retry states can be reproduced.
different-ai/openwork
Manages OpenWork's inference model aliases, discounts and overlays over the upstream OpenRouter catalog, and triggers the GitHub workflow that refreshes base models.
different-ai/openwork
Fires known chat states in the running OpenWork desktop app, such as provider errors, retries and tool steps, so you can check how each renders.
different-ai/openwork
Attaches OpenCode browser tools to the OpenWork Electron dev app through CDP to explore its UI, send a composer task and debug, not to give test verdicts.
different-ai/openwork
Covers Daytona CLI setup, sandbox debugging, keeping a sandbox alive and which credentials the CLI uses, for when Daytona itself is the problem rather than the tests.
Categories
Runs OpenWork's agent-first test specs one at a time, locally or on Daytona, and records exact results, skips and runtime mismatches instead of reporting a loose pass. The skill sets the rules for running OpenWork testkit specs. Tests run against the exact pull request head that will land, with any earlier verdict thrown out after a rebase or cherry-pick, and one test at a time so each failure has a single owner.
OpenWork Test Runner fits situations like: running a single OpenWork spec before a pull request lands; running end-to-end tests locally or on Daytona; investigating why a spec was skipped; checking a change on the runtime the desktop app actually ships with.
Run `npx skills add different-ai/openwork --skill run-tests -a claude-code`. Or copy the skill folder (.opencode/skills/run-tests in different-ai/openwork) into .claude/skills/run-tests in your project. Claude Code loads it when a task matches its description.
Run `npx skills add different-ai/openwork --skill run-tests -a codex`. Or copy the skill folder (.opencode/skills/run-tests in different-ai/openwork) into .agents/skills/run-tests in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add different-ai/openwork --skill run-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-tests, .gemini/skills/run-tests, .github/skills/run-tests and .opencode/skills/run-tests in your project.
Going by SKILL.md and its folder, OpenWork Test Runner needs the command-line tools its instructions call (pnpm and node) and credentials named FREESTYLE_API_KEY. Our summary lists: pnpm and the OpenWork evals install; A Freestyle API key for the core journey test; MySQL at 127.0.0.1:3306 for the local fallback; Daytona access for the Daytona lane.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
OpenWork Test Runner has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.
About 764 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with OpenWork Test Runner: Acceptance Evidence for Deliveries (lobehub/lobehub, 83k stars), Airbyte source-mysql E2E Tests (airbytehq/airbyte, 22k stars), E2E (gronxb/hot-updater, 1.8k stars) and Local CI (module-federation/core, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
different-ai (a GitHub organization) maintains it in different-ai/openwork, which has 23,920 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 7, 2026.
Source: different-ai/openwork on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.