Agent skill

Test Runner

by Prismer-AI in Prismer-AI/PrismerCloud

For a code change, pick the right test tier by change surface, run apc test with baseline diff, and report only NEW reds vs baseline back to the task's acceptance criterion.

MITAuto-check: notes

Install Test Runner

skills CLI
$ npx skills add Prismer-AI/PrismerCloud --skill test-runner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Prismer-AI/PrismerCloud test-runner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Prismer-AI/PrismerCloud.git skills-src && mkdir -p .claude/skills && cp -r skills-src/sdk/apc/skills/test-runner .claude/skills/test-runner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-runner
GitHub stars
1.6k
Token cost
~1.6k tokens
SKILL.md length
476 words
Files
2
Skills in repo
88
Repo updated
First seen
Licence
MIT

At a glance

For a code change, pick the right test tier by change surface, run apc test with baseline diff, and report only NEW reds vs baseline back to the task's acceptance criterion.

  • Works in 4 steps: 先体检(env doctor,取上下文——不是是否跑的最终判据) → 选层跑(apc test) → 解析 TierResult + 判据 → …
  • SKILL.md covers 工具契约(签名以此为准,先核后用), 选层规则(按改动面), Procedure and 输出契约(机器判据按这个复算,别自由发挥格式), plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Test Runner is an agent skill from Prismer-AI/PrismerCloud. For a code change, pick the right test tier by change surface, run apc test with baseline diff, and report only NEW reds vs baseline back to the task's acceptance criterion. Separates envblocked from SUT red so a broken machine never fails the code.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill.json`). Compatibility notes: ["claude-code"]

The licence is MIT.

Example prompts

  • “/test-runner”

Requirements

  • Compatibility (from SKILL.md): ["claude-code"]
  • Pre-approved tools (allowed-tools): Bash

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 先体检(env doctor,取上下文——不是是否跑的最终判据)
  2. 选层跑(apc test)
  3. 解析 TierResult + 判据
  4. 报回 task

What it can do on your machine

Read from SKILL.md and the folder at commit e5d9444. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    ["claude-code"]

    From compatibility in the SKILL.md frontmatter.

Context cost

Test Runner loads about 1.6k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 476 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Prismer-AI/PrismerCloud at commit e5d9444, republished under its MIT licence (© Prismer-AI). 476 words, ~1,570 tokens.

Download SKILL.mdSave it as .claude/skills/test-runner/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
test-runner
description
For a code change, pick the right test tier by change surface, run apc test with baseline diff, and report only NEW reds vs baseline back to the task's acceptance criterion. Separates env_blocked from SUT red so a broken machine never fails the code.
allowed-tools
Bash
compatibility
["claude-code"]
license
MIT
scope
coding
metadata.category
testing

test-runner

对一个改动选层跑测试并把**新增红(vs baseline)**报回 task 的验收 criterion(apc/02 §2 R1 · apc/05 A1 S7)。核心纪律:环境红(env_blocked)与被测代码红(SUT red)分域——机器坏了绝不判代码红。

什么时候用:一个 coding task 改完进 review,需要按改动面选层跑回归、把结果作为 acceptance criterion 的 pass/fail 证据。

工具契约(签名以此为准,先核后用)

命令作用退出码
apc env doctor环境体检(见 env-doctor skill)0 全绿 · 78 env_blocked
apc test [--tier=T0,T1] [--diff] [--json] [--list]全层测试编排(包装 scripts/test203/run.ts)0 绿 · 1 SUT 红/回归 · 2 用法错 · 78 env_blocked
cloud task verify-criterion <task-id> <criterion-id> --outcome <passed|failed|n/a|waived>把一个 criterion 的判定报回 task0 成功
  • apc test 的退出码原样透传,含 78。把 78 折成 1 会让下游去重试一个根本没跑的圈次——绝不这么干。
  • --json 下 stdout 是结构化 test203.run/v1 产物,含每层 TierResult(命令 / 退出码 / failed names / envStatus)+ baseline diff。从 stdout JSON 取结果,人读报告在 stderr。

选层规则(按改动面)

改动面tier
src/lib/**(纯库 / 单元)T0
src/im/**、endpoint / IM 域T1
src/app/** 组件T2
e2e / 跨端(需 cloud:3000)T3

跨 src/lib + endpoint 的改动 → --tier=T0,T1。宁可多选一层,不可漏层。

Procedure

1. 先体检(env doctor,取上下文——不是是否跑的最终判据)
bash
apc env doctor > /tmp/apc-doctor.json; DOC=$?
echo "doctor exit=$DOC"

apc env doctor 是全局体检:它 exit 78 只表示"存在某个环境红",不等于你要跑的 tier 被挡。跑不跑的权威判据是 apc test 自己的退出码——run.ts 按 tier 做逐层 env 门(TIER_ENV_REQUIRES):

tier需要的环境
T0 / T1 / T2 / TD / TA无(纯逻辑,不被任何 infra 红挡)
T3cloud(:3000)
T4mysql + redis + cloud

所以:doctor 红在你选的 tier 的 env 需求之外(典型 toolchain.node pin 漂移之于 T0/T1)→ 记一条故障域备注即可,照跑。别把全局 78 当作"停"——那会因为一条无关的 toolchain 红漏掉整轮回归。真正的 env_blocked 由第 2 步 apc test 的 exit 78 表达(run.ts 只在你选的 tier 的 infra 缺失时才给 78)。

2. 选层跑(apc test)

按上表选层。跨 src/lib + endpoint 的改动:

bash
apc test --tier=T0,T1 --diff --json > /tmp/apc-test.json; T=$?
echo "apc test exit=$T"
3. 解析 TierResult + 判据

从 /tmp/apc-test.json 读结构化产物,真实 schema(字段名以此为准):

  • 顶层:{ schema, doctor, envStatus, tiers:[...], regressions:[...], fixed:[...], exitCode }。
    • envStatus:ok | env_blocked(全局)。
    • regressions[]:新增红 vs baseline(跨所有层的并集)——这是 --diff 判据的核心。
    • fixed[]:本圈由红转绿的用例。
    • exitCode:整轮退出码(与命令退出码一致)。
  • 每个 tiers[] 元素:{ tier, passed, failed, skipped, total, failedNames:[...], envStatus, regressions:[...] }(注意是 failedNames,且 command/exitCode 不在每层——退出码看顶层)。

按 apc test 退出码定 criterion outcome:

exit含义criterion outcome
0全绿,无新增红passed
1有 SUT 红(--diff 下 = 新增红/回归)failed(附 failed names + baseline diff)
78env_blocked(某层环境未就绪)不判 failed——回第 1 步,报环境故障域,不计 SUT 红
2用法错(tier 名非法等)修命令重跑,不报 criterion

只把"新增红"当回归:baseline 里已知的红不是本次回归(--diff 的退出码只看新增红)。

4. 报回 task

拿到本 task 的 criterionId(从 acceptance view),按第 3 步结论上报:

bash
# 绿:
cloud task verify-criterion "$PRISMER_TASK_ID" "<criterion-id>" --outcome passed \
  --note "apc test T0,T1 green; no new reds vs baseline"

# 红(附证据):
cloud task verify-criterion "$PRISMER_TASK_ID" "<criterion-id>" --outcome failed \
  --note "T1 new reds: <failed-name-1>,<failed-name-2>"

verify-criterion 也支持 --run 让服务端下发的 criterion.execution 就地执行后按退出码上报——但那条路的命令串来自服务端,本 skill 走的是"本地 apc test 选层跑 + 手动上报"这条主路。

Show full SKILL.md (187 more words)Show less

输出契约(机器判据按这个复算,别自由发挥格式)

本 skill 的判据不是「报告里出现了 --diff / 78 这些字」,而是判据自己重解析你贴的产物、重算每层算术、并从真 scripts/test203/baseline.json 重推 regressions[](structured-criteria.ts 的 json-claim)。所以报告必须带下面两行 + 一段原样的 JSON:

RUN-EXIT: <apc test 的原始退出码:0 | 1 | 78>
VERDICT: <exit 0 → passed | exit 1 → failed | exit 78 → env-fault>

紧跟着把 --json 的 stdout 一字不改贴进 fenced json 块:

```json
{ "schema": "test203.run/v1", "timestamp": "…", "envStatus": "…", "tiers": [ … ], "regressions": [ … ], "exitCode": 0 }
```

判据会判红的情况(任一):

  • 报告里没有可解析的 fenced JSON(只有散文「跑绿了」);
  • RUN-EXIT: 与产物里的 exitCode 不一致——把红叙述成绿正是这条要抓的;
  • VERDICT: 不是从退出码推出来的那个:78 恒 env-fault,永远不是 failed(环境故障域不计 SUT 红,这是本 skill 的承重纪律,判据里是硬映射不是措辞);
  • 退出码契约被破坏:exit 78 而 envStatus != env_blocked、exit 1 而 regressions 空、exit 0 而 regressions 非空;
  • 某层的 passed+failed+skipped != total,或 failedNames.length != failed——编的计数会露馅;
  • 本 task 钉的 tier 不在 tiers[] 里(跑了别的层顶包);
  • regressions[] 与「用真 baseline.json 重算的新增红集合」不等:把新增红藏进已知红、或把已知红报成回归,两个方向都红;
  • 该层的运行副产物文件不存在,或它的 mtime 与产物自带的 timestamp 相差超过 skew——「我跑了」但磁盘上没有这一刻的痕迹,判红。

诚实边界:freshArtifacts / requireItems 的取值是按本 sample task 钉的(TD 层 + desktop-shell-skip-evidence.txt),换 tier 必须同步换 args;判据也不能证明产物里的数字来自真 vitest——一个肯读 baseline、肯 touch 副产物的伪造者仍能构造自洽产物。它杀的是廉价伪造(编计数、叙述判定、无痕迹)。

产出(副作用 oracle,报告里必须给)

  1. 结构化 TierResult:顶层 {envStatus, exitCode} + 每层 {tier, passed/failed/skipped/total, failedNames, envStatus}(取自 apc test --json stdout)。
  2. 新增红 vs baseline:顶层 regressions[](区分"回归"与"从未通过"——baseline 已知红不进这个数组)。
  3. criterion 上报:cloud task verify-criterion 落库的 outcome 行(可回读 acceptance view 确认)。

不许:把 env_blocked(exit 78)计成 SUT 红去判 criterion failed;把 baseline 已知红当新增红上报;断言聊天文本而非 TierResult 结构字段。

<!-- APC-ACK:v1 -->

调用回执(APC 平台方运营纪律 · 自动注入,勿手改)

执行本 skill 的第一步,先落一条调用回执:

bash
cloud skill ack test-runner --task "$PRISMER_TASK_ID"

按退出码分流(这条命令的退出码是承重信息,禁止用 || 兜底、; true、 set +e 或重定向把它抹掉):

exit含义你要做的
0回执已落库(im_task_logs.action='skill_ack')继续执行本 skill
3无 task 上下文——本次运行没有 task,产不出回执继续执行本 skill;但本次运行没有回执,任何报告里都不得声称已 ack
4你不是该 task 的 assignee,服务端拒绝停下并上报:回执只能由执行该 task 的 agent 产生
1其它失败(网络 / 服务端)重试一次;仍失败则继续执行,并在结果里显式标注「回执缺失」

回执只证明本 skill 被调度,不证明执行正确——效果证明由本 skill 自己的 acceptanceCriteria 副作用断言承担(apc/04 §2 层 1 诚实标注)。

<!-- /APC-ACK:v1 -->

© Prismer-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in sdk/apc/skills/test-runner of Prismer-AI/PrismerCloud.

  • SKILL.md
  • skill.json

Open the folder on GitHubat commit e5d9444

Compare with similar skills

Test Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Runner this skillPrismer-AI/PrismerCloud1.6k—~1.6kAutomated safety check: NotesMIT
Workspace Surface Auditaffaan-m/ECC275k3 repos~1.3kAutomated safety check: NotesMIT
Flutter Cherry Pickflutter/flutter179k—~1.8kAutomated safety check: PassBSD-3-Clause
Agent Test Long Runnerruvnet/ruflo74k2 repos~426Automated safety check: PassMIT
Depot GitHub RunnersPostHog/posthog40k—~2.8kAutomated safety check: PassCustom licence
Add Runner Evalpaperclipai/paperclip99k—~1kAutomated safety check: PassMIT

Similar skills

  • Audit the active repo, MCP servers, plugins, connectors, env surfaces, and harness setup, then recommend the highest-value ECC-native skills, hooks, agents, and operator workflows.

    275k GitHub starsUsed in 3 repos~1.3k tokens
    Agent WorkflowsAuto-check: notes
  • Flutter Cherry Pick

    flutter/flutter

    How to land a formal cherry-pick of a merged PR for the flutter/flutter repo stable or beta channel.

    179k GitHub stars~1.8k tokensUpdated today
    MobileAuto-check passed
  • Agent skill for test-long-runner - invoke with $agent-test-long-runner

    74k GitHub starsUsed in 2 repos~426 tokens
    Agent WorkflowsAuto-check passed
  • Depot GitHub Runners

    PostHog/posthog

    Official

    Configures Depot-managed GitHub Actions runners as a drop-in replacement for GitHub-hosted runners.

    40k GitHub stars~2.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Add Runner Eval

    paperclipai/paperclip

    Add or extend a Paperclip Runner protocol evaluation definition, roster, assertion, or report fixture with provenance and narrow validation.

    99k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Surface

    srid/emanote

    How a downstream app consumes the shared @kolu/surface stack (@kolu/surface · surface-app · surface-nix-host · surface-mcp) — declaring a typed reactive surface, serving it, consuming it (SolidJS…

    963 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed

More from Prismer-AI/PrismerCloud

All 88 skills in this repo
  • Himalaya Email CLI

    Prismer-AI/PrismerCloud

    Operates a mailbox from the terminal with the external Himalaya CLI over IMAP, SMTP, Notmuch or Sendmail, separate from any built-in email gateway adapter.

    1.6k GitHub starsUsed in 3 repos~2.3k tokens
    Auto-check passed
  • Prismer Skill Creator

    Prismer-AI/PrismerCloud

    Walks an agent through creating, importing, editing, validating, testing and publishing Prismer Skills with a fixed workflow and bundled scripts.

    1.6k GitHub stars~2.6k tokensUpdated 8 days ago
    Auto-check: notes
  • Manim Explainer Videos

    Prismer-AI/PrismerCloud

    Produces 3Blue1Brown-style explainer animations with Manim Community Edition for math, algorithms, equations and architecture diagrams, with planning and rendering references.

    1.6k GitHub starsUsed in 2 repos~3.1k tokens
    Auto-check passed
  • Prismer Image Generation

    Prismer-AI/PrismerCloud

    Generates one image from a text prompt with a bundled Node.js helper and delivers it once as the attachment to the current Prismer reply.

    1.6k GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed
  • YouTube Transcript Reformatter

    Prismer-AI/PrismerCloud

    Fetches a YouTube transcript with a helper script and reshapes it into chapters, summaries, X threads, blog posts or timestamped quotes.

    1.6k GitHub starsUsed in 2 repos~905 tokens
    Auto-check passed
  • Prismer Role Builder

    Prismer-AI/PrismerCloud

    Creates or updates Prismer role templates from a persona, SOP or job description, and turns a role into a working agent that runs its first task through a bundled script.

    1.6k GitHub stars~2.3k tokensUpdated 8 days ago
    Auto-check: notes

Questions about Test Runner

What does Test Runner do?

For a code change, pick the right test tier by change surface, run apc test with baseline diff, and report only NEW reds vs baseline back to the task's acceptance criterion. Test Runner is an agent skill from Prismer-AI/PrismerCloud. For a code change, pick the right test tier by change surface, run apc test with baseline diff, and report only NEW reds vs baseline back to the task's acceptance criterion.

How do I install Test Runner in Claude Code?

Run `npx skills add Prismer-AI/PrismerCloud --skill test-runner -a claude-code`. Or copy the skill folder (sdk/apc/skills/test-runner in Prismer-AI/PrismerCloud) into .claude/skills/test-runner in your project. Claude Code loads it when a task matches its description.

How do I install Test Runner in Codex?

Run `npx skills add Prismer-AI/PrismerCloud --skill test-runner -a codex`. Or copy the skill folder (sdk/apc/skills/test-runner in Prismer-AI/PrismerCloud) into .agents/skills/test-runner in your project. Codex loads it when a task matches its description.

Can I use Test Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Prismer-AI/PrismerCloud --skill test-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-runner, .gemini/skills/test-runner, .github/skills/test-runner and .opencode/skills/test-runner in your project.

What does Test Runner need to run?

SKILL.md names no scripts, command-line tools or credentials: Test Runner is instructions for the agent only. Its frontmatter pre-approves these tools: Bash. Compatibility (from SKILL.md): ["claude-code"].

Does Test Runner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Runner safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Test Runner use?

Test Runner is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Runner use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Runner?

Skills that share tags, products or a category with Test Runner: Workspace Surface Audit (affaan-m/ECC, 275k stars), Flutter Cherry Pick (flutter/flutter, 179k stars), Agent Test Long Runner (ruvnet/ruflo, 74k stars) and Depot GitHub Runners (PostHog/posthog, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Runner?

Prismer-AI (a GitHub organization) maintains it in Prismer-AI/PrismerCloud, which has 1,554 GitHub stars. The repository holds 88 skills in this directory. The repository was last updated on September 30, 2026.

Source: Prismer-AI/PrismerCloud on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.