Agent skill

Test Runner

by noumena-labs in noumena-labs/Sipp

Runs the narrowest relevant tests to validate changes. An agent skill from noumena-labs/Sipp.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Test Runner

skills CLI
$ npx skills add noumena-labs/Sipp --skill test-runner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install noumena-labs/Sipp test-runner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/noumena-labs/Sipp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/test-runner .claude/skills/test-runner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-runner
GitHub stars
121
Token cost
~854 tokens
SKILL.md length
261 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs the narrowest relevant tests to validate changes. An agent skill from noumena-labs/Sipp.

  • Works in 7 steps: Broad and automation checks → Rust Native Core (crates/) → Node.js Bindings And Package… → …
  • The user asks to run tests
  • SKILL.md covers Core Rule, Test Targets by Area and Pre-Test Check
  • Calls cargo

What it does

Test Runner is an agent skill from noumena-labs/Sipp. Runs the narrowest relevant tests to validate changes. Use this skill when the user asks to run tests, verify functionality, run checks, or before concluding any change in the repository.

Its SKILL.md is about 850 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires cargo, bun/pnpm, and python testing suites.

It sits in AI & LLM Engineering. It works with Rust, llama.cpp, Python and C++. The repository describes itself as: AI inference, packed simply. A blazing-fast, zero-dependency WebGPU runtime to run GGUF models directly in the browser. Features a symmetric API for seamless local execution and… The licence is Apache-2.0.

When your agent uses it

  • The user asks to run tests
  • Verify functionality
  • Before concluding any change in the repository

Example prompts

  • “/test-runner”

Requirements

  • Python 3
  • Node.js
  • Compatibility (from SKILL.md): Requires cargo, bun/pnpm, and python testing suites.
  • Pre-approved tools (allowed-tools): Bash(cargo:*), Bash(bun:*), Bash(npm:*), Bash(pnpm:*), Bash(pytest:*), Read, Edit

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Broad and automation checks
  2. Rust Native Core (crates/)
  3. Node.js Bindings And Package (bindings/node/, lib/node/)
  4. Browser Package And Demos (lib/web/, demos/)
  5. Python Bindings And Package (bindings/python/, lib/python/)
  6. Browser and holistic smoke checks
  7. Coverage and verification

What it can do on your machine

Read from SKILL.md and the folder at commit 32d5373. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(cargo:*)
    • Bash(bun:*)
    • Bash(npm:*)
    • Bash(pnpm:*)
    • Bash(pytest:*)
    • Read
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires cargo, bun/pnpm, and python testing suites.

    From compatibility in the SKILL.md frontmatter.

Context cost

Test Runner loads about 854 tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 261 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~854

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from noumena-labs/Sipp at commit 32d5373, republished under its Apache-2.0 licence (© noumena-labs). 261 words, ~854 tokens.

Download SKILL.mdSave it as .claude/skills/test-runner/SKILL.md (or your agent's skills folder).
name
test-runner
description
Runs the narrowest relevant tests to validate changes. Use this skill when the user asks to run tests, verify functionality, run checks, or before concluding any change in the repository.
allowed-tools
Bash(cargo:*), Bash(bun:*), Bash(npm:*), Bash(pnpm:*), Bash(pytest:*), Read, Edit
compatibility
Requires cargo, bun/pnpm, and python testing suites.

Test Runner Skill

You are responsible for validating changes in the repository using the appropriate testing framework.

Core Rule

Always run the narrowest relevant test target based on the files you modified. Avoid full-repo checks when a target-specific command covers the change.


Test Targets by Area

1. Broad and automation checks
  • Run every deterministic unit suite:
    bash
    cargo xtask test unit group full
  • Run all white-box unit suites:
    bash
    cargo xtask test unit group whitebox
  • Run xtask-only checks when the change is limited to developer automation:
    bash
    cargo xtask test unit suite xtask
2. Rust Native Core (crates/)
  • Run cataloged Rust unit tests for the affected crate:
    bash
    cargo xtask test unit suite rust-crates --package <crate_name>
  • Example: cargo xtask test unit suite rust-crates --package sipp
3. Node.js Bindings And Package (bindings/node/, lib/node/)
  • Run deterministic Node package API tests:
    bash
    cargo xtask test unit suite node-package --backend cpu
  • Run model-backed Node smoke when local inference behavior changed:
    bash
    cargo xtask test smoke suite example-node --backend cpu
4. Browser Package And Demos (lib/web/, demos/)
  • Run browser package TypeScript tests:
    bash
    cargo xtask test unit suite browser
  • Demo tests are cataloged separately:
    bash
    cargo xtask test unit suite demos
5. Python Bindings And Package (bindings/python/, lib/python/)
  • Run deterministic Python package API tests:
    bash
    cargo xtask test unit suite python-package --backend cpu
  • Run model-backed Python smoke when local inference behavior changed:
    bash
    cargo xtask test smoke suite example-python --backend cpu
Show full SKILL.md (156 more words)Show less
6. Browser and holistic smoke checks
  • Run browser example smoke:
    bash
    cargo xtask test smoke suite example-browser
  • Run browser playground runtime smoke:
    bash
    cargo xtask test smoke suite playground-browser
  • Run CLI, Rust, Node, and Python model-backed smoke:
    bash
    cargo xtask test smoke group local-model --backend cpu
  • Run llama.cpp backend correctness smoke:
    bash
    cargo xtask test smoke suite llama-backend-ops --backend cpu
7. Coverage and verification
  • List the catalog before choosing a target:
    bash
    cargo xtask test list --cases
  • Verify existing coverage artifacts and test structure:
    bash
    cargo xtask test verify --target whitebox
  • Validate changed source files have matching catalog-owned tests:
    bash
    cargo xtask test verify --changed

Pre-Test Check

The xtask test catalog builds required artifacts before suites that need them. Use the build-orchestrator skill first only when you are explicitly compiling or packaging a target outside the test catalog.

Use cargo xtask test list --cases to inspect available suites and discoverable cases before choosing a narrow command.

© noumena-labs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/test-runner of noumena-labs/Sipp.

Open the folder on GitHubat commit 32d5373

Compare with similar skills

Test Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Runner this skillnoumena-labs/Sipp121—~854Automated safety check: PassApache-2.0
Check Runtime Parityayutaz/piper-plus230—~972Automated safety check: PassMIT
Check Loanwordayutaz/piper-plus230—~1.1kAutomated safety check: PassMIT
Run Testsayutaz/piper-plus230—~641Automated safety check: PassMIT
Fory Version Bumpapache/fory4.6k—~1.1kAutomated safety check: PassApache-2.0
Fory Performance Optimizationapache/fory4.6k—~2.2kAutomated safety check: PassApache-2.0

Similar skills

  • Check Runtime Parity

    ayutaz/piper-plus

    推論パスの canonical Python (exportonnx.py / vits/models.py:VitsModel.infer) を変更した PR で、6 ランタイム (Python runtime / Rust / Go / C / C++ / WASM) の inference path が追随しているかを git diff で確認。PR

    230 GitHub stars~972 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Check Loanword

    ayutaz/piper-plus

    ZH-EN code-switching loanword の同期と forward-compat を 1 コマンドで検査。zhenloanword.json を編集したり 5 ランタイムのいずれかに新規エントリを追加する前後に呼ぶ。Python source を canonical とし、Rust×2 / Go / C / WASM / C++ の 6 mirror + Python…

    230 GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Run Tests

    ayutaz/piper-plus

    piper-plus の各言語ランタイムのテストを実行します。引数 python/rust/cs/go/js/cpp/all で対象を選択。未指定なら git diff から自動判定。

    230 GitHub stars~641 tokensUpdated today
    DevelopmentAuto-check passed
  • Bump Apache Fory release or post-release development versions across Java, Kotlin, Scala, Python, Rust, Go, C++, C, Dart, JavaScript, Swift, integration tests, examples, and source docs.

    4.6k GitHub stars~1.1k tokensUpdated yesterday
    MobileAuto-check passed
  • Run profile-driven bottleneck optimization across Apache Fory implementations (Java, C++, Python/Cython, Go, Rust, Swift, C, JavaScript/TypeScript, Dart, Kotlin, Scala).

    4.6k GitHub stars~2.2k tokensUpdated yesterday
    MobileAuto-check passed
  • Crossbind

    crossbind/crossbind

    A skill your agent uses when a user wants to call C++ or Rust from JavaScript or TypeScript; add a native library such as GDAL, SQLite, OpenSSL, GEOS or PROJ to a browser, Node.js, Cloudflare Worker…

    148 GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed

More from noumena-labs/Sipp

  • Style Checker

    noumena-labs/Sipp

    Enforces this monorepo's coding style rules by inspecting git diffs, reading .agents/skills/style-checker/references/styleguidance.md, fixing style violations, and reporting the result.

    121 GitHub stars~1.4k tokensUpdated 19 days ago
    Auto-check passed

Questions about Test Runner

What does Test Runner do?

Runs the narrowest relevant tests to validate changes. An agent skill from noumena-labs/Sipp. Test Runner is an agent skill from noumena-labs/Sipp. Runs the narrowest relevant tests to validate changes.

When should I use Test Runner?

Test Runner fits situations like: the user asks to run tests; verify functionality; before concluding any change in the repository.

How do I install Test Runner in Claude Code?

Run `npx skills add noumena-labs/Sipp --skill test-runner -a claude-code`. Or copy the skill folder (.agents/skills/test-runner in noumena-labs/Sipp) into .claude/skills/test-runner in your project. Claude Code loads it when a task matches its description.

How do I install Test Runner in Codex?

Run `npx skills add noumena-labs/Sipp --skill test-runner -a codex`. Or copy the skill folder (.agents/skills/test-runner in noumena-labs/Sipp) into .agents/skills/test-runner in your project. Codex loads it when a task matches its description.

Can I use Test Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add noumena-labs/Sipp --skill test-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-runner, .gemini/skills/test-runner, .github/skills/test-runner and .opencode/skills/test-runner in your project.

What does Test Runner need to run?

Going by SKILL.md and its folder, Test Runner needs the command-line tools its instructions call (cargo). Our summary lists: Python 3; Node.js. Its frontmatter pre-approves these tools: Bash(cargo:*), Bash(bun:*), Bash(npm:*), Bash(pnpm:*), Bash(pytest:*), Read, Edit. Compatibility (from SKILL.md): Requires cargo, bun/pnpm, and python testing suites..

Does Test Runner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Runner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Runner use?

Test Runner is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Runner use?

About 854 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Runner?

Skills that share tags, products or a category with Test Runner: Check Runtime Parity (ayutaz/piper-plus, 230 stars), Check Loanword (ayutaz/piper-plus, 230 stars), Run Tests (ayutaz/piper-plus, 230 stars) and Fory Version Bump (apache/fory, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Runner?

noumena-labs (a GitHub organization) maintains it in noumena-labs/Sipp, which has 121 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 19, 2026.

Source: noumena-labs/Sipp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.