Agent skill

Testing

by uw-syfi in uw-syfi/vibesys

Write, change, review, or fix tests in any language in this repository.

MITAuto-check passedTesting & QA

Install Testing

skills CLI
$ npx skills add uw-syfi/vibesys --skill testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install uw-syfi/vibesys testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/uw-syfi/vibesys.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/testing .claude/skills/testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing
GitHub stars
103
Token cost
~1.5k tokens
SKILL.md length
793 words
Files
8 (incl. references)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Write, change, review, or fix tests in any language in this repository.

  • Works in 8 steps: Test the public API only. The public API… → No patching or mocking. Do not replace… → Integrate modules with Fakes. A Fake is… → …
  • Editing any test
  • SKILL.md covers Rules and Before handing back
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Testing is an agent skill from uw-syfi/vibesys. Write, change, review, or fix tests in any language in this repository. Use when adding or editing any test, fixing a bug (regression test), tempted to patch or mock, adding or changing a Fake, writing property-based, golden-fixture, or fault-injection tests, testing code that calls agents, clusters, or subprocesses, or diagnosing a flaky test.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `agents/openai.yaml`, `references/fakes-and-contracts.md` and `references/fault-injection.md`).

It sits in Testing & QA, covering Failing and flaky tests and Chaos engineering. The repository describes itself as: Can AI Agents Build Bespoke Systems? The licence is MIT.

When your agent uses it

  • Editing any test
  • Fixing a bug (regression test)
  • Tempted to patch
  • Changing a Fake

Example prompts

  • “/testing”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Test the public API only. The public API is the surface other packages
  2. No patching or mocking. Do not replace code at runtime (monkeypatching
  3. Integrate modules with Fakes. A Fake is an in-memory, faithful
  4. Test properties, not instances. Default to property-based tests and
  5. No flaky tests. Never depend on timeouts, sleeps, wall-clock time,
  6. Bug fixes. Add a regression test at the lowest layer that reproduces the
  7. Test the unhappy paths at boundaries. Canned replies that match what
  8. Test process and signal code in process. Logic that runs in a child

What it can do on your machine

Read from SKILL.md and the folder at commit c7784eb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing loads about 1.5k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 793 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from uw-syfi/vibesys at commit c7784eb, republished under its MIT licence (© uw-syfi). 793 words, ~1,467 tokens.

Download SKILL.mdSave it as .claude/skills/testing/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
testing
description
Write, change, review, or fix tests in any language in this repository. Use when adding or editing any test, fixing a bug (regression test), tempted to patch or mock, adding or changing a Fake, writing property-based, golden-fixture, or fault-injection tests, testing code that calls agents, clusters, or subprocesses, or diagnosing a flaky test.

Testing

Write high-value tests: tests that keep passing through refactors and fail only when observable behavior changes. A test is low value if renaming a private function, splitting a module, or swapping an implementation detail breaks it.

These rules are language-independent. Tool names, commands, and syntax live in a per-language reference; read the one for the language you are editing:

Rules

  1. Test the public API only. The public API is the surface other packages or modules are already allowed to depend on: a library's api entry point, the module API that sibling modules use inside a package, and the application's facade or serialized events at the top. Do not import private modules or assert on private state. The one exception is a test whose point is a tricky implementation detail; mark that site test-isolation: <specific reason> in a comment.
  2. No patching or mocking. Do not replace code at runtime (monkeypatching attributes, module mocking, spies, mock frameworks). It couples the test to how the code is written. Setting inputs is fine: environment variables, working directory, temp directories. If a test seems to need a patch, the code is missing a seam: inject the owning library's interface through a constructor argument or parameter, and pass a Fake implementation.
  3. Integrate modules with Fakes. A Fake is an in-memory, faithful implementation of the same API, not a stub with canned answers. A Fake is never more capable or more forgiving than production: it reports the same supported capabilities (derived from the same definition), validates and loads state as strictly, and enforces the same limits. Whoever changes the real behavior updates the Fake in the same PR. The owning library ships one contract suite per interface; run it against every implementation, Fake and production, without skipping contract cases. See references/fakes-and-contracts.md.
  4. Test properties, not instances. Default to property-based tests and fuzzing for parsers, serializers, validators, pure logic, and state machines. Use single examples only for a named scenario or a regression. Use golden fixtures where behavior reduces to a deterministic state snapshot. Test functional cores with pure properties over generated event sequences, including crash and replay at every durable intent boundary. Use no async, sleeps, or Fakes in core tests. See references/properties-and-goldens.md.
  5. No flaky tests. Never depend on timeouts, sleeps, wall-clock time, scheduling order, or shared state. Inject clocks and simulate timeouts; do not wait for them. Never add retries. See references/flakiness.md.
  6. Bug fixes. Add a regression test at the lowest layer that reproduces the symptom, through the public API, that fails at the merge base. Then add a property test that generalizes the pattern, so the whole class of bug is covered and not just the reported instance. Example: for an uppercase hypothesis id that broke workspace names, test every agent-supplied id with a property test, not only the uppercase one. The class is the mechanism named under the software-design skill's Mechanism checkpoint; a test that replays the one reply or exit path that triggered the bug does not cover it.
  7. Test the unhappy paths at boundaries. Canned replies that match what the code expects test only the happy path. Where code meets an agent, a cluster, an MCP client, or a subprocess, drive it with generated behavior: agent output generated from the declared schema, including options the run does not offer, repeated mistakes, malformed output, and replies slower than the caller's deadline; and scheduled faults on the far side. For a scope that owns a resource, test as a property that an exit at any point (each step, exception, cancellation, signal, stop) releases the resource and ends in a typed terminal state. Composed-system chaos tests inject implementation failures and unknown exceptions; require a typed terminal state, never an implicit success. Test an agent tool server with synthesized tool calls and call sequences, not agents. Assert invariants, not one expected trace. See references/fault-injection.md.
  8. Test process and signal code in process. Logic that runs in a child process or a signal handler is invisible to coverage and slow to test through subprocesses. Put it behind a function the test calls directly with injected I/O interfaces; keep at most one subprocess test per entry point to prove the wiring.
Show full SKILL.md (70 more words)Show less

Before handing back

  • Run the narrowest test target first, then broaden (commands per language).
  • Run the language's test-isolation check if one exists (see its reference). It is a shrink-only ratchet: existing violations are baselined, new ones fail, and removing sites from a baselined file means lowering its recorded counts.
  • Do not migrate unrelated tests to these rules. Apply them to new tests and to the tests you are already changing.

© uw-syfi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in .agents/skills/testing of uw-syfi/vibesys.

  • SKILL.md
  • agents/openai.yaml
  • references/fakes-and-contracts.md
  • references/fault-injection.md
  • references/flakiness.md
  • references/properties-and-goldens.md
  • references/python.md
  • references/typescript.md

Open the folder on GitHubat commit c7784eb

Compare with similar skills

Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing this skilluw-syfi/vibesys103—~1.5kAutomated safety check: PassMIT
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Triage CI FailureDataDog/datadog-agent3.8k—~2.3kAutomated safety check: PassApache-2.0
Dynamo Jira TicketDynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0
Fix Ready PRsfastrepl/anarlog9.5k—~1.4kAutomated safety check: PassMIT
Trx Analysismicrosoft/vstest969—~1.8kAutomated safety check: PassMIT

Similar skills

  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Triage CI Failure

    DataDog/datadog-agent

    Official

    Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

    3.8k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Jira Ticket

    DynamoDS/Dynamo

    Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

    2k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Fix Ready PRs

    fastrepl/anarlog

    Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches.

    9.5k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Trx Analysis

    microsoft/vstest

    Official

    Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest.

    969 GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    200 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed

More from uw-syfi/vibesys

All 15 skills in this repo
  • Neuron Nki Profiling

    uw-syfi/vibesys

    This skill guides using the cli to generate NKI kernel profiles (NEFF + NTFF pairs) to analyze performance on Neuron hardware.

    103 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Neuron Nki Debugging

    uw-syfi/vibesys

    This skill guides debugging NKI compilation errors on Neuron hardware.

    103 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Neuron Nki Docs

    uw-syfi/vibesys

    Research NKI documentation for API lookups, tutorials, error codes, architecture, and optimization guides.

    103 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Query and analyze NKI kernel profile data from neuron-explorer parquet files.

    103 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Neuron Nki Writing

    uw-syfi/vibesys

    Guide for writing and modifying NKI kernels. An agent skill from uw-syfi/vibesys.

    103 GitHub stars~5k tokensUpdated today
    Auto-check passed
  • Triage PRs

    uw-syfi/vibesys

    Triage the open pull requests of the VibeSys repository. An agent skill from uw-syfi/vibesys.

    103 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Categories

Questions about Testing

What does Testing do?

Write, change, review, or fix tests in any language in this repository. Testing is an agent skill from uw-syfi/vibesys. Write, change, review, or fix tests in any language in this repository.

When should I use Testing?

Testing fits situations like: editing any test; fixing a bug (regression test); tempted to patch; changing a Fake.

How do I install Testing in Claude Code?

Run `npx skills add uw-syfi/vibesys --skill testing -a claude-code`. Or copy the skill folder (.agents/skills/testing in uw-syfi/vibesys) into .claude/skills/testing in your project. Claude Code loads it when a task matches its description.

How do I install Testing in Codex?

Run `npx skills add uw-syfi/vibesys --skill testing -a codex`. Or copy the skill folder (.agents/skills/testing in uw-syfi/vibesys) into .agents/skills/testing in your project. Codex loads it when a task matches its description.

Can I use Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add uw-syfi/vibesys --skill testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing, .gemini/skills/testing, .github/skills/testing and .opencode/skills/testing in your project.

What does Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Testing is instructions for the agent only.

Does Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing use?

Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6k tokens, read only when the agent opens those files.

What are the alternatives to Testing?

Skills that share tags, products or a category with Testing: Swig Test (swig/swig, 6.3k stars), Triage CI Failure (DataDog/datadog-agent, 3.8k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars) and Fix Ready PRs (fastrepl/anarlog, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing?

uw-syfi (a GitHub organization) maintains it in uw-syfi/vibesys, which has 103 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.

Source: uw-syfi/vibesys on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.