Agent skill

Tutti Test Audit

by tutti-os in tutti-os/tutti

Audit, design, write, materially revise, or remove Tutti tests.

Apache-2.0Auto-check passedAgent Workflows

Install Tutti Test Audit

skills CLI
$ npx skills add tutti-os/tutti --skill tutti-test-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tutti-os/tutti tutti-test-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tutti-os/tutti.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/tutti-test-audit .claude/skills/tutti-test-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tutti-test-audit
GitHub stars
3.8k
Token cost
~2k tokens
SKILL.md length
997 words
Files
2
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit, design, write, materially revise, or remove Tutti tests.

  • Works in 5 steps: the production owner, public entry… → sibling implementations that share the… → existing tests, fixtures, conformance… → …
  • Repository tests
  • SKILL.md covers Establish the evidence map, Pass the authoring gate, Choose the proof shape and Build sensitive evidence, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tutti Test Audit is an agent skill from tutti-os/tutti. Audit, design, write, materially revise, or remove Tutti tests. Use whenever Codex changes or reviews unit, component, conformance, integration, regression, platform, or repository tests; enforce a protected product contract, credible failure, correct owning boundary, negative-control evidence, overlap review, deterministic setup, and an executing CI lane.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Agent Workflows. The repository describes itself as: Where people and agents build in tune. The licence is Apache-2.0.

When your agent uses it

  • Repository tests
  • Enforce a protected product contract
  • Credible failure
  • Correct owning boundary

Example prompts

  • “/tutti-test-audit”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. the production owner, public entry point, relevant caller and callee;
  2. sibling implementations that share the same invariant;
  3. existing tests, fixtures, conformance suites, and repository checks;
  4. the changed-aware and platform lane that will select the evidence;
  5. relevant history for a regression or a suspicious existing test.

What it can do on your machine

Read from SKILL.md and the folder at commit 207beee. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tutti Test Audit loads about 2k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 997 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tutti-os/tutti at commit 207beee, republished under its Apache-2.0 licence (© tutti-os). 997 words, ~1,992 tokens.

Download SKILL.mdSave it as .claude/skills/tutti-test-audit/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
tutti-test-audit
description
Audit, design, write, materially revise, or remove Tutti tests. Use whenever Codex changes or reviews unit, component, conformance, integration, regression, platform, or repository tests; enforce a protected product contract, credible failure, correct owning boundary, negative-control evidence, overlap review, deterministic setup, and an executing CI lane.

Tutti Test Audit

Make every changed test earn its maintenance cost. A green test, a coverage increase, or a larger test count is not a quality verdict.

Read the root and closest scoped AGENTS.md, then read docs/conventions/unit-testing.md before a non-trivial test change or review. Follow docs/conventions/testing.md#validation-selection for commands and validation scope.

Establish the evidence map

Before editing, inspect:

  1. the production owner, public entry point, relevant caller and callee;
  2. sibling implementations that share the same invariant;
  3. existing tests, fixtures, conformance suites, and repository checks;
  4. the changed-aware and platform lane that will select the evidence;
  5. relevant history for a regression or a suspicious existing test.

For Agent Host lifecycle semantics, also read packages/agent/host/README.md and start with the conformance contract. Do not reimplement lifecycle tests in an adapter.

Pass the authoring gate

Answer all six questions before adding or materially rewriting a test:

  1. Protected contract: What observable behavior, invariant, compatibility promise, or prior failure matters?
  2. Credible failure: What plausible faulty implementation must make the test fail?
  3. Coverage gap: Why would existing evidence not catch that failure? Can an existing table, scenario, or fixture express it without duplication?
  4. Owner and observer: Which module owns the decision, and what is the lowest boundary that can observe the risk without mocking it away?
  5. Production seam: Does the test require an export, flag, wrapper, global, or injection hook that no production caller needs? If so, move to the real boundary or redesign the owner.
  6. Execution: Which blocking lane selects the test, and which native OS must execute it?

A missing answer means the test is not ready. “No new test” is a valid outcome when stronger existing evidence already protects the contract or the proposed assertion only restates a type, literal, implementation detail, or presentation choice.

Choose the proof shape

  • Extend an existing table or conformance scenario when it owns the same contract and fixture.
  • Add a new scenario when the behavior, failure mechanism, setup, or execution lane is independently meaningful.
  • Move upward to a component, narrow integration, conformance, or E2E boundary when a unit double would replace the semantics under test.
  • Use a repository check, not a unit test, when source structure, imports, generated artifacts, manifests, or architecture are the actual contract.
  • Keep visual styling in visual evidence unless styling changes native behavior such as pointer ownership or window dragging.

Prefer one vertical proof through the owning boundary plus narrow tests for important decision partitions over many stub-backed branch tests.

Build sensitive evidence

For a confirmed bug, capture the failing reproduction before the fix whenever possible. The regression test must fail on pre-fix behavior for the intended reason and pass after the owner-boundary repair. If the old revision cannot be run safely, temporarily inject or restore the faulty decision as a local negative control and report that limitation; never commit the mutation.

For new behavior, first demonstrate that the focused test fails because the contract is absent, not because setup, imports, or fixtures are broken.

For a read-only audit, do not edit production code merely to manufacture a negative control. Name the smallest plausible faulty implementation, determine whether the current assertions would detect it, and label that sensitivity unverified when it cannot be executed safely.

Then:

  • assert stable outcomes and important non-effects;
  • include neighboring valid, rejection, cleanup, retry, or preservation cases only when they address a credible risk;
  • use real isolated SQLite, filesystem, Git, HTTP, IPC, parser, or process boundaries when their semantics are the subject;
  • keep external doubles narrow and fail closed on unexpected calls;
  • coordinate concurrency with events, channels, barriers, or latches rather than sleeps or lucky elapsed-time bounds;
  • control clocks, IDs, randomness, ports, user state, credentials, globals, listeners, timers, processes, and temporary resources;
  • assert call counts only when count is itself a contract such as idempotency, single-flight, or at-most-once dispatch.
Show full SKILL.md (359 more words)Show less

Apply Tutti routing

  • Session, Turn, Goal, runtime-operation, and recovery semantics start in packages/agent/host/conformance; consumers provide drivers.
  • Adapter tests own transport, DTO, authorization, query, presentation, and product-policy behavior only.
  • SQLite, filesystem, path, shell, executable, process, signal, permission, and native API behavior cross the real isolated receiver boundary.
  • UI behavior drives accessible keyboard, pointer, focus, and state transitions and observes callbacks or user-visible state.
  • HTTP, IPC, and provider protocol tests cross the real parser, middleware, or codec boundary when receiver behavior is the risk.
  • Platform-dependent behavior runs on the native platform. A serialized request shape or a skipped Linux test is not Windows proof.

Audit existing tests

Treat these as review candidates, not automatic deletions:

  • no meaningful assertion, self-comparison, or only “does not throw”;
  • copied production algorithm, catalog, defaults, or current inventory;
  • source strings or regexes presented as runtime behavior proof;
  • private call-shape or mocked call-graph assertions;
  • static render, class, style, SVG path, or incidental DOM assertions;
  • duplicate coverage weaker than an existing boundary scenario;
  • sleeps, polling delays, pass-on-retry, silent skips, or zero-selected lanes;
  • test-only production exports, wrappers, flags, globals, or dead code.

Before weakening or removing a test, identify the failure it can actually catch, its history, and the stronger evidence that remains. Preserve independent public API, protocol, config, migration, storage, security, platform, generated, release, and architecture contracts even when they are static or slow.

Classify each candidate as:

  • keep — protects an independent contract;
  • strengthen or move — right risk, wrong oracle or boundary;
  • consolidate — useful contract already has a canonical owner;
  • delete — no credible failure remains protected, including any test-only production seam it kept alive.

Validate and hand off

  1. Run the smallest owner and sibling tests that prove the contract.
  2. For authored behavior changes, demonstrate the pre-fix failure or another credible negative control. For a read-only audit, report the failure model without mutating the checkout.
  3. Apply the repository validation-selection policy; verify the intended lane selected and ran non-zero evidence.
  4. Run native-platform proof when required, or name the remaining manual gate.
  5. Inspect the final diff for duplicated setup, weakened assertions, snapshot churn, leaked resources, and unnecessary production seams.

Report this compact evidence for non-trivial test changes:

md
Test evidence

- Protected contract or prior failure:
- Credible faulty implementation / negative control:
- Existing coverage gap:
- Owning seam and test level:
- Observable outcomes and important non-effects:
- Blocking lane and native OS, when relevant:
- Residual risk left to integration, E2E, visual, or manual proof:

© tutti-os, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .codex/skills/tutti-test-audit of tutti-os/tutti.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 207beee

Compare with similar skills

Tutti Test Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tutti Test Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tutti Test Audit this skilltutti-os/tutti3.8k—~2kAutomated safety check: PassApache-2.0
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Hook Development for Claude Code Pluginsanthropics/claude-plugins-official38k10 repos~4.1kAutomated safety check: NotesApache-2.0
Using Superpowersfarm-fe/farm5.6k35 repos~1.4kAutomated safety check: PassMIT
Executing Plans Inlineobra/superpowers297k2 repos~5.1kAutomated safety check: PassMIT
Skill CreatorAzure/azqr79689 repos~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Hook Development for Claude Code Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.

    38k GitHub starsUsed in 10 repos~4.1k tokens
    Agent WorkflowsAuto-check: notes
  • Using Superpowers

    farm-fe/farm

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    5.6k GitHub starsUsed in 35 repos~1.4k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans Inline

    obra/superpowers

    Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.

    297k GitHub starsUsed in 2 repos~5.1k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    796 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed

More from tutti-os/tutti

All 8 skills in this repo
  • Tutti UI System

    tutti-os/tutti

    A skill your agent uses when working with @tutti-os/ui-system components, replacing local UI with shared components, querying component ids or metadata, promoting UI into shared base or business…

    3.8k GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed
  • Analyze Chrome, Chromium, Electron, React DevTools, or Perfetto-compatible JSON traces and audit user-reported profiling findings without loading large artifacts into context; prove…

    3.8k GitHub stars~3.3k tokensUpdated 2 days ago
    Auto-check passed
  • From a Tutti checkout, run, audit, freshly replay, publish, or diagnose Session Replay cassettes that are driven by case-repository scenario scripts (CDP), not by interactive UI recording.

    3.8k GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Create, convert, or repair one Tutti workspace app as either a self-contained publishable package under package/ or a Chrome-style local debug app under .tutti/dev-app/.

    3.8k GitHub stars~5.9k tokensUpdated 2 days ago
    Auto-check passed
  • Build or evolve a complex agent-enabled Tutti workspace app repository.

    3.8k GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Review tutti git diffs for project structure, layering, module ownership, and duplicate event-center infrastructure by planning focused architecture review tasks, then having the main agent…

    3.8k GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Tutti Test Audit

What does Tutti Test Audit do?

Audit, design, write, materially revise, or remove Tutti tests. Tutti Test Audit is an agent skill from tutti-os/tutti. Audit, design, write, materially revise, or remove Tutti tests.

When should I use Tutti Test Audit?

Tutti Test Audit fits situations like: repository tests; enforce a protected product contract; credible failure; correct owning boundary.

How do I install Tutti Test Audit in Claude Code?

Run `npx skills add tutti-os/tutti --skill tutti-test-audit -a claude-code`. Or copy the skill folder (.codex/skills/tutti-test-audit in tutti-os/tutti) into .claude/skills/tutti-test-audit in your project. Claude Code loads it when a task matches its description.

How do I install Tutti Test Audit in Codex?

Run `npx skills add tutti-os/tutti --skill tutti-test-audit -a codex`. Or copy the skill folder (.codex/skills/tutti-test-audit in tutti-os/tutti) into .agents/skills/tutti-test-audit in your project. Codex loads it when a task matches its description.

Can I use Tutti Test Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tutti-os/tutti --skill tutti-test-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tutti-test-audit, .gemini/skills/tutti-test-audit, .github/skills/tutti-test-audit and .opencode/skills/tutti-test-audit in your project.

What does Tutti Test Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Tutti Test Audit is instructions for the agent only.

Does Tutti Test Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tutti Test Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tutti Test Audit use?

Tutti Test Audit is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tutti Test Audit use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tutti Test Audit?

Skills that share tags, products or a category with Tutti Test Audit: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 297k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tutti Test Audit?

tutti-os (a GitHub organization) maintains it in tutti-os/tutti, which has 3,811 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 8, 2026.

Source: tutti-os/tutti on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.