Agent skill

Test Run Rules

by tobihagemann in tobihagemann/turbo

Shared rules for test runs that drive the running app or run a test target: isolation from concurrent runs, stubs for infrastructure that cannot be stood up, provisioning privileged state, approval…

MITAuto-check passed

Install Test Run Rules

skills CLI
$ npx skills add tobihagemann/turbo --skill test-run-rules -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tobihagemann/turbo test-run-rules --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tobihagemann/turbo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/codex/skills/test-run-rules .claude/skills/test-run-rules && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-run-rules
GitHub stars
409
Token cost
~2.9k tokens
SKILL.md length
1,815 words
Files
1
Skills in repo
81
Repo updated
First seen
Licence
MIT

At a glance

Shared rules for test runs that drive the running app or run a test target: isolation from concurrent runs, stubs for infrastructure that cannot be stood up, provisioning privileged state, approval…

  • SKILL.md covers Isolation, Unavailable Infrastructure, Privileged State and Second… and Writes to Shared External…, plus 3 more sections
  • Calls git

What it does

Test Run Rules is an agent skill from tobihagemann/turbo. Shared rules for test runs that drive the running app or run a test target: isolation from concurrent runs, stubs for infrastructure that cannot be stood up, provisioning privileged state, approval before writes to shared external systems, and cleanup of only what the run created. Not typically invoked directly.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Reusable workflows for planning, building, reviewing, and shipping with Claude Code and Codex. The licence is MIT.

Example prompts

  • “/test-run-rules”

What it can do on your machine

Read from SKILL.md and the folder at commit 160a0fa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Run Rules loads about 2.9k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 1,815 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tobihagemann/turbo at commit 160a0fa, republished under its MIT licence (© tobihagemann). 1,815 words, ~2,851 tokens.

Download SKILL.mdSave it as .claude/skills/test-run-rules/SKILL.md (or your agent's skills folder).
name
test-run-rules
description
Shared rules for test runs that drive the running app or run a test target: isolation from concurrent runs, stubs for infrastructure that cannot be stood up, provisioning privileged state, approval before writes to shared external systems, and cleanup of only what the run created. Not typically invoked directly.

Test Run Rules

Apply these rules for the rest of the test run, whether it drives the running app or runs a test target. A scenario they leave blocked could not be set up or run; one they leave inconclusive ran, but its result cannot be read.

Isolation

  • Isolate shared process state so concurrent or sub-agent runs don't collide: bind dev servers and services to unique ports, scope tmux sessions (tmux -L <name>), give each browser session a unique name so cleanup can target only its own, and write screenshots and other scratch state to absolute paths under a uniquely named subdirectory of $TMPDIR. Derive each such identifier once and reuse that exact value in every later command, writing it as a literal or reading it back from a note in that subdirectory. A value recomputed per shell, such as $$, differs between the command that creates a resource and the command that releases it, so cleanup releases something it never created and reports success while the real resource leaks. A port picked as unique may already be held by a concurrent agent, so check it before binding and move to another when it is taken, leaving the incumbent running. When a unique port moves a service off its default address, find the settings elsewhere in the stack that name that default, such as allowed origins and sign-in callback URLs, and bring each in line through runtime overrides, leaving the working tree unchanged: add the new address beside the default in a list, and replace the default only where no process outside this run reads the setting.
  • Reuse a running dev server only when this session started it and has not handed it to the user to try. Otherwise start one on a port this run selected and wait for it to be ready. Confirm it bound to that port before sending it traffic — a failed bind leaves another agent's service answering. Move to another port when the port is taken; report the error and stop when the server itself failed to start.
  • Treat state the app persists at a fixed path, such as a data file or a local database, as shared with every other live instance of the app. When another live instance is using it, give this run's instance a copy of its own under the run's scratch subdirectory: repoint the path there through runtime configuration, or, when the path cannot be repointed and the app resolves it relative to its own tree, run the instance from a copy of the working tree placed there, uncommitted changes included. Leave the working tree unchanged. When neither separates the state, every scenario that writes to it is blocked: name the path.
  • Run each test runner this run starts in its own process group under a timeout enforced from outside the runner.
  • Scope a screen capture or recording this run takes to the surface under test, addressed by its identifier, whenever the capture tool can. A capture of the whole display holds whatever else is open on it and can miss a surface that is covered or not in front.

Unavailable Infrastructure

When a scenario needs infrastructure that cannot be stood up in this session (backend service, auth provider, external dependency), look for a stub before treating the scenario as blocked. When the dependency is reached through a client whose endpoint is runtime configuration, repoint that endpoint at a local stub; the same interception yields an artifact the system emits rather than provides (a token, a session identifier, a single-use link). Confine this to runtime configuration and leave the working tree unchanged. When the code under test makes the call itself, a call-site stub intercepts it without a second process: from the run, install a replacement into the running process that matches one destination and returns the response the scenario needs or delays the real one, gated on a flag the run can flip. Installing it from the run leaves the working tree unchanged as well. It lasts only as long as the process, so re-install it after anything that restarts or reloads it. When no stub applies or the ones that do fail, the scenario is blocked: name what is missing and what was tried.

Privileged State and Second Participants

When a scenario needs privileged state or a second participant (an entitlement or plan tier, an elevated role, seed data, a second concurrent client or session), provision it through a path the project already exposes for development, such as its own development-only endpoint, an administrative command, or a second client this run starts. When the provisioning path writes to a shared external system, carry it through the write sequence under Writes to Shared External Systems. Treat the scenario as blocked only after an attempt to provision failed, naming the precondition and what was tried.

Treat a permission or scope granted mid-run to unblock a scenario as something this run created: before reporting cleanup complete, verify the production code never needs it, then ask for it to be revoked in the report. Name the call sites checked there too.

Writes to Shared External Systems

When a scenario writes to a shared external system and those writes are not cleanly undoable, run it against fixture data with nothing to act on whenever its expected outcome can still be observed that way, so the run still exercises wiring, auth, queries, guards, and failure isolation while writing nothing, and name that fixture data in the result as a substitution. Treat writes as not cleanly undoable whenever restoring the records leaves downstream effects the writes triggered in place.

When the writing path must run, work through it in order. Determine the full write set without executing it: use a dry-run mode when one exists, otherwise trace the code path and enumerate every record it writes, including those reached through triggers, cascades, and hooks. State what the enumeration cannot settle rather than presenting it as complete. Pick the target whose writes are incidental to what the scenario verifies, weighing each candidate's write set against the coverage it adds. Then request approval via request_user_input, presenting the enumeration as what is being consented to, and request it again whenever the enumeration changes. When request_user_input does not reach the user, write nothing: the scenario is blocked, with the approval named as unresolved. Capture a pre-run manifest and write an ordered revert procedure. When approval is declined, the scenario is blocked: name the writes that were withheld.

Show full SKILL.md (738 more words)Show less

Reading Results

  • When a scenario names a control, command, or other affordance the app does not have, establish what the scenario verifies before recording a verdict. When the named mechanism is itself what the scenario verifies, its absence is a failure. When the mechanism is incidental to the outcome the scenario verifies, drive the affordance that delivers that outcome, record the verdict against it, and name the substitution in the result. When which of the two it is cannot be established, the scenario is inconclusive.
  • When a scenario depends on an input mode or device characteristic the browser emulates, confirm the page itself reports that capability before recording a verdict resting on it — a device preset may change only the viewport and the user agent. When the capability cannot be established, the scenario is inconclusive: name the capability.
  • When a scenario's pass condition is that something is visible, record the verdict from an observation rendering can falsify: a capture of the rendered surface, read for that element. The element's presence in a structure snapshot, or a read of its styles or properties, establishes that it exists and what is set on it, and leaves open whether anything shows. When no such observation can be made, the scenario is inconclusive: name the observation that was unavailable.
  • When a scenario fails after an action aimed at an element, establish which element received the action before recording the failure: the element at the point of a click, or the element holding focus for typed input, read as the action ran where the tool allows, since afterward the surface may already have changed. When the run's own aim was off, restore the scenario's starting state, repeat the action the same way on the intended element, and record that result. When the app diverted the action from the intended element, record the failure. When what received the action cannot be established, the scenario is inconclusive: name the action.
  • Verify an observation before reporting it as pre-existing rather than introduced by the change, and state that verification. Read the file at the ref the change under test is measured against with git show <ref>:<path>: HEAD for uncommitted or staged work, the commit the work began from when it is already committed locally, or origin/<base-branch> for a PR, fetched first. Reporting an introduced defect as pre-existing drops it from scope silently, while the opposite error only adds noise.

Logs

  • Tail app logs in a background shell for errors or warnings while testing, so backend failures surface alongside the checks.
  • After the last interaction with the app, perform one additional log read or status check before reporting. Background-shell output that lands after the agent emits final text is not surfaced, so the extra action gives it time to appear in a polled read. Matters most when the run happens inside a sub-agent. When this check targets a server or log process this run did not start, report it as outstanding for the process owner rather than running it yourself; inability to read such a process is outstanding, not a test failure.

Code and Cleanup

  • Leave application code unchanged: every stub above lives in runtime configuration or the running process and is removed on cleanup. Report failures without attempting to fix them.
  • Always clean up: close only the browser sessions this run opened, by name, kill the tmux server of each tmux -L name it scoped, stop the dev servers, services, stubs, and test runners this run started, restore any configuration it repointed or extended, release the sessions, roles, and records it created, and run the revert procedure of each approved write. Capture the PID of each dev server, service, stub, and test runner this run starts and stop it by that PID rather than by a name or command-line pattern, which also matches an identically named process a concurrent agent is running. Stop the process group rather than the captured PID alone — a server started behind a wrapper outlives its parent, and a runner's spawned helpers outlive the runner. Before reporting cleanup complete, confirm each server, service, and stub port released and no process from a stopped group still running, and report by PID any process that could not be stopped. When processes cannot be listed, report that check as unrun and name the groups. Never close all browser sessions at once — concurrent agents may share the browser daemon, so a blanket close is cross-agent destruction.

© tobihagemann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in codex/skills/test-run-rules of tobihagemann/turbo.

Open the folder on GitHubat commit 160a0fa

Compare with similar skills

Test Run Rules next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Run Rules compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Run Rules this skilltobihagemann/turbo409—~2.9kAutomated safety check: PassMIT
Drivedisler/mac-mini-agent272—~1.6kAutomated safety check: PassNone
Feishu Driveopenclaw/openclaw392k—~375Automated safety check: PassMIT
Electron Drive Skillsickn33/agentic-awesome-skills47k1 repos~2.5kAutomated safety check: PassMIT
Gws Drive Uploadgoogleworkspace/cli31k1 repos~322Automated safety check: PassApache-2.0
Google Drivecursor/plugins10k—~1.1kAutomated safety check: PassNone

Similar skills

  • Drive

    disler/mac-mini-agent

    Terminal automation CLI for AI agents. An agent skill from disler/mac-mini-agent.

    272 GitHub stars~1.6k tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • Feishu Drive

    openclaw/openclaw

    Feishu cloud-storage and comment workflows. An agent skill from openclaw/openclaw.

    392k GitHub stars~375 tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Electron Drive Skill

    sickn33/agentic-awesome-skills

    Launch the project's Electron app on a scratch profile and drive it: click, type, screenshot, run renderer or main-process code, read logs.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Testing & QAAuto-check passed
  • Gws Drive Upload

    googleworkspace/cli

    Google Drive: Upload a file with automatic metadata. An agent skill from googleworkspace/cli.

    31k GitHub starsUsed in 1 repo~322 tokens
    Documents & OfficeAuto-check passed
  • Google Drive

    cursor/plugins

    Official

    Work with Google Drive through its MCP tools - search, read, create, organize, share, comment, and export files.

    10k GitHub stars~1.1k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Drive MiMo Code

    XiaomiMiMo/MiMo-Code

    Lets one MiMoCode process drive another, headless with JSON events or interactively through tmux, to test behavior and visual regressions with parseable evidence.

    14k GitHub stars~3.9k tokensUpdated today
    Testing & QAAuto-check passed

More from tobihagemann/turbo

All 81 skills in this repo
  • Consult Oracle

    tobihagemann/turbo

    Consult ChatGPT Pro via ChatGPT browser automation for problems that resist standard approaches.

    409 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Fetch PR Comments

    tobihagemann/turbo

    Fetch and summarize review feedback and conversation from a GitHub PR (unresolved review threads, review bodies, and PR conversation comments) without making changes.

    409 GitHub stars~967 tokensUpdated today
    Auto-check passed
  • Recall Rationale

    tobihagemann/turbo

    Recall why a past change was made by locating the Claude Code transcript that produced it.

    409 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Resolve PR Comments

    tobihagemann/turbo

    Evaluate, fix, answer, and reply to GitHub pull request review comments and conversation comments.

    409 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Resolve PR Comments

    tobihagemann/turbo

    Evaluate, fix, answer, and reply to GitHub pull request review comments and conversation comments.

    409 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Assess Technical Debt

    tobihagemann/turbo

    Assess project-wide structural technical debt: complexity hotspots, deprecated API usage, duplication clusters, architecture rot, and low-value tests.

    409 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Questions about Test Run Rules

What does Test Run Rules do?

Shared rules for test runs that drive the running app or run a test target: isolation from concurrent runs, stubs for infrastructure that cannot be stood up, provisioning privileged state, approval…. Test Run Rules is an agent skill from tobihagemann/turbo. Shared rules for test runs that drive the running app or run a test target: isolation from concurrent runs, stubs for infrastructure that cannot be stood up, provisioning privileged state, approval before writes to shared external systems, and cleanup of only what the run created.

How do I install Test Run Rules in Claude Code?

Run `npx skills add tobihagemann/turbo --skill test-run-rules -a claude-code`. Or copy the skill folder (codex/skills/test-run-rules in tobihagemann/turbo) into .claude/skills/test-run-rules in your project. Claude Code loads it when a task matches its description.

How do I install Test Run Rules in Codex?

Run `npx skills add tobihagemann/turbo --skill test-run-rules -a codex`. Or copy the skill folder (codex/skills/test-run-rules in tobihagemann/turbo) into .agents/skills/test-run-rules in your project. Codex loads it when a task matches its description.

Can I use Test Run Rules in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tobihagemann/turbo --skill test-run-rules -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-run-rules, .gemini/skills/test-run-rules, .github/skills/test-run-rules and .opencode/skills/test-run-rules in your project.

What does Test Run Rules need to run?

Going by SKILL.md and its folder, Test Run Rules needs the command-line tools its instructions call (git).

Does Test Run Rules access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Run Rules safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Run Rules use?

Test Run Rules is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Run Rules use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Run Rules?

Skills that share tags, products or a category with Test Run Rules: Drive (disler/mac-mini-agent, 272 stars), Feishu Drive (openclaw/openclaw, 392k stars), Electron Drive Skill (sickn33/agentic-awesome-skills, 47k stars) and Gws Drive Upload (googleworkspace/cli, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Run Rules?

tobihagemann (a GitHub user) maintains it in tobihagemann/turbo, which has 409 GitHub stars. The repository holds 81 skills in this directory. The repository was last updated on October 9, 2026.

Source: tobihagemann/turbo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.