Agent skill

Regression Hunt

by stella in stella/stella

Track down a behavior that used to work and now fails, changed, or regressed.

Apache-2.0Auto-check passedTesting & QA

Install Regression Hunt

skills CLI
$ npx skills add stella/stella --skill regression-hunt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install stella/stella regression-hunt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/stella/stella.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/regression-hunt .claude/skills/regression-hunt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
regression-hunt
GitHub stars
258
Token cost
~2.5k tokens
SKILL.md length
1,458 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

Track down a behavior that used to work and now fails, changed, or regressed.

  • Works in 10 steps: Restate the regression clearly → Build a feedback loop (this is the skill) → Encode the regression in a test before… → …
  • Tasks that involve QA and bug reports
  • SKILL.md covers Arguments, Instructions and Report back with
  • Calls bun and git

What it does

Regression Hunt is an agent skill from stella/stella. Track down a behavior that used to work and now fails, changed, or regressed. Use this when a bug report points to a recent breakage, especially when the cause is not obvious yet.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering QA and bug reports. The repository describes itself as: Open-source legal workspace. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve QA and bug reports

Example prompts

  • “/regression-hunt”

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Restate the regression clearly
  2. Build a feedback loop (this is the skill)
  3. Encode the regression in a test before fixing it
  4. Generate 3–5 ranked hypotheses before instrumenting
  5. Inspect recent change history
  6. Instrument
  7. Fix it minimally
  8. Verify
  9. Cleanup
  10. Post-mortem

What it can do on your machine

Read from SKILL.md and the folder at commit b225fd8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bun
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Regression Hunt loads about 2.5k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 1,458 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from stella/stella at commit b225fd8, republished under its Apache-2.0 licence (© stella). 1,458 words, ~2,513 tokens.

Download SKILL.mdSave it as .claude/skills/regression-hunt/SKILL.md (or your agent's skills folder).
name
regression-hunt
description
Track down a behavior that used to work and now fails, changed, or regressed. Use this when a bug report points to a recent breakage, especially when the cause is not obvious yet.
argument-hint
[what regressed]

Regression Hunt

Track down a behavior that used to work and now fails, changed, or regressed. Use this when a bug report points to a recent breakage, especially when the cause is not obvious yet.

Arguments

$ARGUMENTS: a short description of what regressed.

Helpful extras when available:

  • failing test name or file
  • error message or log line
  • expected vs actual behavior
  • suspect PR, branch, commit, or file
  • reproduction command, input, route, or endpoint

A plain-English bug report is enough to start.

Instructions

1. Restate the regression clearly
  • what used to work?
  • what is broken now?
  • what is expected instead?
2. Build a feedback loop (this is the skill)

Everything else is mechanical. With a fast, deterministic, agent-runnable pass/fail signal, bisection and hypothesis testing all just consume it. Without one, no amount of staring at code will save you. Spend disproportionate effort here. Be aggressive. Be creative. Refuse to give up.

Resolve the repository's test runner and the flags its scripts wire from its instructions first; the commands below show the Bun shape. Try these in roughly this order:

  1. Run the failing test via the package's test script, e.g. bun run test -- --bail -t "<pattern>", so flags wired into the script (--preload, custom setup) are preserved; calling bun test directly bypasses them. Avoid bun --bun test from a worktree root. Note that bun test positional arguments are file path patterns, not test names; use -t "<pattern>" to filter by test name and --bail to fast-fail.
  2. Curl / HTTP script against the backend dev server. Wrap fetches in AbortSignal.timeout(10_000) so a hung request does not rot the loop.
  3. CLI invocation diffing stdout against a known-good snapshot.
  4. Headless browser (Chrome DevTools MCP, Playwright, or similar) against the frontend dev server.
  5. Replay a captured trace (saved payload, event log) through the code path in isolation.
  6. Throwaway harness exercising just the bug code path with a single function call. For DB-touching code, a fresh test database per loop is usually faster and more honest than mocking.
  7. Differential loop: same input through old-commit vs new-commit, diff outputs.
  8. Bisection harness against the loop above. Drive it through the package script so wired flags survive, e.g.: git bisect run bun run test -- --bail -t "<test-name>". Verify the harness actually fails on a known-bad commit before starting: bun test exits 0 when no test names match, which would silently mark every commit as good and produce the wrong culprit.
  9. Property / fuzz loop if the bug is "sometimes wrong output".

Then treat the loop as a product. Iterate on it:

  • Faster? Cache setup, skip unrelated init, narrow scope.
  • Sharper signal? Assert on the specific symptom, not "didn't crash".
  • More deterministic? Pin time, seed RNG, isolate filesystem, freeze network.

A 30-second flaky loop is barely better than no loop. A 2-second deterministic loop is a debugging superpower.

For non-deterministic bugs the goal is a higher reproduction rate, not a clean repro. Loop the trigger 100×, parallelise, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not. Keep raising the rate until it is.

If you genuinely cannot build a loop, stop and say so explicitly. List what you tried. Ask the user for: environment access, a captured artifact (HAR, log dump, screen recording with timestamps), or permission to add temporary instrumentation. Do not proceed to hypothesise without a loop.

3. Encode the regression in a test before fixing it

Only if a correct seam exists, one where the test exercises the real bug pattern as it occurs at the call site. A test at the wrong seam (too shallow, single-caller test when the bug needs multiple callers, unit test that cannot replicate the trigger chain) gives false confidence.

If no correct seam exists, that itself is the finding. Note it and flag the architectural gap for step 10. The codebase is preventing the bug from being locked down.

Prefer focused integration tests over deep unit tests when the regression crosses layers. If the area has no automated harness, build the smallest reproducible check and say why a proper regression test was not added yet.

4. Generate 3–5 ranked hypotheses before instrumenting

Single-hypothesis generation anchors on the first plausible idea. Each hypothesis must be falsifiable; state its prediction:

If <X> is the cause, then <changing Y> will make it disappear / <changing Z> will make it worse.

If you cannot state a prediction, the hypothesis is a vibe; sharpen or discard. Show the ranked list to the user before testing; domain knowledge often re-ranks instantly ("we just deployed a change to #3"). Don't block on it if the user is AFK; proceed with your own ranking.

5. Inspect recent change history
  • current diff
  • recent commits touching the area
  • suspect PRs or refactors
  • React Compiler: if the regression is a re-render / stale-closure / perf shift in the frontend, check whether the compiler's output for the affected component changed. A lot of "this used to memoize and now it doesn't" bugs live here, not in your hand-written code.

Use git history when helpful, but do not stop at blame; verify the actual cause. For regressions specifically, git bisect run against your step 2 loop is often the fastest path to the offending commit.

Show full SKILL.md (590 more words)Show less
6. Instrument

Each probe must map to a specific prediction from step 4. Change one variable at a time.

Preference:

  1. Bun inspector. bun --inspect-brk is a runtime flag that must attach to the process actually running the code, which means wrapping it in bun run <script> will not propagate to the spawned child. Either prepend --inspect-brk to the test command inside your package script temporarily, or invoke directly while replicating the flags the script wires, e.g. bun --inspect-brk test --preload ./setup.ts <file-path>. Open the printed devtools:// URL in Chrome, set one breakpoint at the suspected fault. One breakpoint beats ten logs.
  2. Targeted logs at the boundaries that distinguish hypotheses.
  3. Never "log everything and grep".

Tag every debug log with a unique prefix, e.g. [DEBUG-a4f2]. Cleanup at the end becomes a single grep: untagged logs survive; tagged logs die.

For performance regressions, logs are usually wrong. Establish a baseline measurement first (prefer Bun.nanoseconds() over performance.now() on the backend (sharper resolution, no Node-portability tax), or a query plan for DB regressions), then bisect. Measure first, fix second.

7. Fix it minimally
  • preserve intended newer behavior where possible
  • do not revert unrelated improvements just to make the symptom disappear
  • identify the exact logic, assumption, or edge case that changed, not just the line that surfaces it
8. Verify
  • regression test now passes (or absence of correct seam is documented)
  • original repro from step 2 no longer reproduces
  • relevant wider checks pass (lint / typecheck / scoped tests)
9. Cleanup
  • All [DEBUG-...] instrumentation removed (grep the prefix)
  • Throwaway harnesses deleted or moved to a clearly-marked debug location
  • The hypothesis that turned out correct is stated in the commit / PR message, so the next debugger learns
10. Post-mortem

Ask: what would have prevented this regression? Make this recommendation after the fix is in; you have more information now than when you started.

Classify the failure before choosing a guard:

  • Representation: invalid or mutually exclusive states were representable. Prefer a branded type, discriminated union, private constructor, or exhaustive transition.
  • Invariant: behavior must hold across a broad input space. Prefer a property, idempotence, fixed-point, metamorphic, or fuzz test.
  • Boundary: malformed or ambiguous data crossed a trust boundary. Parse and validate once, fail early, and return a typed error.
  • Integration: individually valid components disagreed at a seam. Prefer a contract, differential, round-trip, or focused integration test.
  • Observation: the failure was only visible after degradation or deployment. Prefer a metric, baseline, canary, alert, or durable review artifact.
  • Heuristic: behavior is intentionally approximate or distribution-dependent. Prefer a representative benchmark and justified threshold; do not invent a false invariant.

Add the strongest feasible guard for a substantial bug. If the strongest guard needs a broader architectural change, land the immediate fix and create a concrete follow-up rather than merely mentioning the idea.

Additional lenses:

  • Type system could have caught it? Lift the constraint into types (branded types, discriminated unions, exhaustive checks). Patching the finding without lifting the constraint invites the same bug class back.
  • Missing test seam? Note the architectural gap; flag it for a separate refactor.
  • Convention violation? Update the repo's AGENTS.md / conventions so the next agent does not repeat it.
  • Lint could catch it? Consider a custom lint rule for a class of bug that humans and bots keep re-flagging.

Report back with

  • reproduction
  • regression test added (or why no correct seam exists)
  • ranked hypotheses + which one was right
  • root cause
  • fix
  • verification
  • structural prevention:
    • failure class
    • strongest applicable mechanism
    • guard added
    • why a stronger mechanism was not applicable
    • remaining ways this class could recur
  • any remaining uncertainty

© stella, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/regression-hunt of stella/stella.

Open the folder on GitHubat commit b225fd8

Compare with similar skills

Regression Hunt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Regression Hunt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Regression Hunt this skillstella/stella258—~2.5kAutomated safety check: PassApache-2.0
Reproduce Chat Statesdifferent-ai/openwork24k—~673Automated safety check: PassCustom licence
Dynamo Jira TicketDynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0
Moav E2EMotherofallVPNs/MoaV448—~1.9kAutomated safety check: NotesMIT
Creating A Coral TaskHuman-Agent-Society/CORAL1.1k—~2.2kAutomated safety check: PassApache-2.0
Launch Rlmarin-community/marin3.9k—~894Automated safety check: PassApache-2.0

Similar skills

  • Reproduce Chat States

    different-ai/openwork

    Fires known chat states in the running OpenWork desktop app, such as provider errors, retries and tool steps, so you can check how each renders.

    24k GitHub stars~673 tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Jira Ticket

    DynamoDS/Dynamo

    Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

    2k GitHub stars~1.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Moav E2E

    MotherofallVPNs/MoaV

    Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.

    448 GitHub stars~1.9k tokensUpdated 2 days ago
    Testing & QAAuto-check: notes
  • Creating A Coral Task

    Human-Agent-Society/CORAL

    Author a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout…

    1.1k GitHub stars~2.2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Launch Rl

    marin-community/marin

    Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.

    3.9k GitHub stars~894 tokensUpdated today
    Testing & QAAuto-check passed
  • Actionbook Web Test

    actionbook/actionbook

    Run browser-based web tests against websites using Actionbook CLI.

    1.6k GitHub stars~9.7k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from stella/stella

All 24 skills in this repo
  • Plan

    stella/stella

    Create a concise, evidence-backed implementation plan in the repository planning area when the user explicitly asks for a plan.

    258 GitHub stars~917 tokensUpdated today
    Auto-check passed
  • Answer From Sources

    stella/stella

    Answers data-protection (GDPR) questions grounded in the regulation and supervisory guidance, with a citation for every claim.

    258 GitHub stars~735 tokensUpdated today
    Auto-check passed
  • Check Against Rules

    stella/stella

    Reviews a non-disclosure agreement against the firm's NDA checklist and reports findings with citations.

    258 GitHub stars~856 tokensUpdated today
    Auto-check passed
  • Intake To Draft

    stella/stella

    Collects the facts of an unpaid invoice, then drafts a payment demand letter.

    258 GitHub stars~537 tokensUpdated today
    Auto-check passed
  • Conventions Perf

    stella/stella

    Apply when a performance-guard check (network baseline, bundle baseline, DB query count, loader-prefetch lint, RC bailouts) fails or when touching a hot route/endpoint.

    258 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Apply when writing or reviewing React effects in apps/web. An agent skill from stella/stella.

    258 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Regression Hunt

What does Regression Hunt do?

Track down a behavior that used to work and now fails, changed, or regressed. Regression Hunt is an agent skill from stella/stella. Track down a behavior that used to work and now fails, changed, or regressed.

When should I use Regression Hunt?

Regression Hunt fits situations like: tasks that involve QA and bug reports.

How do I install Regression Hunt in Claude Code?

Run `npx skills add stella/stella --skill regression-hunt -a claude-code`. Or copy the skill folder (.agents/skills/regression-hunt in stella/stella) into .claude/skills/regression-hunt in your project. Claude Code loads it when a task matches its description.

How do I install Regression Hunt in Codex?

Run `npx skills add stella/stella --skill regression-hunt -a codex`. Or copy the skill folder (.agents/skills/regression-hunt in stella/stella) into .agents/skills/regression-hunt in your project. Codex loads it when a task matches its description.

Can I use Regression Hunt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add stella/stella --skill regression-hunt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/regression-hunt, .gemini/skills/regression-hunt, .github/skills/regression-hunt and .opencode/skills/regression-hunt in your project.

What does Regression Hunt need to run?

Going by SKILL.md and its folder, Regression Hunt needs the command-line tools its instructions call (bun and git).

Does Regression Hunt access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Regression Hunt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Regression Hunt use?

Regression Hunt is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Regression Hunt use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Regression Hunt?

Skills that share tags, products or a category with Regression Hunt: Reproduce Chat States (different-ai/openwork, 24k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars), Moav E2E (MotherofallVPNs/MoaV, 448 stars) and Creating A Coral Task (Human-Agent-Society/CORAL, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Regression Hunt?

stella (a GitHub organization) maintains it in stella/stella, which has 258 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 9, 2026.

Source: stella/stella on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.