Agent skill

Rigor Regression Sweep

by rigortype in rigortype/rigor

Measure Rigor's baseline drift across the tagged history of a real OSS Ruby project.

MPL-2.0Auto-check passedDevelopment

Install Rigor Regression Sweep

skills CLI
$ npx skills add rigortype/rigor --skill rigor-regression-sweep -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rigortype/rigor rigor-regression-sweep --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rigortype/rigor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/rigor-regression-sweep .claude/skills/rigor-regression-sweep && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rigor-regression-sweep
GitHub stars
106
Token cost
~2.9k tokens
SKILL.md length
1,415 words
Files
3 (incl. scripts)
Skills in repo
36
Repo updated
First seen
Licence
MPL-2.0

At a glance

Measure Rigor's baseline drift across the tagged history of a real OSS Ruby project.

  • Works in 9 steps: When to use → Pick the target and the tag range → Clone the target → …
  • Validating diagnostic realism over a release line
  • SKILL.md covers Phase 0 — When to use, Phase 1 — Pick the target and…, Phase 2 — Clone the target and Phase 3 — Verify the tags exist, plus 7 more sections
  • Runs Shell and Ruby scripts from its folder; calls git, nix and bash; reaches github.com

What it does

Rigor Regression Sweep is an agent skill from rigortype/rigor. Measure Rigor's baseline drift across the tagged history of a real OSS Ruby project. Use when validating diagnostic realism over a release line or growing the survey corpus; not for a single version check or ordinary project onboarding.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/sweep.sh`).

It sits in Development. It works with Ruby. The repository describes itself as: Inference-first static analysis for Ruby. The licence is MPL-2.0.

When your agent uses it

  • Validating diagnostic realism over a release line
  • Growing the survey corpus
  • Not for a single version check
  • Ordinary project onboarding

Example prompts

  • “/rigor-regression-sweep”

Requirements

  • A Bash shell

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. When to use
  2. Pick the target and the tag range
  3. Clone the target
  4. Verify the tags exist
  5. Write the frozen config
  6. Baseline at the first tag
  7. Sweep the tags
  8. Tabulate
  9. Interpret and record

What it can do on your machine

Read from SKILL.md and the folder at commit 57a67cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and Ruby), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • nix
    • bash
    • make
    • bundle
    • mise

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rigor Regression Sweep loads about 2.9k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 1,415 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from rigortype/rigor at commit 57a67cf, republished under its MPL-2.0 licence (© rigortype). 1,415 words, ~2,899 tokens.

Download SKILL.mdSave it as .claude/skills/rigor-regression-sweep/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
rigor-regression-sweep
description
Measure Rigor's baseline drift across the tagged history of a real OSS Ruby project. Use when validating diagnostic realism over a release line or growing the survey corpus; not for a single version check or ordinary project onboarding.
metadata.internal
true

Rigor Regression Sweep

A contributor workflow for measuring how Rigor's diagnostics behave over a real project's development flow. Pick an OSS Ruby project, baseline it at one version, then run rigor check against every later tag with that baseline + config frozen. The per-tag surfaced diagnostic count is the "error increase" a team adopting Rigor in acknowledge mode would have seen.

It validates two things at once: the baseline mechanism (ADR-22) and the realism of the rigor-project-init acknowledge- mode environment — and grows an empirical corpus under docs/notes/.

First worked run: docs/notes/20260521-mastodon-v4.5-regression-sweep.md (Mastodon, 16 tags v4.5.0-beta.1 → v4.5.10).

Phase 0 — When to use

Trigger when asked to "sweep Rigor across versions of X", "check how the error count grows over a release line", "validate the baseline against real churn", or to add a project to the regression corpus.

Do NOT use for: a single-version check of one project (just run rigor check); authoring a plugin (rigor-plugin-author); the 22-library single-snapshot survey style of docs/notes/20260519-oss-library-survey.md (that is breadth; this is depth-over-time for one project).

Phase 1 — Pick the target and the tag range

Choose a real OSS Ruby project with git tags.

The sampling decision is the single most important one — and it governs what the sweep can prove.

Released tags are a post-spec-gate population. A mature project's CI catches obvious errors before merge, and genuine bugs are fixed before a tag is cut. So a sweep over released tags — whether a patch series or a feature-spanning minor/major range — measures:

  • baseline stability — does ordinary maintenance churn produce false regressions? (a real, valuable question), and
  • standing diagnostics — does newly-added released code carry diagnostics the baseline did not already cover?

It does not measure Rigor's bug-detection power. surfaced = 0 over released tags is the expected result for a healthy project, not a Rigor weakness: the bugs Rigor would catch are the same class the project's specs already caught and removed before the tag existed. (Confirmed by the Mastodon v4.5.x run — flat at 0 across 16 tags.)

To actually test bug-detection, sample where unfixed bugs still exist — finer than release tags:

  • per-commit on main between two releases (or a window of it);
  • PR-head commits (pre-merge state);
  • bug-introducing commits — for a fixed set of known bug fixes, sweep the commit before each fix and check whether Rigor surfaced the defect.

Pick the sampling to match the question:

QuestionSample
Does maintenance churn cause false regressions?released tags (this is the easy, default run)
Does new released code add standing diagnostics?feature-spanning released-tag range
Does Rigor catch real bugs?per-commit / PR-head / bug-introducing commits

List the chosen revisions in chronological order (betas / RCs included — they are part of the dev flow). Record the sampling choice and its consequence in the survey note.

The later phases say "tag" for brevity, but every step works on any checkout-able revision — git checkout and the scripts' TAGS list accept commit SHAs just as well as tag names.

Phase 2 — Clone the target

Clone into the survey area, not the rigor repo:

sh
cd ~/repo/ruby/rigor-survey
git clone --filter=blob:none https://github.com/<org>/<name>.git <name>

--filter=blob:none (blobless partial clone) avoids fetching all historical blobs up front; each tag checkout fetches what it needs.

Phase 3 — Verify the tags exist

sh
cd ~/repo/ruby/rigor-survey/<name>
for t in <tag-list>; do
  git rev-parse -q --verify "refs/tags/$t" >/dev/null \
    && echo "OK   $t" || echo "MISS $t"
done

Drop or substitute any missing tag before sweeping; note omissions.

Phase 4 — Write the frozen config

Write .rigor.dist.yml into the target's root (untracked there, so git checkout between tags never disturbs it). Approximate what rigor-project-init would generate for the project's stack, then freeze it — identical config across every tag is what makes the surfaced-count delta attributable to the project's code.

yaml
# Frozen — held identical across every tag in the sweep.
paths: [app, lib]            # the project's source roots
exclude: [vendor, tmp]
severity_profile: lenient    # acknowledge-mode default for a large project
plugins: [rigor-activesupport-core-ext]
cache:
  path: /abs/path/to/rigor-survey/_<name>-sweep/cache

Rules that keep the sweep clean:

  • Artefacts live OUTSIDE the target tree. Put the baseline, cache, and per-tag reports under a sibling _<name>-sweep/ directory. cache.path is set absolute for the same reason. Then git checkout <tag> only ever changes the project's own files.
  • Plugins. Bundled plugins ship inside rigortype; activate them with plugins: in the frozen config and record which were and were not active.
  • The content-hashed cache is safe to share across tags — a changed file misses, an unchanged file hits. Keep it; it makes the sweep fast.

Phase 5 — Baseline at the first tag

sh
cd ~/repo/ruby/rigor-survey/<name> && git checkout -q -f <first-tag>

Generate the baseline (run from the rigor repo so the Nix flake resolves; cd into the target inside the command):

sh
nix develop --command \
  bash -c 'cd ~/repo/ruby/rigor-survey/<name> && \
    BUNDLE_GEMFILE=<rigor>/Gemfile bundle exec <rigor>/exe/rigor \
    baseline generate --output=<rigor>/../rigor-survey/_<name>-sweep/baseline.yml'

Record the bucket / diagnostic count it reports — that is the sweep's zero point.

Phase 6 — Sweep the tags

Use scripts/sweep.sh — set RIGOR, TARGET, SWEEP, and TAGS at the top, then:

sh
nix develop --command \
  bash .claude/skills/rigor-regression-sweep/scripts/sweep.sh

It checks out each tag (git checkout -q -f), runs rigor check --baseline=… --no-stats --format json, and saves reports/<tag>.json (stdout) + reports/<tag>.err (stderr — carries the "N diagnostic(s) silenced by baseline" line). Long sweeps: run it backgrounded.

Phase 7 — Tabulate

Use scripts/tabulate.rb (set the same paths

  • TAGS):
sh
nix develop --command ruby \
  .claude/skills/rigor-regression-sweep/scripts/tabulate.rb

Per tag it prints raw / silenced / surfaced and the severity + rule breakdown. raw = surfaced + silenced. surfaced is the headline metric — diagnostics beyond the frozen baseline envelope, i.e. the error increase a team would have seen on that tag.

Show full SKILL.md (606 more words)Show less

Phase 8 — Interpret and record

Read the curve, then write a docs/notes/<date>-<project>-…-regression-sweep.md survey note (mirror the Mastodon one). Cover:

  • The surfaced curve — flat at 0 means ordinary development never breached the baseline (baseline stability validated); a rising curve means real regressions or churn artefacts to inspect.
  • Churn cross-check — always measure git diff --stat <first> <last> -- 'app/**/*.rb' 'lib/**/*.rb'. A flat curve is only meaningful if real files changed; report the changed-file count so surfaced = 0 cannot be mistaken for "nothing moved".
  • The rename caveat. The baseline keys on (file, rule, count). A renamed file with a baselined diagnostic shows as the old bucket going :cleared and a new (file, rule) surfacing — a surfaced > 0 that is a churn artefact, not a regression. When surfaced jumps, diff the tag and rule out renames before calling it a regression.
  • Cold-cache spot check. Re-run the last tag with --no-cache --no-baseline and confirm raw matches the swept value — proves the shared cache masked nothing.
  • The parse-error floor. A constant, non-zero surfaced at every tag — including the baseline tag itself — is the signature of rule-less diagnostics: parse errors and internal-analyzer errors carry no rule, so Baseline never buckets them and they surface forever. The usual source is a Rails generator .rb template (an ERB file with a .rb extension — <%= … %> fails to parse). Confirm by reading the rule-less rows; the fix is config-side — exclude: the generator-template directory — not a baseline concern. (Seen in the Redmine sweep: 22 parse errors from one lib/generators/.../templates/migration.rb.)
  • Feed findings back. Validated behaviour → cite in the relevant ADR. New false positives or a surprising curve → a queued engine/plugin item, or a regression spec under spec/.

Invocation gotcha

nix develop resolves the flake from the current directory. The rigor flake is in the rigor repo, the target is elsewhere — so do not cd into the target before nix develop. Run nix develop --command bash -c 'cd <target> && …' from the rigor repo, and pass BUNDLE_GEMFILE=<rigor>/Gemfile so bundle exec <rigor>/exe/rigor runs with rigor's gem environment while the target is the working directory (diagnostic paths then resolve target-relative, keeping baseline keys stable across tags).

Corpus-gating gotchas

  • A worktree isolates the ENGINE, never a PLUGIN. exe/rigor unshifts its own tree's lib/, so $worktree/exe/rigor measures that worktree's engine — but plugins/ still resolves from the main repo, so a worktree "before" run silently loads the modified plugin and the diff comes out empty. It presents as a passing gate. For a plugin change swap the directory instead: git checkout origin/master -- plugins/rigor-<name> (plus rm any file the change adds), run the before side, then git checkout HEAD -- plugins/rigor-<name>. Verify with rigor plugins, which prints each manifest's version.
  • make verify gates only lib + plugins — the external corpus still exposes false positives from receivers Rigor mistypes. Never widen a union / nilable-receiver diagnostic (or promote any default) on a clean make verify alone; run this sweep, or at minimum a before/after corpus diff, first.
  • Adjudicate against the framework's own source, not the symptom. Treat a GENUINE verdict on a hot production code path as a presumptive FP until the library's source confirms it.
  • dump_type-via-check is the ground truth — single-file dump_type probes are wrong for cross-file symbols. Analyze the whole directory.
  • Prove the path is exercised before trusting a green run. A clean corpus result proves nothing until you count what the changed layer actually did. Measure the layer, not the aggregate: a --depth 1 clone collapses every commit to one author/date, and a vendored directory inflates grep counts.
  • Survey checkouts: mise.toml needs mise trust first; git stash push -- <tracked files> (an untracked pathspec errors and stashes nothing); never run two rigor processes against one target — cache-lock contention corrupts the run.

© rigortype, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in .claude/skills/rigor-regression-sweep of rigortype/rigor.

  • SKILL.md
  • scripts/sweep.sh
  • scripts/tabulate.rb

Open the folder on GitHubat commit 57a67cf

Compare with similar skills

Rigor Regression Sweep next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rigor Regression Sweep compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rigor Regression Sweep this skillrigortype/rigor106—~2.9kAutomated safety check: PassMPL-2.0
Gumroad Prod Consoleantiwork/gumroad9.8k—~2.9kAutomated safety check: NotesMIT
Fastlane Pull Request Reviewfastlane/fastlane42k—~550Automated safety check: PassMIT
Dependency UpdaterAsvarox/allkaraoke2614 repos~3.5kAutomated safety check: PassMIT
Wise APIlineofflight/frankfurter2k—~1.1kAutomated safety check: PassMIT
Write RbsDataDog/dd-trace-rb417—~805Automated safety check: PassCustom licence

Similar skills

  • Gumroad Prod Console

    antiwork/gumroad

    Execute read-only Ruby/Rails commands against Gumroad's production database for debugging and investigation.

    9.8k GitHub stars~2.9k tokensUpdated today
    DevelopmentAuto-check: notes
  • Reviews a fastlane pull request against its linked issue and the project guides, separating blocking from non-blocking findings and handling vulnerabilities privately.

    42k GitHub stars~550 tokensUpdated today
    DevelopmentAuto-check passed
  • Dependency Updater

    Asvarox/allkaraoke

    Smart dependency management for any language. An agent skill from Asvarox/allkaraoke.

    261 GitHub starsUsed in 4 repos~3.5k tokens
    DevelopmentAuto-check passed
  • Wise API

    lineofflight/frankfurter

    A skill your agent uses when querying Wise for exchange rates (real-time or historical), validating Frankfurter rates against Wise mid-market, debugging rate discrepancies, or when the user mentions…

    2k GitHub stars~1.1k tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Write Rbs

    DataDog/dd-trace-rb

    Official

    A skill your agent uses when writing, reviewing, or modifying RBS type signatures (sig//.rbs, vendor/rbs//.rbs, or inline : annotations) or running Steep – e.g.

    417 GitHub stars~805 tokensUpdated today
    DevelopmentAuto-check passed
  • Addresses unresolved pull request review threads and suppressed (low-confidence) Copilot review comments on the current branch, folds each fix into the…

    1.8k GitHub stars~657 tokensUpdated 7 days ago
    DevelopmentAuto-check passed

More from rigortype/rigor

All 36 skills in this repo
  • Adjudicate a rigor unused report safely before proposing dead-code removal.

    106 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Rigor Baseline Reduce

    rigortype/rigor

    Reduce an existing .rigor-baseline.yml rule by rule by triaging sites, fixing or intentionally suppressing them, and regenerating the baseline.

    106 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Rigor Doctor

    rigortype/rigor

    Validate that a project's Rigor configuration, plugins, paths, and baseline are actually healthy.

    106 GitHub stars~767 tokensUpdated yesterday
    Auto-check passed
  • Rigor Plugin Author

    rigortype/rigor

    Author a new Rigor plugin, choosing plugins/ for production support or examples/ for a contract walkthrough.

    106 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check: notes
  • Rigor Plugin Author

    rigortype/rigor

    Author a Rigor plugin in an adopting project or standalone rigor- gem for a DSL, framework, or metaprogramming pattern.

    106 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Rigor Plugin Review

    rigortype/rigor

    Audit an existing Rigor plugin against the current authoring contract and produce a prioritized upgrade path.

    106 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Rigor Regression Sweep

What does Rigor Regression Sweep do?

Measure Rigor's baseline drift across the tagged history of a real OSS Ruby project. Rigor Regression Sweep is an agent skill from rigortype/rigor. Measure Rigor's baseline drift across the tagged history of a real OSS Ruby project.

When should I use Rigor Regression Sweep?

Rigor Regression Sweep fits situations like: validating diagnostic realism over a release line; growing the survey corpus; not for a single version check; ordinary project onboarding.

How do I install Rigor Regression Sweep in Claude Code?

Run `npx skills add rigortype/rigor --skill rigor-regression-sweep -a claude-code`. Or copy the skill folder (.claude/skills/rigor-regression-sweep in rigortype/rigor) into .claude/skills/rigor-regression-sweep in your project. Claude Code loads it when a task matches its description.

How do I install Rigor Regression Sweep in Codex?

Run `npx skills add rigortype/rigor --skill rigor-regression-sweep -a codex`. Or copy the skill folder (.claude/skills/rigor-regression-sweep in rigortype/rigor) into .agents/skills/rigor-regression-sweep in your project. Codex loads it when a task matches its description.

Can I use Rigor Regression Sweep in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rigortype/rigor --skill rigor-regression-sweep -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rigor-regression-sweep, .gemini/skills/rigor-regression-sweep, .github/skills/rigor-regression-sweep and .opencode/skills/rigor-regression-sweep in your project.

What does Rigor Regression Sweep need to run?

Going by SKILL.md and its folder, Rigor Regression Sweep needs a shell and Ruby for the scripts in its folder and the command-line tools its instructions call (git, nix, bash, make, bundle and mise). Our summary lists: A Bash shell.

Does Rigor Regression Sweep access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Rigor Regression Sweep safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Rigor Regression Sweep use?

Rigor Regression Sweep is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rigor Regression Sweep use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Rigor Regression Sweep?

Skills that share tags, products or a category with Rigor Regression Sweep: Gumroad Prod Console (antiwork/gumroad, 9.8k stars), Fastlane Pull Request Review (fastlane/fastlane, 42k stars), Dependency Updater (Asvarox/allkaraoke, 261 stars) and Wise API (lineofflight/frankfurter, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rigor Regression Sweep?

rigortype (a GitHub organization) maintains it in rigortype/rigor, which has 106 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on October 8, 2026.

Source: rigortype/rigor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.