Agent skill

Gating

by oaustegard in oaustegard/claude-skills

Build and audit deterministic verification gates — a check that blocks a pipeline and can be shown to go red.

MITAuto-check passedBusiness, Finance & HR

Install Gating

skills CLI
$ npx skills add oaustegard/claude-skills --skill gating -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills gating --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/gating .claude/skills/gating && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gating
GitHub stars
150
Token cost
~3.1k tokens
SKILL.md length
1,846 words
Files
7 (incl. scripts, references)
Skills in repo
69
Repo updated
First seen
Licence
MIT

At a glance

Build and audit deterministic verification gates — a check that blocks a pipeline and can be shown to go red.

  • Writing a calibration gate
  • SKILL.md covers When NOT to use this skill, The three obligations, Building a gate and Auditing an existing check suite, plus 4 more sections
  • Runs Python scripts from its folder; calls python3
  • Validation script

What it does

Gating is an agent skill from oaustegard/claude-skills. Build and audit deterministic verification gates — a check that blocks a pipeline and can be shown to go red. Use when writing a calibration gate, CI check, validation script or pre-publication check for a numeric or empirical result; when a plausible-but-wrong value would survive review; when asking whether an existing test, linter rule or check could actually fail; and when a suite passes first try, passes suspiciously often, or was written by whatever produced the thing it checks. Triggers on "can this check…

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `CHANGELOG.md`, `README.md` and `references/anchors.md`).

It sits in Business, Finance & HR, covering Performance reviews and Linting and formatting. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • Writing a calibration gate
  • Validation script
  • Pre-publication check for a numeric
  • Empirical result

Example prompts

  • “can this check fail”
  • “known-bad”
  • “negative control”
  • “/gating”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 6fc82b8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gating loads about 3.1k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 163 tokens; SKILL.md has 1,846 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~163
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 6fc82b8, republished under its MIT licence (© oaustegard). 1,846 words, ~3,144 tokens.

Download SKILL.mdSave it as .claude/skills/gating/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
gating
description
Build and audit deterministic verification gates — a check that blocks a pipeline and can be shown to go red. Use when writing a calibration gate, CI check, validation script or pre-publication check for a numeric or empirical result; when a plausible-but-wrong value would survive review; when asking whether an existing test, linter rule or check could actually fail; and when a suite passes first try, passes suspiciously often, or was written by whatever produced the thing it checks. Triggers on "can this check fail", "known-bad", "negative control", "calibration gate", "sanity check my results", "is this test actually testing anything".
metadata.version
0.3.0

gating

A gate is a check that blocks. Its only job is to go red when it should.

The characteristic failure is not a wrong check — a wrong check gets noticed. It is a check that cannot fail, which reports PASS forever and is indistinguishable from a working one from the outside. That is what makes this different from ordinary testing: the object under suspicion is the check.

When NOT to use this skill

Scope is ONE check and whether it can be made to fail.

SituationUse
Sequence several steps with branches and retriesflowing
Run the repo's existing suiterun it
Decide what to test at allthis skill has no opinion; that is design

A gate is a thing that goes red. If nothing here can go red, there is no gate to audit.

The three obligations

Every gate owes these. A gate missing any of them is not yet a gate.

1. An anchor outside your own code. Something the check compares against that your implementation did not produce: a published constant, a closed-form answer, a conservation law, a degenerate case with a known result, an independent implementation. A check that compares this run to the last run only ever tells you the code still does what it did. See references/anchors.md.

2. A known-bad it demonstrably rejects. Break the subject the way it would plausibly break, run the gate, confirm red. Until you have done this you have not shown the gate works — you have shown it runs. This is the obligation people skip, because a passing gate feels like evidence.

Two things about known-bads that are easy to get wrong:

  • Validate it at the configuration it will run in. A case tuned on a small or fast setting can stop being bad at full size. An "untrained" grid built from one Lloyd iteration was genuinely zero-gain at m=2/K=16 and earned a real +0.10 dB at m=8/K=65536, where one iteration relocates ~63,000 empty cells toward the mode. It passed the fast gate and certified nothing about the real one. This matters more than it sounds, because mutate.py needs a fast gate variant and it is tempting to validate everything there.
  • Measure its reach. One known-bad is the floor, not the goal. Name which checks it exercises (known_bad(..., covers=(...))); the harness prints the checks no known-bad reaches. An audited gate had a single known-bad covering 1 of 8 checks — and the check its whole result rested on accepted the same bad case.

3. A written statement of what it cannot catch. Coverage holes are invisible from inside a green run: the gate is silent about the thing it does not look at, in exactly the same tone it uses for the thing it looked at and approved. The author has to assert the hole; nothing else will.

scripts/gate.py enforces obligations 2 and 3 mechanically — it returns exit code 2 (INCONCLUSIVE, not PASS) when a gate registers no known-bad or no coverage limit.

Building a gate

Work in this order. The first step is the one that determines whether the rest is worth anything.

Name the wrong conclusion, not the component. Not "check the quantizer is correct" but "prevent shipping scalar wins at high bit rates when that would really be an optimizer artifact." A gate aimed at a conclusion knows what counts as a near-miss; a gate aimed at a component just exercises the code.

Find an anchor. references/anchors.md lists the kinds, in rough order of strength, with the questions that find each one.

Prefer brackets to point checks. Assert a value lies strictly between two things it cannot legitimately pass: better than a baseline, worse than a theoretical bound. A one-sided check passes for a result that collapsed as readily as for one that is right — which is how an implementation that silently does nothing gets certified.

Derive the tolerance from measured noise. Run the thing several times, see how much it moves, put the threshold outside that. A tolerance picked for comfort tends to land wider than the defect you are trying to catch, and then it swallows it.

Then check it is not too tight to mean anything. The opposite failure is real and less obvious: a margin can be statistically impeccable and practically empty. A paired estimator — scoring both arms on one shared sample so the common fluctuation cancels — is the right way to measure a difference, and its standard error shrinks as the two arms converge. So "beats the baseline by 3 se" degenerates: a codebook perturbed by N(0, 1e-3) gained +0.0001 dB against a 3-se margin of 1.2e-06 and was accepted, while real ones gained 0.35–1.41 dB. The check certified the effect is real, not the effect is worth having. Those are different assertions and need different thresholds — and the second one has to come from an anchor (there, a published lattice codebook), never from the estimator, which knows nothing about what magnitude would matter.

Build the known-bad and confirm red. Then run scripts/mutate.py for the failures you did not think of.

Wire it to a non-zero exit and run it before the thing it gates, not after. A gate that runs after the results are written is a report.

Auditing an existing check suite

Given tests, a linter config, a CI job, or a gate someone already wrote, the question is not "do these pass" but "can these fail". Full procedure in references/auditing.md; the fast version:

  • For each assertion, name a concrete input that makes it fail. If you cannot, it is decoration — delete it or fix it.
  • Check each oracle's range against the range you actually operate in. A published table that stops short of your regime is a hole with a green light on it.
  • Run scripts/mutate.py against the code the suite covers. Every survivor is a behaviour nothing checks.
  • Look for assertions whose truth does not depend on the subject at all.
  • Look for tolerances wider than the effect being measured.

Scripts

bash
# harness: refuses to report PASS without a known-bad and a coverage limit
python3 scripts/gate.py           # importable; see the module docstring

# mutation pass: which single-token changes does the gate NOT notice?
python3 scripts/mutate.py --target src/codec.py -- python3 calibrate.py
python3 scripts/mutate.py --target grids.py --max 40 -- pytest -q

mutate.py requires the gate to pass on unmutated code first and refuses to run otherwise, because survivor counts against an already-red gate mean nothing. It restores the file even on interrupt, and uses tokenize so string literals and comments are never corrupted. It is the zero-dependency pass that works against any gate command; once it stops finding survivors, mutmut or cosmic-ray go deeper on Python test suites specifically.

Show full SKILL.md (783 more words)Show less

Anti-patterns

Each of these has shipped a wrong result somewhere. They are ordered by how convincingly they impersonate a working gate.

Anti-patternWhy it survives review
Slack wider than the defectA tolerance chosen for comfort. The gate passes the real thing and the broken thing, and reports PASS for both. Derive the threshold from noise, then confirm the known-bad falls outside it.
An oracle with a coverage holePublished anchors end somewhere. If the defect is past the end of the table, the check is structurally incapable of catching it and looks fine. State the range the anchor covers.
An assertion whose truth doesn't depend on the subject"The output has at least N distinct colours" is equally true of an unchanged frame. Prefer differential checks: the state must change when it should, and a toggle applied twice must return to the byte-identical original.
Confirming the check ran, not that it can fail"Invoke it and confirm the step appears in the output" catches a check that was never wired up. It says nothing about a check that is wired up and toothless.
Comparing against your own previous outputRegenerated goldens ratify drift. If the golden came from the code under test, it is a changelog, not an oracle.
A cache keyed on the problem rather than the methodcache[(m, K)] cannot notice that the code producing the value changed. Version-stamp the artifact and delete on mismatch instead of trusting.
A self-matching predicateuntil ! pgrep -f trainer never exits, because the watching shell's own argv contains trainer. Worse, a malformed variant exits immediately and reports the job finished while it runs. Wait on a PID.
A margin that is significant but not meaningfulA threshold derived purely from estimator noise certifies that an effect is real, not that it is worth having — and a paired estimator's noise shrinks as the arms converge, so the margin can approach zero. Pair every noise-derived floor with a magnitude an anchor says would matter.
A strict bracket at an attainable optimumA theoretical bound is often reachable, and reaching it is the best possible outcome. A strict edge then goes red on a perfect result and blocks real work. Ask of each edge whether the subject can legitimately sit exactly there.
A gate written by whatever produced the artifactShared assumptions produce shared blind spots, and the convention both inherited is the one neither questions. Anchors are the defence, because an anchor is the one input the producer did not choose.

Division of labour

This skill is for results and pipelines where the failure mode is a plausible wrong number that would survive a careful read.

UseWhen the risk is
gating (this)A number or empirical result is about to be published or acted on, and a wrong-but-reasonable value would pass unnoticed. Output: a gate that blocks.
challengingAn artifact would draw a specific objection from a skeptical reader — prose, analysis, a recommendation, a diff. LLM judgement against a persona. Output: findings and a SHIP/REVISE/RETHINK verdict.
verifying-claimsDocumentation says something about code that is no longer true. Output: prose-vs-code disagreements.
A test suite / TDDCode you wrote does not behave as specified. Output: red tests.

challenging asks would a careful reader object? gating asks can this check go red? They are complements and they miss different things: an adversarial reviewer will not recompute your constants, and a gate will not notice that your framing is wrong.

A gate checks correctness, and cannot check comparability. If two arms of a comparison are each individually correct but not comparably implemented, no anchor and no mutant will see it: nothing is broken, so nothing goes red. A published ablation reported one transform 11–24× slower than another and carried a caveat saying the number was implementation-bound; both arms passed every correctness check, and the real finding — one arm was a tuned BLAS call and the other an interpreted loop — was found by a human reviewer months of gate-work later. Related: performance claims have no anchor in this framework at all. There is no published constant for how fast your code should be. Wall-clock belongs to benchmarking discipline (matched implementation effort, min-of-trials, stated hardware), not to gating.

One caution about pairing them. A same-model reviewer is an independent context, not an independent reviewer — it shares your priors, so the convention you did not question is the one it will not question either. Where that matters, an anchor beats a reviewer, because an anchor is not negotiable.

References

  • references/anchors.md — kinds of oracle, strongest first, and how to find one when nothing published exists.
  • references/auditing.md — the full "can this fail?" pass over an existing suite, including how to read a mutation report.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in gating of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • references/anchors.md
  • references/auditing.md
  • scripts/gate.py
  • scripts/mutate.py

Open the folder on GitHubat commit 6fc82b8

Compare with similar skills

Gating next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gating compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gating this skilloaustegard/claude-skills150—~3.1kAutomated safety check: PassMIT
Authoring Skillsfriday-platform/friday-studio104—~2.4kAutomated safety check: PassCustom licence
AI Indexmizchi/skills356—~3.6kAutomated safety check: PassNone
Jev Lint Repomizchi/jev-lint119—~942Automated safety check: PassMIT
System Onemagnus919/agent-skills113—~3.7kAutomated safety check: PassMIT
Nextjs React ExpertDokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI507—~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Authoring Skills

    friday-platform/friday-studio

    Authors new agent skills that follow the Anthropic + agentskills.io specification.

    104 GitHub stars~2.4k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • AI Index

    mizchi/skills

    Method and tooling for measuring how AI-generated a piece of prose reads, in Japanese or English.

    356 GitHub stars~3.6k tokensUpdated 6 days ago
    Business, Finance & HRAuto-check passed
  • Jev Lint Repo

    mizchi/jev-lint

    A skill your agent uses when changing jev-lint ITSELF — editing src/, shipped rule suites under rules/<language/<id/, or recorded runs in docs/data/.

    119 GitHub stars~942 tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • System One

    magnus919/agent-skills

    Design, integrate, evaluate, self-host, and troubleshoot typed System One decision models including TypeSafe Jev, Convai Innovations Laya, CLM, and experimental Strands Decider.

    113 GitHub stars~3.7k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Nextjs React Expert

    Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI

    React and Next.js performance optimization from Vercel Engineering.

    507 GitHub stars~1.7k tokensUpdated 3 mo ago
    Business, Finance & HRAuto-check passed
  • Deep Analysis

    nicepkg/auto-company

    Analytical thinking patterns for comprehensive evaluation, code audits, security analysis, and performance reviews.

    192 GitHub starsUsed in 2 repos~2.7k tokens
    Business, Finance & HRAuto-check: notes

More from oaustegard/claude-skills

All 69 skills in this repo
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated today
    Auto-check passed
  • Forecasting Reverso

    oaustegard/claude-skills

    Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference).

    150 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated today
    Auto-check passed

Questions about Gating

What does Gating do?

Build and audit deterministic verification gates — a check that blocks a pipeline and can be shown to go red. Gating is an agent skill from oaustegard/claude-skills. Build and audit deterministic verification gates — a check that blocks a pipeline and can be shown to go red.

When should I use Gating?

Gating fits situations like: writing a calibration gate; validation script; pre-publication check for a numeric; empirical result.

How do I install Gating in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill gating -a claude-code`. Or copy the skill folder (gating in oaustegard/claude-skills) into .claude/skills/gating in your project. Claude Code loads it when a task matches its description.

How do I install Gating in Codex?

Run `npx skills add oaustegard/claude-skills --skill gating -a codex`. Or copy the skill folder (gating in oaustegard/claude-skills) into .agents/skills/gating in your project. Codex loads it when a task matches its description.

Can I use Gating in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill gating -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gating, .gemini/skills/gating, .github/skills/gating and .opencode/skills/gating in your project.

What does Gating need to run?

Going by SKILL.md and its folder, Gating needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Gating access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gating safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gating use?

Gating is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gating use?

About 3.1k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.7k tokens, read only when the agent opens those files.

What are the alternatives to Gating?

Skills that share tags, products or a category with Gating: Authoring Skills (friday-platform/friday-studio, 104 stars), AI Index (mizchi/skills, 356 stars), Jev Lint Repo (mizchi/jev-lint, 119 stars) and System One (magnus919/agent-skills, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gating?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 69 skills in this directory. The repository was last updated on October 8, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.