Agent skill

Skill Autobench

by garrytan in garrytan/gbrain

Author an eval for an existing skill from its REAL usage history, not its spec.

MITAuto-check passedAI & LLM Engineering

Install Skill Autobench

skills CLI
$ npx skills add garrytan/gbrain --skill skill-autobench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install garrytan/gbrain skill-autobench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/garrytan/gbrain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skill-autobench .claude/skills/skill-autobench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-autobench
GitHub stars
31k
Token cost
~3.2k tokens
SKILL.md length
1,397 words
Files
2
Skills in repo
47
Repo updated
First seen
Licence
MIT

At a glance

Author an eval for an existing skill from its REAL usage history, not its spec.

  • Works in 3 steps: MINE — extract real invocation windows → SYNTH — turn windows into a proposed eval → STAGE — human gate, always
  • AI & LLM Engineering work in your project
  • SKILL.md covers Pipeline, The loop (after approval), Panel integrity — trust no… and Fail-improve taxonomy — what…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Skill Autobench is an agent skill from garrytan/gbrain. Author an eval for an existing skill from its REAL usage history, not its spec. Mine invocations from the brain's conversation archive (conversations/) and per-harness session transcripts — a user correction after an invocation is the gold signal — then synthesize an evalcontract plus 4-8 replayable cases with honesty labels (SPEC-DERIVED vs HISTORY-IMPLIED) and stage the result at skills/<name/eval/autobench-<date.md as PENDING-HUMAN-APPROVAL. Never rewrites SKILL.md. Ships two guard companions: panel integrity…

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in AI & LLM Engineering. The repository describes itself as: Garry's Opinionated OpenClaw/Hermes Agent Brain. The licence is MIT.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/skill-autobench”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. MINE — extract real invocation windows
  2. SYNTH — turn windows into a proposed eval
  3. STAGE — human gate, always

What it can do on your machine

Read from SKILL.md and the folder at commit f250a51. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Autobench loads about 3.2k tokens when it runs. Until then it costs about 175 tokens; SKILL.md has 1,397 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~175
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from garrytan/gbrain at commit f250a51, republished under its MIT licence (© garrytan). 1,397 words, ~3,158 tokens.

Download SKILL.mdSave it as .claude/skills/skill-autobench/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
skill-autobench
description
Author an eval for an existing skill from its REAL usage history, not its spec. Mine invocations from the brain's conversation archive (conversations/) and per-harness session transcripts — a user correction after an invocation is the gold signal — then synthesize an eval_contract plus 4-8 replayable cases with honesty labels (SPEC-DERIVED vs HISTORY-IMPLIED) and stage the result at skills/<name>/eval/autobench-<date>.md as PENDING-HUMAN-APPROVAL. Never rewrites SKILL.md. Ships two guard companions: panel integrity (multi-model judging must prove each provider actually responded) and the fail-improve taxonomy (logged LLM-fallback cases convert to deterministic code over time).
version
1.0.0
triggers
skill autobench, autobench, write the eval from usage history, synthesize an eval for this skill, mine how this skill is actually used, build a benchmark from…
requires
dir:conversations/
mutating
true
writes_pages
false
writes_to
skills/<name>/eval/
upstream
skill-autobench@fc834ee + panel-integrity@fc834ee + fail-improve-loop@fc834ee (taxonomy only)

skill-autobench — write the eval from lived usage

Convention: see conventions/brain-first.md — mining starts in the brain. Search the conversation archive before touching raw transcript files, and never declare "no history" without having queried the brain first.

Convention: see conventions/model-routing.md — mining and synthesis run on the cheap tier by default. The full multi-model judging pass is an explicit opt-in (see Contract).

The self-improving loop has three legs: an eval, a variant generator (SkillOpt), and a replay + judge harness (gbrain eval cross-modal). The generator and the judge ship with gbrain. The persistently missing leg is the eval author — someone has to WRITE the eval, and a spec-derived benchmark only tests what the skill promised, not what users actually asked for or what actually went wrong. This skill writes the eval from reality instead of imagination.

Pipeline

1. MINE — extract real invocation windows

Substrates, in priority order:

  1. Brain conversation archive — pages under conversations/, populated by the conversation-archive skill (hard dependency for this substrate: if it hasn't ingested your history yet, run it first). Search for the target skill's name, trigger phrases, and output shapes:

    bash
    gbrain search "<skill-name>"
    gbrain query "when did I use <skill-name> and what did I ask for"
  2. Per-harness session transcripts, as available — use what the harness exposes; do not assume a layout. gbrain transcripts recent --full reads the configured local transcript corpus (local-only by design). Claude Code keeps per-project session JSONL under ~/.claude/projects/; other harnesses have their own session stores. Absent stores are simply skipped.

From each hit, extract an invocation window: the user ask before the invocation, the invocation turn itself, and the 2 turns after — because that is where corrections live. A user correction after an invocation is the gold signal: it is a real, observed failure mode, and it becomes a hard_fail plus a replayable case.

FAIL-CLOSED: if no substrate yields a single real invocation of the target skill, emit an honest no-history report (substrates checked, queries run, windows scanned, zero matches) and stop. Do NOT invent "typical" invocations. For a skill with no history, the right tool is gbrain skillopt <name> --bootstrap-from-skill (spec-derived, and honest about it) — see Dedup.

2. SYNTH — turn windows into a proposed eval

From the mined windows plus the current SKILL.md, produce:

  • A proposed eval_contract: goal, dimensions, hard_fails. Dimensions come from observed asks; hard_fails encode observed corrections.
  • 4-8 replayable cases, each shaped {input, expected_behavior, failure_mode_to_catch} — realistic input, a checkable expected behavior, and the named failure mode the case exists to catch.
  • Spec-vs-usage gaps: "the spec says X, users consistently ask Y."

HONESTY LABELS are mandatory. Every dimension and every case is labeled:

  • HISTORY-IMPLIED — a real mined window backs it; cite which one.
  • SPEC-DERIVED — inferred from SKILL.md only; no usage evidence.

Never conflate the two. If history is thin or off-target, say so prominently at the top of the staged file ("GROUNDING WARNING: only N windows found, none exercised the core path") instead of padding with fabricated evidence.

Privacy scrub before staging: staged evals live in the skill repo and are distributable. Mined windows contain real names, companies, and deals — rewrite every case onto placeholder slugs (alice-example, acme-example) before writing the file. A history-grounded case keeps its shape and failure mode, never its real entities.

3. STAGE — human gate, always

Write skills/<name>/eval/autobench-<date>.md with frontmatter status: PENDING-HUMAN-APPROVAL.

This skill NEVER rewrites SKILL.md — not the eval_contract, not the body, not the triggers. Merging the staged eval is the human's decision. (This is a workflow contract the agent must honor, not a mechanically-enforced gate.)

The loop (after approval)

  1. Human reviews, edits, and approves the staged eval; the approved eval_contract is merged into the skill's frontmatter explicitly.

  2. Convert approved cases into skills/<name>/skillopt-benchmark.jsonl lines and run gbrain skillopt <name> — this is the SkillOpt surface extension: a history-grounded benchmark replacing the spec-derived bootstrap.

  3. Judge outputs through the native gate:

    bash
    gbrain eval cross-modal --task "<what the output was meant to achieve>" --output <path>
  4. Re-run autobench after more usage accumulates; diff against the prior staged baseline.

Panel integrity — trust no aggregate

Multi-model judging is only as good as the panel being real. The silent failure class: a model-id normalizer strips provider prefixes and every "different model" call lands on one host, so a "3-frontier consensus" is one model's opinion in a trench coat. Before trusting any multi-model verdict, assert over the result object:

  1. Each named model returned a non-empty response. An empty or failed slot is the first tell of a collapse.
  2. Responses came from DISTINCT provider endpoints. If three "different models" all report the same provider, the panel collapsed to one host.
  3. No two "different models" returned byte-identical output. If two differently-named models return the same bytes, they are the same model. This catches a collapse even when provider metadata is missing or faked.

gbrain eval cross-modal already exits 2 (INCONCLUSIVE) when fewer than 2/3 models return parseable scores; the byte-identical duplicate check and the distinct-endpoint check are the independent backstops this skill layers on top. Run them over the receipt JSON (written to the receipt dir) before treating a PASS/FAIL as authoritative. The integrity check is pure assertion logic over an existing result — it never calls a model itself: no network, no cost.

Show full SKILL.md (564 more words)Show less

Fail-improve taxonomy — what mined failures become

Classify each mined correction/failure case by its cheapest durable fix:

ClassSignal in historyDurable fix
DETERMINISTIC-CODIFIABLEAn LLM fallback repeatedly handles the same input shape (regex, parsing, slugs, dates)Convert to deterministic code + a permanent test case. The LLM is not the solution; it is the training-data generator for the code that replaces it.
PROMPT-FIXABLEThe correction targets tone, format, or an omission the SKILL.md could specifyAn eval case + a gbrain skillopt run
SPEC-GAPUsers consistently ask for something the spec never promisedA spec-vs-usage gap observation for the human
ROUTING-MISSThe skill fired on the wrong ask, or failed to fireA routing-eval.jsonl case, not a benchmark case

Direction of travel: every fixed failure becomes a permanent test, the deterministic share rises, and the LLM-fallback share falls. Log and improve; never silently drop a mined failure.

Contract

  • Input: a skill name that exists under skills/.
  • Substrates: conversations/ archive pages (brain-first), then per-harness session transcripts as available (gbrain transcripts recent is local-only by design). No usable substrate → honest no-history report, never fabricated evidence.
  • Output: one staged file at skills/<name>/eval/autobench-<date>.md, status: PENDING-HUMAN-APPROVAL — or the no-history report. Never an edit to SKILL.md, triggers, or any routing surface.
  • Cost posture: cheap-model default for mining and synthesis. The full multi-model judging pass (3 provider slots per cycle) is an explicit opt-in, and judging goes through native gbrain eval cross-modal — no bespoke judging harness.
  • Honesty: every dimension and case carries a SPEC-DERIVED or HISTORY-IMPLIED label; thin history is flagged, not papered over.
  • Privacy: mined cases are rewritten onto placeholder entities before staging.

Output Format

markdown
---
skill: <name>
status: PENDING-HUMAN-APPROVAL
generated: <date>
substrate: { conversation_pages: N, transcript_files: M, windows: K, corrections: C }
---

# Autobench: <name> — <date>

## Grounding
<one paragraph: how much real history backs this eval; GROUNDING WARNING if thin>

## Proposed eval_contract
goal / dimensions (each labeled HISTORY-IMPLIED|SPEC-DERIVED) / hard_fails

## Cases (4-8)
### case-01 [HISTORY-IMPLIED — window ref]
input: ...
expected_behavior: ...
failure_mode_to_catch: ...

## Spec-vs-usage gaps
- spec says X; users ask Y (windows: ...)

## Fail-improve classification
- case-03 → DETERMINISTIC-CODIFIABLE (same date-format fallback, 4 windows)

When it fails

Follow the agent operator protocol for any gbrain error code, exit code, [AGENT] block or notice block. Specific to this skill:

  • No substrate yields a real invocation of the skill: FAIL-CLOSED. Report it and write no eval; never synthesize cases from imagination.
  • gbrain eval cross-modal or the judge run lacks a key or hits no_pricing: say the eval was not judged; leave the draft as PENDING-HUMAN-APPROVAL.

Anti-Patterns

  • ❌ Auto-merging a synthesized eval into SKILL.md. The human gate is the contract.
  • ❌ Presenting SPEC-DERIVED dimensions as history-grounded — fabricating usage evidence is the cardinal sin.
  • ❌ Mining nothing and still emitting a confident eval. Fail loudly or label honestly.
  • ❌ Trusting "3 models scored it 8/10" without checking that three providers actually returned distinct, non-identical responses.
  • ❌ Calling a model inside the panel-integrity check — it is pure assertion logic over a result object.
  • ❌ Building a bespoke judging harness when gbrain eval cross-modal is the native gate.
  • ❌ Staging mined cases with real people/companies in them. Placeholders only.

Dedup (sharp boundaries)

  • skill-optimizer (SkillOpt, host-side) — optimizes a skill's body against an EXISTING benchmark; its --bootstrap-from-skill derives tasks from the spec. THIS skill authors the benchmark from lived usage and feeds it into skills/<name>/skillopt-benchmark.jsonl — it extends SkillOpt's surface, never duplicates it. Read the skill-optimizer SKILL.md (on the host, where its engine lives) before the handoff. No history at all → use --bootstrap-from-skill, not this.
  • BrainBench (gbrain eval suites) — evals the ENGINE (retrieval, memory conformance, calibration). This evals SKILLS over their history.
  • skills/skillify/SKILL.md / skills/skill-creator/SKILL.md — create skills from descriptions; they don't mine lived usage.
  • skills/cross-modal-review/SKILL.md / gbrain eval cross-modal — RUN judging panels; they don't author evals. The panel-integrity assertions here verify their panels were real.
  • skills/skillpack-check/SKILL.md — audits skill structure/conformance, not behavior quality.
  • routing-eval.jsonl — tests dispatch (does the right skill fire); autobench tests behavior after dispatch. ROUTING-MISS findings route there.

© garrytan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/skill-autobench of garrytan/gbrain.

  • SKILL.md
  • routing-eval.jsonl

Open the folder on GitHubat commit f250a51

Compare with similar skills

Skill Autobench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Autobench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Autobench this skillgarrytan/gbrain31k—~3.2kAutomated safety check: PassMIT
Agent BuildershareAI-lab/learn-claude-code78k4 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.9k14 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 4 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.9k GitHub starsUsed in 14 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed

More from garrytan/gbrain

All 47 skills in this repo
  • Traces a factual error the user points out back to its source (a brain page, a memory file, SOUL.md or USER.md, or a hallucination) and fixes that source instead of just noting the correction.

    31k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Searches and writes a company-wide knowledge brain through the gbrain CLI, so durable decisions and facts about people, projects and history stay findable beyond one session.

    31k GitHub stars~875 tokensUpdated today
    Auto-check passed
  • Idea Ingest

    garrytan/gbrain

    Ingest links, articles, tweets, and ideas into the brain. An agent skill from garrytan/gbrain.

    31k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Sends what your notes already know about a topic to Perplexity, so the cited web search reports only what is new, such as entity updates or deal changes.

    31k GitHub stars~2k tokensUpdated today
    Auto-check: notes
  • Schema Unify

    garrytan/gbrain

    Migrate a brain from gbrain-base (or any pack) to gbrain-base-v2's 14-canonical-type taxonomy via gbrain onboard --check + the unify-types Minion handler.

    31k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Skillpack Check

    garrytan/gbrain

    Run gbrain skillpack-check to produce an agent-readable JSON health report for the gbrain install.

    31k GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Skill Autobench

What does Skill Autobench do?

Author an eval for an existing skill from its REAL usage history, not its spec. Skill Autobench is an agent skill from garrytan/gbrain. Author an eval for an existing skill from its REAL usage history, not its spec.

When should I use Skill Autobench?

Skill Autobench fits situations like: AI & LLM Engineering work in your project.

How do I install Skill Autobench in Claude Code?

Run `npx skills add garrytan/gbrain --skill skill-autobench -a claude-code`. Or copy the skill folder (skills/skill-autobench in garrytan/gbrain) into .claude/skills/skill-autobench in your project. Claude Code loads it when a task matches its description.

How do I install Skill Autobench in Codex?

Run `npx skills add garrytan/gbrain --skill skill-autobench -a codex`. Or copy the skill folder (skills/skill-autobench in garrytan/gbrain) into .agents/skills/skill-autobench in your project. Codex loads it when a task matches its description.

Can I use Skill Autobench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add garrytan/gbrain --skill skill-autobench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-autobench, .gemini/skills/skill-autobench, .github/skills/skill-autobench and .opencode/skills/skill-autobench in your project.

What does Skill Autobench need to run?

SKILL.md names no scripts, command-line tools or credentials: Skill Autobench is instructions for the agent only.

Does Skill Autobench access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skill Autobench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skill Autobench use?

Skill Autobench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Autobench use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skill Autobench?

Skills that share tags, products or a category with Skill Autobench: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Autobench?

garrytan (a GitHub user) maintains it in garrytan/gbrain, which has 30,736 GitHub stars. The repository holds 47 skills in this directory. The repository was last updated on October 10, 2026.

Source: garrytan/gbrain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.