Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
Author an eval for an existing skill from its REAL usage history, not its spec.
$ npx skills add garrytan/gbrain --skill skill-autobench -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install garrytan/gbrain skill-autobench --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/garrytan/gbrain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skill-autobench .claude/skills/skill-autobench && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skill-autobench" agent skill from https://github.com/garrytan/gbrain/tree/master/skills/skill-autobench into .claude/skills/skill-autobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-autobench", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/garrytan/gbrain/tree/master/skills/skill-autobenchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add garrytan/gbrain --skill skill-autobench -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install garrytan/gbrain skill-autobench --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/garrytan/gbrain.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/skill-autobench .agents/skills/skill-autobench && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skill-autobench" agent skill from https://github.com/garrytan/gbrain/tree/master/skills/skill-autobench into .agents/skills/skill-autobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-autobench", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add garrytan/gbrain --skill skill-autobench -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install garrytan/gbrain skill-autobench --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/garrytan/gbrain.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/skill-autobench .cursor/skills/skill-autobench && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skill-autobench" agent skill from https://github.com/garrytan/gbrain/tree/master/skills/skill-autobench into .cursor/skills/skill-autobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-autobench", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/garrytan/gbrain.git --path skills/skill-autobench--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add garrytan/gbrain --skill skill-autobench -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install garrytan/gbrain skill-autobench --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/garrytan/gbrain.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/skill-autobench .gemini/skills/skill-autobench && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skill-autobench" agent skill from https://github.com/garrytan/gbrain/tree/master/skills/skill-autobench into .gemini/skills/skill-autobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-autobench", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install garrytan/gbrain skill-autobenchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add garrytan/gbrain --skill skill-autobench -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/garrytan/gbrain.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/skill-autobench .github/skills/skill-autobench && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skill-autobench" agent skill from https://github.com/garrytan/gbrain/tree/master/skills/skill-autobench into .github/skills/skill-autobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-autobench", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add garrytan/gbrain --skill skill-autobench -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install garrytan/gbrain skill-autobench --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/garrytan/gbrain.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/skill-autobench .opencode/skills/skill-autobench && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skill-autobench" agent skill from https://github.com/garrytan/gbrain/tree/master/skills/skill-autobench into .opencode/skills/skill-autobench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-autobench", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skill-autobenchAuthor an eval for an existing skill from its REAL usage history, not its spec.
Skill Autobench is an agent skill from garrytan/gbrain. Author an eval for an existing skill from its REAL usage history, not its spec. Mine invocations from the brain's conversation archive (conversations/) and per-harness session transcripts — a user correction after an invocation is the gold signal — then synthesize an evalcontract plus 4-8 replayable cases with honesty labels (SPEC-DERIVED vs HISTORY-IMPLIED) and stage the result at skills/<name/eval/autobench-<date.md as PENDING-HUMAN-APPROVAL. Never rewrites SKILL.md. Ships two guard companions: panel integrity…
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.
It sits in AI & LLM Engineering. The repository describes itself as: Garry's Opinionated OpenClaw/Hermes Agent Brain. The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit f250a51. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skill Autobench loads about 3.2k tokens when it runs. Until then it costs about 175 tokens; SKILL.md has 1,397 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from garrytan/gbrain at commit f250a51, republished under its MIT licence (© garrytan). 1,397 words, ~3,158 tokens.
.claude/skills/skill-autobench/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Convention: see conventions/brain-first.md — mining starts in the brain. Search the conversation archive before touching raw transcript files, and never declare "no history" without having queried the brain first.
Convention: see conventions/model-routing.md — mining and synthesis run on the cheap tier by default. The full multi-model judging pass is an explicit opt-in (see Contract).
The self-improving loop has three legs: an eval, a variant generator
(SkillOpt), and a replay + judge harness (gbrain eval cross-modal). The
generator and the judge ship with gbrain. The persistently missing leg is the
eval author — someone has to WRITE the eval, and a spec-derived benchmark
only tests what the skill promised, not what users actually asked for or what
actually went wrong. This skill writes the eval from reality instead of
imagination.
Substrates, in priority order:
Brain conversation archive — pages under conversations/, populated by
the conversation-archive skill (hard dependency for this substrate: if it
hasn't ingested your history yet, run it first). Search for the target
skill's name, trigger phrases, and output shapes:
gbrain search "<skill-name>"
gbrain query "when did I use <skill-name> and what did I ask for"Per-harness session transcripts, as available — use what the harness
exposes; do not assume a layout. gbrain transcripts recent --full reads
the configured local transcript corpus (local-only by design). Claude Code
keeps per-project session JSONL under ~/.claude/projects/; other
harnesses have their own session stores. Absent stores are simply skipped.
From each hit, extract an invocation window: the user ask before the
invocation, the invocation turn itself, and the 2 turns after — because that
is where corrections live. A user correction after an invocation is the
gold signal: it is a real, observed failure mode, and it becomes a
hard_fail plus a replayable case.
FAIL-CLOSED: if no substrate yields a single real invocation of the
target skill, emit an honest no-history report (substrates checked, queries
run, windows scanned, zero matches) and stop. Do NOT invent "typical"
invocations. For a skill with no history, the right tool is
gbrain skillopt <name> --bootstrap-from-skill (spec-derived, and honest
about it) — see Dedup.
From the mined windows plus the current SKILL.md, produce:
{input, expected_behavior, failure_mode_to_catch} — realistic input, a
checkable expected behavior, and the named failure mode the case exists to
catch.HONESTY LABELS are mandatory. Every dimension and every case is labeled:
HISTORY-IMPLIED — a real mined window backs it; cite which one.SPEC-DERIVED — inferred from SKILL.md only; no usage evidence.Never conflate the two. If history is thin or off-target, say so prominently at the top of the staged file ("GROUNDING WARNING: only N windows found, none exercised the core path") instead of padding with fabricated evidence.
Privacy scrub before staging: staged evals live in the skill repo and are
distributable. Mined windows contain real names, companies, and deals —
rewrite every case onto placeholder slugs (alice-example, acme-example)
before writing the file. A history-grounded case keeps its shape and failure
mode, never its real entities.
Write skills/<name>/eval/autobench-<date>.md with frontmatter
status: PENDING-HUMAN-APPROVAL.
This skill NEVER rewrites SKILL.md — not the eval_contract, not the body, not the triggers. Merging the staged eval is the human's decision. (This is a workflow contract the agent must honor, not a mechanically-enforced gate.)
Human reviews, edits, and approves the staged eval; the approved eval_contract is merged into the skill's frontmatter explicitly.
Convert approved cases into skills/<name>/skillopt-benchmark.jsonl lines
and run gbrain skillopt <name> — this is the SkillOpt surface extension:
a history-grounded benchmark replacing the spec-derived bootstrap.
Judge outputs through the native gate:
gbrain eval cross-modal --task "<what the output was meant to achieve>" --output <path>Re-run autobench after more usage accumulates; diff against the prior staged baseline.
Multi-model judging is only as good as the panel being real. The silent failure class: a model-id normalizer strips provider prefixes and every "different model" call lands on one host, so a "3-frontier consensus" is one model's opinion in a trench coat. Before trusting any multi-model verdict, assert over the result object:
gbrain eval cross-modal already exits 2 (INCONCLUSIVE) when fewer than 2/3
models return parseable scores; the byte-identical duplicate check and the
distinct-endpoint check are the independent backstops this skill layers on
top. Run them over the receipt JSON (written to the receipt dir) before
treating a PASS/FAIL as authoritative. The integrity check is pure assertion
logic over an existing result — it never calls a model itself: no network,
no cost.
Classify each mined correction/failure case by its cheapest durable fix:
| Class | Signal in history | Durable fix |
|---|---|---|
| DETERMINISTIC-CODIFIABLE | An LLM fallback repeatedly handles the same input shape (regex, parsing, slugs, dates) | Convert to deterministic code + a permanent test case. The LLM is not the solution; it is the training-data generator for the code that replaces it. |
| PROMPT-FIXABLE | The correction targets tone, format, or an omission the SKILL.md could specify | An eval case + a gbrain skillopt run |
| SPEC-GAP | Users consistently ask for something the spec never promised | A spec-vs-usage gap observation for the human |
| ROUTING-MISS | The skill fired on the wrong ask, or failed to fire | A routing-eval.jsonl case, not a benchmark case |
Direction of travel: every fixed failure becomes a permanent test, the deterministic share rises, and the LLM-fallback share falls. Log and improve; never silently drop a mined failure.
skills/.conversations/ archive pages (brain-first), then
per-harness session transcripts as available (gbrain transcripts recent
is local-only by design). No usable substrate → honest no-history report,
never fabricated evidence.skills/<name>/eval/autobench-<date>.md,
status: PENDING-HUMAN-APPROVAL — or the no-history report. Never an edit
to SKILL.md, triggers, or any routing surface.gbrain eval cross-modal — no
bespoke judging harness.---
skill: <name>
status: PENDING-HUMAN-APPROVAL
generated: <date>
substrate: { conversation_pages: N, transcript_files: M, windows: K, corrections: C }
---
# Autobench: <name> — <date>
## Grounding
<one paragraph: how much real history backs this eval; GROUNDING WARNING if thin>
## Proposed eval_contract
goal / dimensions (each labeled HISTORY-IMPLIED|SPEC-DERIVED) / hard_fails
## Cases (4-8)
### case-01 [HISTORY-IMPLIED — window ref]
input: ...
expected_behavior: ...
failure_mode_to_catch: ...
## Spec-vs-usage gaps
- spec says X; users ask Y (windows: ...)
## Fail-improve classification
- case-03 → DETERMINISTIC-CODIFIABLE (same date-format fallback, 4 windows)Follow the agent operator protocol for any gbrain error code, exit code, [AGENT] block or notice block. Specific to this skill:
gbrain eval cross-modal or the judge run lacks a key or hits no_pricing: say the eval was not judged; leave the draft as PENDING-HUMAN-APPROVAL.gbrain eval cross-modal is the
native gate.--bootstrap-from-skill derives tasks
from the spec. THIS skill authors the benchmark from lived usage and feeds
it into skills/<name>/skillopt-benchmark.jsonl — it extends SkillOpt's
surface, never duplicates it. Read the skill-optimizer SKILL.md (on the
host, where its engine lives) before the handoff. No history at all → use
--bootstrap-from-skill, not this.gbrain eval suites) — evals the ENGINE (retrieval,
memory conformance, calibration). This evals SKILLS over their history.skills/skillify/SKILL.md / skills/skill-creator/SKILL.md — create
skills from descriptions; they don't mine lived usage.skills/cross-modal-review/SKILL.md / gbrain eval cross-modal — RUN
judging panels; they don't author evals. The panel-integrity assertions
here verify their panels were real.skills/skillpack-check/SKILL.md — audits skill structure/conformance,
not behavior quality.routing-eval.jsonl — tests dispatch (does the right skill fire);
autobench tests behavior after dispatch. ROUTING-MISS findings route there.© garrytan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/skill-autobench of garrytan/gbrain.
Open the folder on GitHubat commit f250a51
Skill Autobench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skill Autobench this skillgarrytan/gbrain | 31k | — | ~3.2k | Automated safety check: Pass | MIT | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 4 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.9k | 14 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
jarrodwatts/claude-code-config
Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.
garrytan/gbrain
Traces a factual error the user points out back to its source (a brain page, a memory file, SOUL.md or USER.md, or a hallucination) and fixes that source instead of just noting the correction.
garrytan/gbrain
Searches and writes a company-wide knowledge brain through the gbrain CLI, so durable decisions and facts about people, projects and history stay findable beyond one session.
garrytan/gbrain
Ingest links, articles, tweets, and ideas into the brain. An agent skill from garrytan/gbrain.
garrytan/gbrain
Sends what your notes already know about a topic to Perplexity, so the cited web search reports only what is new, such as entity updates or deal changes.
garrytan/gbrain
Migrate a brain from gbrain-base (or any pack) to gbrain-base-v2's 14-canonical-type taxonomy via gbrain onboard --check + the unify-types Minion handler.
garrytan/gbrain
Run gbrain skillpack-check to produce an agent-readable JSON health report for the gbrain install.
Categories
Author an eval for an existing skill from its REAL usage history, not its spec. Skill Autobench is an agent skill from garrytan/gbrain. Author an eval for an existing skill from its REAL usage history, not its spec.
Skill Autobench fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add garrytan/gbrain --skill skill-autobench -a claude-code`. Or copy the skill folder (skills/skill-autobench in garrytan/gbrain) into .claude/skills/skill-autobench in your project. Claude Code loads it when a task matches its description.
Run `npx skills add garrytan/gbrain --skill skill-autobench -a codex`. Or copy the skill folder (skills/skill-autobench in garrytan/gbrain) into .agents/skills/skill-autobench in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add garrytan/gbrain --skill skill-autobench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-autobench, .gemini/skills/skill-autobench, .github/skills/skill-autobench and .opencode/skills/skill-autobench in your project.
SKILL.md names no scripts, command-line tools or credentials: Skill Autobench is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Skill Autobench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Skill Autobench: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
garrytan (a GitHub user) maintains it in garrytan/gbrain, which has 30,736 GitHub stars. The repository holds 47 skills in this directory. The repository was last updated on October 10, 2026.
Source: garrytan/gbrain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.