Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
Evolve an agent harness through Gear using measured seed and held-out evaluations.
$ npx skills add rsi-gear/gear --skill refine -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install rsi-gear/gear refine --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/rsi-gear/gear.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/refine .claude/skills/refine && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "refine" agent skill from https://github.com/rsi-gear/gear/tree/main/skills/refine into .claude/skills/refine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "refine", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/rsi-gear/gear/tree/main/skills/refineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add rsi-gear/gear --skill refine -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install rsi-gear/gear refine --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rsi-gear/gear.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/refine .agents/skills/refine && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "refine" agent skill from https://github.com/rsi-gear/gear/tree/main/skills/refine into .agents/skills/refine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "refine", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rsi-gear/gear --skill refine -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install rsi-gear/gear refine --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rsi-gear/gear.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/refine .cursor/skills/refine && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "refine" agent skill from https://github.com/rsi-gear/gear/tree/main/skills/refine into .cursor/skills/refine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "refine", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/rsi-gear/gear.git --path skills/refine--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add rsi-gear/gear --skill refine -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install rsi-gear/gear refine --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rsi-gear/gear.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/refine .gemini/skills/refine && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "refine" agent skill from https://github.com/rsi-gear/gear/tree/main/skills/refine into .gemini/skills/refine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "refine", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install rsi-gear/gear refineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add rsi-gear/gear --skill refine -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/rsi-gear/gear.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/refine .github/skills/refine && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "refine" agent skill from https://github.com/rsi-gear/gear/tree/main/skills/refine into .github/skills/refine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "refine", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rsi-gear/gear --skill refine -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install rsi-gear/gear refine --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rsi-gear/gear.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/refine .opencode/skills/refine && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "refine" agent skill from https://github.com/rsi-gear/gear/tree/main/skills/refine into .opencode/skills/refine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "refine", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
refineEvolve an agent harness through Gear using measured seed and held-out evaluations.
Refine is an agent skill from rsi-gear/gear. Evolve an agent harness through Gear using measured seed and held-out evaluations. Use when the user asks to refine, evolve, optimize, compare, continue, inspect, publish, or roll back a Gear-managed harness from Codex, Claude Code, DSH, or another Agent Skills-compatible harness.
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/dsh-target-harness.md` and `references/protocol.md`).
It sits in AI & LLM Engineering. The repository describes itself as: Gear (General Evolution Architecture) is RSI infrastructure for AI agents. Use Gear to refine agent harnesses. The licence is MIT.
9 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 05f3cf6. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (JavaScript), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Refine loads about 2.8k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 1,447 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from rsi-gear/gear at commit 05f3cf6, republished under its MIT licence (© rsi-gear). 1,447 words, ~2,834 tokens.
.claude/skills/refine/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Use Gear as the authority for evolution state, candidate workspaces, evaluation, selection, and promotion. This skill is the Meta Agent entrypoint; the Target Agent runs separately through the rollout provider configured in Gear.
Use the native refine_request tool when it is available. It carries the same
protocol as the CLI and binds the current DSH session's client and configured
runtime/skill identity after verifying the packaged skill was loaded; pass only
the method and its ordinary parameters, with no clientId or identity.
If the bridge reports missing instructions (for example after compaction),
reload refine with DSH's native skill tool before retrying. In Code Mode,
finish the skill-loading call before making a separate Refine request.
This does not change or attest the host session's other tools, history, or OS
permissions; those remain the host's responsibility.
Otherwise require all of the following before starting or claiming work:
gear-refine executable;GEAR_REFINE_SOCKET, unless the user supplies --socket;Do not guess an identity or silently change it. Gear seals it into the evolution spec and rejects a different harness on claim or resume.
Read references/protocol.md before constructing calls or handling an active assignment. Before diagnosing evidence or changing a candidate, also read references/target-harness-editing.md. When the candidate is Gear's DSH carrier, also read references/dsh-target-harness.md before choosing an artifact or hook. These references are the complete method contract, general editing guide, and version-specific DSH authoring guide; do not infer missing field names or harness APIs from errors.
Inspect existing status before creating a new evolution when the user's wording could mean continue. A plain refine request creates a new evolution; continuation must name the evolution explicitly.
Call control.start or control.continue, then poll control.status and
meta.claim. Baseline evaluation can finish before an assignment appears.
Translate optimization preferences into control.start.objective.terms and
optional explicit constraints. For example, equal pass/process weights are
two terms with weight 0.5; cost/token penalties use negative weights and fixed
scales in their original units. Do not put a different objective only in the
prompt, invent a metric, infer passes from positive rewards, or alter the
objective during search. Omission means strict pass rate. Check the returned
resolvedObjective; inspect available raw metrics if admission rejects a term.
Treat the returned lease id, token, session id, candidate id, and (for CLI clients) client id as one inseparable capability. Never reuse them for another assignment.
Use only the candidate file and meta.call methods exposed by Gear. The
direct candidate methods are exactly candidate.tree, candidate.read,
candidate.write, candidate.edit, and candidate.remove. Invoke
harness.current, harness.read, seed_tasks.load, trajectory.query,
experience.query, experience.read, hitch.status, candidate.diff,
candidate.check, candidate.finalize, and candidate.decline only as
meta.call capabilities; they are not top-level request methods. Read the
active candidate files before selecting an edit. Do not discover or modify
Gear state, Git metadata, held-out data, credentials, or host paths directly.
Review the assignment's optional experienceContext. It contains the
actual paired seed result of the direct parent's last edit when available,
plus only relevant bounded history cards. Treat rationale and expected
outcome as the earlier proposer's claims; treat task reward changes as
observations, not causal or statistically significant conclusions. Use
experience.query and experience.read for focused history. Historical
experienceRef values never belong in evidenceRefs and never satisfy the
active baseline diagnosis requirement.
Review the baseline summary. Also read baseline.rawMetrics and
baseline.objectiveScore when present:
the raw scores, units and contributions explain the frozen objective. A
successful but expensive task can still warrant an improvement. State which
measured term the proposal is expected to improve and verify its explicit
constraints. Never replace a missing metric with zero or change the weights.
If the assignment contains workplanDelivery,
consume its hypothesis, parent identity, sourced dossier excerpt, scope guards,
generation budget, and seed-only findings first. Keep modifications within
workplan.modificationPaths (paths relative to the harness root), or explicitly
disclose the wider effect; a narrow scope does not qualify wider changes for
promotion. The claimed assignment records workplan-dossier-consumed, not
personal diagnosis of all parent failures. When
evidencePolicy.diagnoseEveryFailedRunBeforeProposal is false, the full-failure
reading loop below does not apply: read extra seed evidence only when needed,
and use candidate.check to confirm finalization readiness. Do not request or
use held-out promotion details to plan candidate changes.
For assignments without a workplan, start with trajectory.query without arguments
on every assignment: it restores validated diagnostics from earlier attempts
of this candidate and returns their summaries plus current-session detail
refs. If diagnosisRecovery.remaining is positive, repeat that query to
receive the rest. Use diagnosisProgress.remainingRunIds to choose new reads.
Recovery does not restore candidate edits; inspect the current tree. Changed
baseline, trajectory, verifier, or sanitization policy requires reading the
affected evidence again.
generationBudget reports attempt/round deadlines and remaining time. Its
finalization reserve is advisory time for editing, checking, and sealing; it
does not extend the deadline. Reconnecting does not reset time. A
DIAGNOSIS_BUDGET_AT_RISK warning means the observed pace may leave too little
time for the remaining work; do not skip required diagnoses or assume that
continuing an old evolution adopts new budget settings.
Query trajectory.query with refs for every remaining
failed baseline run before proposing a change. It returns a compact
diagnostic card containing the task, outcome, verifier failure summary,
and the chronological message transcript without raw chunk noise. The card
keeps the last 80,000 transcript characters and previews each tool result at
up to 2,000 characters. Follow earlierRef for messages before that window
and detailRef for a complete long result.
Treat result_only as missing verifier logs, not as a complete failure
explanation. When the card includes a detailRef, pass it back to
trajectory.query; pass a returned nextRef back as the next detailRef,
or add find to search that long content. Keep cited evidence limited to
references actually returned for the active seed baseline.
Start with one diagnostic card and size later reads from its actual output.
Process the evidence before fetching more; use targeted detail reads instead
of accumulating transcripts that repeatedly force context compaction.
Identify an evidenced harness gap before choosing an intervention. A failed
task or omitted check alone does not establish that gap. Use relevant
accessible seed comparisons to test the explanation, including successful
runs when useful. Choose the smallest supported mechanism from the editing
guide, with an applicability boundary that transfers beyond the observed
tasks. New files are allowed but must be wired into the existing load graph.
Prompt changes, skills, hooks, and workflows need the same causal support;
do not force any artifact type. An implementation mistake can still expose
a reusable prevention or detection opportunity; existing guidance alone
does not show that an effective procedure or check exists. Apply the editing
guide's diagnosis before either choosing a prompt edit or declining.
Record the evidence, mechanism choice,
applicability, and main uncertainty in rationale, and an observable
behavioral prediction in expectedOutcome.
Inspect candidate.diff, remove accidental or task-specific changes, and
run candidate.check before finalizing and require
finalizationReadiness.ready: true. Inspect each runtime stage: loading,
prompt assembly and Skill reads cover different paths; they do not prove
that arbitrary tool/hook bodies, routing, compaction or workflows executed.
Record remaining behavior checks in the proposal. Use candidate.decline when the
evidence does not justify a harness change or the required fix is outside
the editable substrate.
After finalization, stop using that lease and poll status. Claim and complete every subsequent candidate or round until the requested batch reaches a terminal state.
Do not call control.publish or control.rollback unless the user explicitly
requests that state change. Automatic per-evolution promotion remains governed
by Gear's sealed promotion policy.
control.rerun only for repairable evaluation slots reported by status.searchPendingEvidence, control.search-repair takes
evolutionId, roundId, a stable repairId, and evidenceDigest. It repairs
original invalid slots and resumes the same round; it is not a request for a
new candidate or a new experiment. Leave held-out repair to the operator flow,
and do not feed its results into a candidate lease.accepted:false, recoverable:true, execute
nextAction, then remainingActions, and retry the same operation with the
same arguments. The lease remains active until accepted:true.accepted:false, recoverable:false, report the exact
operatorAction and do not repeat the same tool call. Verifier evidence may
require an operator/configuration change. TRAJECTORY_EVIDENCE_UNAVAILABLE
means Hitch could not construct bounded analysis for the listed blockedRuns;
Hitch or that stored trajectory must be repaired before rereading the card.© rsi-gear, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/refine of rsi-gear/gear.
Open the folder on GitHubat commit 05f3cf6
Refine next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Refine this skillrsi-gear/gear | 139 | — | ~2.8k | Automated safety check: Pass | MIT | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 4 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.9k | 14 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
jarrodwatts/claude-code-config
Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.
Categories
Evolve an agent harness through Gear using measured seed and held-out evaluations. Refine is an agent skill from rsi-gear/gear. Evolve an agent harness through Gear using measured seed and held-out evaluations.
Refine fits situations like: the user asks to refine; roll back a Gear-managed harness from Codex; another Agent Skills-compatible harness.
Run `npx skills add rsi-gear/gear --skill refine -a claude-code`. Or copy the skill folder (skills/refine in rsi-gear/gear) into .claude/skills/refine in your project. Claude Code loads it when a task matches its description.
Run `npx skills add rsi-gear/gear --skill refine -a codex`. Or copy the skill folder (skills/refine in rsi-gear/gear) into .agents/skills/refine in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rsi-gear/gear --skill refine -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/refine, .gemini/skills/refine, .github/skills/refine and .opencode/skills/refine in your project.
Going by SKILL.md and its folder, Refine needs JavaScript for the scripts in its folder. Our summary lists: Node.js.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Refine is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Refine: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
rsi-gear (a GitHub organization) maintains it in rsi-gear/gear, which has 139 GitHub stars. The repository was last updated on October 10, 2026.
Source: rsi-gear/gear on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.