Agent skill

Refine

by rsi-gear in rsi-gear/gear

Evolve an agent harness through Gear using measured seed and held-out evaluations.

MITAuto-check passedAI & LLM Engineering

Install Refine

skills CLI
$ npx skills add rsi-gear/gear --skill refine -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rsi-gear/gear refine --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rsi-gear/gear.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/refine .claude/skills/refine && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
refine
GitHub stars
139
Token cost
~2.8k tokens
SKILL.md length
1,447 words
Files
6 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Evolve an agent harness through Gear using measured seed and held-out evaluations.

  • Works in 9 steps: Inspect existing status before creating… → Call control.start or control.continue,… → Treat the returned lease id, token,… → …
  • The user asks to refine
  • SKILL.md covers Connect, Run an evolution and Failure handling
  • Runs JavaScript scripts from its folder

What it does

Refine is an agent skill from rsi-gear/gear. Evolve an agent harness through Gear using measured seed and held-out evaluations. Use when the user asks to refine, evolve, optimize, compare, continue, inspect, publish, or roll back a Gear-managed harness from Codex, Claude Code, DSH, or another Agent Skills-compatible harness.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/dsh-target-harness.md` and `references/protocol.md`).

It sits in AI & LLM Engineering. The repository describes itself as: Gear (General Evolution Architecture) is RSI infrastructure for AI agents. Use Gear to refine agent harnesses. The licence is MIT.

When your agent uses it

  • The user asks to refine
  • Roll back a Gear-managed harness from Codex
  • Another Agent Skills-compatible harness

Example prompts

  • “/refine”

Requirements

  • Node.js

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Inspect existing status before creating a new evolution when the user's
  2. Call control.start or control.continue, then poll control.status and
  3. Treat the returned lease id, token, session id, candidate id, and (for CLI
  4. Use only the candidate file and meta.call methods exposed by Gear. The
  5. Review the assignment's optional experienceContext. It contains the
  6. Review the baseline summary. Also read baseline.rawMetrics and
  7. Identify an evidenced harness gap before choosing an intervention. A failed
  8. Inspect candidate.diff, remove accidental or task-specific changes, and
  9. After finalization, stop using that lease and poll status. Claim and complete

What it can do on your machine

Read from SKILL.md and the folder at commit 05f3cf6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Refine loads about 2.8k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 1,447 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~23k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from rsi-gear/gear at commit 05f3cf6, republished under its MIT licence (© rsi-gear). 1,447 words, ~2,834 tokens.

Download SKILL.mdSave it as .claude/skills/refine/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
refine
description
Evolve an agent harness through Gear using measured seed and held-out evaluations. Use when the user asks to refine, evolve, optimize, compare, continue, inspect, publish, or roll back a Gear-managed harness from Codex, Claude Code, DSH, or another Agent Skills-compatible harness.

Refine

Use Gear as the authority for evolution state, candidate workspaces, evaluation, selection, and promotion. This skill is the Meta Agent entrypoint; the Target Agent runs separately through the rollout provider configured in Gear.

Connect

Use the native refine_request tool when it is available. It carries the same protocol as the CLI and binds the current DSH session's client and configured runtime/skill identity after verifying the packaged skill was loaded; pass only the method and its ordinary parameters, with no clientId or identity. If the bridge reports missing instructions (for example after compaction), reload refine with DSH's native skill tool before retrying. In Code Mode, finish the skill-loading call before making a separate Refine request. This does not change or attest the host session's other tools, history, or OS permissions; those remain the host's responsibility.

Otherwise require all of the following before starting or claiming work:

  • the gear-refine executable;
  • GEAR_REFINE_SOCKET, unless the user supplies --socket;
  • a stable client id for this harness session;
  • the exact runtime, skill, model, and sampling identity configured for the evolution.

Do not guess an identity or silently change it. Gear seals it into the evolution spec and rejects a different harness on claim or resume.

Read references/protocol.md before constructing calls or handling an active assignment. Before diagnosing evidence or changing a candidate, also read references/target-harness-editing.md. When the candidate is Gear's DSH carrier, also read references/dsh-target-harness.md before choosing an artifact or hook. These references are the complete method contract, general editing guide, and version-specific DSH authoring guide; do not infer missing field names or harness APIs from errors.

Run an evolution

  1. Inspect existing status before creating a new evolution when the user's wording could mean continue. A plain refine request creates a new evolution; continuation must name the evolution explicitly.

  2. Call control.start or control.continue, then poll control.status and meta.claim. Baseline evaluation can finish before an assignment appears. Translate optimization preferences into control.start.objective.terms and optional explicit constraints. For example, equal pass/process weights are two terms with weight 0.5; cost/token penalties use negative weights and fixed scales in their original units. Do not put a different objective only in the prompt, invent a metric, infer passes from positive rewards, or alter the objective during search. Omission means strict pass rate. Check the returned resolvedObjective; inspect available raw metrics if admission rejects a term.

  3. Treat the returned lease id, token, session id, candidate id, and (for CLI clients) client id as one inseparable capability. Never reuse them for another assignment.

  4. Use only the candidate file and meta.call methods exposed by Gear. The direct candidate methods are exactly candidate.tree, candidate.read, candidate.write, candidate.edit, and candidate.remove. Invoke harness.current, harness.read, seed_tasks.load, trajectory.query, experience.query, experience.read, hitch.status, candidate.diff, candidate.check, candidate.finalize, and candidate.decline only as meta.call capabilities; they are not top-level request methods. Read the active candidate files before selecting an edit. Do not discover or modify Gear state, Git metadata, held-out data, credentials, or host paths directly.

  5. Review the assignment's optional experienceContext. It contains the actual paired seed result of the direct parent's last edit when available, plus only relevant bounded history cards. Treat rationale and expected outcome as the earlier proposer's claims; treat task reward changes as observations, not causal or statistically significant conclusions. Use experience.query and experience.read for focused history. Historical experienceRef values never belong in evidenceRefs and never satisfy the active baseline diagnosis requirement.

  6. Review the baseline summary. Also read baseline.rawMetrics and baseline.objectiveScore when present: the raw scores, units and contributions explain the frozen objective. A successful but expensive task can still warrant an improvement. State which measured term the proposal is expected to improve and verify its explicit constraints. Never replace a missing metric with zero or change the weights. If the assignment contains workplanDelivery, consume its hypothesis, parent identity, sourced dossier excerpt, scope guards, generation budget, and seed-only findings first. Keep modifications within workplan.modificationPaths (paths relative to the harness root), or explicitly disclose the wider effect; a narrow scope does not qualify wider changes for promotion. The claimed assignment records workplan-dossier-consumed, not personal diagnosis of all parent failures. When evidencePolicy.diagnoseEveryFailedRunBeforeProposal is false, the full-failure reading loop below does not apply: read extra seed evidence only when needed, and use candidate.check to confirm finalization readiness. Do not request or use held-out promotion details to plan candidate changes.

    For assignments without a workplan, start with trajectory.query without arguments on every assignment: it restores validated diagnostics from earlier attempts of this candidate and returns their summaries plus current-session detail refs. If diagnosisRecovery.remaining is positive, repeat that query to receive the rest. Use diagnosisProgress.remainingRunIds to choose new reads. Recovery does not restore candidate edits; inspect the current tree. Changed baseline, trajectory, verifier, or sanitization policy requires reading the affected evidence again. generationBudget reports attempt/round deadlines and remaining time. Its finalization reserve is advisory time for editing, checking, and sealing; it does not extend the deadline. Reconnecting does not reset time. A DIAGNOSIS_BUDGET_AT_RISK warning means the observed pace may leave too little time for the remaining work; do not skip required diagnoses or assume that continuing an old evolution adopts new budget settings. Query trajectory.query with refs for every remaining failed baseline run before proposing a change. It returns a compact diagnostic card containing the task, outcome, verifier failure summary, and the chronological message transcript without raw chunk noise. The card keeps the last 80,000 transcript characters and previews each tool result at up to 2,000 characters. Follow earlierRef for messages before that window and detailRef for a complete long result. Treat result_only as missing verifier logs, not as a complete failure explanation. When the card includes a detailRef, pass it back to trajectory.query; pass a returned nextRef back as the next detailRef, or add find to search that long content. Keep cited evidence limited to references actually returned for the active seed baseline. Start with one diagnostic card and size later reads from its actual output. Process the evidence before fetching more; use targeted detail reads instead of accumulating transcripts that repeatedly force context compaction.

  7. Identify an evidenced harness gap before choosing an intervention. A failed task or omitted check alone does not establish that gap. Use relevant accessible seed comparisons to test the explanation, including successful runs when useful. Choose the smallest supported mechanism from the editing guide, with an applicability boundary that transfers beyond the observed tasks. New files are allowed but must be wired into the existing load graph. Prompt changes, skills, hooks, and workflows need the same causal support; do not force any artifact type. An implementation mistake can still expose a reusable prevention or detection opportunity; existing guidance alone does not show that an effective procedure or check exists. Apply the editing guide's diagnosis before either choosing a prompt edit or declining. Record the evidence, mechanism choice, applicability, and main uncertainty in rationale, and an observable behavioral prediction in expectedOutcome.

  8. Inspect candidate.diff, remove accidental or task-specific changes, and run candidate.check before finalizing and require finalizationReadiness.ready: true. Inspect each runtime stage: loading, prompt assembly and Skill reads cover different paths; they do not prove that arbitrary tool/hook bodies, routing, compaction or workflows executed. Record remaining behavior checks in the proposal. Use candidate.decline when the evidence does not justify a harness change or the required fix is outside the editable substrate.

  9. After finalization, stop using that lease and poll status. Claim and complete every subsequent candidate or round until the requested batch reaches a terminal state.

Show full SKILL.md (213 more words)Show less

Do not call control.publish or control.rollback unless the user explicitly requests that state change. Automatic per-evolution promotion remains governed by Gear's sealed promotion policy.

Failure handling

  • Stop using a lease immediately when Gear reports it stale or invalid.
  • On a timeout or failed round, inspect status; do not create a replacement evolution unless the user asked for a new one.
  • Use control.rerun only for repairable evaluation slots reported by status.
  • For staged-search searchPendingEvidence, control.search-repair takes evolutionId, roundId, a stable repairId, and evidenceDigest. It repairs original invalid slots and resumes the same round; it is not a request for a new candidate or a new experiment. Leave held-out repair to the operator flow, and do not feed its results into a candidate lease.
  • If finalize or decline returns accepted:false, recoverable:true, execute nextAction, then remainingActions, and retry the same operation with the same arguments. The lease remains active until accepted:true.
  • If it returns accepted:false, recoverable:false, report the exact operatorAction and do not repeat the same tool call. Verifier evidence may require an operator/configuration change. TRAJECTORY_EVIDENCE_UNAVAILABLE means Hitch could not construct bounded analysis for the listed blockedRuns; Hitch or that stored trajectory must be repaired before rereading the card.
  • Never bypass a failed compiler check, evidence requirement, identity check, or promotion decision by editing state files.

© rsi-gear, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/refine of rsi-gear/gear.

  • SKILL.md
  • agents/openai.yaml
  • references/dsh-target-harness.md
  • references/protocol.md
  • references/target-harness-editing.md
  • scripts/transport.mjs

Open the folder on GitHubat commit 05f3cf6

Compare with similar skills

Refine next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Refine compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Refine this skillrsi-gear/gear139—~2.8kAutomated safety check: PassMIT
Agent BuildershareAI-lab/learn-claude-code78k4 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.9k14 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 4 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.9k GitHub starsUsed in 14 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed

Questions about Refine

What does Refine do?

Evolve an agent harness through Gear using measured seed and held-out evaluations. Refine is an agent skill from rsi-gear/gear. Evolve an agent harness through Gear using measured seed and held-out evaluations.

When should I use Refine?

Refine fits situations like: the user asks to refine; roll back a Gear-managed harness from Codex; another Agent Skills-compatible harness.

How do I install Refine in Claude Code?

Run `npx skills add rsi-gear/gear --skill refine -a claude-code`. Or copy the skill folder (skills/refine in rsi-gear/gear) into .claude/skills/refine in your project. Claude Code loads it when a task matches its description.

How do I install Refine in Codex?

Run `npx skills add rsi-gear/gear --skill refine -a codex`. Or copy the skill folder (skills/refine in rsi-gear/gear) into .agents/skills/refine in your project. Codex loads it when a task matches its description.

Can I use Refine in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rsi-gear/gear --skill refine -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/refine, .gemini/skills/refine, .github/skills/refine and .opencode/skills/refine in your project.

What does Refine need to run?

Going by SKILL.md and its folder, Refine needs JavaScript for the scripts in its folder. Our summary lists: Node.js.

Does Refine access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Refine safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Refine use?

Refine is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Refine use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.

What are the alternatives to Refine?

Skills that share tags, products or a category with Refine: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Refine?

rsi-gear (a GitHub organization) maintains it in rsi-gear/gear, which has 139 GitHub stars. The repository was last updated on October 10, 2026.

Source: rsi-gear/gear on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.