Agent skill

Octocode Benchmark Runner

by bgauryy in bgauryy/octocode

Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

MITAuto-check passedAI & LLM Engineering

Install Octocode Benchmark Runner

skills CLI
$ npx skills add bgauryy/octocode --skill octocode-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bgauryy/octocode octocode-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/octocode-benchmark/skills/octocode-benchmark .claude/skills/octocode-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
octocode-benchmark
GitHub stars
949
Token cost
~2.1k tokens
SKILL.md length
721 words
Files
24 (incl. scripts, references)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

  • Works in 4 steps: Preflight — verify + pin every arm; a… → Answer — 2 isolated runners (anchor +… → Judge — after both sections exist, one… → …
  • Comparing Octocode against a gh-based baseline on a research question
  • SKILL.md covers Flow (4 phases), How it is MEASURED…, How it is SCORED (blind judge,… and Quickstart (copy-paste, one…, plus 5 more sections
  • Calls python3, npx and gh

What it does

Each matchup pits Octocode as the anchor against one baseline (gh plus RTK, gh plus Headroom, or plain gh) on the same question. Two isolated runner agents answer independently, a blind judge grades the two answers labelled X and Y in randomized order so it cannot tell which tool produced which, and at least three passes are run before the rollup combines every matchup.

The scored metric is total characters through the model: the command strings and arguments the model wrote plus its final answer, and the tool output pulled back into context, both read from an instrumented per-call log rather than trusted as a hand count. Every research command goes through a thin wrapper specific to its arm that shells the real CLI unchanged and appends one log row per call, and dedicated scripts recompute the per-question total and validate the whole campaign byte-for-byte rather than accepting a self-reported number.

When your agent uses it

  • Comparing Octocode against a gh-based baseline on a research question
  • Running a blind-judged benchmark pass across several baselines
  • Validating that a benchmark campaign's character counts are correct

Example prompts

  • “Run the Octocode vs plain gh matchup on the questions in compare/github-questions.”
  • “Judge this pass's answers for question 4 and record the verdict.”
  • “Validate the latest benchmark campaign's logs byte-for-byte.”

Requirements

  • The gh CLI and the octocode, rtk and Headroom tools being compared
  • Python scripts in compare/bin/ for logging and validation

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Preflight — verify + pin every arm; a failure invalidates the run.
  2. Answer — 2 isolated runners (anchor + baseline) per question/pass, leanest-legal path, each appends a ## Q section to answers/-p.md.
  3. Judge — after both sections exist, one blind judge reasons to a verdict, then scores.
  4. Summarize — validate logs, aggregate paired stats, update the rollup.

What it can do on your machine

Read from SKILL.md and the folder at commit c265e3f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • npx
    • gh
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Octocode Benchmark Runner loads about 2.1k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 137 tokens; SKILL.md has 721 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bgauryy/octocode at commit c265e3f, republished under its MIT licence (© bgauryy). 721 words, ~2,067 tokens.

Download SKILL.mdSave it as .claude/skills/octocode-benchmark/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.
name
octocode-benchmark
description
Use when planning, running, grading, or reporting the by-hand Octocode research benchmark — pairwise matchups (Octocode anchor vs one baseline: gh+RTK, gh+Headroom, or plain gh) over markdown questions, with a fresh isolated runner agent per (question, arm, pass), one blind judge per question grading two answers X/Y in randomized order, and an orchestrator that summarizes accuracy/quality/workflow/characters. Results measured in total characters through the model (model-in delivered + model-out commands/args + final answer).

Octocode benchmark

Plain-markdown, run-by-hand CLI research comparison. Octocode is the anchor; each baseline is a separate pairwise matchup (octocode vs rtk | headroom | gh). Per question, per pass: two isolated runners answer, one blind judge grades them X / Y (randomized per question). Run ≥3 passes; the rollup shows every matchup together. No harness, no JSON.

Paths below are relative to the package root packages/octocode-benchmark/. Shared tooling lives in compare/bin/; questions in compare/github-questions/; reports in results/.

Flow (4 phases)

  1. Preflight — verify + pin every arm; a failure invalidates the run.
  2. Answer — 2 isolated runners (anchor + baseline) per question/pass, leanest-legal path, each appends a ## Q<n> section to answers/<arm>-p<pass>.md.
  3. Judge — after both sections exist, one blind judge reasons to a verdict, then scores.
  4. Summarize — validate logs, aggregate paired stats, update the rollup.

How it is MEASURED (characters, never self-reported)

The metric is total_chars = model-in + model-out in Unicode code points, from an instrumented log — the tool transcript only (excludes system prompt, tool schemas, model reasoning; the fixed per-arm primer is excluded by rule; any later help/schema/failed call is counted).

  • model-out = the command string + args the model wrote, plus the final answer.
  • model-in = the tool output pulled back into context (for Headroom, the compressed output).

Mechanism: every research command runs through its arm's thin wrapper, which shells the real CLI unchanged, prints output verbatim, and appends one JSONL row per call:

ArmWrapperRunsLog env
octocode (local build)compare/bin/octocnpx octocode tools …OCTO_LOG
octocode (published pin)compare/bin/octoc1822npx -y octocode@18.2.2 tools …OCTO_LOG
gh+RTKcompare/bin/rtkmrtk gh …RTK_LOG
gh+Headroomcompare/bin/ghcgh … → Headroom compressGHC_LOG (+ HR_PY)
plain ghcompare/bin/ghmgh … (read-only)—

The final answer is logged as pure model-out via compare/bin/record_answer.py. Per-question total = compare/bin/sumlog.py --strict <log>; the whole campaign is checked byte-faithfully by compare/bin/validate_campaign.py. Never trust a hand-counted number — recompute from the JSONL. Only elapsed_ms (octocode/rtk/gh) is captured for time; it is not a fair latency metric (npx bootstrap per call, no Headroom timing) — do not headline it.

How it is SCORED (blind judge, correctness-first)

One blind judge per question grades the two answers as X / Y (order randomized per question, tool identity redacted). It reasons to ground truth first, then scores each answer: correctness 0–10, research depth 1–5, workflow 1–5 (rubric in references/JUDGING.md).

Decision per pairing: if one arm is net strictly more correct (paired sign test) it wins — a confidently-wrong answer never wins on footprint. If correctness is statistically tied, characters decide by the geometric-mean of per-question ratios (baseline ÷ octocode) + median + leaner win-rate + bootstrap CI — never a pooled sum alone. Aggregate paired, per question, over ≥3 passes. Method + worked example: references/aggregation-and-stats.md.

Show full SKILL.md (292 more words)Show less

Quickstart (copy-paste, one matchup, one pass)

bash
cd packages/octocode-benchmark
export HR_PY="$HOME/.local/share/uv/tools/headroom-ai/bin/python"   # Headroom arm only

# Phase 0 — preflight (non-zero exit = fix before running)
bash skills/octocode-benchmark/scripts/check-prereqs.sh 18.2.2

# set up a campaign dir
CAMP="campaigns/run-$(date -u +%H%M%S)-$(date -u +%Y-%m-%d)"; mkdir -p "$CAMP/answers" "$CAMP/judge"

# Phase 1 — answer (spawn ONE isolated agent per arm; never mix arms in an agent).
# Each research call sets its per-question log, e.g. octocode Q4:
OCTO_LOG="$CAMP/octocode-p1-Q4.jsonl" ./compare/bin/octoc1822 ghGetFileContent \
  --queries '{"owner":"axios","repo":"axios","path":"lib/adapters/http.js","matchString":"follow-redirects"}'
# baseline (gh+RTK) Q4:
RTK_LOG="$CAMP/rtk-p1-Q4.jsonl" ./compare/bin/rtkm search code --repo axios/axios follow-redirects --limit 20
# log the final answer, then append a "## Q4" section (Answer + Research steps) to answers/<arm>-p1.md
python3 compare/bin/record_answer.py --log "$CAMP/octocode-p1-Q4.jsonl" --question Q4 --file answer.txt

# Phase 2 — judge: build the blind packet, then one reasoning-first verdict per question
python3 compare/bin/build_blind_packet.py --help          # X/Y randomized, tool identity redacted

# Phase 3 — validate + aggregate + report
python3 compare/bin/sumlog.py --strict "$CAMP/octocode-p1-Q4.jsonl"
python3 compare/bin/validate_campaign.py "$CAMP" --question-count 30
python3 compare/bin/per_question_summary.py --out results/PER_QUESTION_SUMMARY.md --json results/per_question_summary.json

Runner/judge briefing packets, spawn scaling (batch Q1-15/Q16-30 within one arm), and output layout: references/run-with-agents.md → run-preflight.md + run-phases.md.

Hard gates (skip one and the run is worthless)

  • Isolation — a fresh agent per (question, arm, pass) + a separate judge; no shared transcript, no answer key. Batch questions within an arm, never mix arms.
  • Fairness — leanest-legal path on every arm; no whole-tree/whole-file dump where a targeted read/search answers (inflates chars, invalidates the ratio). sumlog.py emits advisory FAIRNESS: lines for recursive=1 dumps / oversized reads — review them.
  • Blind + reasoned — grade X/Y in randomized order; the judge reasons before scoring, correctness-first; a confidently-wrong answer never wins.
  • Measured, not self-reported — total_chars = model-in + model-out from the instrumented log; recompute, never hand-count.
  • Honest stats — geometric-mean char ratio (never a pooled sum) + bootstrap CI; ≥3 passes; the public set is orientation, not a shipping gate.

Routes — load only what the step needs

WhenLoad
understand the designreferences/BENCHMARK.md
run a matchupreferences/INSTRUCTIONS.md then references/run-with-agents.md
brief a runnerreferences/RUNNER.md + references/RUNNER_TOOL_CONTEXT.md (+ the arm's primer-*.md)
judge a questionreferences/JUDGING.md + references/example-verdict.md
score + aggregatereferences/SCORING.md then references/aggregation-and-stats.md
write the reportreferences/REPORT_TEMPLATE.md
author a matchup READMEreferences/matchup-readme.md

Scripts + tooling

  • scripts/check-prereqs.sh — Phase 0 gate (all arms + questions + primers).
  • scripts/measure.sh — fallback char wrapper for an arm without a dedicated bin/ wrapper.
  • compare/bin/: octoc · octoc1822 · rtkm · ghc · ghm (arm wrappers) · instrument_command.py / hr_compress.py (char capture) · record_answer.py · sumlog.py (per-question total, --strict) · build_blind_packet.py (X/Y packet) · validate_campaign.py (byte-faithful campaign check) · per_question_summary.py (per-question + overall chars & correctness across all 4 arms) · test_instrumentation.py.

Stop when

The matchup's questions are answered, judged, and aggregated across ≥3 passes with CIs — or a preflight/fairness violation blocks the run; fix before continuing.

Add a question

Copy an existing Q<n>.md, bump the number, edit title / id / ## Question only. GitHub → compare/github-questions/; corpus-local → that matchup's questions/; add its README row.

© bgauryy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 23 other files (scripts, references) in packages/octocode-benchmark/skills/octocode-benchmark of bgauryy/octocode.

  • SKILL.md
  • README.md
  • references/BENCHMARK.md
  • references/INSTRUCTIONS.md
  • references/JUDGING.md
  • references/REPORT_TEMPLATE.md
  • references/RUNNER.md
  • references/RUNNER_TOOL_CONTEXT.md
  • references/SCORING.md
  • references/aggregation-and-stats.md
  • references/example-verdict.md
  • references/matchup-readme.md
  • references/primer-gh-headroom.md
  • references/primer-gh-rtk.md
  • references/primer-gh.md
  • references/primer-octocode.md
  • references/run-phases.md
  • references/run-preflight.md
  • references/run-with-agents.md
  • scripts
  • … and 4 more

Open the folder on GitHubat commit c265e3f

Compare with similar skills

Octocode Benchmark Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Octocode Benchmark Runner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Octocode Benchmark Runner this skillbgauryy/octocode949—~2.1kAutomated safety check: PassMIT
SWE Benchmark Task Adderory/lumen307—~497Automated safety check: PassCustom licence
AI Project Copilotsun461941-hub/ai-project-copilot97—~3kAutomated safety check: PassMIT
Managed Deep Agentslangchain-ai/langchain-skills1.3k—~8.7kAutomated safety check: NotesMIT
Waza Interactivemicrosoft/waza1.4k—~1.3kAutomated safety check: PassMIT
Benchmark Agentsvercel/vercel-plugin301—~3.6kAutomated safety check: PassCustom licence

Similar skills

  • Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

    307 GitHub stars~497 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • AI Project Copilot

    sun461941-hub/ai-project-copilot

    A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

    97 GitHub stars~3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Managed Deep Agents

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.

    1.3k GitHub stars~8.7k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • Waza Interactive

    microsoft/waza

    Official

    Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

    1.4k GitHub stars~1.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Benchmark Agents

    vercel/vercel-plugin

    Official

    Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

    301 GitHub stars~3.6k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Windmill AI Evals

    windmill-labs/windmill

    Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

    18k GitHub stars~969 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from bgauryy/octocode

All 12 skills in this repo
  • Octocode Code Research

    bgauryy/octocode

    Researches code with evidence: traces callers, imports and cross-repo links, diagnoses failures and reports findings with exact file and line references and a confidence label.

    949 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Writes, repairs and copyedits project docs against the Google developer documentation style guide, verifying claims in the repository before stating them.

    949 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Octocode Mannequin

    bgauryy/octocode

    Poses and animates a 22-bone anatomical humanoid rig with joint range-of-motion limits, using a Node CLI, a Three.js viewer and WebMCP tools an agent can drive live.

    949 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Octocode Skills Manager

    bgauryy/octocode

    Finds, rates, reviews, creates, improves, installs and syncs Agent Skill folders from local workspaces, registries or remote sources, with a user gate before any write.

    949 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Octocode Brainstorming

    bgauryy/octocode

    Walks an idea through framing, diverging into options, researching evidence and stress-testing before converging on a build, prototype, narrow or park decision.

    949 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Octocode Chrome Devtools

    bgauryy/octocode

    A skill your agent uses when a live page needs Chrome DevTools/CDP evidence: network failures, console errors, performance, DOM/CSS actionability, screenshots/PDF, cookies/storage…

    949 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed

Questions about Octocode Benchmark Runner

What does Octocode Benchmark Runner do?

Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report. Each matchup pits Octocode as the anchor against one baseline (gh plus RTK, gh plus Headroom, or plain gh) on the same question. Two isolated runner agents answer independently, a blind judge grades the two answers labelled X and Y in randomized order so it cannot tell which tool produced which, and at least three passes are run before the rollup combines every matchup.

When should I use Octocode Benchmark Runner?

Octocode Benchmark Runner fits situations like: comparing Octocode against a gh-based baseline on a research question; running a blind-judged benchmark pass across several baselines; validating that a benchmark campaign's character counts are correct.

How do I install Octocode Benchmark Runner in Claude Code?

Run `npx skills add bgauryy/octocode --skill octocode-benchmark -a claude-code`. Or copy the skill folder (packages/octocode-benchmark/skills/octocode-benchmark in bgauryy/octocode) into .claude/skills/octocode-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Octocode Benchmark Runner in Codex?

Run `npx skills add bgauryy/octocode --skill octocode-benchmark -a codex`. Or copy the skill folder (packages/octocode-benchmark/skills/octocode-benchmark in bgauryy/octocode) into .agents/skills/octocode-benchmark in your project. Codex loads it when a task matches its description.

Can I use Octocode Benchmark Runner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bgauryy/octocode --skill octocode-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/octocode-benchmark, .gemini/skills/octocode-benchmark, .github/skills/octocode-benchmark and .opencode/skills/octocode-benchmark in your project.

What does Octocode Benchmark Runner need to run?

Going by SKILL.md and its folder, Octocode Benchmark Runner needs the command-line tools its instructions call (python3, npx, gh and bash). Our summary lists: The gh CLI and the octocode, rtk and Headroom tools being compared; Python scripts in compare/bin/ for logging and validation.

Does Octocode Benchmark Runner access the network?

SKILL.md contains no URLs. Its commands use npx and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Octocode Benchmark Runner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Octocode Benchmark Runner use?

Octocode Benchmark Runner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Octocode Benchmark Runner use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to Octocode Benchmark Runner?

Skills that share tags, products or a category with Octocode Benchmark Runner: SWE Benchmark Task Adder (ory/lumen, 307 stars), AI Project Copilot (sun461941-hub/ai-project-copilot, 97 stars), Managed Deep Agents (langchain-ai/langchain-skills, 1.3k stars) and Waza Interactive (microsoft/waza, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Octocode Benchmark Runner?

bgauryy (a GitHub user) maintains it in bgauryy/octocode, which has 949 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 9, 2026.

Source: bgauryy/octocode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.