SWE Benchmark Task Adder
ory/lumen
Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.
Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.
$ npx skills add bgauryy/octocode --skill octocode-benchmark -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install bgauryy/octocode octocode-benchmark --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/octocode-benchmark/skills/octocode-benchmark .claude/skills/octocode-benchmark && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "octocode-benchmark" agent skill from https://github.com/bgauryy/octocode/tree/main/packages/octocode-benchmark/skills/octocode-benchmark into .claude/skills/octocode-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-benchmark", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/bgauryy/octocode/tree/main/packages/octocode-benchmark/skills/octocode-benchmarkType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add bgauryy/octocode --skill octocode-benchmark -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install bgauryy/octocode octocode-benchmark --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .agents/skills && cp -r skills-src/packages/octocode-benchmark/skills/octocode-benchmark .agents/skills/octocode-benchmark && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "octocode-benchmark" agent skill from https://github.com/bgauryy/octocode/tree/main/packages/octocode-benchmark/skills/octocode-benchmark into .agents/skills/octocode-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-benchmark", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bgauryy/octocode --skill octocode-benchmark -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install bgauryy/octocode octocode-benchmark --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/packages/octocode-benchmark/skills/octocode-benchmark .cursor/skills/octocode-benchmark && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "octocode-benchmark" agent skill from https://github.com/bgauryy/octocode/tree/main/packages/octocode-benchmark/skills/octocode-benchmark into .cursor/skills/octocode-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-benchmark", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/bgauryy/octocode.git --path packages/octocode-benchmark/skills/octocode-benchmark--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add bgauryy/octocode --skill octocode-benchmark -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install bgauryy/octocode octocode-benchmark --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/packages/octocode-benchmark/skills/octocode-benchmark .gemini/skills/octocode-benchmark && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "octocode-benchmark" agent skill from https://github.com/bgauryy/octocode/tree/main/packages/octocode-benchmark/skills/octocode-benchmark into .gemini/skills/octocode-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-benchmark", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install bgauryy/octocode octocode-benchmarkInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add bgauryy/octocode --skill octocode-benchmark -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .github/skills && cp -r skills-src/packages/octocode-benchmark/skills/octocode-benchmark .github/skills/octocode-benchmark && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "octocode-benchmark" agent skill from https://github.com/bgauryy/octocode/tree/main/packages/octocode-benchmark/skills/octocode-benchmark into .github/skills/octocode-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-benchmark", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bgauryy/octocode --skill octocode-benchmark -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install bgauryy/octocode octocode-benchmark --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/packages/octocode-benchmark/skills/octocode-benchmark .opencode/skills/octocode-benchmark && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "octocode-benchmark" agent skill from https://github.com/bgauryy/octocode/tree/main/packages/octocode-benchmark/skills/octocode-benchmark into .opencode/skills/octocode-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-benchmark", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
octocode-benchmarkRuns blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.
Each matchup pits Octocode as the anchor against one baseline (gh plus RTK, gh plus Headroom, or plain gh) on the same question. Two isolated runner agents answer independently, a blind judge grades the two answers labelled X and Y in randomized order so it cannot tell which tool produced which, and at least three passes are run before the rollup combines every matchup.
The scored metric is total characters through the model: the command strings and arguments the model wrote plus its final answer, and the tool output pulled back into context, both read from an instrumented per-call log rather than trusted as a hand count. Every research command goes through a thin wrapper specific to its arm that shells the real CLI unchanged and appends one log row per call, and dedicated scripts recompute the per-question total and validate the whole campaign byte-for-byte rather than accepting a self-reported number.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c265e3f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
python3npxghbashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx and gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Octocode Benchmark Runner loads about 2.1k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 137 tokens; SKILL.md has 721 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from bgauryy/octocode at commit c265e3f, republished under its MIT licence (© bgauryy). 721 words, ~2,067 tokens.
.claude/skills/octocode-benchmark/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.Plain-markdown, run-by-hand CLI research comparison. Octocode is the anchor; each baseline
is a separate pairwise matchup (octocode vs rtk | headroom | gh). Per question,
per pass: two isolated runners answer, one blind judge grades them X / Y (randomized per
question). Run ≥3 passes; the rollup shows every matchup together. No harness, no JSON.
Paths below are relative to the package root packages/octocode-benchmark/. Shared tooling
lives in compare/bin/; questions in compare/github-questions/; reports in results/.
## Q<n> section to answers/<arm>-p<pass>.md.The metric is total_chars = model-in + model-out in Unicode code points, from an
instrumented log — the tool transcript only (excludes system prompt, tool schemas, model
reasoning; the fixed per-arm primer is excluded by rule; any later help/schema/failed call
is counted).
Mechanism: every research command runs through its arm's thin wrapper, which shells the real CLI unchanged, prints output verbatim, and appends one JSONL row per call:
| Arm | Wrapper | Runs | Log env |
|---|---|---|---|
| octocode (local build) | compare/bin/octoc | npx octocode tools … | OCTO_LOG |
| octocode (published pin) | compare/bin/octoc1822 | npx -y octocode@18.2.2 tools … | OCTO_LOG |
| gh+RTK | compare/bin/rtkm | rtk gh … | RTK_LOG |
| gh+Headroom | compare/bin/ghc | gh … → Headroom compress | GHC_LOG (+ HR_PY) |
| plain gh | compare/bin/ghm | gh … (read-only) | — |
The final answer is logged as pure model-out via compare/bin/record_answer.py. Per-question
total = compare/bin/sumlog.py --strict <log>; the whole campaign is checked byte-faithfully
by compare/bin/validate_campaign.py. Never trust a hand-counted number — recompute from
the JSONL. Only elapsed_ms (octocode/rtk/gh) is captured for time; it is not a fair latency
metric (npx bootstrap per call, no Headroom timing) — do not headline it.
One blind judge per question grades the two answers as X / Y (order randomized per
question, tool identity redacted). It reasons to ground truth first, then scores each
answer: correctness 0–10, research depth 1–5, workflow 1–5 (rubric in references/JUDGING.md).
Decision per pairing: if one arm is net strictly more correct (paired sign test) it wins —
a confidently-wrong answer never wins on footprint. If correctness is statistically tied,
characters decide by the geometric-mean of per-question ratios (baseline ÷ octocode) +
median + leaner win-rate + bootstrap CI — never a pooled sum alone. Aggregate paired, per
question, over ≥3 passes. Method + worked example: references/aggregation-and-stats.md.
cd packages/octocode-benchmark
export HR_PY="$HOME/.local/share/uv/tools/headroom-ai/bin/python" # Headroom arm only
# Phase 0 — preflight (non-zero exit = fix before running)
bash skills/octocode-benchmark/scripts/check-prereqs.sh 18.2.2
# set up a campaign dir
CAMP="campaigns/run-$(date -u +%H%M%S)-$(date -u +%Y-%m-%d)"; mkdir -p "$CAMP/answers" "$CAMP/judge"
# Phase 1 — answer (spawn ONE isolated agent per arm; never mix arms in an agent).
# Each research call sets its per-question log, e.g. octocode Q4:
OCTO_LOG="$CAMP/octocode-p1-Q4.jsonl" ./compare/bin/octoc1822 ghGetFileContent \
--queries '{"owner":"axios","repo":"axios","path":"lib/adapters/http.js","matchString":"follow-redirects"}'
# baseline (gh+RTK) Q4:
RTK_LOG="$CAMP/rtk-p1-Q4.jsonl" ./compare/bin/rtkm search code --repo axios/axios follow-redirects --limit 20
# log the final answer, then append a "## Q4" section (Answer + Research steps) to answers/<arm>-p1.md
python3 compare/bin/record_answer.py --log "$CAMP/octocode-p1-Q4.jsonl" --question Q4 --file answer.txt
# Phase 2 — judge: build the blind packet, then one reasoning-first verdict per question
python3 compare/bin/build_blind_packet.py --help # X/Y randomized, tool identity redacted
# Phase 3 — validate + aggregate + report
python3 compare/bin/sumlog.py --strict "$CAMP/octocode-p1-Q4.jsonl"
python3 compare/bin/validate_campaign.py "$CAMP" --question-count 30
python3 compare/bin/per_question_summary.py --out results/PER_QUESTION_SUMMARY.md --json results/per_question_summary.jsonRunner/judge briefing packets, spawn scaling (batch Q1-15/Q16-30 within one arm), and output
layout: references/run-with-agents.md → run-preflight.md + run-phases.md.
sumlog.py emits advisory FAIRNESS: lines for recursive=1 dumps / oversized reads — review them.total_chars = model-in + model-out from the instrumented log; recompute, never hand-count.| When | Load |
|---|---|
| understand the design | references/BENCHMARK.md |
| run a matchup | references/INSTRUCTIONS.md then references/run-with-agents.md |
| brief a runner | references/RUNNER.md + references/RUNNER_TOOL_CONTEXT.md (+ the arm's primer-*.md) |
| judge a question | references/JUDGING.md + references/example-verdict.md |
| score + aggregate | references/SCORING.md then references/aggregation-and-stats.md |
| write the report | references/REPORT_TEMPLATE.md |
| author a matchup README | references/matchup-readme.md |
scripts/check-prereqs.sh — Phase 0 gate (all arms + questions + primers).scripts/measure.sh — fallback char wrapper for an arm without a dedicated bin/ wrapper.compare/bin/: octoc · octoc1822 · rtkm · ghc · ghm (arm wrappers) ·
instrument_command.py / hr_compress.py (char capture) · record_answer.py ·
sumlog.py (per-question total, --strict) · build_blind_packet.py (X/Y packet) ·
validate_campaign.py (byte-faithful campaign check) · per_question_summary.py
(per-question + overall chars & correctness across all 4 arms) · test_instrumentation.py.The matchup's questions are answered, judged, and aggregated across ≥3 passes with CIs — or a preflight/fairness violation blocks the run; fix before continuing.
Copy an existing Q<n>.md, bump the number, edit title / id / ## Question only. GitHub →
compare/github-questions/; corpus-local → that matchup's questions/; add its README row.
© bgauryy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 23 other files (scripts, references) in packages/octocode-benchmark/skills/octocode-benchmark of bgauryy/octocode.
Open the folder on GitHubat commit c265e3f
Octocode Benchmark Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Octocode Benchmark Runner this skillbgauryy/octocode | 949 | — | ~2.1k | Automated safety check: Pass | MIT | |
| SWE Benchmark Task Adderory/lumen | 307 | — | ~497 | Automated safety check: Pass | Custom licence | |
| AI Project Copilotsun461941-hub/ai-project-copilot | 97 | — | ~3k | Automated safety check: Pass | MIT | |
| Managed Deep Agentslangchain-ai/langchain-skills | 1.3k | — | ~8.7k | Automated safety check: Notes | MIT | |
| Waza Interactivemicrosoft/waza | 1.4k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Benchmark Agentsvercel/vercel-plugin | 301 | — | ~3.6k | Automated safety check: Pass | Custom licence |
ory/lumen
Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.
sun461941-hub/ai-project-copilot
A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…
langchain-ai/langchain-skills
INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.
microsoft/waza
Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.
vercel/vercel-plugin
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.
windmill-labs/windmill
Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.
bgauryy/octocode
Researches code with evidence: traces callers, imports and cross-repo links, diagnoses failures and reports findings with exact file and line references and a confidence label.
bgauryy/octocode
Writes, repairs and copyedits project docs against the Google developer documentation style guide, verifying claims in the repository before stating them.
bgauryy/octocode
Poses and animates a 22-bone anatomical humanoid rig with joint range-of-motion limits, using a Node CLI, a Three.js viewer and WebMCP tools an agent can drive live.
bgauryy/octocode
Finds, rates, reviews, creates, improves, installs and syncs Agent Skill folders from local workspaces, registries or remote sources, with a user gate before any write.
bgauryy/octocode
Walks an idea through framing, diverging into options, researching evidence and stress-testing before converging on a build, prototype, narrow or park decision.
bgauryy/octocode
A skill your agent uses when a live page needs Chrome DevTools/CDP evidence: network failures, console errors, performance, DOM/CSS actionability, screenshots/PDF, cookies/storage…
Works with
Categories
Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report. Each matchup pits Octocode as the anchor against one baseline (gh plus RTK, gh plus Headroom, or plain gh) on the same question. Two isolated runner agents answer independently, a blind judge grades the two answers labelled X and Y in randomized order so it cannot tell which tool produced which, and at least three passes are run before the rollup combines every matchup.
Octocode Benchmark Runner fits situations like: comparing Octocode against a gh-based baseline on a research question; running a blind-judged benchmark pass across several baselines; validating that a benchmark campaign's character counts are correct.
Run `npx skills add bgauryy/octocode --skill octocode-benchmark -a claude-code`. Or copy the skill folder (packages/octocode-benchmark/skills/octocode-benchmark in bgauryy/octocode) into .claude/skills/octocode-benchmark in your project. Claude Code loads it when a task matches its description.
Run `npx skills add bgauryy/octocode --skill octocode-benchmark -a codex`. Or copy the skill folder (packages/octocode-benchmark/skills/octocode-benchmark in bgauryy/octocode) into .agents/skills/octocode-benchmark in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bgauryy/octocode --skill octocode-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/octocode-benchmark, .gemini/skills/octocode-benchmark, .github/skills/octocode-benchmark and .opencode/skills/octocode-benchmark in your project.
Going by SKILL.md and its folder, Octocode Benchmark Runner needs the command-line tools its instructions call (python3, npx, gh and bash). Our summary lists: The gh CLI and the octocode, rtk and Headroom tools being compared; Python scripts in compare/bin/ for logging and validation.
SKILL.md contains no URLs. Its commands use npx and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Octocode Benchmark Runner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Octocode Benchmark Runner: SWE Benchmark Task Adder (ory/lumen, 307 stars), AI Project Copilot (sun461941-hub/ai-project-copilot, 97 stars), Managed Deep Agents (langchain-ai/langchain-skills, 1.3k stars) and Waza Interactive (microsoft/waza, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
bgauryy (a GitHub user) maintains it in bgauryy/octocode, which has 949 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 9, 2026.
Source: bgauryy/octocode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.