Benchmark Agents
vercel/vercel-plugin
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.
Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.
$ npx skills add bgauryy/octocode --skill octocode-graph-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install bgauryy/octocode octocode-graph-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/octocode-graph-eval .claude/skills/octocode-graph-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "octocode-graph-eval" agent skill from https://github.com/bgauryy/octocode/tree/main/skills/octocode-graph-eval into .claude/skills/octocode-graph-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-graph-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/bgauryy/octocode/tree/main/skills/octocode-graph-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add bgauryy/octocode --skill octocode-graph-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install bgauryy/octocode octocode-graph-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/octocode-graph-eval .agents/skills/octocode-graph-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "octocode-graph-eval" agent skill from https://github.com/bgauryy/octocode/tree/main/skills/octocode-graph-eval into .agents/skills/octocode-graph-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-graph-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bgauryy/octocode --skill octocode-graph-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install bgauryy/octocode octocode-graph-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/octocode-graph-eval .cursor/skills/octocode-graph-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "octocode-graph-eval" agent skill from https://github.com/bgauryy/octocode/tree/main/skills/octocode-graph-eval into .cursor/skills/octocode-graph-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-graph-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/bgauryy/octocode.git --path skills/octocode-graph-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add bgauryy/octocode --skill octocode-graph-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install bgauryy/octocode octocode-graph-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/octocode-graph-eval .gemini/skills/octocode-graph-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "octocode-graph-eval" agent skill from https://github.com/bgauryy/octocode/tree/main/skills/octocode-graph-eval into .gemini/skills/octocode-graph-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-graph-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install bgauryy/octocode octocode-graph-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add bgauryy/octocode --skill octocode-graph-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/octocode-graph-eval .github/skills/octocode-graph-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "octocode-graph-eval" agent skill from https://github.com/bgauryy/octocode/tree/main/skills/octocode-graph-eval into .github/skills/octocode-graph-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-graph-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bgauryy/octocode --skill octocode-graph-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install bgauryy/octocode octocode-graph-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/octocode-graph-eval .opencode/skills/octocode-graph-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "octocode-graph-eval" agent skill from https://github.com/bgauryy/octocode/tree/main/skills/octocode-graph-eval into .opencode/skills/octocode-graph-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "octocode-graph-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
octocode-graph-evalRuns a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.
The skill structures evaluation as a fixed flow: error-analyze, frame a goal into a KPI, measure a baseline, loop a change, judge the result, capture a lesson, verify on held-out data, then evolve the test suite. It refuses to proceed without a goal linked to a measurable KPI and a runnable sensor to check it, and rejects accepting a change on narrative alone or by editing the harness or test cases just to make them pass.
Its rules call for deterministic graders over binary or LLM judgment where possible, a change accepted only when the primary metric improves on held-out data and guardrail metrics still hold, and a counter-metric guardrail for every primary KPI so it cannot be gamed by tuning alone. A verifier sharing the same context as the agent that made the change does not count as independent; fresh context is required first. Multi-agent workflows are checked for true independence between steps, and the harness is frozen during an experiment. Reference files cover benchmarking, error analysis and failure modes, loaded only as each step needs them.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c265e3f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Octocode Graph Eval Loop loads about 1.6k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 742 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from bgauryy/octocode at commit c265e3f, republished under its MIT licence (© bgauryy). 742 words, ~1,591 tokens.
.claude/skills/octocode-graph-eval/SKILL.md (or your agent's skills folder). This skill also uses 30 other files; get the full folder from GitHub.Evaluate outcomes and run improvement loops with evidence, not vibes — for one loop or a graph of loops.
Flow: ERROR-ANALYZE → FRAME(goal→KPI) → BASELINE → LOOP → JUDGE → CAPTURE → VERIFY → SUITE-EVOLVE.
Modes: ErrorAnalyze · Define · Run · Suite · Benchmark · Audit.
references/error-analysis.md; when connecting intent to measures load references/goal-kpi-cascade.md, then fill references/kpi-contract.md — make success and budget explicit.references/nested-loops.md; before the first iteration load references/feedback-loops.md, then for the inner keep/discard cycle load references/agent-loop.md — no workable sensor, no loop.references/graph-of-loops.md — run edge detection first, require anchor nodes, check verifier independence, name Goodhart guardrails, then set primary KPI at the graph boundary with per-node sensors.references/graph-failure-modes.md; add a suite case on a mode's first trace appearance.references/subagent-cookbook.md first for the ownership split; spawn mechanics stay in octocode-subagent.references/subagent-protocol.md for the frozen FRAME→verdict protocol; when choosing worker and graph-boundary metrics, load references/subagent-kpis.md so spawn cost is measured, not invisible.references/subagent-communication.md — bad channels create false certainty and unattributable failures; when choosing the topology itself, load references/subagent-approaches.md because the pattern decides which KPIs and checks matter.references/nested-loops.md for bilevel escalation, then references/karpathy-patterns.md for the Bilevel Autoresearch pattern.references/eval-techniques.md; when grading agent tool-call sequences or multi-turn trajectories load references/trajectory-grading.md; when trusting public/private suites load references/benchmarking.md — match evidence strength to the decision.references/eval-harness.md; before acceptance load references/held-out-and-guards.md — prevent leakage, overfitting, and greenwashing.references/karpathy-patterns.md — anchor techniques in proven loops.references/routing.md; when closing a meta improvement cycle load references/improve-loop.md — transfer ownership without losing the decision rule.references/output.md and run scripts/loop-report.mjs — require goal, baseline, result, and verdict.octocode-research for evidence under test; octocode-brainstorming before evaluating an unresolved idea; octocode-rfc-generator for a design KPI contract.octocode-subagent to fan out parallel hypotheses or benchmark trials within one iteration — measurement, keep/discard, graders, and the subagent cookbook (references/subagent-cookbook.md) stay frozen here.octocode-prompt-optimizer for wording after the KPI is fixed; octocode-skills for folder edits after ACCEPT.scripts/check-description.mjs then scripts/eval-eval.mjs --self-test and a matching --case — catch trigger and self-routing regressions; cases live in evals/ (cases.json, trigger-cases.json, kpi-contract.json).© bgauryy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 30 other files (scripts, references) in skills/octocode-graph-eval of bgauryy/octocode.
Open the folder on GitHubat commit c265e3f
Octocode Graph Eval Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Octocode Graph Eval Loop this skillbgauryy/octocode | 949 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Benchmark Agentsvercel/vercel-plugin | 301 | — | ~3.6k | Automated safety check: Pass | Custom licence | |
| Waza Interactivemicrosoft/waza | 1.4k | — | ~1.3k | Automated safety check: Pass | MIT | |
| AI Project Copilotsun461941-hub/ai-project-copilot | 97 | — | ~3k | Automated safety check: Pass | MIT | |
| SWE Benchmark Task Adderory/lumen | 307 | — | ~497 | Automated safety check: Pass | Custom licence | |
| Managed Deep Agentslangchain-ai/langchain-skills | 1.3k | — | ~8.7k | Automated safety check: Notes | MIT |
vercel/vercel-plugin
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.
microsoft/waza
Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.
sun461941-hub/ai-project-copilot
A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…
ory/lumen
Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.
langchain-ai/langchain-skills
INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.
greyhaven-ai/autocontext
Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.
bgauryy/octocode
Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.
bgauryy/octocode
Researches code with evidence: traces callers, imports and cross-repo links, diagnoses failures and reports findings with exact file and line references and a confidence label.
bgauryy/octocode
Writes, repairs and copyedits project docs against the Google developer documentation style guide, verifying claims in the repository before stating them.
bgauryy/octocode
Poses and animates a 22-bone anatomical humanoid rig with joint range-of-motion limits, using a Node CLI, a Three.js viewer and WebMCP tools an agent can drive live.
bgauryy/octocode
Finds, rates, reviews, creates, improves, installs and syncs Agent Skill folders from local workspaces, registries or remote sources, with a user gate before any write.
bgauryy/octocode
Walks an idea through framing, diverging into options, researching evidence and stress-testing before converging on a build, prototype, narrow or park decision.
Works with
Categories
Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification. The skill structures evaluation as a fixed flow: error-analyze, frame a goal into a KPI, measure a baseline, loop a change, judge the result, capture a lesson, verify on held-out data, then evolve the test suite. It refuses to proceed without a goal linked to a measurable KPI and a runnable sensor to check it, and rejects accepting a change on narrative alone or by editing the harness or test cases just to make them pass.
Octocode Graph Eval Loop fits situations like: setting up a measurable improvement loop for an agent or system change; deciding whether a proposed change actually improved a held-out metric; checking that a verification step is truly independent of the change it is judging.
Run `npx skills add bgauryy/octocode --skill octocode-graph-eval -a claude-code`. Or copy the skill folder (skills/octocode-graph-eval in bgauryy/octocode) into .claude/skills/octocode-graph-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add bgauryy/octocode --skill octocode-graph-eval -a codex`. Or copy the skill folder (skills/octocode-graph-eval in bgauryy/octocode) into .agents/skills/octocode-graph-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bgauryy/octocode --skill octocode-graph-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/octocode-graph-eval, .gemini/skills/octocode-graph-eval, .github/skills/octocode-graph-eval and .opencode/skills/octocode-graph-eval in your project.
SKILL.md names no scripts, command-line tools or credentials: Octocode Graph Eval Loop is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Octocode Graph Eval Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Octocode Graph Eval Loop: Benchmark Agents (vercel/vercel-plugin, 301 stars), Waza Interactive (microsoft/waza, 1.4k stars), AI Project Copilot (sun461941-hub/ai-project-copilot, 97 stars) and SWE Benchmark Task Adder (ory/lumen, 307 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
bgauryy (a GitHub user) maintains it in bgauryy/octocode, which has 949 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 9, 2026.
Source: bgauryy/octocode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.