Agent skill

Octocode Graph Eval Loop

by bgauryy in bgauryy/octocode

Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

MITAuto-check passedAgent Workflows

Install Octocode Graph Eval Loop

skills CLI
$ npx skills add bgauryy/octocode --skill octocode-graph-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bgauryy/octocode octocode-graph-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bgauryy/octocode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/octocode-graph-eval .claude/skills/octocode-graph-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
octocode-graph-eval
GitHub stars
949
Token cost
~1.6k tokens
SKILL.md length
742 words
Files
31 (incl. scripts, references)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

  • Works in 4 steps: Error-analyze traces into a failure… → Measure a fixed-budget baseline; make… → Judge grader quality, fairness,… → …
  • Setting up a measurable improvement loop for an agent or system change
  • SKILL.md covers Lobby rules, Workflow, Smart routes — load only what… and Related routes and verification
  • Deciding whether a proposed change actually improved a held-out metric

What it does

The skill structures evaluation as a fixed flow: error-analyze, frame a goal into a KPI, measure a baseline, loop a change, judge the result, capture a lesson, verify on held-out data, then evolve the test suite. It refuses to proceed without a goal linked to a measurable KPI and a runnable sensor to check it, and rejects accepting a change on narrative alone or by editing the harness or test cases just to make them pass.

Its rules call for deterministic graders over binary or LLM judgment where possible, a change accepted only when the primary metric improves on held-out data and guardrail metrics still hold, and a counter-metric guardrail for every primary KPI so it cannot be gamed by tuning alone. A verifier sharing the same context as the agent that made the change does not count as independent; fresh context is required first. Multi-agent workflows are checked for true independence between steps, and the harness is frozen during an experiment. Reference files cover benchmarking, error analysis and failure modes, loaded only as each step needs them.

When your agent uses it

  • Setting up a measurable improvement loop for an agent or system change
  • Deciding whether a proposed change actually improved a held-out metric
  • Checking that a verification step is truly independent of the change it is judging

Example prompts

  • “Frame a goal and KPI for reducing our agent's tool-call error rate, then baseline it.”
  • “Run one improvement loop on this retrieval change and judge it against the held-out set.”
  • “Check whether our evaluation harness was edited just to make the new version pass.”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Error-analyze traces into a failure taxonomy; frame success, primary/leading metrics, guardrails, and decision rule.
  2. Measure a fixed-budget baseline; make the smallest subject change; keep or discard from comparable results.
  3. Judge grader quality, fairness, capability versus regression, and contamination; capture one durable lesson.
  4. Verify held-out results and required checks; then add new failure cases between experiments.

What it can do on your machine

Read from SKILL.md and the folder at commit c265e3f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Octocode Graph Eval Loop loads about 1.6k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 742 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bgauryy/octocode at commit c265e3f, republished under its MIT licence (© bgauryy). 742 words, ~1,591 tokens.

Download SKILL.mdSave it as .claude/skills/octocode-graph-eval/SKILL.md (or your agent's skills folder). This skill also uses 30 other files; get the full folder from GitHub.
name
octocode-graph-eval
description
Use when you need a measurable keep/discard loop — goal→KPI, baseline vs target, held-out checks, eval suites, or don't-stop-till-done against a runnable sensor. Not for ordinary ship checks where 'tests passed' is enough.

Octocode Graph Eval

Evaluate outcomes and run improvement loops with evidence, not vibes — for one loop or a graph of loops. Flow: ERROR-ANALYZE → FRAME(goal→KPI) → BASELINE → LOOP → JUDGE → CAPTURE → VERIFY → SUITE-EVOLVE. Modes: ErrorAnalyze · Define · Run · Suite · Benchmark · Audit.

Lobby rules

  • No goal→KPI link → STOP. No measurable primary → STOP. No runnable sensor → build one before looping.
  • Narrative-only accept → REJECT. Editing harness/cases/graders to pass → REJECT.
  • ACCEPT only if primary moves on held-out and guardrails hold.
  • Prefer deterministic graders; binary/LLM next; humans calibrate. Grade outcomes over paths.
  • TDD for agents: write or select a failing case / KPI check before mutating the subject; green only after the change (red → green → keep|discard).
  • Public benches orient; private failure suites gate ships. Distrust saturated/contaminated boards.
  • Freeze the harness during an experiment; evolve the suite only between experiments.
  • Graph check: before evaluating a multi-agent workflow, run edge detection — if no two nodes are independent (every step reads the prior step's output), it is a loop, not a graph. Build a loop.
  • Goodhart guard: every primary KPI must have a counter-metric guardrail the agent cannot tune. Primary improving + guardrail degrading → reframe the goal, not the loop.
  • Verifier independence: a verifier sharing the executor's context is not independent. Require fresh context before calling a result verified.
  • Anchor requirement: every graph must have at least one node whose output cannot be argued with (tests that ran, build exit codes, type errors). No anchors → build one before trusting the graph.

Workflow

  1. Error-analyze traces into a failure taxonomy; frame success, primary/leading metrics, guardrails, and decision rule.
  2. Measure a fixed-budget baseline; make the smallest subject change; keep or discard from comparable results.
  3. Judge grader quality, fairness, capability versus regression, and contamination; capture one durable lesson.
  4. Verify held-out results and required checks; then add new failure cases between experiments. Stop when goal/KPI is undefined, checks did not run, the harness changed to pass, or another loop cannot change the verdict.

Smart routes — load only what the current step needs

  • When deriving failures, load references/error-analysis.md; when connecting intent to measures load references/goal-kpi-cascade.md, then fill references/kpi-contract.md — make success and budget explicit.
  • When choosing experiment, suite, or meta scope, load references/nested-loops.md; before the first iteration load references/feedback-loops.md, then for the inner keep/discard cycle load references/agent-loop.md — no workable sensor, no loop.
  • When the subject is a multi-agent workflow (graph of loops), load references/graph-of-loops.md — run edge detection first, require anchor nodes, check verifier independence, name Goodhart guardrails, then set primary KPI at the graph boundary with per-node sensors.
  • When auditing that graph for structural failure risk before trusting its green lights — shared context, opaque state, no checkpoint/resume, unbounded tool permissions, missing human gates — load references/graph-failure-modes.md; add a suite case on a mode's first trace appearance.
  • When managing or measuring subagents under eval, load references/subagent-cookbook.md first for the ownership split; spawn mechanics stay in octocode-subagent.
  • When running an evaluated multi-agent iteration, load references/subagent-protocol.md for the frozen FRAME→verdict protocol; when choosing worker and graph-boundary metrics, load references/subagent-kpis.md so spawn cost is measured, not invisible.
  • When defining how parent and workers talk during an evaluated run, load references/subagent-communication.md — bad channels create false certainty and unattributable failures; when choosing the topology itself, load references/subagent-approaches.md because the pattern decides which KPIs and checks matter.
  • When inner loop is flat and no new hypothesis exists, suspect stuck search priors — load references/nested-loops.md for bilevel escalation, then references/karpathy-patterns.md for the Bilevel Autoresearch pattern.
  • When selecting graders or statistical checks, load references/eval-techniques.md; when grading agent tool-call sequences or multi-turn trajectories load references/trajectory-grading.md; when trusting public/private suites load references/benchmarking.md — match evidence strength to the decision.
  • When creating cases and runners, load references/eval-harness.md; before acceptance load references/held-out-and-guards.md — prevent leakage, overfitting, and greenwashing.
  • When grounding methods in primary patterns, load references/karpathy-patterns.md — anchor techniques in proven loops.
  • When a result needs another skill or durable capture, load references/routing.md; when closing a meta improvement cycle load references/improve-loop.md — transfer ownership without losing the decision rule.
  • When reporting, load references/output.md and run scripts/loop-report.mjs — require goal, baseline, result, and verdict.
Show full SKILL.md (86 more words)Show less
  • Use octocode-research for evidence under test; octocode-brainstorming before evaluating an unresolved idea; octocode-rfc-generator for a design KPI contract.
  • Use octocode-subagent to fan out parallel hypotheses or benchmark trials within one iteration — measurement, keep/discard, graders, and the subagent cookbook (references/subagent-cookbook.md) stay frozen here.
  • Use octocode-prompt-optimizer for wording after the KPI is fixed; octocode-skills for folder edits after ACCEPT.
  • When changing this skill, run scripts/check-description.mjs then scripts/eval-eval.mjs --self-test and a matching --case — catch trigger and self-routing regressions; cases live in evals/ (cases.json, trigger-cases.json, kpi-contract.json).

© bgauryy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 30 other files (scripts, references) in skills/octocode-graph-eval of bgauryy/octocode.

  • SKILL.md
  • README.md
  • evals/cases.json
  • evals/kpi-contract.json
  • evals/trigger-cases.json
  • references/agent-loop.md
  • references/benchmarking.md
  • references/error-analysis.md
  • references/eval-harness.md
  • references/eval-techniques.md
  • references/feedback-loops.md
  • references/goal-kpi-cascade.md
  • references/graph-failure-modes.md
  • references/graph-of-loops.md
  • references/held-out-and-guards.md
  • references/improve-loop.md
  • references/karpathy-patterns.md
  • references/kpi-contract.md
  • references/nested-loops.md
  • … and 12 more

Open the folder on GitHubat commit c265e3f

Compare with similar skills

Octocode Graph Eval Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Octocode Graph Eval Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Octocode Graph Eval Loop this skillbgauryy/octocode949—~1.6kAutomated safety check: PassMIT
Benchmark Agentsvercel/vercel-plugin301—~3.6kAutomated safety check: PassCustom licence
Waza Interactivemicrosoft/waza1.4k—~1.3kAutomated safety check: PassMIT
AI Project Copilotsun461941-hub/ai-project-copilot97—~3kAutomated safety check: PassMIT
SWE Benchmark Task Adderory/lumen307—~497Automated safety check: PassCustom licence
Managed Deep Agentslangchain-ai/langchain-skills1.3k—~8.7kAutomated safety check: NotesMIT

Similar skills

  • Benchmark Agents

    vercel/vercel-plugin

    Official

    Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

    301 GitHub stars~3.6k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Waza Interactive

    microsoft/waza

    Official

    Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

    1.4k GitHub stars~1.3k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • AI Project Copilot

    sun461941-hub/ai-project-copilot

    A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

    97 GitHub stars~3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

    307 GitHub stars~497 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Managed Deep Agents

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.

    1.3k GitHub stars~8.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed

More from bgauryy/octocode

All 12 skills in this repo
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Octocode Code Research

    bgauryy/octocode

    Researches code with evidence: traces callers, imports and cross-repo links, diagnoses failures and reports findings with exact file and line references and a confidence label.

    949 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Writes, repairs and copyedits project docs against the Google developer documentation style guide, verifying claims in the repository before stating them.

    949 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Octocode Mannequin

    bgauryy/octocode

    Poses and animates a 22-bone anatomical humanoid rig with joint range-of-motion limits, using a Node CLI, a Three.js viewer and WebMCP tools an agent can drive live.

    949 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Octocode Skills Manager

    bgauryy/octocode

    Finds, rates, reviews, creates, improves, installs and syncs Agent Skill folders from local workspaces, registries or remote sources, with a user gate before any write.

    949 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Octocode Brainstorming

    bgauryy/octocode

    Walks an idea through framing, diverging into options, researching evidence and stress-testing before converging on a build, prototype, narrow or park decision.

    949 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed

Questions about Octocode Graph Eval Loop

What does Octocode Graph Eval Loop do?

Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification. The skill structures evaluation as a fixed flow: error-analyze, frame a goal into a KPI, measure a baseline, loop a change, judge the result, capture a lesson, verify on held-out data, then evolve the test suite. It refuses to proceed without a goal linked to a measurable KPI and a runnable sensor to check it, and rejects accepting a change on narrative alone or by editing the harness or test cases just to make them pass.

When should I use Octocode Graph Eval Loop?

Octocode Graph Eval Loop fits situations like: setting up a measurable improvement loop for an agent or system change; deciding whether a proposed change actually improved a held-out metric; checking that a verification step is truly independent of the change it is judging.

How do I install Octocode Graph Eval Loop in Claude Code?

Run `npx skills add bgauryy/octocode --skill octocode-graph-eval -a claude-code`. Or copy the skill folder (skills/octocode-graph-eval in bgauryy/octocode) into .claude/skills/octocode-graph-eval in your project. Claude Code loads it when a task matches its description.

How do I install Octocode Graph Eval Loop in Codex?

Run `npx skills add bgauryy/octocode --skill octocode-graph-eval -a codex`. Or copy the skill folder (skills/octocode-graph-eval in bgauryy/octocode) into .agents/skills/octocode-graph-eval in your project. Codex loads it when a task matches its description.

Can I use Octocode Graph Eval Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bgauryy/octocode --skill octocode-graph-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/octocode-graph-eval, .gemini/skills/octocode-graph-eval, .github/skills/octocode-graph-eval and .opencode/skills/octocode-graph-eval in your project.

What does Octocode Graph Eval Loop need to run?

SKILL.md names no scripts, command-line tools or credentials: Octocode Graph Eval Loop is instructions for the agent only.

Does Octocode Graph Eval Loop access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Octocode Graph Eval Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Octocode Graph Eval Loop use?

Octocode Graph Eval Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Octocode Graph Eval Loop use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Octocode Graph Eval Loop?

Skills that share tags, products or a category with Octocode Graph Eval Loop: Benchmark Agents (vercel/vercel-plugin, 301 stars), Waza Interactive (microsoft/waza, 1.4k stars), AI Project Copilot (sun461941-hub/ai-project-copilot, 97 stars) and SWE Benchmark Task Adder (ory/lumen, 307 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Octocode Graph Eval Loop?

bgauryy (a GitHub user) maintains it in bgauryy/octocode, which has 949 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 9, 2026.

Source: bgauryy/octocode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.