Copilot Session Failure Analysis
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
Run isolated eval and grading calls using CC 2.1.81 --bare mode.
$ npx skills add yonatangross/orchestkit --skill bare-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install yonatangross/orchestkit bare-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/bare-eval .claude/skills/bare-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bare-eval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/bare-eval into .claude/skills/bare-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bare-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/yonatangross/orchestkit/tree/main/src/skills/bare-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add yonatangross/orchestkit --skill bare-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install yonatangross/orchestkit bare-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/skills/bare-eval .agents/skills/bare-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bare-eval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/bare-eval into .agents/skills/bare-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bare-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill bare-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install yonatangross/orchestkit bare-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/skills/bare-eval .cursor/skills/bare-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bare-eval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/bare-eval into .cursor/skills/bare-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bare-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/yonatangross/orchestkit.git --path src/skills/bare-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add yonatangross/orchestkit --skill bare-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install yonatangross/orchestkit bare-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/skills/bare-eval .gemini/skills/bare-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bare-eval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/bare-eval into .gemini/skills/bare-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bare-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install yonatangross/orchestkit bare-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add yonatangross/orchestkit --skill bare-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/skills/bare-eval .github/skills/bare-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bare-eval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/bare-eval into .github/skills/bare-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bare-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill bare-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install yonatangross/orchestkit bare-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/skills/bare-eval .opencode/skills/bare-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bare-eval" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/bare-eval into .opencode/skills/bare-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bare-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bare-evalRun isolated eval and grading calls using CC 2.1.81 --bare mode.
Bare Eval is an agent skill from yonatangross/orchestkit. Run isolated eval and grading calls using CC 2.1.81 --bare mode. Constructs claude -p --bare invocations for skill evaluation, trigger testing, and LLM grading without plugin/hook interference. Use when running eval pipelines, grading skill outputs, benchmarking prompt quality, or testing trigger accuracy in isolation.
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `references/grading-schemas.md`, `references/invocation-patterns.md` and `references/troubleshooting.md`). Compatibility notes: Claude Code 2.1.277+
It sits in Agent Workflows, covering Agent evaluation and testing. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.
Read from SKILL.md and the folder at commit 0ef71d2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (JavaScript), which the agent can run.
Shell commands in SKILL.md call:
claudejqnpmbashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
ANTHROPIC_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Claude Code 2.1.277+
From compatibility in the SKILL.md frontmatter.
Bare Eval loads about 2.2k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 873 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from yonatangross/orchestkit at commit 0ef71d2, republished under its MIT licence (© yonatangross). 873 words, ~2,242 tokens.
.claude/skills/bare-eval/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.Run claude -p --bare for fast, clean eval/grading without plugin overhead.
CC 2.1.81 required. The --bare flag skips hooks, LSP, plugin sync, and skill directory walks.
-p call that doesn't need plugins--plugin-dir)# --bare requires ANTHROPIC_API_KEY (OAuth/keychain disabled)
export ANTHROPIC_API_KEY="sk-ant-..."
# Verify CC version
claude --version # Must be >= 2.1.81| Call Type | Command Pattern |
|---|---|
| Grading | claude -p "$prompt" --bare --max-turns 1 --output-format text |
| Trigger | claude -p "$prompt" --bare --json-schema "$schema" --output-format json |
| Streaming grade | claude -p "$prompt" --bare --max-turns 1 --output-format stream-json |
| Optimize | echo "$prompt" | claude -p --bare --max-turns 1 --output-format text |
| Force-skill | claude -p "$prompt" --bare --print --append-system-prompt "$content" |
| @-file in prompt | claude -p "grade @fixtures/case-1.md against rubric" --bare (CC 2.1.113 Remote Control autocomplete) |
Long harness runs (CC 2.1.199+): set
CLAUDE_CODE_RETRY_WATCHDOG=1for unattended eval batches — it raises the default retry count for non-capacity transient errors to 300 and lifts the cap of 15 onCLAUDE_CODE_MAX_RETRIES, so an overnight grading run survives transient API blips instead of dying mid-batch.
--output-format stream-jsonNewline-delimited JSON events (one per token/tool-call) — lets a runner score partial output or abort early on a failing probe without waiting for the full response.
claude -p "$prompt" --bare --max-turns 1 --output-format stream-json \
| while IFS= read -r line; do
# line is a single JSON event; inspect $.type == "content_block_delta"
jq -r 'select(.type == "content_block_delta") | .delta.text' <<< "$line"
doneUse stream-json over json when:
ork:eval-runner),Load detailed patterns and examples:
Read("references/invocation-patterns.md")JSON schemas for structured eval output:
Read("references/grading-schemas.md")OrchestKit's eval scripts (npm run eval:skill) auto-detect bare mode:
# eval-common.sh detects ANTHROPIC_API_KEY → sets BARE_MODE=true
# Scripts add --bare to all non-plugin calls automaticallyBare calls: Trigger classification, force-skill, baseline, all grading.
Never bare: run_with_skill (needs plugin context for routing tests).
--print honors agent tools: / disallowedTools: (M122)Before CC 2.1.119, --print mode ran with the full default tool set regardless of the agent's frontmatter tools: and disallowedTools:. Bare-eval grading was effectively ungated — graders could call any tool they wanted, even if the agent definition restricted them.
As of 2.1.119, --print enforces the agent's declared tool surface. Implications for eval design:
| Consequence | Action |
|---|---|
| Eval graders that relied on unrestricted tool access may now fail | Audit grader prompts for tools they actually need; whitelist explicitly via the agent's tools: frontmatter |
| Eval results match interactive runs | Reproducibility improves — grading what the model can actually do, not what it could do in an unsandboxed --print |
--agent <name> also honors permissionMode in --print | Permission-gated tools (Bash, Edit) require either permissionMode: acceptEdits or explicit allowlists in the agent definition |
Migration test:
# Run an eval against an agent with a deliberately tight tools: list.
# Graders that previously called Read/Bash freely will now fail unless those
# tools are declared on the agent.
claude -p "$prompt" --bare --print --agent grader-testIf the grader fails with a "tool not permitted" error, add the required tool to the agent's tools: frontmatter and re-run.
CLAUDE_CODE_FORK_SUBAGENT=1 for grader determinism (#1545)Before CC 2.1.121, the env var only worked in interactive sessions. As of 2.1.121, non-interactive paths (claude -p, SDK) honor it too — each grader invocation gets a fresh forked subagent context.
The cross-eval state-leak problem this fixes:
Without forking, sequential claude -p --bare graders inherit harness state:
| Inherited | Symptom |
|---|---|
| memory MCP query cache | grader sees stale hit from previous run; same fixture grades differently |
.claude/chain/*.json on disk | grader for "implement" thinks "explore" already ran (file is from previous test) |
| ToolSearch deferred-tool cache | first grader's MCP loads bleed into next grader's tool registry |
| model picker pref | grader N inherits --model=opus from grader N-1 |
This produced ~5–10% retry rate and non-reproducible scores — the eval baseline drifted between runs, engineers chased phantom regressions.
Fix: tests/evals/scripts/lib/eval-common.sh exports CLAUDE_CODE_FORK_SUBAGENT=1, so every script that sources it (run-trigger-eval, run-quality-eval, run-agent-eval, optimize-description, etc.) gets forked graders automatically. The CI workflow .github/workflows/skill-eval.yml also sets it at the workflow level (its predecessor orchestkit-eval.yml was retired 2026-08-01 — its grading phases rated description prose, not behavior). Older CC silently ignores the env var (no-op).
Determinism contract: running the same grader on the same fixture twice in a row produces the same score. Verified by tests/evals/scripts/test-grader-determinism.sh.
| Scenario | Without --bare | With --bare | Savings |
|---|---|---|---|
| Single grading call | ~3-5s startup | ~0.5-1s | 2-4x |
| Trigger (per prompt) | ~3-5s | ~0.5-1s | 2-4x |
| Full eval (50 calls) | ~150-250s overhead | ~25-50s | 3-5x |
Read("rules/_sections.md")Read("references/troubleshooting.md")workflows/skill-fitness.js is a runnable dynamic-workflow template — the workflow-backed complement to the static conformance grader (scripts/eval/conformance-check.mjs). It fans out one isolated-context agent per skill to score fitness (freshness / router-clarity / structure) and synthesizes a ranked scorecard, catching qualitative drift a static grep can't (description/body count mismatches, duplicate headings, install-specific absolute paths, version drift). Run it with the Workflow tool:
Workflow({ scriptPath: "${CLAUDE_SKILL_DIR}/workflows/skill-fitness.js",
args: ["assess", "commit", "doctor"] })Treat it as a template, not a verbatim script — adapt the SKILLS list and rubric per use. Cost is real (~50k tokens/skill; scoring all ~112 is ~6M tokens), so pass an explicit batch via args. Static-first: run conformance-check.mjs (zero tokens) to pre-filter, then this harness for the judgment grep can't make.
The holdout-promotion gate grades a champion and a challenger SKILL.md over the same sealed holdout via bare-mode forked graders — the canonical consumer of the determinism contract above: identical grader + identical ork-rubric/1.0 + identical sealed set, with CLAUDE_CODE_FORK_SUBAGENT=1 so the only variable is the version under test. Both --bare constraints apply (requires ANTHROPIC_API_KEY, bills tokens directly → on-demand / CI only). Run it with bash tests/evals/scripts/run-skill-eval.sh --holdout-promote <skill>.
eval:skill npm script — unified skill evaluation runnereval:trigger — trigger accuracy testingeval:quality — A/B quality comparisonoptimize-description.sh — iterative description improvementdoctor/references/version-compatibility.md© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files (references) in src/skills/bare-eval of yonatangross/orchestkit.
Open the folder on GitHubat commit 0ef71d2
Bare Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bare Eval this skillyonatangross/orchestkit | 289 | — | ~2.2k | Automated safety check: Pass | MIT | |
| Copilot Session Failure Analysisdotnet/maui | 23k | — | ~3.4k | Automated safety check: Pass | MIT | |
| Autocontext for Hermesgreyhaven-ai/autocontext | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Agentic Harness Design and ReviewNateBJones-Projects/OB1 | 4.7k | — | ~1.8k | Automated safety check: Pass | Custom licence | |
| Octocode Graph Eval Loopbgauryy/octocode | 949 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Write Skilldruxt/druxt.js | 114 | — | ~926 | Automated safety check: Pass | MIT |
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
greyhaven-ai/autocontext
Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.
NateBJones-Projects/OB1
Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.
bgauryy/octocode
Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.
druxt/druxt.js
Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it.
affaan-m/ECC
Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.
yonatangross/orchestkit
API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.
yonatangross/orchestkit
ADR templates in the Nygard format with context, decision, consequences, and alternatives.
yonatangross/orchestkit
Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.
yonatangross/orchestkit
Structured review processes, conventional comments, language-specific checklists, and feedback templates.
yonatangross/orchestkit
Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.
yonatangross/orchestkit
Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.
Categories
Run isolated eval and grading calls using CC 2.1.81 --bare mode. Bare Eval is an agent skill from yonatangross/orchestkit.81 --bare mode.
Bare Eval fits situations like: LLM grading without plugin/hook interference; running eval pipelines; grading skill outputs; benchmarking prompt quality.
Run `npx skills add yonatangross/orchestkit --skill bare-eval -a claude-code`. Or copy the skill folder (src/skills/bare-eval in yonatangross/orchestkit) into .claude/skills/bare-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add yonatangross/orchestkit --skill bare-eval -a codex`. Or copy the skill folder (src/skills/bare-eval in yonatangross/orchestkit) into .agents/skills/bare-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill bare-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bare-eval, .gemini/skills/bare-eval, .github/skills/bare-eval and .opencode/skills/bare-eval in your project.
Going by SKILL.md and its folder, Bare Eval needs JavaScript for the scripts in its folder, the command-line tools its instructions call (claude, jq, npm and bash) and credentials named ANTHROPIC_API_KEY. Our summary lists: Node.js; A credential in ANTHROPIC_API_KEY. Compatibility (from SKILL.md): Claude Code 2.1.277+.
SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bare Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Bare Eval: Copilot Session Failure Analysis (dotnet/maui, 23k stars), Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars), Agentic Harness Design and Review (NateBJones-Projects/OB1, 4.7k stars) and Octocode Graph Eval Loop (bgauryy/octocode, 949 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 289 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.
Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.