Agent skill

Bare Eval

by yonatangross in yonatangross/orchestkit

Run isolated eval and grading calls using CC 2.1.81 --bare mode.

MITAuto-check passedAgent Workflows

Install Bare Eval

skills CLI
$ npx skills add yonatangross/orchestkit --skill bare-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit bare-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/bare-eval .claude/skills/bare-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bare-eval
GitHub stars
289
Token cost
~2.2k tokens
SKILL.md length
873 words
Files
10 (incl. references)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Run isolated eval and grading calls using CC 2.1.81 --bare mode.

  • LLM grading without plugin/hook interference
  • SKILL.md covers When to Use, When NOT to Use, Prerequisites and Quick Reference, plus 9 more sections
  • Runs JavaScript scripts from its folder; calls claude, jq and npm; needs ANTHROPIC_API_KEY
  • Running eval pipelines

What it does

Bare Eval is an agent skill from yonatangross/orchestkit. Run isolated eval and grading calls using CC 2.1.81 --bare mode. Constructs claude -p --bare invocations for skill evaluation, trigger testing, and LLM grading without plugin/hook interference. Use when running eval pipelines, grading skill outputs, benchmarking prompt quality, or testing trigger accuracy in isolation.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `references/grading-schemas.md`, `references/invocation-patterns.md` and `references/troubleshooting.md`). Compatibility notes: Claude Code 2.1.277+

It sits in Agent Workflows, covering Agent evaluation and testing. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • LLM grading without plugin/hook interference
  • Running eval pipelines
  • Grading skill outputs
  • Benchmarking prompt quality

Example prompts

  • “/bare-eval”

Requirements

  • Node.js
  • A credential in ANTHROPIC_API_KEY
  • Compatibility (from SKILL.md): Claude Code 2.1.277+

What it can do on your machine

Read from SKILL.md and the folder at commit 0ef71d2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • claude
    • jq
    • npm
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+

    From compatibility in the SKILL.md frontmatter.

Context cost

Bare Eval loads about 2.2k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 873 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit 0ef71d2, republished under its MIT licence (© yonatangross). 873 words, ~2,242 tokens.

Download SKILL.mdSave it as .claude/skills/bare-eval/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
bare-eval
description
Run isolated eval and grading calls using CC 2.1.81 --bare mode. Constructs claude -p --bare invocations for skill evaluation, trigger testing, and LLM grading without plugin/hook interference. Use when running eval pipelines, grading skill outputs, benchmarking prompt quality, or testing trigger accuracy in isolation.
compatibility
Claude Code 2.1.277+
user-invocable
false
context
inherit
effort
low
metadata.version
1.1.0
metadata.author
OrchestKit
metadata.complexity
medium
metadata.tags
eval, bare, grading, pipeline, testing, ci

Bare Eval — Isolated Evaluation Calls

Run claude -p --bare for fast, clean eval/grading without plugin overhead.

CC 2.1.81 required. The --bare flag skips hooks, LSP, plugin sync, and skill directory walks.

When to Use

  • Grading skill outputs against assertions
  • Trigger classification (which skill matches a prompt)
  • Description optimization iterations
  • Any scripted -p call that doesn't need plugins

When NOT to Use

  • Testing skill routing (needs --plugin-dir)
  • Testing agent orchestration (needs full plugin context)
  • Interactive sessions

Prerequisites

bash
# --bare requires ANTHROPIC_API_KEY (OAuth/keychain disabled)
export ANTHROPIC_API_KEY="sk-ant-..."

# Verify CC version
claude --version  # Must be >= 2.1.81

Quick Reference

Call TypeCommand Pattern
Gradingclaude -p "$prompt" --bare --max-turns 1 --output-format text
Triggerclaude -p "$prompt" --bare --json-schema "$schema" --output-format json
Streaming gradeclaude -p "$prompt" --bare --max-turns 1 --output-format stream-json
Optimizeecho "$prompt" | claude -p --bare --max-turns 1 --output-format text
Force-skillclaude -p "$prompt" --bare --print --append-system-prompt "$content"
@-file in promptclaude -p "grade @fixtures/case-1.md against rubric" --bare (CC 2.1.113 Remote Control autocomplete)

Long harness runs (CC 2.1.199+): set CLAUDE_CODE_RETRY_WATCHDOG=1 for unattended eval batches — it raises the default retry count for non-capacity transient errors to 300 and lifts the cap of 15 on CLAUDE_CODE_MAX_RETRIES, so an overnight grading run survives transient API blips instead of dying mid-batch.

--output-format stream-json

Newline-delimited JSON events (one per token/tool-call) — lets a runner score partial output or abort early on a failing probe without waiting for the full response.

bash
claude -p "$prompt" --bare --max-turns 1 --output-format stream-json \
  | while IFS= read -r line; do
      # line is a single JSON event; inspect $.type == "content_block_delta"
      jq -r 'select(.type == "content_block_delta") | .delta.text' <<< "$line"
    done

Use stream-json over json when:

  • grading long outputs and you want incremental scoring,
  • piping into another CLI step-by-step (e.g. ork:eval-runner),
  • you need per-token timing data alongside the content.

Invocation Patterns

Load detailed patterns and examples:

Read("references/invocation-patterns.md")

Grading Schemas

JSON schemas for structured eval output:

Read("references/grading-schemas.md")

Pipeline Integration

OrchestKit's eval scripts (npm run eval:skill) auto-detect bare mode:

bash
# eval-common.sh detects ANTHROPIC_API_KEY → sets BARE_MODE=true
# Scripts add --bare to all non-plugin calls automatically

Bare calls: Trigger classification, force-skill, baseline, all grading. Never bare: run_with_skill (needs plugin context for routing tests).

CC 2.1.119: --print honors agent tools: / disallowedTools: (M122)

Before CC 2.1.119, --print mode ran with the full default tool set regardless of the agent's frontmatter tools: and disallowedTools:. Bare-eval grading was effectively ungated — graders could call any tool they wanted, even if the agent definition restricted them.

As of 2.1.119, --print enforces the agent's declared tool surface. Implications for eval design:

ConsequenceAction
Eval graders that relied on unrestricted tool access may now failAudit grader prompts for tools they actually need; whitelist explicitly via the agent's tools: frontmatter
Eval results match interactive runsReproducibility improves — grading what the model can actually do, not what it could do in an unsandboxed --print
--agent <name> also honors permissionMode in --printPermission-gated tools (Bash, Edit) require either permissionMode: acceptEdits or explicit allowlists in the agent definition

Migration test:

bash
# Run an eval against an agent with a deliberately tight tools: list.
# Graders that previously called Read/Bash freely will now fail unless those
# tools are declared on the agent.
claude -p "$prompt" --bare --print --agent grader-test

If the grader fails with a "tool not permitted" error, add the required tool to the agent's tools: frontmatter and re-run.

Show full SKILL.md (431 more words)Show less
CC 2.1.121: CLAUDE_CODE_FORK_SUBAGENT=1 for grader determinism (#1545)

Before CC 2.1.121, the env var only worked in interactive sessions. As of 2.1.121, non-interactive paths (claude -p, SDK) honor it too — each grader invocation gets a fresh forked subagent context.

The cross-eval state-leak problem this fixes:

Without forking, sequential claude -p --bare graders inherit harness state:

InheritedSymptom
memory MCP query cachegrader sees stale hit from previous run; same fixture grades differently
.claude/chain/*.json on diskgrader for "implement" thinks "explore" already ran (file is from previous test)
ToolSearch deferred-tool cachefirst grader's MCP loads bleed into next grader's tool registry
model picker prefgrader N inherits --model=opus from grader N-1

This produced ~5–10% retry rate and non-reproducible scores — the eval baseline drifted between runs, engineers chased phantom regressions.

Fix: tests/evals/scripts/lib/eval-common.sh exports CLAUDE_CODE_FORK_SUBAGENT=1, so every script that sources it (run-trigger-eval, run-quality-eval, run-agent-eval, optimize-description, etc.) gets forked graders automatically. The CI workflow .github/workflows/skill-eval.yml also sets it at the workflow level (its predecessor orchestkit-eval.yml was retired 2026-08-01 — its grading phases rated description prose, not behavior). Older CC silently ignores the env var (no-op).

Determinism contract: running the same grader on the same fixture twice in a row produces the same score. Verified by tests/evals/scripts/test-grader-determinism.sh.

Performance

ScenarioWithout --bareWith --bareSavings
Single grading call~3-5s startup~0.5-1s2-4x
Trigger (per prompt)~3-5s~0.5-1s2-4x
Full eval (50 calls)~150-250s overhead~25-50s3-5x

Rules

Read("rules/_sections.md")

Troubleshooting

Read("references/troubleshooting.md")

Dynamic-workflow harness (template-in-skill)

workflows/skill-fitness.js is a runnable dynamic-workflow template — the workflow-backed complement to the static conformance grader (scripts/eval/conformance-check.mjs). It fans out one isolated-context agent per skill to score fitness (freshness / router-clarity / structure) and synthesizes a ranked scorecard, catching qualitative drift a static grep can't (description/body count mismatches, duplicate headings, install-specific absolute paths, version drift). Run it with the Workflow tool:

Workflow({ scriptPath: "${CLAUDE_SKILL_DIR}/workflows/skill-fitness.js",
           args: ["assess", "commit", "doctor"] })

Treat it as a template, not a verbatim script — adapt the SKILLS list and rubric per use. Cost is real (~50k tokens/skill; scoring all ~112 is ~6M tokens), so pass an explicit batch via args. Static-first: run conformance-check.mjs (zero tokens) to pre-filter, then this harness for the judgment grep can't make.

Holdout Bake-Off Grading

The holdout-promotion gate grades a champion and a challenger SKILL.md over the same sealed holdout via bare-mode forked graders — the canonical consumer of the determinism contract above: identical grader + identical ork-rubric/1.0 + identical sealed set, with CLAUDE_CODE_FORK_SUBAGENT=1 so the only variable is the version under test. Both --bare constraints apply (requires ANTHROPIC_API_KEY, bills tokens directly → on-demand / CI only). Run it with bash tests/evals/scripts/run-skill-eval.sh --holdout-promote <skill>.

  • eval:skill npm script — unified skill evaluation runner
  • eval:trigger — trigger accuracy testing
  • eval:quality — A/B quality comparison
  • optimize-description.sh — iterative description improvement
  • Version compatibility: doctor/references/version-compatibility.md

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in src/skills/bare-eval of yonatangross/orchestkit.

  • SKILL.md
  • references/grading-schemas.md
  • references/invocation-patterns.md
  • references/troubleshooting.md
  • rules/_sections.md
  • rules/bare-grading-only.md
  • rules/bare-plugin-conflict.md
  • rules/bare-requires-api-key.md
  • test-cases.json
  • workflows/skill-fitness.js

Open the folder on GitHubat commit 0ef71d2

Compare with similar skills

Bare Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bare Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bare Eval this skillyonatangross/orchestkit289—~2.2kAutomated safety check: PassMIT
Copilot Session Failure Analysisdotnet/maui23k—~3.4kAutomated safety check: PassMIT
Autocontext for Hermesgreyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.0
Agentic Harness Design and ReviewNateBJones-Projects/OB14.7k—~1.8kAutomated safety check: PassCustom licence
Octocode Graph Eval Loopbgauryy/octocode949—~1.6kAutomated safety check: PassMIT
Write Skilldruxt/druxt.js114—~926Automated safety check: PassMIT

Similar skills

  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Agentic Harness Design and Review

    NateBJones-Projects/OB1

    Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.

    4.7k GitHub stars~1.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Octocode Graph Eval Loop

    bgauryy/octocode

    Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

    949 GitHub stars~1.6k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • Write Skill

    druxt/druxt.js

    Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it.

    114 GitHub stars~926 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    275k GitHub stars~1.5k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed

More from yonatangross/orchestkit

All 108 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    289 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    289 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    289 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    289 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    289 GitHub stars~4.5k tokensUpdated today
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    289 GitHub stars~3.9k tokensUpdated today
    Auto-check: notes

Questions about Bare Eval

What does Bare Eval do?

Run isolated eval and grading calls using CC 2.1.81 --bare mode. Bare Eval is an agent skill from yonatangross/orchestkit.81 --bare mode.

When should I use Bare Eval?

Bare Eval fits situations like: LLM grading without plugin/hook interference; running eval pipelines; grading skill outputs; benchmarking prompt quality.

How do I install Bare Eval in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill bare-eval -a claude-code`. Or copy the skill folder (src/skills/bare-eval in yonatangross/orchestkit) into .claude/skills/bare-eval in your project. Claude Code loads it when a task matches its description.

How do I install Bare Eval in Codex?

Run `npx skills add yonatangross/orchestkit --skill bare-eval -a codex`. Or copy the skill folder (src/skills/bare-eval in yonatangross/orchestkit) into .agents/skills/bare-eval in your project. Codex loads it when a task matches its description.

Can I use Bare Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill bare-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bare-eval, .gemini/skills/bare-eval, .github/skills/bare-eval and .opencode/skills/bare-eval in your project.

What does Bare Eval need to run?

Going by SKILL.md and its folder, Bare Eval needs JavaScript for the scripts in its folder, the command-line tools its instructions call (claude, jq, npm and bash) and credentials named ANTHROPIC_API_KEY. Our summary lists: Node.js; A credential in ANTHROPIC_API_KEY. Compatibility (from SKILL.md): Claude Code 2.1.277+.

Does Bare Eval access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bare Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bare Eval use?

Bare Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bare Eval use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Bare Eval?

Skills that share tags, products or a category with Bare Eval: Copilot Session Failure Analysis (dotnet/maui, 23k stars), Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars), Agentic Harness Design and Review (NateBJones-Projects/OB1, 4.7k stars) and Octocode Graph Eval Loop (bgauryy/octocode, 949 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bare Eval?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 289 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.