Agent skill

Agent Evaluation Operations

by TheGoat395 in TheGoat395/Codex-Skills

Evaluate agent and skill behavior or routing. An agent skill from TheGoat395/Codex-Skills.

MITAuto-check passedAgent Workflows

Install Agent Evaluation Operations

skills CLI
$ npx skills add TheGoat395/Codex-Skills --skill agent-evaluation-operations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TheGoat395/Codex-Skills agent-evaluation-operations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TheGoat395/Codex-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-evaluation-operations .claude/skills/agent-evaluation-operations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-evaluation-operations
GitHub stars
126
Token cost
~1.5k tokens
SKILL.md length
732 words
Files
10 (incl. scripts, references)
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Evaluate agent and skill behavior or routing. An agent skill from TheGoat395/Codex-Skills.

  • Tasks that involve Agent evaluation and testing
  • SKILL.md covers Choose the evaluation mode, Specify before measuring, Score observable behavior and Test collisions, not only…, plus 3 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Agent Evaluation Operations is an agent skill from TheGoat395/Codex-Skills. Evaluate agent and skill behavior or routing.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/fixture-execution.md` and `references/fixtures/draft-send.json`).

It sits in Agent Workflows, covering Agent evaluation and testing. The repository describes itself as: Codex-first Agent Skills library for premium frontend, website, motion, accessibility, QA, and handoff workflows. The licence is MIT.

When your agent uses it

  • Tasks that involve Agent evaluation and testing

Example prompts

  • “/agent-evaluation-operations”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit feff772. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Evaluation Operations loads about 1.5k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 18 tokens; SKILL.md has 732 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from TheGoat395/Codex-Skills at commit feff772, republished under its MIT licence (© TheGoat395). 732 words, ~1,468 tokens.

Download SKILL.mdSave it as .claude/skills/agent-evaluation-operations/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
agent-evaluation-operations
description
Evaluate agent and skill behavior or routing.

Agent Evaluation Operations

Evaluate the claim the change is supposed to support, not the amount of new prompt text or the number of passing examples. Routine copyediting of a test comment, ordinary execution of one existing test, or a task that merely mentions agents does not require an evaluation program.

Choose the evaluation mode

  • Agent workflow: test the real prompt, model, tools, approvals, state, and failure paths.
  • Skill behavior: test whether the skill triggers on the right requests, stays out of unrelated requests, cooperates with adjacent skills, and improves outcomes without excessive context or rigidity.
  • Release comparison: hold the harness constant and compare the current baseline with the proposed change.

For evaluating a skill creation or behavior-bearing modification, read skill behavior evaluation. Use the regression corpus when testing synthetic public agent and Codex operating behavior. Validate the corpus with python3 "${CODEX_HOME:-$HOME/.codex}/skills/agent-evaluation-operations/scripts/validate_regression_corpus.py".

Specify before measuring

Record the evaluation claim, tested system, model/reasoning setting, prompt and skill versions, tool access, side-effect policy, attempt budget, acceptance threshold, and what would falsify the claim. Do not compare two runs that silently differ on these dimensions.

Build cases from real work: ordinary success, ambiguous input, missing data, conflicting evidence, missing access, unavailable tools, duplicate events, unsafe external actions, escalation, recovery, and every confirmed historical failure. Keep a small smoke set plus a growing regression set.

Trace run ID, tested-system version, model/reasoning setting, prompt and skill versions, tools, input class, structured result, error, latency, cost, and approval path. Redact or avoid sensitive payload capture by default.

Test retrieval and application separately. A skill may fail to trigger even when its rules are sound, or it may trigger and still fail to change behavior. Keep a small matched baseline and treatment set with identical prompts, inputs, model settings, tools, and viewports; score first attempts blind when subjective judgment matters.

Score observable behavior

Prefer deterministic assertions for file state, structured fields, tool calls, authorization boundaries, and exact completion status. Use written rubrics for judgment. Calibrate model graders against examples and human review; do not let the candidate skill be the sole judge of its own success.

Measure:

  • factual grounding and evidence quality;
  • scope coverage and completion-state accuracy;
  • tool and skill routing, including false-positive triggers;
  • action correctness, policy/approval adherence, and reversibility;
  • correction quality after contradictory evidence;
  • latency, tokens/cost, retries, and human review;
  • privacy, recoverability, and sensitive-data handling.

Test collisions, not only positive triggers

Run positive, negative, adjacent, overlap, precedence, and co-invocation cases. Use python3 "${CODEX_HOME:-$HOME/.codex}/skills/agent-evaluation-operations/scripts/analyze_skill_collisions.py" --root "${CODEX_HOME:-$HOME/.codex}/skills" only to generate candidate pairs; lexical similarity is discovery evidence, not proof of a semantic collision.

Reject a skill change when it attracts unrelated work, duplicates an existing owner without a routing reason, weakens a capability floor, improves only the curated examples, or adds more context cost than demonstrated value.

Show full SKILL.md (274 more words)Show less

Gate releases

  • Run baseline and treatment under the same harness.
  • Require the proposed change to fix its target regressions without material degradation elsewhere.
  • Fail the release when a required scenario, authorization boundary, cost budget, or quality threshold is missed.
  • Test external calls in a sandbox, fixture, or dry-run path before production.
  • Separate local structural validation, simulated behavior, and live production evidence.
  • Preserve run identifiers, versions, aggregate scores, failures, and reviewer overrides while redacting sensitive payloads.

Do not infer safety, production readiness, or broad behavioral improvement from one successful demonstration.

Correction loop

For recurring review feedback, preserve the original artifact, the correction, its source, the proposed destination, exceptions, and a holdout case. Treat collector, reviewer and maintainer as logical roles, using separate staffing only when authorized and justified. Gather raw evidence, verify and group it, then decide whether it becomes guidance, a component or token, a deterministic check, an exemplar, an evaluation fixture, a coverage gap, or no change. Rerun affected cases after accepted changes and watch whether the same complaint actually becomes less common.

The corpus validator checks structure only. Use the concrete synthetic dossiers and local observer example in fixture execution to connect cases to real inputs and observed task state; an unexecuted case is not a model result.

Use Promptfoo or another project-local harness only when its telemetry, credentials, remote execution, and configuration have been reviewed. Route agent-system architecture to $agent-orchestration-architecture; route repository security to $repository-release-security.

Optional specialists

For repository release checks, use repository-release-security only if installed; otherwise run the repository-required checks, inspect the patch for secrets and dependency changes, and preserve rollback. Its absence does not block an otherwise verified release.

© TheGoat395, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references) in skills/agent-evaluation-operations of TheGoat395/Codex-Skills.

  • SKILL.md
  • agents/openai.yaml
  • references/fixture-execution.md
  • references/fixtures/draft-send.json
  • references/regression-corpus.json
  • references/skill-behavior-evaluation.md
  • scripts/analyze_skill_collisions.py
  • scripts/observe_draft_fixture.py
  • scripts/test_observe_draft_fixture.py
  • scripts/validate_regression_corpus.py

Open the folder on GitHubat commit feff772

Compare with similar skills

Agent Evaluation Operations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Evaluation Operations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Evaluation Operations this skillTheGoat395/Codex-Skills126—~1.5kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers297k3 repos~1.7kAutomated safety check: PassMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch66k—~1kAutomated safety check: PassMIT
CodeGraph Agent Evalcolbymchenry/codegraph74k—~950Automated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    297k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    66k GitHub stars~1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    74k GitHub stars~950 tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from TheGoat395/Codex-Skills

All 23 skills in this repo
  • Premium Visual Reference Library

    TheGoat395/Codex-Skills

    Find references, Viktor Oddy and MotionSites. An agent skill from TheGoat395/Codex-Skills.

    126 GitHub stars~867 tokensUpdated 28 days ago
    Auto-check passed
  • Creativity

    TheGoat395/Codex-Skills

    Generate original ideas or escape fixation. An agent skill from TheGoat395/Codex-Skills.

    126 GitHub stars~1.2k tokensUpdated 28 days ago
    Auto-check passed
  • Final Client Handoff

    TheGoat395/Codex-Skills

    Prepare verified ownership and delivery notes. An agent skill from TheGoat395/Codex-Skills.

    126 GitHub stars~477 tokensUpdated 28 days ago
    Auto-check passed
  • Visual Regression Lab

    TheGoat395/Codex-Skills

    Compare captures; stitch lazy/animated pages. An agent skill from TheGoat395/Codex-Skills.

    126 GitHub stars~516 tokensUpdated 28 days ago
    Auto-check passed
  • Web Reference Research

    TheGoat395/Codex-Skills

    Analyze selected sources: Viktor or MotionSites. An agent skill from TheGoat395/Codex-Skills.

    126 GitHub stars~475 tokensUpdated 28 days ago
    Auto-check passed
  • Code Change Safety Checkpoint

    TheGoat395/Codex-Skills

    Preserve rollback before materially risky edits. An agent skill from TheGoat395/Codex-Skills.

    126 GitHub stars~654 tokensUpdated 28 days ago
    Auto-check passed

Categories

Questions about Agent Evaluation Operations

What does Agent Evaluation Operations do?

Evaluate agent and skill behavior or routing. An agent skill from TheGoat395/Codex-Skills. Agent Evaluation Operations is an agent skill from TheGoat395/Codex-Skills. Evaluate agent and skill behavior or routing.

When should I use Agent Evaluation Operations?

Agent Evaluation Operations fits situations like: tasks that involve Agent evaluation and testing.

How do I install Agent Evaluation Operations in Claude Code?

Run `npx skills add TheGoat395/Codex-Skills --skill agent-evaluation-operations -a claude-code`. Or copy the skill folder (skills/agent-evaluation-operations in TheGoat395/Codex-Skills) into .claude/skills/agent-evaluation-operations in your project. Claude Code loads it when a task matches its description.

How do I install Agent Evaluation Operations in Codex?

Run `npx skills add TheGoat395/Codex-Skills --skill agent-evaluation-operations -a codex`. Or copy the skill folder (skills/agent-evaluation-operations in TheGoat395/Codex-Skills) into .agents/skills/agent-evaluation-operations in your project. Codex loads it when a task matches its description.

Can I use Agent Evaluation Operations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TheGoat395/Codex-Skills --skill agent-evaluation-operations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-evaluation-operations, .gemini/skills/agent-evaluation-operations, .github/skills/agent-evaluation-operations and .opencode/skills/agent-evaluation-operations in your project.

What does Agent Evaluation Operations need to run?

Going by SKILL.md and its folder, Agent Evaluation Operations needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Agent Evaluation Operations access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Evaluation Operations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Evaluation Operations use?

Agent Evaluation Operations is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Evaluation Operations use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.5k tokens, read only when the agent opens those files.

What are the alternatives to Agent Evaluation Operations?

Skills that share tags, products or a category with Agent Evaluation Operations: MCP Server Builder (anthropics/skills, 180k stars), Diagnosing Superpowers Sessions (obra/superpowers, 297k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars) and Skill Release Gate (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Evaluation Operations?

TheGoat395 (a GitHub user) maintains it in TheGoat395/Codex-Skills, which has 126 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on September 11, 2026.

Source: TheGoat395/Codex-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.