Arize Evaluator
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Evaluate a single retort experiment run. An agent skill from adrianco/retort.
$ npx skills add adrianco/retort --skill evaluate-run -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install adrianco/retort evaluate-run --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/adrianco/retort.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluate-run .claude/skills/evaluate-run && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluate-run" agent skill from https://github.com/adrianco/retort/tree/main/skills/evaluate-run into .claude/skills/evaluate-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-run", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/adrianco/retort/tree/main/skills/evaluate-runType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add adrianco/retort --skill evaluate-run -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install adrianco/retort evaluate-run --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adrianco/retort.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/evaluate-run .agents/skills/evaluate-run && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluate-run" agent skill from https://github.com/adrianco/retort/tree/main/skills/evaluate-run into .agents/skills/evaluate-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-run", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add adrianco/retort --skill evaluate-run -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install adrianco/retort evaluate-run --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adrianco/retort.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/evaluate-run .cursor/skills/evaluate-run && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluate-run" agent skill from https://github.com/adrianco/retort/tree/main/skills/evaluate-run into .cursor/skills/evaluate-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-run", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/adrianco/retort.git --path skills/evaluate-run--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add adrianco/retort --skill evaluate-run -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install adrianco/retort evaluate-run --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adrianco/retort.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/evaluate-run .gemini/skills/evaluate-run && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluate-run" agent skill from https://github.com/adrianco/retort/tree/main/skills/evaluate-run into .gemini/skills/evaluate-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-run", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install adrianco/retort evaluate-runInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add adrianco/retort --skill evaluate-run -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/adrianco/retort.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/evaluate-run .github/skills/evaluate-run && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-run" agent skill from https://github.com/adrianco/retort/tree/main/skills/evaluate-run into .github/skills/evaluate-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-run", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add adrianco/retort --skill evaluate-run -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install adrianco/retort evaluate-run --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/adrianco/retort.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/evaluate-run .opencode/skills/evaluate-run && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-run" agent skill from https://github.com/adrianco/retort/tree/main/skills/evaluate-run into .opencode/skills/evaluate-run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-run", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluate-runEvaluate a single retort experiment run. An agent skill from adrianco/retort.
Evaluate Run is an agent skill from adrianco/retort. Evaluate a single retort experiment run. Score the generated code against the task's TASK.md requirements, run its build and tests, compute metrics, and emit a structured evaluation report plus a machine-readable findings file.
Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `evaluate-run.py`).
The repository describes itself as: Platform Evolution Engine. Distill the best from the combinatorial mess. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1f75769. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
python3sqlite3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluate Run loads about 4.2k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 1,328 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from adrianco/retort at commit 1f75769, republished under its Apache-2.0 licence (© adrianco). 1,328 words, ~4,166 tokens.
.claude/skills/evaluate-run/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.A retort run produces a workspace directory — generated source code for one factor-level combination — archived under <experiment>/runs/<cell>/rep<N>/. This skill evaluates that workspace against the task spec it was asked to implement, captures quantitative and qualitative findings, and writes results in a format comparable across runs.
This is the per-run counterpart to pourpoise's evaluate-attempt, adapted for retort's DoE structure: instead of ad-hoc attempts, each run is a point in a design matrix.
experiment-1/runs/language=rust_model=opus_tooling=beads/rep2/{run_dir}/evaluation.md): Where to write the human-readable report{run_dir}/findings.jsonl): Where to write structured findings (one JSON object per line) suitable for file-run-issuesEach run_dir is laid out by retort's LocalRunner and contains:
| File | Purpose |
|---|---|
TASK.md | Task spec — the prompt the agent received. This is the "requirements" source of truth. |
stack.json | {"language": ..., "agent": ..., "framework": ...} — the factor levels for this run |
| All generated source files | Exactly as the agent left them |
Possibly .beads/ | Only if tooling=beads was in effect — the agent used bd for tracking |
| Possibly build artifacts | node_modules/, target/, __pycache__/, etc. |
The retort database (experiment-<N>/retort.db) also holds this run's ExperimentRun + RunResult rows. You MAY query it read-only for cross-checking scores; you MUST NOT write to it.
test -d "{run_dir}" || { echo "run_dir missing"; exit 1; }
test -f "{run_dir}/TASK.md" || { echo "TASK.md missing — not a retort workspace"; exit 1; }
test -f "{run_dir}/stack.json" || { echo "stack.json missing"; exit 1; }Constraints:
run_dir. Run all commands read-only or in a temp copy.-failed on the directory). Evaluate what exists; note the failure up front.retort.dbDo NOT re-run the build, tests, or linter. retort's scorers already ran them for this run during scoring and stored the results — re-running the toolchain (especially compiled/JVM languages) is the slowest part of evaluation and is pure duplication. Read the stored scores instead.
Fastest source — {run_dir}/scores.json. When the eval runs inline as a
gate during retort run, the run isn't in retort.db yet, so the runner drops
the just-computed mechanical scores into scores.json in the archive. If it
exists, read it and skip the DB query:
[ -f "{run_dir}/scores.json" ] && cat "{run_dir}/scores.json" # {"test_coverage": 1.0, "code_quality": 0.83, ...}If scores.json is absent (e.g. retroactive retort evaluate), fall back to the
database.
The database is at <experiment>/retort.db. run_dir is runs/<cell>/<rep>,
so walk up until you find retort.db (don't hard-code a level count — the
nesting can vary). Match this run by the factors in stack.json plus the
replicate (the trailing repN of run_dir, also in _meta.json):
db=""; d="{run_dir}"
for _ in 1 2 3 4 5; do d="$(cd "$d/.." && pwd)"; [ -f "$d/retort.db" ] && { db="$d/retort.db"; break; }; done
lang=$(python3 -c "import json;print(json.load(open('{run_dir}/stack.json'))['language'])")
model=$(python3 -c "import json;print(json.load(open('{run_dir}/stack.json')).get('model',''))")
tooling=$(python3 -c "import json;print(json.load(open('{run_dir}/stack.json')).get('tooling',''))")
rep=$(basename "{run_dir}" | sed -E 's/rep([0-9]+).*/\1/')
# A resumed/retried cell can have BOTH a stale `failed` row (test_coverage=0)
# and the real `completed` row for the same (factors, replicate). Pull scores
# from the single most-recent matching run, preferring the archive's own state:
# a `-failed` run_dir -> the failed row, otherwise the completed row.
want_status=completed
case "{run_dir}" in *-failed) want_status=failed;; esac
sqlite3 -readonly "$db" "
SELECT rr.metric_name, rr.value
FROM run_results rr
WHERE rr.run_id = (
SELECT er.id FROM experiment_runs er
WHERE json_extract(er.run_config_json,'\$.language')='$lang'
AND json_extract(er.run_config_json,'\$.model')='$model'
AND json_extract(er.run_config_json,'\$.tooling')='$tooling'
AND er.replicate=$rep AND er.status='$want_status'
ORDER BY er.finished_at DESC LIMIT 1)
AND rr.metric_name IN ('test_coverage','code_quality','defect_rate',
'maintainability','idiomatic','token_efficiency');"Interpret the stored scores (all 0–1) — these stand in for re-running:
test_coverage — coverage / pass-rate. 1.0 ⇒ build + all tests passed; 0.0 ⇒ tests did not execute (build or import failure — the test gate). Use this as the build+test signal.code_quality — lint/quality score. Use it for the Lint line.defect_rate — 1.0 ⇒ build+test succeeded.Constraints:
unavailable, not failed.First: prefer a pinned requirement list. Per-run requirement extraction is
non-deterministic (the same task yields different counts on different runs,
which makes requirement_coverage non-comparable). So if the experiment ships a
fixed list, you MUST use it verbatim. Walk up from run_dir (as you did for
retort.db) to find REQUIREMENTS.json:
req=""; d="{run_dir}"
for _ in 1 2 3 4 5; do d="$(cd "$d/.." && pwd)"; [ -f "$d/REQUIREMENTS.json" ] && { req="$d/REQUIREMENTS.json"; break; }; doneIf REQUIREMENTS.json exists, its requirements[] array IS the checklist —
use those exact ids and requirement texts, in that order, as the complete
and only list. Do NOT add, drop, merge, or re-number any. The denominator
(total) is fixed at len(requirements) for every run of this task. Skip
the extraction below entirely; go straight to step 4. (how_to_verify on each
entry tells you what evidence to look for.)
Otherwise (no pinned list), extract requirements as below.
The run must conform to the full prompt the agent was actually given. retort
assembles that prompt as: "Read TASK.md … implement everything it asks for" +
(a tooling instruction) + (only when a prompt factor was set) the contents of
prompts/<level>.md. So there are up to two requirement sources:
TASK.md — the task spec, always present. Parse into a checklist (R1, R2, …). Typical patterns:1. Implement ...), "must"/"should" bullets, code-fenced API signatures.stack.json has prompt set to something other than none/absent. Then read prompts/<prompt>.md from the experiment dir (where workspace.yaml lives — walk up from run_dir like you did for retort.db). Extract its additional, checkable instructions as prompt requirements (P1, P2, …) and verify the code/output followed them.Ignore prompts.txt — it is a benchmark-template placeholder (it literally begins with #ignore this file), NOT the prompt retort gave the agent. Do not derive requirements from it.
Constraints:
R<N> for TASK.md, P<N> for prompt-factor instructions) so comparisons across runs align.prompt factor, so the P* list is usually empty — that's fine; TASK.md is then the whole spec.This is the conformance gate: a run that doesn't implement the spec (and follow the prompt) is a failure, so be accurate — cite evidence, don't guess.
For each R<N> (TASK.md) and each P<N> (prompt), classify as one of:
implemented — code clearly satisfies it, tests exercise itpartial — code attempts it but is incomplete or untestedmissing — no evidence in the codebasecannot-verify — you genuinely can't tell from the code (rare). Use sparingly with evidence.Tests are non-negotiable: if test_coverage == 0 (tests did not run), the run already FAILS the test gate — that is always a failure, full stop. Note it up front and don't dress it up as cannot-verify.
Base the assessment on:
test_coverage from Step 2 (1.0 ⇒ build + all tests pass; 0.0 ⇒ tests did not execute, so treat unverified requirements as cannot-verify)Constraints:
implemented solely because it has a stub function.Skips inflate pass rates without verifying behavior. Count them:
# Python
grep -rE "pytest\.skip|@pytest\.mark\.skip|xfail" tests/ --include="*.py" 2>/dev/null | wc -l
# Go
grep -rE "t\.Skip\(|t\.Skipf\(" . --include="*.go" 2>/dev/null | wc -l
# Rust
grep -rE "#\[ignore\]|#\[cfg\(ignore\)\]" . --include="*.rs" 2>/dev/null | wc -l
# TypeScript (jest/vitest)
grep -rE "\.skip\(|xit\(|xdescribe\(|it\.todo\(" . --include="*.ts" --include="*.js" 2>/dev/null | wc -lConstraints:
effective_tests = passed + failed (skipped excluded).skipped_test finding for each skip, even if the skip looks "reasonable" — the signal matters for cross-run comparison.# Lines of code (exclude build artifacts)
cloc . --exclude-dir=node_modules,target,__pycache__,.git,dist,build 2>/dev/null | tail -20
# File count
find . -type f \
-not -path "*/node_modules/*" -not -path "*/target/*" \
-not -path "*/__pycache__/*" -not -path "*/.git/*" \
| wc -l
# Dependency count (language-appropriate)
case $lang in
python) wc -l requirements.txt pyproject.toml 2>/dev/null ;;
typescript) node -e "const p=require('./package.json');console.log(Object.keys({...p.dependencies,...p.devDependencies}).length)" 2>/dev/null ;;
go) grep -c "^\s*\S" go.sum 2>/dev/null ;;
rust) grep -cE "^\S+ = " Cargo.toml 2>/dev/null ;;
esacIf cloc isn't available, fall back to a simple wc -l loop over source files for the language's extensions only. Never include node_modules, target, etc.
Delegate architecture analysis to the run-summary skill:
summarize codebase {run_dir} to {run_dir}/summary/This produces structured markdown under {run_dir}/summary/ covering modules, interfaces, and flow. Reference it from the final report rather than duplicating its content.
One JSON object per line, one object per finding. Schema:
{"id": "R3", "kind": "requirement_missing", "severity": "high", "title": "No pagination support on GET /books", "evidence": "src/app.py:42 returns full list unconditionally", "suggestion": "Add ?limit and ?offset query params"}
{"id": "test-skip-1", "kind": "skipped_test", "severity": "medium", "title": "test_concurrent_writes is skipped", "evidence": "tests/test_app.py:87 @pytest.mark.skip", "suggestion": "Implement the concurrency check or delete the test"}
{"id": "build-fail", "kind": "build_failure", "severity": "critical", "title": "cargo build fails with E0308", "evidence": "src/main.rs:23 — mismatched types", "suggestion": "Fix the type signature before this run can be scored"}Allowed kind values:
requirement_missing, requirement_partialbuild_failure, test_failureskipped_test, disabled_testlint_warning, security_concerndoc_missing, enhancementAllowed severity: critical, high, medium, low, info.
Constraints:
evidence — the file + line or command + output snippet that backs the claim.Use the template in Output Format below. The human-readable report links to findings.jsonl and summary/index.md rather than inlining them.
# Evaluation: {cell_name} · rep {replicate}
## Summary
- **Factors:** language={lang}, model={model}, tooling={tooling} (plus any extras)
- **Status:** ok | failed ({reason}) | cannot-verify ({reason})
- **Requirements:** {implemented}/{total} implemented, {partial} partial, {missing} missing
- **Tests:** {passed} passed / {failed} failed / {skipped} skipped ({effective} effective)
- **Build:** {pass|fail|unavailable} — {duration}s
- **Lint:** {pass|fail|unavailable} — {warning_count} warnings
- **Architecture:** see `summary/index.md`
- **Findings:** {n} items in `findings.jsonl` ({critical} critical, {high} high, ...)
## Requirements
| ID | Requirement (short) | Status | Evidence |
|----|----|----|----|
| R1 | ... | ✓ implemented | `src/app.py:Book` |
| R2 | ... | ~ partial | `src/app.py:list_books` — no pagination |
| R3 | ... | ✗ missing | no search endpoint found |
## Build & Test
```text
{build command}
{first 40 lines of output, elided if long}{test command}
{test summary + failures}| Metric | Value |
|---|---|
| Lines of code (source only) | {n} |
| Files | {n} |
| Dependencies | {n} |
| Tests total | {n} |
| Tests effective | {n} |
| Skip ratio | {pct}% |
| Build duration | {s}s |
Top 5 by severity (full list in findings.jsonl):
cd {run_dir}
{exact commands used above, in order}
## Interaction with retort
- The retort CLI invokes this skill after each successful run (see `cli.py:_evaluate_run`). You SHOULD assume the archive already exists when this skill is called.
- Evaluation failures MUST NOT abort the experiment — the skill exits with stderr written but always exit code 0 so the run loop continues.
- Results are cached per run — if `evaluation.md` already exists and is newer than all source files in `run_dir`, the skill MAY exit early (idempotent re-invocation).
## Constraints Summary
- You MUST NOT modify files in `run_dir` except under `summary/`, and MUST create `evaluation.md` and `findings.jsonl` inside `run_dir`.
- You MUST NOT write to `retort.db` or any file outside `run_dir`.
- You MUST finish in under 5 minutes wall-clock. If you can't, emit whatever you have and return.
- You MUST cite file:line evidence for every finding.
- You MUST keep the output deterministic enough that re-running against the same workspace produces the same requirement IDs and the same findings (order may differ).
## Troubleshooting
**Toolchain missing (e.g. `cargo: command not found`)**
- Mark build/test as `unavailable`.
- Add a finding `toolchain_missing` (severity: info) so cross-run comparison knows why this run wasn't verified.
**TASK.md looks generic / doesn't list discrete requirements**
- Extract one requirement per imperative sentence in the prompt.
- Emit a `doc_missing` info finding noting that the task spec is under-specified.
**`run-summary` skill fails**
- Continue without it. Note in evaluation.md under Architecture: "summary skill unavailable".
- Do not let summary failure prevent the evaluation report from being written.© adrianco, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/evaluate-run of adrianco/retort.
Open the folder on GitHubat commit 1f75769
Evaluate Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluate Run this skilladrianco/retort | 207 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 1 repos | ~8.1k | Automated safety check: Notes | MIT | |
| EvaluatorsArize-ai/phoenix | 12k | — | ~1.7k | Automated safety check: Pass | Custom licence | |
| LLM Evaluationdavila7/claude-code-templates | 32k | 12 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Agent Evaluationsickn33/agentic-awesome-skills | 47k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Claw Scoreopenclaw/openclaw | 392k | — | ~2.5k | Automated safety check: Pass | MIT |
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Arize-ai/phoenix
Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
sickn33/agentic-awesome-skills
Evaluate agent behavior with versioned cases and explicit verifiers.
openclaw/openclaw
Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.
PostHog/posthog
Resolves a PostHog experiment reference from natural language to a concrete experiment ID by browsing experiment-list (not feature-flag tools), with disambiguation when multiple experiments match.
adrianco/retort
Compare evaluated runs in a retort experiment along factor dimensions.
adrianco/retort
Determine the TRUE cause of a failed retort run before attributing it.
adrianco/retort
Aggregate a retort run's findings.jsonl into a machine-readable assessment.json summary with severity counts, penalty score, requirement coverage, and top findings.
adrianco/retort
Summarize the architecture of code generated by a single retort run.
adrianco/retort
Refresh the data tables in optimal-blog.md from master.db. An agent skill from adrianco/retort.
Evaluate a single retort experiment run. An agent skill from adrianco/retort. Evaluate Run is an agent skill from adrianco/retort. Evaluate a single retort experiment run.
Run `npx skills add adrianco/retort --skill evaluate-run -a claude-code`. Or copy the skill folder (skills/evaluate-run in adrianco/retort) into .claude/skills/evaluate-run in your project. Claude Code loads it when a task matches its description.
Run `npx skills add adrianco/retort --skill evaluate-run -a codex`. Or copy the skill folder (skills/evaluate-run in adrianco/retort) into .agents/skills/evaluate-run in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adrianco/retort --skill evaluate-run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-run, .gemini/skills/evaluate-run, .github/skills/evaluate-run and .opencode/skills/evaluate-run in your project.
Going by SKILL.md and its folder, Evaluate Run needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and sqlite3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evaluate Run is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evaluate Run: Arize Evaluator (github/awesome-copilot, 40k stars), Evaluators (Arize-ai/phoenix, 12k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars) and Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
adrianco (a GitHub user) maintains it in adrianco/retort, which has 207 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 9, 2026.
Source: adrianco/retort on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.