LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Quality and performance evaluation with baseline comparison.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/epicsagas/epic-harness/skills/eval .claude/skills/eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/epicsagas/epic-harness/skills/eval into .claude/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/epicsagas/epic-harness/skills/evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/epicsagas/epic-harness/skills/eval .agents/skills/eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/epicsagas/epic-harness/skills/eval into .agents/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/epicsagas/epic-harness/skills/eval .cursor/skills/eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/epicsagas/epic-harness/skills/eval into .cursor/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/hashgraph-online/awesome-codex-plugins.git --path plugins/epicsagas/epic-harness/skills/eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/epicsagas/epic-harness/skills/eval .gemini/skills/eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/epicsagas/epic-harness/skills/eval into .gemini/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install hashgraph-online/awesome-codex-plugins evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add hashgraph-online/awesome-codex-plugins --skill eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/epicsagas/epic-harness/skills/eval .github/skills/eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/epicsagas/epic-harness/skills/eval into .github/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add hashgraph-online/awesome-codex-plugins --skill eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install hashgraph-online/awesome-codex-plugins eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/epicsagas/epic-harness/skills/eval .opencode/skills/eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval" agent skill from https://github.com/hashgraph-online/awesome-codex-plugins/tree/main/plugins/epicsagas/epic-harness/skills/eval into .opencode/skills/eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evalQuality and performance evaluation with baseline comparison.
Eval is an agent skill from hashgraph-online/awesome-codex-plugins. Quality and performance evaluation with baseline comparison. Sub-modes: correctness, performance, quality, regression. Outputs PASS/WARN/FAIL per dimension. Use for pre-ship evaluation or regression checks.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. It works with Rust. The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 78497e5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitmakeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval loads about 2k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 753 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from hashgraph-online/awesome-codex-plugins at commit 78497e5, republished under its Apache-2.0 licence (© hashgraph-online). 753 words, ~2,012 tokens.
.claude/skills/eval/SKILL.md (or your agent's skills folder).CRITICAL: Run HARNESS_DIR=$(epic path) first. Never use .harness/ in the project directory.
/ship creates a PR (automatic if eval.yaml exists)/go completes a feature/eval commandmake eval or epic eval --json4 dimensions run in parallel where possible:
HARNESS_DIR=$(epic path)If $HARNESS_DIR/eval/eval.yaml does not exist, run scaffold:
epic eval --initRead the config:
cat $HARNESS_DIR/eval/eval.yamlIf eval.yaml has benchmarks: [] and no benchmark files are found in the project:
Generate stub files using the CLI:
epic eval --scaffoldSupported stacks (auto-detected from project markers):
| Stack | Detected by | Generated file | Output format |
|---|---|---|---|
| Rust | Cargo.toml | benches/eval_harness.rs | criterion (exit code) |
| Python | pyproject.toml / setup.py | benchmarks/eval_runner.py | JSON composite |
| TypeScript | tsconfig.json | benchmarks/eval.ts | JSON composite |
| Node.js | package.json | benchmarks/eval.mjs | JSON composite |
| Go | go.mod | benchmarks/eval_test.go | JSON composite |
| Java | pom.xml / build.gradle | benchmarks/EvalBenchmark.java | exit code |
| Kotlin | build.gradle.kts | benchmarks/EvalBenchmark.kt | exit code |
| Ruby | Gemfile | benchmarks/eval_benchmark.rb | JSON composite |
| PHP | composer.json | benchmarks/eval_benchmark.php | JSON composite |
| C# | *.csproj / *.sln | Benchmarks/EvalBenchmark.cs | JSON composite |
| Swift | Package.swift | benchmarks/EvalBenchmark.swift | JSON composite |
| Elixir | mix.exs | benchmarks/eval_benchmark.exs | JSON composite |
| C++ | CMakeLists.txt | benchmarks/eval_benchmark.cpp | exit code |
Customize the generated file — every file has # TODO / // TODO markers:
If --scaffold can't generate a useful stub (domain too complex, custom evaluation logic needed), generate a custom benchmark with LLM assistance:
{"composite": 0.0–1.0, ...}benchmarks/eval_runner.{ext} matching the project languageWire into eval.yaml:
benchmarks:
- name: eval_runner
command: python3 benchmarks/eval_runner.py full
result_type: composite # parse composite field from JSON stdoutUse result_type: exit_code for frameworks (criterion, JMH, BenchmarkDotNet) that manage their own output.
Execute the structured evaluation via the Rust binary:
epic eval --jsonThis runs all enabled dimensions and outputs a JSON result. Capture the output.
If the CLI reports llm_judge: SKIPPED (no LLM available in CLI mode), proceed to Step 2 for LLM-as-judge. Otherwise, skip to Step 3.
If the quality dimension has llm_judge: true and CLI marked it SKIPPED:
Sample 3–5 changed files from the current branch:
git diff --name-only $(git merge-base HEAD main)For each sampled file, evaluate on a 1-10 rubric:
Average scores across files. Map to 0.0–1.0 scale.
Record results alongside CLI output.
cat $HARNESS_DIR/eval/baselines/latest.jsonIf no baseline exists, the current run BECOMES the first baseline. Save it:
epic eval --baseline-updateReport: "First baseline established. Future runs will compare against this."
Combine CLI output + LLM-as-judge results into a single report:
## Eval Report
- Branch: {branch}
- Commit: {commit_short}
### Correctness: [PASS/WARN/FAIL] — score: {score}
- Tests: {passed}/{total} passing ({pass_rate}%)
- Mutation score: {mutation_score}% (if enabled)
- Delta vs baseline: {+/-delta}
### Performance: [PASS/WARN/FAIL] — score: {score} (if enabled)
- Avg latency: {latency}ms (delta: {+/-delta})
- Throughput: {throughput} (delta: {+/-delta})
### Quality: [PASS/WARN/FAIL] — score: {score}
- Lint errors: {count}
- LLM judge: {score}/10 (if enabled)
### Regression: [PASS/FAIL]
| Dimension | Baseline | Current | Delta | Verdict |
|-----------|----------|---------|-------|---------|
| correctness | {prev} | {cur} | {delta} | {pass/fail} |
| quality | {prev} | {cur} | {delta} | {pass/fail} |
### Overall: [PASS/WARN/FAIL] — {overall_score}/ship to create a PR."/go, then re-run /eval."epic eval --baseline-update # if user approves this as new baselineResults auto-saved to $HARNESS_DIR/eval/results/EVAL-{timestamp}.json.
| Excuse | Rebuttal | What to do instead |
|---|---|---|
| "Tests pass, no need for eval" | Tests pass today but regress tomorrow without baselines | Run eval and establish a baseline |
| "Performance testing is premature" | Latency regressions are invisible until users complain | Enable performance dimension, run benchmarks now |
| "Mutation testing is too slow" | Slow mutation catches bugs fast tests miss | Run on changed modules only (--dimension correctness) |
| "LLM-as-judge is subjective" | Subjective beats absent — fixed rubric + averaging reduces variance | Use the 4-axis rubric, average across 3+ files |
| "We can add eval later" | Later never comes; regressions accumulate silently | Start with correctness+quality, add dimensions incrementally |
| "CI will catch regressions" | CI only catches build/test failures, not quality drift | Eval measures what CI misses: mutation score, LLM quality |
epic eval --json output captured (all enabled dimensions scored)$HARNESS_DIR/eval/results/epic eval output© hashgraph-online, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/epicsagas/epic-harness/skills/eval of hashgraph-online/awesome-codex-plugins.
Open the folder on GitHubat commit 78497e5
Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval this skillhashgraph-online/awesome-codex-plugins | 1.2k | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
hashgraph-online/awesome-codex-plugins
Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.
hashgraph-online/awesome-codex-plugins
Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).
hashgraph-online/awesome-codex-plugins
A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…
hashgraph-online/awesome-codex-plugins
Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…
hashgraph-online/awesome-codex-plugins
Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.
hashgraph-online/awesome-codex-plugins
Analyze nonfiction manuscripts for reader engagement signals, including heading-level word counts, slow starts, long slogs, weak takeaway titles, value pacing, beta-reader comment dropoff, and…
Works with
Categories
Quality and performance evaluation with baseline comparison. Eval is an agent skill from hashgraph-online/awesome-codex-plugins. Quality and performance evaluation with baseline comparison.
Eval fits situations like: pre-ship evaluation; regression checks.
Run `npx skills add hashgraph-online/awesome-codex-plugins --skill eval -a claude-code`. Or copy the skill folder (plugins/epicsagas/epic-harness/skills/eval in hashgraph-online/awesome-codex-plugins) into .claude/skills/eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add hashgraph-online/awesome-codex-plugins --skill eval -a codex`. Or copy the skill folder (plugins/epicsagas/epic-harness/skills/eval in hashgraph-online/awesome-codex-plugins) into .agents/skills/eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval, .gemini/skills/eval, .github/skills/eval and .opencode/skills/eval in your project.
Going by SKILL.md and its folder, Eval needs the command-line tools its instructions call (git and make). Our summary lists: Python 3; Node.js.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,242 GitHub stars. The repository holds 686 skills in this directory. The repository was last updated on October 8, 2026.
Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.