LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…
$ npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ProfSynapse/nexus nexus-eval-harness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.skills/nexus-eval-harness .claude/skills/nexus-eval-harness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nexus-eval-harness" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-eval-harness into .claude/skills/nexus-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-eval-harness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-eval-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ProfSynapse/nexus nexus-eval-harness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.skills/nexus-eval-harness .agents/skills/nexus-eval-harness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nexus-eval-harness" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-eval-harness into .agents/skills/nexus-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-eval-harness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ProfSynapse/nexus nexus-eval-harness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.skills/nexus-eval-harness .cursor/skills/nexus-eval-harness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nexus-eval-harness" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-eval-harness into .cursor/skills/nexus-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-eval-harness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ProfSynapse/nexus.git --path .skills/nexus-eval-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ProfSynapse/nexus nexus-eval-harness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.skills/nexus-eval-harness .gemini/skills/nexus-eval-harness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nexus-eval-harness" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-eval-harness into .gemini/skills/nexus-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-eval-harness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ProfSynapse/nexus nexus-eval-harnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .github/skills && cp -r skills-src/.skills/nexus-eval-harness .github/skills/nexus-eval-harness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nexus-eval-harness" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-eval-harness into .github/skills/nexus-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-eval-harness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ProfSynapse/nexus nexus-eval-harness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.skills/nexus-eval-harness .opencode/skills/nexus-eval-harness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nexus-eval-harness" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-eval-harness into .opencode/skills/nexus-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-eval-harness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nexus-eval-harnessWork on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…
Nexus Eval Harness is an agent skill from ProfSynapse/nexus. Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing, everything-fails, or numbers that disagree. Use when an eval scenario is wrong, a run behaves oddly, or the harness itself needs to change. To grade a model rather than change the harness, use nexus-model-eval.
Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts and reference files (for example `agents/openai.yaml`, `protocols/add-a-scenario.md` and `protocols/configure-a-run.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit eb20895. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nexus Eval Harness loads about 1k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 466 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ProfSynapse/nexus at commit eb20895, republished under its MIT licence (© ProfSynapse). 466 words, ~1,018 tokens.
.claude/skills/nexus-eval-harness/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Context: the harness under tests/eval/ drives the real production path —
StreamingOrchestrator plus tool continuation — with a mock or live tool
executor swapped in, and grades the tool calls a model emits. It is a fixture
system, and almost every surprising result is the fixture talking, not the
model. This skill owns the fixtures, the configs, the executors and the
reports.
Pick the job and open its protocol. Work from the protocol; this router names procedures, it does not contain them.
| Job | Protocol |
|---|---|
| Add a scenario, or fix one that grades wrongly | protocols/add-a-scenario.md |
| Write or change a config, choose targets, mode, retries | protocols/configure-a-run.md |
| A run produced nothing, all-fails, a hang, or odd numbers | protocols/debug-a-run.md |
| Change the executors, assertions, loader or reports | protocols/extend-the-harness.md |
Derive every list from the tree, never from this skill. It names no scenarios, no configs, no models and no env-var table on purpose, and you MUST NOT add one — the harness gains knobs faster than a document survives.
ls tests/eval/scenarios/ tests/eval/configs/
grep -rhoE "get(Number|List)?Env\('[A-Z_]+'\)|process\.env\.[A-Z_]+" tests/eval/ \
| grep -oE "[A-Z][A-Z_]{3,}" | sort -u # every knob, including the
# ones ConfigLoader mediatesBefore calling any scenario change done, run the checker from the repo root and fix everything it prints:
python3 .claude/skills/nexus-eval-harness/scripts/check_scenarios.pyNEVER trust jest's exit code as the verdict on a run, and never report a
pass rate you read from stdout. The saved reports under the configured
artifacts dir are the only source of truth, and they are written even when
the run times out — see references/run-behavior.md.
At the end of a session that used this skill, run protocols/self-refine.md.
protocols/ the procedures named in step 1, plus self-refine.md.references/ read on demand: harness-map.md (what each file owns and how a
run is assembled), scenario-contract.md (what a scenario fixture means and
the traps in it), run-behavior.md (config resolution, concurrency, retries,
artifacts, what the numbers count).scripts/check_scenarios.py the mechanical check from step 3. Run it; do not
reimplement it.refinement-log.md what past sessions changed here and why.agents/openai.yaml an interface manifest several nexus-* skills carry.
Not a subagent prompt; nothing in this skill reads it.nexus-model-eval — grading models. Which models to run, whether a slug
resolves, how to read a leaderboard, and whether a failure indicts the model.
That skill consumes the harness; this one changes it. If the question is "how
good is model X", stop here and use it.nexus-testing — the gate that keeps this suite from running (and billing) in
CI, and how to watch a run in flight.nexus-agents — the real getTools/useTools contract the fixtures imitate.
When a fixture and production disagree, production wins, and that skill says
what production does.nexus-llm-adapters — provider adapters. A run that fails inside streaming
for one provider only is an adapter problem, not a harness problem.nexus-tool-schemas — the live tool catalog, for checking a fixture's tool
slugs against the registry instead of guessing.© ProfSynapse, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (scripts, references) in .skills/nexus-eval-harness of ProfSynapse/nexus.
Open the folder on GitHubat commit eb20895
Nexus Eval Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nexus Eval Harness this skillProfSynapse/nexus | 157 | — | ~1k | Automated safety check: Pass | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
ProfSynapse/nexus
How to add, change and verify a Nexus agent or tool, and the contract every tool must satisfy.
ProfSynapse/nexus
Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's.
ProfSynapse/nexus
Add, change or verify a Nexus LLM model definition — the registry entry, the provider default, and proof the model id actually works against the live endpoint.
ProfSynapse/nexus
Cut a Nexus release — bump the version with the repo's own machinery, get the docs and generated sources right, push a tag the GitHub Actions workflow will actually pick up, and recover when it does…
ProfSynapse/nexus
How to change, persist, migrate and recover Nexus data without losing it.
ProfSynapse/nexus
Verify a Nexus change — pick a Jest lane, write a test that can actually fail, run the in-app Obsidian CLI loop, drive the eval harness, or fix a shipped-docs drift failure.
Categories
Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…. Nexus Eval Harness is an agent skill from ProfSynapse/nexus. Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing, everything-fails, or numbers that disagree.
Nexus Eval Harness fits situations like: an eval scenario is wrong; A run behaves oddly; the harness itself needs to change.
Run `npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a claude-code`. Or copy the skill folder (.skills/nexus-eval-harness in ProfSynapse/nexus) into .claude/skills/nexus-eval-harness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a codex`. Or copy the skill folder (.skills/nexus-eval-harness in ProfSynapse/nexus) into .agents/skills/nexus-eval-harness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nexus-eval-harness, .gemini/skills/nexus-eval-harness, .github/skills/nexus-eval-harness and .opencode/skills/nexus-eval-harness in your project.
Going by SKILL.md and its folder, Nexus Eval Harness needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Nexus Eval Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Nexus Eval Harness: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ProfSynapse (a GitHub user) maintains it in ProfSynapse/nexus, which has 157 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 8, 2026.
Source: ProfSynapse/nexus on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.