Agent skill

Nexus Eval Harness

by ProfSynapse in ProfSynapse/nexus

Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…

MITAuto-check passedAI & LLM Engineering

Install Nexus Eval Harness

skills CLI
$ npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ProfSynapse/nexus nexus-eval-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.skills/nexus-eval-harness .claude/skills/nexus-eval-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nexus-eval-harness
GitHub stars
157
Token cost
~1k tokens
SKILL.md length
466 words
Files
12 (incl. scripts, references)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…

  • Works in 5 steps: Pick the job and open its protocol. Work… → Derive every list from the tree, never… → Before calling any scenario change done,… → …
  • An eval scenario is wrong
  • SKILL.md covers Workflow, Map and Siblings — name them, do not…
  • Runs Python scripts from its folder; calls python3

What it does

Nexus Eval Harness is an agent skill from ProfSynapse/nexus. Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing, everything-fails, or numbers that disagree. Use when an eval scenario is wrong, a run behaves oddly, or the harness itself needs to change. To grade a model rather than change the harness, use nexus-model-eval.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts and reference files (for example `agents/openai.yaml`, `protocols/add-a-scenario.md` and `protocols/configure-a-run.md`).

It sits in AI & LLM Engineering, covering LLM evaluation. The licence is MIT.

When your agent uses it

  • An eval scenario is wrong
  • A run behaves oddly
  • The harness itself needs to change

Example prompts

  • “/nexus-eval-harness”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Pick the job and open its protocol. Work from the protocol; this router
  2. Derive every list from the tree, never from this skill. It names no
  3. Before calling any scenario change done, run the checker from the repo root
  4. NEVER trust jest's exit code as the verdict on a run, and never report a
  5. At the end of a session that used this skill, run protocols/self-refine.md.

What it can do on your machine

Read from SKILL.md and the folder at commit eb20895. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nexus Eval Harness loads about 1k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 466 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ProfSynapse/nexus at commit eb20895, republished under its MIT licence (© ProfSynapse). 466 words, ~1,018 tokens.

Download SKILL.mdSave it as .claude/skills/nexus-eval-harness/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
nexus-eval-harness
description
Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing, everything-fails, or numbers that disagree. Use when an eval scenario is wrong, a run behaves oddly, or the harness itself needs to change. To grade a model rather than change the harness, use nexus-model-eval.

Nexus Eval Harness

Context: the harness under tests/eval/ drives the real production path — StreamingOrchestrator plus tool continuation — with a mock or live tool executor swapped in, and grades the tool calls a model emits. It is a fixture system, and almost every surprising result is the fixture talking, not the model. This skill owns the fixtures, the configs, the executors and the reports.

Workflow

  1. Pick the job and open its protocol. Work from the protocol; this router names procedures, it does not contain them.

    JobProtocol
    Add a scenario, or fix one that grades wronglyprotocols/add-a-scenario.md
    Write or change a config, choose targets, mode, retriesprotocols/configure-a-run.md
    A run produced nothing, all-fails, a hang, or odd numbersprotocols/debug-a-run.md
    Change the executors, assertions, loader or reportsprotocols/extend-the-harness.md
  2. Derive every list from the tree, never from this skill. It names no scenarios, no configs, no models and no env-var table on purpose, and you MUST NOT add one — the harness gains knobs faster than a document survives.

    bash
    ls tests/eval/scenarios/ tests/eval/configs/
    grep -rhoE "get(Number|List)?Env\('[A-Z_]+'\)|process\.env\.[A-Z_]+" tests/eval/ \
      | grep -oE "[A-Z][A-Z_]{3,}" | sort -u    # every knob, including the
                                                # ones ConfigLoader mediates
  3. Before calling any scenario change done, run the checker from the repo root and fix everything it prints:

    bash
    python3 .claude/skills/nexus-eval-harness/scripts/check_scenarios.py
  4. NEVER trust jest's exit code as the verdict on a run, and never report a pass rate you read from stdout. The saved reports under the configured artifacts dir are the only source of truth, and they are written even when the run times out — see references/run-behavior.md.

  5. At the end of a session that used this skill, run protocols/self-refine.md.

Show full SKILL.md (220 more words)Show less

Map

  • protocols/ the procedures named in step 1, plus self-refine.md.
  • references/ read on demand: harness-map.md (what each file owns and how a run is assembled), scenario-contract.md (what a scenario fixture means and the traps in it), run-behavior.md (config resolution, concurrency, retries, artifacts, what the numbers count).
  • scripts/check_scenarios.py the mechanical check from step 3. Run it; do not reimplement it.
  • refinement-log.md what past sessions changed here and why.
  • agents/openai.yaml an interface manifest several nexus-* skills carry. Not a subagent prompt; nothing in this skill reads it.

Siblings — name them, do not duplicate them

  • nexus-model-eval — grading models. Which models to run, whether a slug resolves, how to read a leaderboard, and whether a failure indicts the model. That skill consumes the harness; this one changes it. If the question is "how good is model X", stop here and use it.
  • nexus-testing — the gate that keeps this suite from running (and billing) in CI, and how to watch a run in flight.
  • nexus-agents — the real getTools/useTools contract the fixtures imitate. When a fixture and production disagree, production wins, and that skill says what production does.
  • nexus-llm-adapters — provider adapters. A run that fails inside streaming for one provider only is an adapter problem, not a harness problem.
  • nexus-tool-schemas — the live tool catalog, for checking a fixture's tool slugs against the registry instead of guessing.

© ProfSynapse, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts, references) in .skills/nexus-eval-harness of ProfSynapse/nexus.

  • SKILL.md
  • agents/openai.yaml
  • protocols/add-a-scenario.md
  • protocols/configure-a-run.md
  • protocols/debug-a-run.md
  • protocols/extend-the-harness.md
  • protocols/self-refine.md
  • references/harness-map.md
  • references/run-behavior.md
  • references/scenario-contract.md
  • refinement-log.md
  • scripts/check_scenarios.py

Open the folder on GitHubat commit eb20895

Compare with similar skills

Nexus Eval Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nexus Eval Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nexus Eval Harness this skillProfSynapse/nexus157—~1kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from ProfSynapse/nexus

All 12 skills in this repo
  • Nexus Agents

    ProfSynapse/nexus

    How to add, change and verify a Nexus agent or tool, and the contract every tool must satisfy.

    157 GitHub stars~843 tokensUpdated yesterday
    Auto-check passed
  • Nexus Model Eval

    ProfSynapse/nexus

    Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's.

    157 GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Nexus Model Updates

    ProfSynapse/nexus

    Add, change or verify a Nexus LLM model definition — the registry entry, the provider default, and proof the model id actually works against the live endpoint.

    157 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Nexus Release

    ProfSynapse/nexus

    Cut a Nexus release — bump the version with the repo's own machinery, get the docs and generated sources right, push a tag the GitHub Actions workflow will actually pick up, and recover when it does…

    157 GitHub stars~944 tokensUpdated yesterday
    Auto-check passed
  • Nexus Storage

    ProfSynapse/nexus

    How to change, persist, migrate and recover Nexus data without losing it.

    157 GitHub stars~749 tokensUpdated yesterday
    Auto-check passed
  • Nexus Testing

    ProfSynapse/nexus

    Verify a Nexus change — pick a Jest lane, write a test that can actually fail, run the in-app Obsidian CLI loop, drive the eval harness, or fix a shipped-docs drift failure.

    157 GitHub stars~994 tokensUpdated yesterday
    Auto-check passed

Questions about Nexus Eval Harness

What does Nexus Eval Harness do?

Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…. Nexus Eval Harness is an agent skill from ProfSynapse/nexus. Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing, everything-fails, or numbers that disagree.

When should I use Nexus Eval Harness?

Nexus Eval Harness fits situations like: an eval scenario is wrong; A run behaves oddly; the harness itself needs to change.

How do I install Nexus Eval Harness in Claude Code?

Run `npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a claude-code`. Or copy the skill folder (.skills/nexus-eval-harness in ProfSynapse/nexus) into .claude/skills/nexus-eval-harness in your project. Claude Code loads it when a task matches its description.

How do I install Nexus Eval Harness in Codex?

Run `npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a codex`. Or copy the skill folder (.skills/nexus-eval-harness in ProfSynapse/nexus) into .agents/skills/nexus-eval-harness in your project. Codex loads it when a task matches its description.

Can I use Nexus Eval Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ProfSynapse/nexus --skill nexus-eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nexus-eval-harness, .gemini/skills/nexus-eval-harness, .github/skills/nexus-eval-harness and .opencode/skills/nexus-eval-harness in your project.

What does Nexus Eval Harness need to run?

Going by SKILL.md and its folder, Nexus Eval Harness needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Nexus Eval Harness access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Nexus Eval Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Nexus Eval Harness use?

Nexus Eval Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nexus Eval Harness use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.7k tokens, read only when the agent opens those files.

What are the alternatives to Nexus Eval Harness?

Skills that share tags, products or a category with Nexus Eval Harness: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nexus Eval Harness?

ProfSynapse (a GitHub user) maintains it in ProfSynapse/nexus, which has 157 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 8, 2026.

Source: ProfSynapse/nexus on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.