Agent skill

Nexus Model Eval

by ProfSynapse in ProfSynapse/nexus

Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's.

MITAuto-check passedAI & LLM Engineering

Install Nexus Model Eval

skills CLI
$ npx skills add ProfSynapse/nexus --skill nexus-model-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ProfSynapse/nexus nexus-model-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.skills/nexus-model-eval .claude/skills/nexus-model-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nexus-model-eval
GitHub stars
154
Token cost
~1k tokens
SKILL.md length
509 words
Files
10 (incl. scripts, references)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's.

  • Works in 5 steps: Get current truth before running… → Run the grade:… → You MUST attribute every failure before… → …
  • Compare models on Nexus tool use
  • SKILL.md covers Workflow, Map and Siblings
  • Runs Python scripts from its folder; calls python3

What it does

Nexus Model Eval is an agent skill from ProfSynapse/nexus. Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's. Use when asked to grade, benchmark, rank or compare models on Nexus tool use, when picking a default model, or when an eval report needs interpreting.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `protocols/attribute-failures.md`, `protocols/grade-models.md` and `protocols/self-refine.md`).

It sits in AI & LLM Engineering. The licence is MIT.

When your agent uses it

  • Compare models on Nexus tool use
  • Picking a default model
  • An eval report needs interpreting

Example prompts

  • “s fault or the harness”
  • “/nexus-model-eval”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Get current truth before running anything. A model cannot be graded on a
  2. Run the grade: protocols/grade-models.md. Read it before you start; a
  3. You MUST attribute every failure before you report a number
  4. Report both numbers — raw pass rate and the attributed rate that charges only
  5. At the end of a session that used this skill, run protocols/self-refine.md.

What it can do on your machine

Read from SKILL.md and the folder at commit 6cb7bce. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nexus Model Eval loads about 1k tokens when it runs, and up to ~3.5k if it reads all its reference files. Until then it costs about 79 tokens; SKILL.md has 509 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ProfSynapse/nexus at commit 6cb7bce, republished under its MIT licence (© ProfSynapse). 509 words, ~1,008 tokens.

Download SKILL.mdSave it as .claude/skills/nexus-model-eval/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
nexus-model-eval
description
Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's. Use when asked to grade, benchmark, rank or compare models on Nexus tool use, when picking a default model, or when an eval report needs interpreting.

Nexus model eval

Context: the harness in tests/eval/ shows a model the same two tools the app does — getTools for discovery, useTools for execution — and grades the calls it makes, not the prose it writes. This skill owns the verdict: which models to run, and what a FAIL actually means. Running, configuring and extending the harness itself belongs to nexus-eval-harness. This file routes; detail loads when you take the path.

Workflow

  1. Get current truth before running anything. A model cannot be graded on a fixture no model can satisfy, and the fixture set moves:
    bash
    ls tests/eval/scenarios/ tests/eval/configs/
    python3 .claude/skills/nexus-eval-harness/scripts/check_scenarios.py
    python3 .claude/skills/nexus-model-eval/scripts/check_advertised_tools.py
    A non-zero exit from the scenario checker means some scenario can never pass; resolve that first, and the fix belongs to nexus-eval-harness, not to this run. The advertised-tools gap is not a defect — it is the list of correct model behaviors this harness punishes, and you will need it in step 3.
  2. Run the grade: protocols/grade-models.md. Read it before you start; a summarized procedure is one you will improvise, and every scenario in the matrix costs live, billed API calls.
  3. You MUST attribute every failure before you report a number: protocols/attribute-failures.md. The harness fails models for things the model did not do, so a raw pass rate with unread failures is not a grade. scripts/summarize_eval.py --labels refuses to sign off while any failure is unlabelled.
  4. Report both numbers — raw pass rate and the attributed rate that charges only model-failure verdicts — plus what the excluded failures actually were. One number alone is either unfair to the model or unfair to the reader.
  5. At the end of a session that used this skill, run protocols/self-refine.md.
Show full SKILL.md (237 more words)Show less

Map

  • protocols/ the procedures: grade-models.md (target list → run → artifacts), attribute-failures.md (FAIL → verdict → defensible grade), self-refine.md.
  • references/ read on demand: what-is-graded.md (what makes a scenario pass, what a "turn" counts, how retries and exclusions move the number), harness-artifacts.md (symptom → cause → proof for failures the model did not cause — read this before blaming any model).
  • scripts/ run them, do not reimplement:
    • scripts/check_advertised_tools.py — the commands the eval system prompt tells the model to use that the executor cannot run, so obeying the prompt scores as a hallucination.
    • scripts/preflight_models.py — do these slugs exist, before the run spends money proving they do not.
    • scripts/summarize_eval.py — report JSON → per-model rollup, bucketed failures, and an attribution that is checked rather than asserted.
  • refinement-log.md what past sessions changed here and why.

Siblings

The boundary with nexus-eval-harness: it owns the instrument, this skill owns the verdict. Anything that changes the harness or its inputs — env knobs, target syntax, live mode and the headless vault, config YAML, scenario authoring, harness code — is that skill's. Anything that changes what you conclude about a model is this one's. When a run reveals a fixture defect, hand it over rather than fixing it here.

Also: nexus-model-updates owns provider model definitions and whether a model ID works at all (grade nothing until it does); nexus-testing owns Jest lanes and what a mock can prove; nexus-agents owns the two-tool contract the harness is imitating; nexus-llm-adapters owns the adapter a stream error comes from.

© ProfSynapse, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references) in .skills/nexus-model-eval of ProfSynapse/nexus.

  • SKILL.md
  • protocols/attribute-failures.md
  • protocols/grade-models.md
  • protocols/self-refine.md
  • references/harness-artifacts.md
  • references/what-is-graded.md
  • refinement-log.md
  • scripts/check_advertised_tools.py
  • scripts/preflight_models.py
  • scripts/summarize_eval.py

Open the folder on GitHubat commit 6cb7bce

Compare with similar skills

Nexus Model Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nexus Model Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nexus Model Eval this skillProfSynapse/nexus154—~1kAutomated safety check: PassMIT
Agent BuildershareAI-lab/learn-claude-code78k6 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.8k13 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 6 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.8k GitHub starsUsed in 13 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from ProfSynapse/nexus

All 12 skills in this repo
  • Nexus Agents

    ProfSynapse/nexus

    How to add, change and verify a Nexus agent or tool, and the contract every tool must satisfy.

    154 GitHub stars~843 tokensUpdated 5 days ago
    Auto-check passed
  • Nexus Eval Harness

    ProfSynapse/nexus

    Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…

    154 GitHub stars~1k tokensUpdated 5 days ago
    Auto-check passed
  • Nexus Model Updates

    ProfSynapse/nexus

    Add, change or verify a Nexus LLM model definition — the registry entry, the provider default, and proof the model id actually works against the live endpoint.

    154 GitHub stars~1.1k tokensUpdated 5 days ago
    Auto-check passed
  • Nexus Release

    ProfSynapse/nexus

    Cut a Nexus release — bump the version with the repo's own machinery, get the docs and generated sources right, push a tag the GitHub Actions workflow will actually pick up, and recover when it does…

    154 GitHub stars~944 tokensUpdated 5 days ago
    Auto-check passed
  • Nexus Storage

    ProfSynapse/nexus

    How to change, persist, migrate and recover Nexus data without losing it.

    154 GitHub stars~749 tokensUpdated 5 days ago
    Auto-check passed
  • Nexus Testing

    ProfSynapse/nexus

    Verify a Nexus change — pick a Jest lane, write a test that can actually fail, run the in-app Obsidian CLI loop, drive the eval harness, or fix a shipped-docs drift failure.

    154 GitHub stars~994 tokensUpdated 5 days ago
    Auto-check passed

Questions about Nexus Model Eval

What does Nexus Model Eval do?

Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's. Nexus Model Eval is an agent skill from ProfSynapse/nexus. Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's.

When should I use Nexus Model Eval?

Nexus Model Eval fits situations like: compare models on Nexus tool use; picking a default model; an eval report needs interpreting.

How do I install Nexus Model Eval in Claude Code?

Run `npx skills add ProfSynapse/nexus --skill nexus-model-eval -a claude-code`. Or copy the skill folder (.skills/nexus-model-eval in ProfSynapse/nexus) into .claude/skills/nexus-model-eval in your project. Claude Code loads it when a task matches its description.

How do I install Nexus Model Eval in Codex?

Run `npx skills add ProfSynapse/nexus --skill nexus-model-eval -a codex`. Or copy the skill folder (.skills/nexus-model-eval in ProfSynapse/nexus) into .agents/skills/nexus-model-eval in your project. Codex loads it when a task matches its description.

Can I use Nexus Model Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ProfSynapse/nexus --skill nexus-model-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nexus-model-eval, .gemini/skills/nexus-model-eval, .github/skills/nexus-model-eval and .opencode/skills/nexus-model-eval in your project.

What does Nexus Model Eval need to run?

Going by SKILL.md and its folder, Nexus Model Eval needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Nexus Model Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Nexus Model Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Nexus Model Eval use?

Nexus Model Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nexus Model Eval use?

About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Nexus Model Eval?

Skills that share tags, products or a category with Nexus Model Eval: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nexus Model Eval?

ProfSynapse (a GitHub user) maintains it in ProfSynapse/nexus, which has 154 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 2, 2026.

Source: ProfSynapse/nexus on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.