Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's.
$ npx skills add ProfSynapse/nexus --skill nexus-model-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ProfSynapse/nexus nexus-model-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.skills/nexus-model-eval .claude/skills/nexus-model-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nexus-model-eval" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-model-eval into .claude/skills/nexus-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-model-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-model-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ProfSynapse/nexus --skill nexus-model-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ProfSynapse/nexus nexus-model-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.skills/nexus-model-eval .agents/skills/nexus-model-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nexus-model-eval" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-model-eval into .agents/skills/nexus-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-model-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ProfSynapse/nexus --skill nexus-model-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ProfSynapse/nexus nexus-model-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.skills/nexus-model-eval .cursor/skills/nexus-model-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nexus-model-eval" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-model-eval into .cursor/skills/nexus-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-model-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ProfSynapse/nexus.git --path .skills/nexus-model-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ProfSynapse/nexus --skill nexus-model-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ProfSynapse/nexus nexus-model-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.skills/nexus-model-eval .gemini/skills/nexus-model-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nexus-model-eval" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-model-eval into .gemini/skills/nexus-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-model-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ProfSynapse/nexus nexus-model-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ProfSynapse/nexus --skill nexus-model-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .github/skills && cp -r skills-src/.skills/nexus-model-eval .github/skills/nexus-model-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nexus-model-eval" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-model-eval into .github/skills/nexus-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-model-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ProfSynapse/nexus --skill nexus-model-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ProfSynapse/nexus nexus-model-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ProfSynapse/nexus.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.skills/nexus-model-eval .opencode/skills/nexus-model-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nexus-model-eval" agent skill from https://github.com/ProfSynapse/nexus/tree/main/.skills/nexus-model-eval into .opencode/skills/nexus-model-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nexus-model-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nexus-model-evalGrade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's.
Nexus Model Eval is an agent skill from ProfSynapse/nexus. Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's. Use when asked to grade, benchmark, rank or compare models on Nexus tool use, when picking a default model, or when an eval report needs interpreting.
Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `protocols/attribute-failures.md`, `protocols/grade-models.md` and `protocols/self-refine.md`).
It sits in AI & LLM Engineering. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 6cb7bce. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nexus Model Eval loads about 1k tokens when it runs, and up to ~3.5k if it reads all its reference files. Until then it costs about 79 tokens; SKILL.md has 509 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ProfSynapse/nexus at commit 6cb7bce, republished under its MIT licence (© ProfSynapse). 509 words, ~1,008 tokens.
.claude/skills/nexus-model-eval/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.Context: the harness in tests/eval/ shows a model the same two tools the app
does — getTools for discovery, useTools for execution — and grades the calls
it makes, not the prose it writes. This skill owns the verdict: which models to
run, and what a FAIL actually means. Running, configuring and extending the
harness itself belongs to nexus-eval-harness. This file routes; detail loads
when you take the path.
ls tests/eval/scenarios/ tests/eval/configs/
python3 .claude/skills/nexus-eval-harness/scripts/check_scenarios.py
python3 .claude/skills/nexus-model-eval/scripts/check_advertised_tools.pynexus-eval-harness, not to this
run. The advertised-tools gap is not a defect — it is the list of correct
model behaviors this harness punishes, and you will need it in step 3.protocols/grade-models.md. Read it before you start; a
summarized procedure is one you will improvise, and every scenario in the
matrix costs live, billed API calls.protocols/attribute-failures.md. The harness fails models for things the
model did not do, so a raw pass rate with unread failures is not a grade.
scripts/summarize_eval.py --labels refuses to sign off while any failure is
unlabelled.model-failure verdicts — plus what the excluded failures actually were. One
number alone is either unfair to the model or unfair to the reader.protocols/self-refine.md.protocols/ the procedures: grade-models.md (target list → run → artifacts),
attribute-failures.md (FAIL → verdict → defensible grade), self-refine.md.references/ read on demand: what-is-graded.md (what makes a scenario pass,
what a "turn" counts, how retries and exclusions move the number),
harness-artifacts.md (symptom → cause → proof for failures the model did not
cause — read this before blaming any model).scripts/ run them, do not reimplement:scripts/check_advertised_tools.py — the commands the eval system prompt
tells the model to use that the executor cannot run, so obeying the prompt
scores as a hallucination.scripts/preflight_models.py — do these slugs exist, before the run spends
money proving they do not.scripts/summarize_eval.py — report JSON → per-model rollup, bucketed
failures, and an attribution that is checked rather than asserted.refinement-log.md what past sessions changed here and why.The boundary with nexus-eval-harness: it owns the instrument, this skill owns
the verdict. Anything that changes the harness or its inputs — env knobs,
target syntax, live mode and the headless vault, config YAML, scenario authoring,
harness code — is that skill's. Anything that changes what you conclude about a
model is this one's. When a run reveals a fixture defect, hand it over rather
than fixing it here.
Also: nexus-model-updates owns provider model definitions and whether a model
ID works at all (grade nothing until it does); nexus-testing owns Jest lanes
and what a mock can prove; nexus-agents owns the two-tool contract the harness
is imitating; nexus-llm-adapters owns the adapter a stream error comes from.
© ProfSynapse, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files (scripts, references) in .skills/nexus-model-eval of ProfSynapse/nexus.
Open the folder on GitHubat commit 6cb7bce
Nexus Model Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nexus Model Eval this skillProfSynapse/nexus | 154 | — | ~1k | Automated safety check: Pass | MIT | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 6 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| Peft Fine TuningOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.1k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.8k | 13 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
ProfSynapse/nexus
How to add, change and verify a Nexus agent or tool, and the contract every tool must satisfy.
ProfSynapse/nexus
Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…
ProfSynapse/nexus
Add, change or verify a Nexus LLM model definition — the registry entry, the provider default, and proof the model id actually works against the live endpoint.
ProfSynapse/nexus
Cut a Nexus release — bump the version with the repo's own machinery, get the docs and generated sources right, push a tag the GitHub Actions workflow will actually pick up, and recover when it does…
ProfSynapse/nexus
How to change, persist, migrate and recover Nexus data without losing it.
ProfSynapse/nexus
Verify a Nexus change — pick a Jest lane, write a test that can actually fail, run the in-app Obsidian CLI loop, drive the eval harness, or fix a shipped-docs drift failure.
Categories
Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's. Nexus Model Eval is an agent skill from ProfSynapse/nexus. Grade how well a model drives the Nexus two-tool protocol (getTools/useTools) and decide whether a low score is the model's fault or the harness's.
Nexus Model Eval fits situations like: compare models on Nexus tool use; picking a default model; an eval report needs interpreting.
Run `npx skills add ProfSynapse/nexus --skill nexus-model-eval -a claude-code`. Or copy the skill folder (.skills/nexus-model-eval in ProfSynapse/nexus) into .claude/skills/nexus-model-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ProfSynapse/nexus --skill nexus-model-eval -a codex`. Or copy the skill folder (.skills/nexus-model-eval in ProfSynapse/nexus) into .agents/skills/nexus-model-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ProfSynapse/nexus --skill nexus-model-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nexus-model-eval, .gemini/skills/nexus-model-eval, .github/skills/nexus-model-eval and .opencode/skills/nexus-model-eval in your project.
Going by SKILL.md and its folder, Nexus Model Eval needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Nexus Model Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Nexus Model Eval: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ProfSynapse (a GitHub user) maintains it in ProfSynapse/nexus, which has 154 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 2, 2026.
Source: ProfSynapse/nexus on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.