LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Scaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner.
$ npx skills add undefined-ui/second-brain-os --skill evals-bootstrap -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install undefined-ui/second-brain-os evals-bootstrap --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/agents-course/skills/evals-bootstrap .claude/skills/evals-bootstrap && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evals-bootstrap" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/evals-bootstrap into .claude/skills/evals-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-bootstrap", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/evals-bootstrapType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add undefined-ui/second-brain-os --skill evals-bootstrap -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install undefined-ui/second-brain-os evals-bootstrap --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/agents-course/skills/evals-bootstrap .agents/skills/evals-bootstrap && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evals-bootstrap" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/evals-bootstrap into .agents/skills/evals-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-bootstrap", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add undefined-ui/second-brain-os --skill evals-bootstrap -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install undefined-ui/second-brain-os evals-bootstrap --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/agents-course/skills/evals-bootstrap .cursor/skills/evals-bootstrap && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evals-bootstrap" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/evals-bootstrap into .cursor/skills/evals-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-bootstrap", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/undefined-ui/second-brain-os.git --path plugins/agents-course/skills/evals-bootstrap--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add undefined-ui/second-brain-os --skill evals-bootstrap -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install undefined-ui/second-brain-os evals-bootstrap --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/agents-course/skills/evals-bootstrap .gemini/skills/evals-bootstrap && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evals-bootstrap" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/evals-bootstrap into .gemini/skills/evals-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-bootstrap", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install undefined-ui/second-brain-os evals-bootstrapInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add undefined-ui/second-brain-os --skill evals-bootstrap -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/agents-course/skills/evals-bootstrap .github/skills/evals-bootstrap && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evals-bootstrap" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/evals-bootstrap into .github/skills/evals-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-bootstrap", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add undefined-ui/second-brain-os --skill evals-bootstrap -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install undefined-ui/second-brain-os evals-bootstrap --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/undefined-ui/second-brain-os.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/agents-course/skills/evals-bootstrap .opencode/skills/evals-bootstrap && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evals-bootstrap" agent skill from https://github.com/undefined-ui/second-brain-os/tree/main/plugins/agents-course/skills/evals-bootstrap into .opencode/skills/evals-bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evals-bootstrap", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evals-bootstrapScaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner.
Evals Bootstrap is an agent skill from undefined-ui/second-brain-os. Scaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner. Use when the user wants evals, regression tests for an agent, a golden set, or asks how to know a prompt/model change did not break things. Do NOT use for a single task's done-check (goal-test) or for auditing context (context-audit).
Its SKILL.md is about 790 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: An AI second brain that maintains itself. Full guide, starter vault, agent skills and scripts for a self-organizing knowledge base in Claude Code and Obsidian. The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d6861cc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
undefined-ui.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evals Bootstrap loads about 793 tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 378 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from undefined-ui/second-brain-os at commit d6861cc, republished under its MIT licence (© undefined-ui). 378 words, ~793 tokens.
.claude/skills/evals-bootstrap/SKILL.md (or your agent's skills folder).Theory: Two kinds of checks and the full walkthrough in Evals practice. An eval suite is the same test after every change. Behavioural checks read the steps of a trace; end-to-end checks read only the result. Start behavioural: they are deterministic, run in seconds, and diagnose instead of just scoring.
traces/<id>.json —
a list of events including tool calls. If the user's harness is Claude
Code, the transcript already is the trace; wire up whatever copies or
converts it. No trace, no behavioural checks.cases.yaml. One entry per failure:- id: refund_1042
input: "Refund order #1042, customer says it arrived broken"
expect: looks up the order before replying; asks approval before refund
check:
- trace has get_order before send_reply
- trace has approval_request before refund The expect line is for humans; the check lines are the test. Keep the
rule language tiny: trace has X, trace has X before Y, trace lacks X.
5. Generate check_traces.py. A small runner: load cases.yaml, parse
each rule with a regex, walk the tool-call list, print one line per case,
exit non-zero on any failure. Keep it dependency-light (pyyaml only) and
fast — the whole suite should run in seconds, with no model calls.
6. Run it and hand over the flywheel. Show the pass/fail lines. Then
leave the loop in writing at the end of your report: read fresh traces
weekly, add every new failure as a case, fix the biggest cluster, re-run.
python check_traces.py) so it can gate a
CI job or a pre-release habit without ceremony.© undefined-ui, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/agents-course/skills/evals-bootstrap of undefined-ui/second-brain-os.
Open the folder on GitHubat commit d6861cc
Evals Bootstrap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evals Bootstrap this skillundefined-ui/second-brain-os | 999 | — | ~793 | Automated safety check: Pass | MIT | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
undefined-ui/second-brain-os
Audit an agent's context layout against the four places: system prompt, tools, history, tail.
undefined-ui/second-brain-os
Find the decisions in a pipeline that do not need the expensive model and propose the gate for each: a rule, a classic classifier, or a small model, with fail-closed routing.
undefined-ui/second-brain-os
Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.
undefined-ui/second-brain-os
Audit an agent's harness against the four rings: containment, guides, sensors, permissions.
undefined-ui/second-brain-os
Triage an exported chat history into what to ingest, what to archive and what to delete, with a privacy pass first.
undefined-ui/second-brain-os
Analyse the vault's link graph and report on its shape: orphan rate, average degree, components, hubs, bridges and clusters, with what each number means for retrieval.
Categories
Scaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner. Evals Bootstrap is an agent skill from undefined-ui/second-brain-os. Scaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner.
Evals Bootstrap fits situations like: the user wants evals; regression tests for an agent; asks how to know a prompt/model change did not break things; A single tasks done-check (goal-test).
Run `npx skills add undefined-ui/second-brain-os --skill evals-bootstrap -a claude-code`. Or copy the skill folder (plugins/agents-course/skills/evals-bootstrap in undefined-ui/second-brain-os) into .claude/skills/evals-bootstrap in your project. Claude Code loads it when a task matches its description.
Run `npx skills add undefined-ui/second-brain-os --skill evals-bootstrap -a codex`. Or copy the skill folder (plugins/agents-course/skills/evals-bootstrap in undefined-ui/second-brain-os) into .agents/skills/evals-bootstrap in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add undefined-ui/second-brain-os --skill evals-bootstrap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-bootstrap, .gemini/skills/evals-bootstrap, .github/skills/evals-bootstrap and .opencode/skills/evals-bootstrap in your project.
Going by SKILL.md and its folder, Evals Bootstrap needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: undefined-ui.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evals Bootstrap is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 793 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evals Bootstrap: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
undefined-ui (a GitHub user) maintains it in undefined-ui/second-brain-os, which has 999 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on September 29, 2026.
Source: undefined-ui/second-brain-os on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.