LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Turn the current conversation's workflow into a reusable agent skill.
$ npx skills add Undertone0809/rudder --skill conversation-to-skill -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Undertone0809/rudder conversation-to-skill --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Undertone0809/rudder.git skills-src && mkdir -p .claude/skills && cp -r skills-src/server/resources/bundled-skills/conversation-to-skill .claude/skills/conversation-to-skill && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "conversation-to-skill" agent skill from https://github.com/Undertone0809/rudder/tree/main/server/resources/bundled-skills/conversation-to-skill into .claude/skills/conversation-to-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conversation-to-skill", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Undertone0809/rudder/tree/main/server/resources/bundled-skills/conversation-to-skillType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Undertone0809/rudder --skill conversation-to-skill -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Undertone0809/rudder conversation-to-skill --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Undertone0809/rudder.git skills-src && mkdir -p .agents/skills && cp -r skills-src/server/resources/bundled-skills/conversation-to-skill .agents/skills/conversation-to-skill && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "conversation-to-skill" agent skill from https://github.com/Undertone0809/rudder/tree/main/server/resources/bundled-skills/conversation-to-skill into .agents/skills/conversation-to-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conversation-to-skill", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Undertone0809/rudder --skill conversation-to-skill -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Undertone0809/rudder conversation-to-skill --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Undertone0809/rudder.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/server/resources/bundled-skills/conversation-to-skill .cursor/skills/conversation-to-skill && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "conversation-to-skill" agent skill from https://github.com/Undertone0809/rudder/tree/main/server/resources/bundled-skills/conversation-to-skill into .cursor/skills/conversation-to-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conversation-to-skill", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Undertone0809/rudder.git --path server/resources/bundled-skills/conversation-to-skill--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Undertone0809/rudder --skill conversation-to-skill -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Undertone0809/rudder conversation-to-skill --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Undertone0809/rudder.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/server/resources/bundled-skills/conversation-to-skill .gemini/skills/conversation-to-skill && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "conversation-to-skill" agent skill from https://github.com/Undertone0809/rudder/tree/main/server/resources/bundled-skills/conversation-to-skill into .gemini/skills/conversation-to-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conversation-to-skill", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Undertone0809/rudder conversation-to-skillInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Undertone0809/rudder --skill conversation-to-skill -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Undertone0809/rudder.git skills-src && mkdir -p .github/skills && cp -r skills-src/server/resources/bundled-skills/conversation-to-skill .github/skills/conversation-to-skill && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "conversation-to-skill" agent skill from https://github.com/Undertone0809/rudder/tree/main/server/resources/bundled-skills/conversation-to-skill into .github/skills/conversation-to-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conversation-to-skill", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Undertone0809/rudder --skill conversation-to-skill -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Undertone0809/rudder conversation-to-skill --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Undertone0809/rudder.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/server/resources/bundled-skills/conversation-to-skill .opencode/skills/conversation-to-skill && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "conversation-to-skill" agent skill from https://github.com/Undertone0809/rudder/tree/main/server/resources/bundled-skills/conversation-to-skill into .opencode/skills/conversation-to-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conversation-to-skill", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
conversation-to-skillTurn the current conversation's workflow into a reusable agent skill.
Conversation To Skill is an agent skill from Undertone0809/rudder. Turn the current conversation's workflow into a reusable agent skill. Use this whenever the user wants to make a workflow reusable, standardize a successful thread, package an agent capability, or convert an ad hoc process into a repeatable skill. Read the thread first, extract the stable pattern, decide whether the skill should live in ~/.agents/skills/<name or <project-path/.agents/skills/<name, write the skill, and when quality matters add lightweight evals and iteration instead of just transcribing the chat.
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including scripts, reference files and assets (for example `agents/analyzer.md`, `agents/comparator.md` and `agents/grader.md`).
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Open-source local Agent harness for self-improving agent teams: run agents, review work, and turn feedback into reusable skills. The licence is Apache-2.0.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e2ba0f1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 4 files in scripts/ (Python, from the files we listed), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Conversation To Skill loads about 3.6k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 137 tokens; SKILL.md has 1,930 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from Undertone0809/rudder at commit e2ba0f1, republished under its Apache-2.0 licence (© Undertone0809). 1,930 words, ~3,563 tokens.
.claude/skills/conversation-to-skill/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.This skill turns the work happening in the current conversation into a reusable agent skill.
Its job is not just to write SKILL.md.
Its job is to identify the durable workflow, separate it from one-off thread
noise, decide the right packaging and placement, and produce a skill that will
actually help a future agent perform better.
When useful, this skill should borrow the practical methods of skill-creator:
good descriptions, clean skill structure, eval-friendly organization, and an
improve-via-feedback loop. When this skill owns evaluation, bundle the relevant
toolchain locally under agents/, assets/, eval-viewer/, scripts/, and
references/ so it stays self-contained instead of depending on another skill
directory at runtime.
Use this skill when the user is trying to:
Typical prompts:
Do not use this skill when the user mainly wants:
If the conversation does not yet reveal a stable workflow, say that plainly and help the user clarify the reusable part first.
The skill should capture the repeatable value, not the accidental details.
A good abstraction preserves:
A bad abstraction copies:
Prefer instructions that explain why a step matters. Avoid brittle mandates unless the workflow truly requires them.
If you find yourself writing a long list of rigid commands with no reasoning, you are probably transcribing the thread instead of building a skill.
Do not overbuild the skill. Use the smallest structure that preserves the capability:
SKILL.md only, when the workflow is mostly reasoning and sequencingSKILL.md plus references/, when the skill needs domain guidanceSKILL.md plus scripts/, when deterministic repeated work should be bundledSKILL.md plus evals/, when the skill benefits from repeatable testingPick the skill location before creating files so the paths stay stable:
~/.agents/skills/<skill-name><project-path>/.agents/skills/<skill-name>If the user wants a global skill to be discoverable by Codex immediately, also create:
~/.codex/skills/<skill-name> as a symlink to the global skill directoryIf you plan to run evals, place the workspace next to the skill directory as:
<skill-name>-workspace/Follow this sequence unless the user already provided enough structure.
Read the current conversation first. Pull out the real workflow before asking the user to restate everything.
Capture:
Classify each detail into one of three buckets:
Useful heuristic:
Do not ask the user to restate the whole workflow if the thread already tells you most of it. Only ask for the missing pieces that affect the resulting skill:
If examples, edge cases, dependencies, or adjacent skills matter, gather that context before writing the final version.
Before generating the final skill, write a short abstraction brief for the user to review unless they already said to just build it.
Use this structure:
## Skill Intent
- Name:
- Goal:
- Why this should exist:
## Trigger
- Use when:
- Do not use when:
## Inputs
- Required inputs:
- Optional inputs:
## Outputs
- Main deliverable:
- Secondary artifacts:
## Workflow
1. ...
2. ...
3. ...
## Judgment Rules
- What must stay true:
- What to avoid:
## Open Questions
- ...If the conversation already settles these points, keep the brief short and move on.
Do not act like a passive stenographer. If the proposed skill is overfit, under-scoped, or missing the real judgment logic, say so and correct it.
Common failure modes to call out:
Make these decisions before writing:
SKILL.md alone is enoughreferences/, scripts/, assets/, or evals/Default location rules:
~/.agents/skills/<skill-name><project-path>/.agents/skills/<skill-name>If updating an existing skill, preserve the directory name and frontmatter name unless the user asked for a rename.
When writing SKILL.md, include:
name and a trigger-oriented descriptionBring in the skill-creator quality bar here:
Prefer this structure when it helps:
skill-name/
├── SKILL.md
├── references/
├── scripts/
├── assets/
└── evals/Use progressive disclosure:
SKILL.md should explain the workflow clearlyWhen the skill supports multiple variants or domains, organize references by variant and tell the future agent which file to read for which case.
If the user wants more than a draft, or explicitly asks for testing, benchmarking, or trigger tuning, add local references that capture the evaluation workflow instead of leaving that logic implicit.
If the workflow needs actual tooling, prefer bundling it inside this skill rather than pointing at another repo's copy.
If multiple runs of the workflow would obviously repeat the same deterministic
steps, package that work into scripts/ instead of forcing future agents to
reinvent it every time.
Good candidates:
Do not add scripts just because you can. Only bundle work that is repeated, stable, and cheaper to reuse than to re-derive.
Not every conversation-derived skill needs evals. But if the skill produces objectively testable outputs, if the user asks for benchmarking, or if you are iterating on quality instead of just drafting, do not stop at a hand-wavy "light eval."
When you choose to evaluate, use the full evaluation suite:
evals/evals.json<skill-name>-workspace/ for iteration outputswith_skill against without_skill or an old snapshotThe detailed procedure lives in:
references/evaluation-suite.md for test execution, grading, benchmark aggregation, feedback, and iterationreferences/description-optimization.md for trigger-query generation and description tuningreferences/compatibility.md and references/schemas.md for host differences and file formatsThe local support toolchain lives in:
agents/ for grader, comparator, and analyst instructionsassets/ for review UI assetseval-viewer/ for viewer generationscripts/ for aggregation, optimization, validation, and packagingIf you decide evals are needed, read those reference files before proceeding.
Prefer qualitative review for subjective skills. Prefer assertions and benchmarks for objective skills.
If the first draft feels narrow, ambiguous, or weakly triggered, improve it. Useful improvement passes include:
Do not force a full benchmark loop if the user only wants a draft. But do not pretend the first draft is final if it clearly is not.
After creating or revising the skill, report:
Choose names that are short, clear, and capability-oriented.
Prefer names like:
conversation-to-skillworkflow-standardizertask-to-playbookAvoid names that depend on this thread's temporary wording unless the user explicitly wants that.
If updating an existing skill, preserve the existing directory name and frontmatter name unless the user asked for a rename.
Unless the user wants files written immediately, start with:
If the user asks to proceed, then write the files.
When the user already said "build it" or "just make it", go straight from the brief into file creation in the same turn.
If you also set up evals, mention:
The resulting skill should make a future agent meaningfully better at the task.
That usually means it captures at least one of these:
Strong skills often also have at least one of these:
If it captures none of those, it is probably not a real skill yet.
Do not create misleading, hostile, or surprise-heavy skills. The skill should do what its description honestly suggests.
Do not package instructions that facilitate unauthorized access, harmful automation, or disguised exfiltration.
Roleplay, stylistic framing, and benign workflow abstraction are fine.
© Undertone0809, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 21 other files (scripts, references, assets) in server/resources/bundled-skills/conversation-to-skill of Undertone0809/rudder.
Open the folder on GitHubat commit e2ba0f1
Conversation To Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Conversation To Skill this skillUndertone0809/rudder | 292 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
Undertone0809/rudder
A skill your agent uses when starting the current Rudder checkout as a temporary managed local preview with a stable URL, readiness check, logs, stop command, and cleanup path for manual inspection…
Undertone0809/rudder
A skill your agent uses when the user explicitly asks to stop, restart, kill, or clean Rudder repo-local pnpm dev processes or local dev runtime residue, including “把 pnpm dev 停了”, “重启 dev”, or “清掉…
Undertone0809/rudder
Conducts enterprise-grade research with multi-source synthesis, citation tracking, and verification.
Undertone0809/rudder
A skill your agent uses to audit or clean Rudder worktrees, generated artifacts, logs, caches, and repo-owned processes without deleting active work, user data, or unrelated machine state.
Undertone0809/rudder
Create safe inline visual explanations in Rudder Chat. An agent skill from Undertone0809/rudder.
Undertone0809/rudder
A skill your agent uses when Rudder development work needs first-principles advisor analysis plus independent reviewer rounds: proposals, UI/product decisions, architecture, release readiness…
Categories
Turn the current conversation's workflow into a reusable agent skill. Conversation To Skill is an agent skill from Undertone0809/rudder. Turn the current conversation's workflow into a reusable agent skill.
Conversation To Skill fits situations like: wants to make a workflow reusable; standardize a successful thread; package an agent capability; convert an ad hoc process into a repeatable skill.
Run `npx skills add Undertone0809/rudder --skill conversation-to-skill -a claude-code`. Or copy the skill folder (server/resources/bundled-skills/conversation-to-skill in Undertone0809/rudder) into .claude/skills/conversation-to-skill in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Undertone0809/rudder --skill conversation-to-skill -a codex`. Or copy the skill folder (server/resources/bundled-skills/conversation-to-skill in Undertone0809/rudder) into .agents/skills/conversation-to-skill in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Undertone0809/rudder --skill conversation-to-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/conversation-to-skill, .gemini/skills/conversation-to-skill, .github/skills/conversation-to-skill and .opencode/skills/conversation-to-skill in your project.
Going by SKILL.md and its folder, Conversation To Skill needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Conversation To Skill is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Conversation To Skill: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Undertone0809 (a GitHub user) maintains it in Undertone0809/rudder, which has 292 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on October 9, 2026.
Source: Undertone0809/rudder on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.