Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
A skill your agent uses when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks.
$ npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Mesh-LLM/mesh-llm skippy-bench --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/skippy-bench .claude/skills/skippy-bench && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skippy-bench" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/skippy-bench into .claude/skills/skippy-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skippy-bench", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/skippy-benchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Mesh-LLM/mesh-llm skippy-bench --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/skippy-bench .agents/skills/skippy-bench && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skippy-bench" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/skippy-bench into .agents/skills/skippy-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skippy-bench", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Mesh-LLM/mesh-llm skippy-bench --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/skippy-bench .cursor/skills/skippy-bench && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skippy-bench" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/skippy-bench into .cursor/skills/skippy-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skippy-bench", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Mesh-LLM/mesh-llm.git --path .agents/skills/skippy-bench--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Mesh-LLM/mesh-llm skippy-bench --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/skippy-bench .gemini/skills/skippy-bench && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skippy-bench" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/skippy-bench into .gemini/skills/skippy-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skippy-bench", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Mesh-LLM/mesh-llm skippy-benchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/skippy-bench .github/skills/skippy-bench && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skippy-bench" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/skippy-bench into .github/skills/skippy-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skippy-bench", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Mesh-LLM/mesh-llm skippy-bench --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mesh-LLM/mesh-llm.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/skippy-bench .opencode/skills/skippy-bench && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skippy-bench" agent skill from https://github.com/Mesh-LLM/mesh-llm/tree/main/.agents/skills/skippy-bench into .opencode/skills/skippy-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skippy-bench", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skippy-benchA skill your agent uses when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks.
Skippy Bench is an agent skill from Mesh-LLM/mesh-llm. Use this skill when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks.
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering. The repository describes itself as: Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit aaf5a6c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
cargojustjqpythondockerFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
pypi.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
EVAL_LLM_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skippy Bench loads about 2.2k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 976 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Mesh-LLM/mesh-llm at commit aaf5a6c, republished under its Apache-2.0 licence (© Mesh-LLM). 976 words, ~2,168 tokens.
.claude/skills/skippy-bench/SKILL.md (or your agent's skills folder).Use this skill for performance, orchestration, and report-oriented checks.
Use skippy-correctness when the question is pass/fail exactness.
All reportable benchmark runs need metrics-server. run, focused-runtime,
and local-single start a collector by default; endpoint-driving commands such
as chat-corpus and eval run require --metrics-http to point at an
already-running metrics-server and should use --metrics-run-id matching the
target endpoint's Skippy run id.
Benchmark-managed Skippy server runs must use a release skippy-serving build.
Run just release-build before run, focused-runtime, local-single, or
local split binary benchmarks, and use target/release/skippy-server (the
SkippyBench default). Do not use target/debug/skippy-server for performance or
full-corpus validation; SkippyBench rejects that path because debug builds can
create false timeout and throughput failures.
Standalone skippy-bench may not be present in this mesh checkout yet. Confirm
available packages before using old source-repo commands:
cargo metadata --no-deps --format-version 1 | jq -r '.packages[].name' | sortUseful current checks:
cargo test -p skippy-serving --lib
cargo test -p mesh-llm-host-runtime --lib inference::skippyWhen benchmark harnesses are imported, keep reporting separate from request-path serving. Stage runtimes emit telemetry; benchmark/report tooling owns reports.
Use skippy-bench eval for external agent/coding benchmark harnesses. The
local SkippyBench corpora are for runtime behavior, cache behavior, transport
stress, and perf regression; they are not the source of agent benchmark claims.
Core pack:
skippy-bench eval list
skippy-bench eval info terminal-bench
skippy-bench eval sync --pack core
skippy-bench eval doctor
skippy-bench eval run speed-bench \
--base-url http://127.0.0.1:9337/v1 \
--model org/repo:Q4_K_M \
--endpoint-concurrency 1 \
--metrics-http http://127.0.0.1:18080 \
--metrics-run-id run-local-qwen--timeout-secs is passed to the native harness as its request/task timeout
where supported. It is not a full-run dataset limit. Use
--harness-timeout-secs only when you need a hard wall-clock cap for an
operator/debug run; omit it for canonical full-dataset validation.
--endpoint-concurrency must match the target endpoint's
serve-openai --generation-concurrency value. SkippyBench keeps native harness
request concurrency equal to that value; adapter-specific request concurrency
overrides such as SWE_BENCH_PRO_NUM_WORKERS and
MCP_ATLAS_COMPLETION_CONCURRENCY must match it or eval run fails before
starting the upstream harness. Do not run multiple LLM workers against a
single-lane Skippy endpoint when validating full corpora.
Core eval ids:
speed-bench — llama.cpp SPEED-Bench client for OpenAI-compatible serving
latency/throughput. Run the upstream qualitative benchmark across all
categories with no Skippy-owned sample limit.terminal-bench — Terminal-Bench CLI via terminal-bench-core==0.1.1.swe-bench-pro — Scale SWE-Bench Pro OS repo; uses the upstream data and
SWE-agent patch generation/evaluation flow rather than a Skippy-owned mini
benchmark.mcp-atlas — Scale MCP-Atlas native harness. eval run starts the
MCP agent environment and completion service when their localhost ports are
not already live, then runs the upstream completion script with --no-filter
so all Hugging Face dataset rows are attempted, plus the upstream scoring
path, without Skippy-specific task limits or tool_choice overrides.Use-case routing:
| Need | Eval | Why |
|---|---|---|
| OpenAI-compatible serving latency, tok/s, and full SPEED-Bench traffic | speed-bench | Native SPEED-Bench client over the upstream dataset selection. |
| Terminal agent behavior, shell/task execution, Docker sandbox readiness | terminal-bench | Exercises an agent loop that has to operate in a real terminal task environment. |
| Coding-agent patch generation and issue-resolution style prompts | swe-bench-pro | Uses upstream SWE-agent instance generation, patch gathering, and swe_bench_pro_eval.py. |
| MCP tool-use benchmark flow | mcp-atlas | Uses upstream MCP-Atlas completion and scoring scripts with the full Hugging Face dataset. |
| Cache, runtime, transport, split, or mesh performance regression | Built-in SkippyBench run, focused-runtime, local-single, or chat-corpus | These are Skippy/runtime benchmarks, not external agent-quality claims. |
Optional future packs are intentionally not wired yet:
repo-generation: NL2RepoBench.tool-expanded: Toolathlon / Tool-Decathlon.Keep sync/install opt-in. Do not make normal just build or cargo build
download external harnesses, datasets, or Docker images.
eval sync checks out the fetched upstream ref directly, and eval run records
the resolved harness SHA as harness_commit in run.json. Preserve both
behaviors so benchmark evidence remains reproducible even when definitions use
floating upstream refs.
Terminal-Bench should be installed with uv tool install --python 3.12 terminal-bench; Python 3.14 currently breaks the tb Typer CLI. Treat Docker
as ready only when skippy-bench eval doctor reports that the daemon can start
a container; docker info alone is insufficient. skippy-bench eval run
performs the same prerequisite checks before launching a native harness. Do not
add Skippy-owned task filters, dataset limits, compatibility shims,
response-format substitutions, or tool_choice overrides to external evals
unless the user explicitly asks for a noncanonical experiment.
For MCP-Atlas scoring, the wrapper defaults EVAL_LLM_MODEL,
EVAL_LLM_BASE_URL, and EVAL_LLM_API_KEY to the same local endpoint/model
used for completion, while preserving caller-provided EVAL_LLM_* overrides
for judge-model runs. When validating with a very small local Skippy model, run
completion against the normal compatibility endpoint and point EVAL_LLM_* at
a separate strict structured-output scorer endpoint, for example a second
skippy-serving serve-openai --guardrails enforce process; do not patch
or post-process the MCP scorer. For resumed operator runs, set
MCP_ATLAS_COMPLETION_OUTPUT_NAME to an existing upstream
completion_results/*.csv basename so the native completion script can reuse
its own processed-row skip behavior, and use MCP_ATLAS_SCORE_CONCURRENCY for
the scorer's native --concurrency setting.
For SWE-Bench Pro, the wrapper defaults to the official Docker image namespace
with local Docker deployment and local Docker evaluation so the core pack can
run without Modal credentials. It still runs upstream
helper_code/generate_sweagent_instances.py for the full dataset, then supplies
SWE-agent with a native expert_file instance file for local Docker platform,
entrypoint settings and SWE-agent's standalone Python/SWE-Rex Docker runtime.
Local Docker runs install SWE-agent into a dedicated venv and default
SWE_BENCH_PRO_SWEREX_SPEC to swe-rex[modal]==1.4.0, which keeps the native
SWE-ReX Docker runtime but includes the upstream
python:3.11.9-slim-bookworm builder fix. Modal remains an explicit
environment override and uses the Scale SWE-ReX patch flow. Some official
SWE-Pro base images point pip at an unavailable localhost package mirror; local
Docker runs default SWE_BENCH_PRO_SWEREX_PIP_INDEX_URL to
https://pypi.org/simple for the derived-image SWE-ReX install. Use
SWE_BENCH_PRO_PARSE_FUNCTION=thought_action for local OpenAI-compatible models
that do not emit OpenAI tool calls; this is the upstream SWE-agent local-model
path, not a Skippy dataset or harness rewrite.
For TTFT/FTTT, use metrics-server correlation rather than harness-only timing.
skippy-bench eval run and skippy-bench chat-corpus create/finalize a
metrics-server run. eval run keeps harness success independent from a
finalization/export failure and records telemetry as unavailable; chat-corpus
still fails when its metrics report cannot be exported. The target endpoint
must be emitting OTLP for the same run id. Debug telemetry is required for
per-token spans such as stage.openai_decode_token; without it, the JSON report
will still include a telemetry block explaining why TTFT/FTTT was unavailable.
© Mesh-LLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/skippy-bench of Mesh-LLM/mesh-llm.
Open the folder on GitHubat commit aaf5a6c
Skippy Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skippy Bench this skillMesh-LLM/mesh-llm | 3.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 5 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.9k | 14 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
jarrodwatts/claude-code-config
Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.
Mesh-LLM/mesh-llm
A skill your agent uses when validating a MeshLLM release candidate or current HEAD against the last GitHub release, assembling the canonical feature/fix/modification inventory, testing locally…
Mesh-LLM/mesh-llm
A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…
Mesh-LLM/mesh-llm
A skill your agent uses when adding, renaming, removing, validating, or exposing mesh-llm config settings, including built-in settings, plugin config schemas, owner-control apply behavior, CLI…
Mesh-LLM/mesh-llm
A skill your agent uses when connecting agent tools or OpenAI clients to mesh-llm — launching or configuring Goose, Claude Code, OpenCode, Pi, curl, or any OpenAI-compatible client against a local…
Mesh-LLM/mesh-llm
A skill your agent uses when converting Hugging Face SafeTensors checkpoints into split BF16 GGUF model repos with skippy-quantize on Hugging Face Jobs or a local machine, then publishing the…
Mesh-LLM/mesh-llm
A skill your agent uses when creating, monitoring, validating, or documenting low-memory Hugging Face Jobs or local runs that quantize split BF16/FP16 GGUF model repos into custom quant GGUF repos…
Categories
A skill your agent uses when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks. Skippy Bench is an agent skill from Mesh-LLM/mesh-llm. Use this skill when running benchmark orchestration, local single-stage or split benchmarks, benchmark report flow, or performance-oriented skippy runtime checks.
Skippy Bench fits situations like: running benchmark orchestration; local single-stage; split benchmarks; benchmark report flow.
Run `npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a claude-code`. Or copy the skill folder (.agents/skills/skippy-bench in Mesh-LLM/mesh-llm) into .claude/skills/skippy-bench in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a codex`. Or copy the skill folder (.agents/skills/skippy-bench in Mesh-LLM/mesh-llm) into .agents/skills/skippy-bench in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mesh-LLM/mesh-llm --skill skippy-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skippy-bench, .gemini/skills/skippy-bench, .github/skills/skippy-bench and .opencode/skills/skippy-bench in your project.
Going by SKILL.md and its folder, Skippy Bench needs the command-line tools its instructions call (cargo, just, jq, python and docker) and credentials named EVAL_LLM_API_KEY. Our summary lists: Python 3; Docker; A credential in EVAL_LLM_API_KEY.
SKILL.md names 1 domain. In commands or code: pypi.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Skippy Bench is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Skippy Bench: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Mesh-LLM (a GitHub organization) maintains it in Mesh-LLM/mesh-llm, which has 3,487 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 9, 2026.
Source: Mesh-LLM/mesh-llm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.