Agent skill

Bench

by ddalcu in ddalcu/mlx-serve

mlx-serve benchmarking methodology — bench.sh/llmprobe usage, comparison-trap rules (same-methodology cells only, spec-decode variance, thermal lies, engine naming), perf-claim etiquette.

Custom licenceAuto-check passedMedia & Creative

Install Bench

skills CLI
$ npx skills add ddalcu/mlx-serve --skill bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ddalcu/mlx-serve bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ddalcu/mlx-serve.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/bench .claude/skills/bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bench
GitHub stars
1.8k
Used in
1 other repo
Token cost
~1.1k tokens
SKILL.md length
580 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
Custom licence

At a glance

mlx-serve benchmarking methodology — bench.sh/llmprobe usage, comparison-trap rules (same-methodology cells only, spec-decode variance, thermal lies, engine naming), perf-claim etiquette.

  • Media & Creative work in your project
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Bench is an agent skill from ddalcu/mlx-serve. mlx-serve benchmarking methodology — bench.sh/llmprobe usage, comparison-trap rules (same-methodology cells only, spec-decode variance, thermal lies, engine naming), perf-claim etiquette. Use before running benchmarks or making any performance claim.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative. It works with macOS. The repository describes itself as: Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Zig backend, Swift frontend macOS app with chat, music, voice, video generation.

When your agent uses it

  • Media & Creative work in your project

Example prompts

  • “/bench”

What it can do on your machine

Read from SKILL.md and the folder at commit 2ebcbdb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bench loads about 1.1k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 580 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 580 words (~1,067 tokens).

“llmprobe is the measurement layer. tests/bench.sh boots mlx-serve (one model at a time: boot, probe, kill, settle) and llmprobe takes every number via --bench-only. We do not hand-roll timing loops — llmprobe discards a warmup per scenario, reports median-of-3 as…”

— opening of SKILL.md by ddalcu, Custom licence
name
bench

Read the full SKILL.md on GitHub

Files

Just SKILL.md in .claude/skills/bench of ddalcu/mlx-serve.

Open the folder on GitHubat commit 2ebcbdb

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ddalcu/mlx-serve, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bench this skillddalcu/mlx-serve1.8k1 repos~1.1kAutomated safety check: PassCustom licence
Kkclawkk43994/kkclaw174—~576Automated safety check: PassMIT
Voxtype Installpchalasani/claude-code-tools2k—~857Automated safety check: NotesMIT
Vlog Auto Editznyupup/ai-video-editing-skill144—~6.8kAutomated safety check: PassMIT
Sayitcallebtc/sayit159—~2.9kAutomated safety check: PassMIT
Gpt Imagenikships/droidproxy122—~2.5kAutomated safety check: PassMIT

Similar skills

  • Kkclaw

    kk43994/kkclaw

    给你的 AI Agent 一个桌面身体 — Setup Wizard、14情绪球体、语音克隆、歌词窗、Doctor 自检、跨平台支持(Windows + macOS)

    174 GitHub stars~576 tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Voxtype Install

    pchalasani/claude-code-tools

    Guide the user through installing, configuring, and launching voxtype — local on-device voice dictation (speech-to-text that types wherever the cursor is).

    2k GitHub stars~857 tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • Vlog Auto Edit

    znyupup/ai-video-editing-skill

    AI Agent自动剪辑旅行Vlog的完整工作流。从原始素材到成品视频,系统级只需ffmpeg,其余在Python venv内完成。by nyx研究所 (GitHub @znyupup · B站/小红书 @nyx研究所)

    144 GitHub stars~6.8k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Sayit

    callebtc/sayit

    Live, low-latency spoken narration for hands-free agent sessions through the sayit TTS command.

    159 GitHub stars~2.9k tokensUpdated 14 days ago
    Media & CreativeAuto-check passed
  • Gpt Image

    nikships/droidproxy

    Generate or edit images via GPT Image 2.5 Flare or Sunburst through DroidProxy Codex OAuth (no OPENAIAPIKEY).

    122 GitHub stars~2.5k tokensUpdated today
    Media & CreativeAuto-check passed
  • Agentvibes

    paulpreibisch/AgentVibes

    🎤 AgentVibes Voice Management - Manage your text-to-speech voices across multiple providers (Piper TTS, Piper, macOS Say).

    155 GitHub stars~1.5k tokensUpdated 20 days ago
    Media & CreativeAuto-check passed

More from ddalcu/mlx-serve

  • Mlx Serve

    ddalcu/mlx-serve

    Hook an app, game or script up to the local mlx-serve server for LLM chat, embeddings, image, speech, music, sound effect, video and 3D generation, and Laya/Kev/Clef typed decisions.

    1.8k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Release

    ddalcu/mlx-serve

    mlx-serve pre-release validation checklist, CalVer versioning, release steps, and CHANGELOG style.

    1.8k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • PR Charts

    ddalcu/mlx-serve

    Turn llmprobe reports into the charts a perf PR embeds (one line per arm across context size, every point labelled, machine and model in the caption) and host them so the PR body renders them.

    1.8k GitHub stars~594 tokensUpdated today
    Auto-check passed

Works with

Questions about Bench

What does Bench do?

mlx-serve benchmarking methodology — bench.sh/llmprobe usage, comparison-trap rules (same-methodology cells only, spec-decode variance, thermal lies, engine naming), perf-claim etiquette. Bench is an agent skill from ddalcu/mlx-serve.sh/llmprobe usage, comparison-trap rules (same-methodology cells only, spec-decode variance, thermal lies, engine naming), perf-claim etiquette.

When should I use Bench?

Bench fits situations like: media & Creative work in your project.

How do I install Bench in Claude Code?

Run `npx skills add ddalcu/mlx-serve --skill bench -a claude-code`. Or copy the skill folder (.claude/skills/bench in ddalcu/mlx-serve) into .claude/skills/bench in your project. Claude Code loads it when a task matches its description.

How do I install Bench in Codex?

Run `npx skills add ddalcu/mlx-serve --skill bench -a codex`. Or copy the skill folder (.claude/skills/bench in ddalcu/mlx-serve) into .agents/skills/bench in your project. Codex loads it when a task matches its description.

Can I use Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ddalcu/mlx-serve --skill bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bench, .gemini/skills/bench, .github/skills/bench and .opencode/skills/bench in your project.

What does Bench need to run?

SKILL.md names no scripts, command-line tools or credentials: Bench is instructions for the agent only.

Does Bench access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bench use?

Bench has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Bench use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bench?

Skills that share tags, products or a category with Bench: Kkclaw (kk43994/kkclaw, 174 stars), Voxtype Install (pchalasani/claude-code-tools, 2k stars), Vlog Auto Edit (znyupup/ai-video-editing-skill, 144 stars) and Sayit (callebtc/sayit, 159 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bench?

ddalcu (a GitHub user) maintains it in ddalcu/mlx-serve, which has 1,782 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.

Source: ddalcu/mlx-serve on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.