Agent skill

Model Benchmarks

by theopenco in theopenco/llmgateway

Run and report repository model or provider-mapping benchmarks.

Custom licenceAuto-check: notesAI & LLM Engineering

Install Model Benchmarks

skills CLI
$ npx skills add theopenco/llmgateway --skill model-benchmarks -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install theopenco/llmgateway model-benchmarks --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/theopenco/llmgateway.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/model-benchmarks .claude/skills/model-benchmarks && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-benchmarks
GitHub stars
1.7k
Token cost
~1.1k tokens
SKILL.md length
449 words
Files
3 (incl. scripts)
Skills in repo
19
Repo updated
First seen
Licence
Custom licence

At a glance

Run and report repository model or provider-mapping benchmarks.

  • Asked to benchmark a model
  • SKILL.md covers Profiles and Against a local gateway
  • Runs JavaScript scripts from its folder; calls node and python3; needs LLM_GATEWAY_API_KEY and LLMGATEWAY_API_KEY
  • Compare all mappings

What it does

Model Benchmarks is an agent skill from theopenco/llmgateway. Run and report repository model or provider-mapping benchmarks. Use when asked to benchmark a model, compare all mappings or regions, collect latency or throughput metrics, measure agentic coding or multi-turn tool-calling performance, run the smoke, standard, coding, load, capability, quality, or performance suites, or inspect an existing benchmark JSON result.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts.

It sits in AI & LLM Engineering, covering Structured output and tool calling. The repository describes itself as: Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.

When your agent uses it

  • Asked to benchmark a model
  • Compare all mappings
  • Collect latency
  • Throughput metrics

Example prompts

  • “/model-benchmarks”

Requirements

  • Python 3
  • Node.js
  • A credential in LLM_GATEWAY_API_KEY
  • A credential in LLMGATEWAY_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 4fd76f3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LLM_GATEWAY_API_KEY
    • LLMGATEWAY_API_KEY
    • LOCAL_GATEWAY_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Benchmarks loads about 1.1k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 449 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:28
    `.env` file.
  • NoteMentions a .env fileSKILL.md:68
    _KEY` for this — the CLI loads the root `.env`, which
  • NoteMentions a .env fileSKILL.md:81
    node --env-file=../../.env dist/bench-serve.js > /tmp/gateway.log 2>&1 < /dev/null & )

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 449 words (~1,112 tokens).

“Run paid benchmarks only when the user explicitly asks. Use the wrapper so raw results survive and failed evaluations still retain their timing data:”

— opening of SKILL.md by theopenco, Custom licence
name
model-benchmarks

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files (scripts) in .agents/skills/model-benchmarks of theopenco/llmgateway.

  • SKILL.md
  • scripts/render-results.mjs
  • scripts/run-benchmark.mjs

Open the folder on GitHubat commit 4fd76f3

Compare with similar skills

Model Benchmarks next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Benchmarks compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Benchmarks this skilltheopenco/llmgateway1.7k—~1.1kAutomated safety check: NotesCustom licence
Planning With Filesjarrodwatts/claude-code-config1.1k5 repos~967Automated safety check: PassNone
Tool Use Data Synthesissunny-glow/Auto-BenchMax1.3k—~3.3kAutomated safety check: PassNone
Agent Harness ConstructionKartikLabhshetwar/mind-mentor1477 repos~500Automated safety check: PassApache-2.0
Prompt Engineering Patternswshobson/agents40k—~1.3kAutomated safety check: PassMIT
Agent Prompt Quality Barmastra-ai/mastra29k—~2kAutomated safety check: PassCustom licence

Similar skills

  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed
  • Tool Use Data Synthesis

    sunny-glow/Auto-BenchMax

    Synthesize training data for ANY tool-use / agentic benchmark, in ANY repo.

    1.3k GitHub stars~3.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agent Harness Construction

    KartikLabhshetwar/mind-mentor

    Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates.

    147 GitHub starsUsed in 7 repos~500 tokens
    AI & LLM EngineeringAuto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Agent Prompt Quality Bar

    mastra-ai/mastra

    Universal quality bar and final audit rubric for any agent system prompt.

    29k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • AI Tools

    Dbxstudio/dbx-studio

    Reference for all AI tools available in DBX Studio's AI chat system.

    103 GitHub starsUsed in 2 repos~643 tokens
    AI & LLM EngineeringAuto-check passed

More from theopenco/llmgateway

All 19 skills in this repo
  • Blog

    theopenco/llmgateway

    Write and validate an LLM Gateway marketing blog post in the repository's current house style, including structured frontmatter and a gpt-image-2 OpenGraph image.

    1.7k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Skill Authoring

    theopenco/llmgateway

    Create, update, and audit repository-local agent skills against the current LLM Gateway codebase.

    1.7k GitHub stars~662 tokensUpdated today
    Auto-check passed
  • Security Audit

    theopenco/llmgateway

    Security best practices, vulnerability review, and full security audits for LLM Gateway.

    1.7k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Add Model

    theopenco/llmgateway

    Add a model or provider mapping to the catalogue, or verify one that was already written — pricing, capability and reasoning metadata, scoped e2e, and playground options for image/video models.

    1.7k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Changelog

    theopenco/llmgateway

    Write a new LLM Gateway changelog entry. An agent skill from theopenco/llmgateway.

    1.7k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Knowledge Base

    theopenco/llmgateway

    Write a new LLM Gateway docs Knowledge base page under apps/docs/content/(product)/learn with light and dark dashboard screenshots.

    1.7k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Questions about Model Benchmarks

What does Model Benchmarks do?

Run and report repository model or provider-mapping benchmarks. Model Benchmarks is an agent skill from theopenco/llmgateway. Run and report repository model or provider-mapping benchmarks.

When should I use Model Benchmarks?

Model Benchmarks fits situations like: asked to benchmark a model; compare all mappings; collect latency; throughput metrics.

How do I install Model Benchmarks in Claude Code?

Run `npx skills add theopenco/llmgateway --skill model-benchmarks -a claude-code`. Or copy the skill folder (.agents/skills/model-benchmarks in theopenco/llmgateway) into .claude/skills/model-benchmarks in your project. Claude Code loads it when a task matches its description.

How do I install Model Benchmarks in Codex?

Run `npx skills add theopenco/llmgateway --skill model-benchmarks -a codex`. Or copy the skill folder (.agents/skills/model-benchmarks in theopenco/llmgateway) into .agents/skills/model-benchmarks in your project. Codex loads it when a task matches its description.

Can I use Model Benchmarks in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add theopenco/llmgateway --skill model-benchmarks -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-benchmarks, .gemini/skills/model-benchmarks, .github/skills/model-benchmarks and .opencode/skills/model-benchmarks in your project.

What does Model Benchmarks need to run?

Going by SKILL.md and its folder, Model Benchmarks needs JavaScript for the scripts in its folder, the command-line tools its instructions call (node and python3) and credentials named LLM_GATEWAY_API_KEY, LLMGATEWAY_API_KEY and LOCAL_GATEWAY_KEY. Our summary lists: Python 3; Node.js; A credential in LLM_GATEWAY_API_KEY; A credential in LLMGATEWAY_API_KEY.

Does Model Benchmarks access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Model Benchmarks safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Model Benchmarks use?

Model Benchmarks has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Model Benchmarks use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Benchmarks?

Skills that share tags, products or a category with Model Benchmarks: Planning With Files (jarrodwatts/claude-code-config, 1.1k stars), Tool Use Data Synthesis (sunny-glow/Auto-BenchMax, 1.3k stars), Agent Harness Construction (KartikLabhshetwar/mind-mentor, 147 stars) and Prompt Engineering Patterns (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Benchmarks?

theopenco (a GitHub organization) maintains it in theopenco/llmgateway, which has 1,674 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 7, 2026.

Source: theopenco/llmgateway on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.