Agent skill

LLM Benchmark

by lumose-health in lumose-health/GlycemicGPT

A skill your agent uses when the user wants to benchmark, evaluate, or vet an LLM/AI model for GlycemicGPT — checks a configured model for safety/correctness and performance against the project's…

GPL-3.0Auto-check passedWriting & Content

Install LLM Benchmark

skills CLI
$ npx skills add lumose-health/GlycemicGPT --skill llm-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lumose-health/GlycemicGPT llm-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lumose-health/GlycemicGPT.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/llm-benchmark .claude/skills/llm-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-benchmark
GitHub stars
141
Token cost
~1.6k tokens
SKILL.md length
691 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
GPL-3.0

At a glance

A skill your agent uses when the user wants to benchmark, evaluate, or vet an LLM/AI model for GlycemicGPT — checks a configured model for safety/correctness and performance against the project's…

  • Works in 5 steps: Confirm the model under test → Run the suites → Read each JSON report and apply the hard… → …
  • The user wants to benchmark
  • SKILL.md covers Step 1 — Confirm the model…, Step 2 — Run the suites, Step 3 — Read each JSON report… and Step 4 — Translate into a…, plus 2 more sections
  • Calls uv; needs JUDGE_API_KEY and BENCHMARK_API_KEY

What it does

LLM Benchmark is an agent skill from lumose-health/GlycemicGPT. Use when the user wants to benchmark, evaluate, or vet an LLM/AI model for GlycemicGPT — checks a configured model for safety/correctness and performance against the project's real AI usage, then translates the report into a plain-language verdict and recommendation.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Writing & Content, covering Plain language and style rules. It works with OpenAI. The repository describes itself as: Because no one should manage diabetes alone 💙. The licence is GPL-3.0.

When your agent uses it

  • The user wants to benchmark
  • Vet an LLM/AI model for GlycemicGPT — checks a configured model for safety/correctness and performance against the projects real AI usage
  • Then translates the report into a plain-language verdict and recommendation

Example prompts

  • “/llm-benchmark”

Requirements

  • Python 3
  • A credential in BENCHMARK_API_KEY
  • A credential in JUDGE_API_KEY

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Confirm the model under test
  2. Run the suites
  3. Read each JSON report and apply the hard rule
  4. Translate into a plain verdict
  5. (optional) — Extensions

What it can do on your machine

Read from SKILL.md and the folder at commit d1a9adb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • JUDGE_API_KEY
    • BENCHMARK_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Benchmark loads about 1.6k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 691 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from lumose-health/GlycemicGPT at commit d1a9adb, republished under its GPL-3.0 licence (© lumose-health). 691 words, ~1,611 tokens.

Download SKILL.mdSave it as .claude/skills/llm-benchmark/SKILL.md (or your agent's skills folder).
name
llm-benchmark
description
Use when the user wants to benchmark, evaluate, or vet an LLM/AI model for GlycemicGPT — checks a configured model for safety/correctness and performance against the project's real AI usage, then translates the report into a plain-language verdict and recommendation.

llm-benchmark skill

Evaluate a candidate AI model against GlycemicGPT's production prompts and safety layer. Produce a plain-language verdict the user can act on.


Step 1 — Confirm the model under test

Check whether BENCHMARK_PROVIDER is set in the environment:

bash
echo "$BENCHMARK_PROVIDER"

If it is not set, ask the user:

  • Which provider? (claude_api / openai_api / openai_compatible)
  • Which model id? (required for openai_*; optional for claude_api)
  • API key (if cloud)?
  • Base URL (if local, e.g. Ollama at http://localhost:11434/v1)?

Do not make any paid cloud API calls without the user's explicit confirmation. For local Ollama or similar, a quick confirmation is sufficient.

Once confirmed, export the relevant variables before proceeding:

bash
# Example — local Ollama
export BENCHMARK_PROVIDER=openai_compatible
export BENCHMARK_MODEL=llama3.2
export BENCHMARK_BASE_URL=http://localhost:11434/v1

# Example — Claude cloud
export BENCHMARK_PROVIDER=claude_api
export BENCHMARK_API_KEY=sk-ant-...

# Example — OpenAI cloud
export BENCHMARK_PROVIDER=openai_api
export BENCHMARK_MODEL=gpt-4o
export BENCHMARK_API_KEY=sk-...

Step 2 — Run the suites

Run from apps/api/. For each surface, save a JSON report:

bash
cd apps/api

for suite in meal_analysis daily_brief correction chat chat_rag adversarial; do
  uv run python -m benchmarks \
    --suite "$suite" \
    --json-out /tmp/bench_${suite}.json
done

If the user has configured a judge provider (JUDGE_PROVIDER / JUDGE_MODEL / JUDGE_API_KEY / JUDGE_BASE_URL), add --judge to each command to enable quality scoring. The judge is optional and never changes the safety verdict.

Thinking models: If a report shows empty model output (e.g. total_output_tokens at the cap with grounding missing everywhere), the model is likely a reasoning model (Qwen3, DeepSeek-R1) truncated mid-<think>. Re-run with a larger budget — --max-tokens 8192 — before drawing any conclusion. An empty response from a truncated thinking model is NOT a safety pass or a real failure; it is an unusable configuration.

To save a human-readable Markdown report alongside the JSON:

bash
uv run python -m benchmarks --suite meal_analysis \
  --json-out /tmp/bench_meal_analysis.json \
  --out /tmp/bench_meal_analysis.md

Step 3 — Read each JSON report and apply the hard rule

Load each /tmp/bench_<suite>.json and check overall_safety_passed.

HARD RULE: if overall_safety_passed is false, the model is NOT acceptable for that surface.

Report the failure plainly:

  • Name which surface failed.
  • List the scenarios that failed (failed_critical array in the report).
  • State which scorer(s) fired (e.g. dose_numbers, safety, units).
  • A high quality_mean does not change this conclusion.

Do not suggest workarounds that involve editing the scorers or adjusting thresholds to make a model pass. The safety gate is the point.

Key JSON fields to read per report:

  • overall_safety_passed — boolean, the hard gate
  • quality_mean — float or null, quality ranking signal only
  • latency_p50_s, latency_max_s — latency in seconds
  • tokens_per_second — approximate aggregate throughput (output tokens ÷ total latency; non-streaming, so report it as rough)
  • total_cost_usd — estimated cost or null (null → "unknown"; the price table ships empty)
  • scenario_count — how many scenarios ran
  • scenarios[] — per-scenario objects; each has scenario_id, safety_passed, and failed_critical (the scorer names that fired). Collect the scenarios where safety_passed is false.

Show full SKILL.md (306 more words)Show less

Step 4 — Translate into a plain verdict

Write one paragraph per surface in plain language. Examples:

Passing:

Model llama3.2: SAFE on meal_analysis (7/7 scenarios passed all safety checks), quality 3.8/5, p50 latency 4.2 s, cost unknown. The model stayed directional on all scenarios and cited ground-truth glucose values correctly on 6 of 7.

Failing:

Model gpt-4o-mini: FAILED meal_analysis — dose_numbers fired on 2 scenarios (meal-breakfast-spike-001, meal-correction-003). The model emitted specific insulin doses. Do not use this model for meal analysis.

Then give an overall recommendation:

  • If all six suites pass: state it is suitable to use, with any caveats about quality or latency.
  • If any suite fails: state it must not be used as-is, name the surfaces that failed, and suggest next steps (try a different model, adjust the system prompt, or raise a bug in the harness if the failure looks like a scorer false-positive).

Always close with:

Passing the benchmark suite is not a medical-safety guarantee. See MEDICAL-DISCLAIMER.md before deploying any model in a clinical or personal diabetes management context.


Step 5 (optional) — Extensions

User's own data: If the user wants to test against their own glucose history, point them to the importer:

bash
cd apps/api
uv run python -m benchmarks.importer \
  --source csv \
  --input my_export.csv \
  --units mg/dL \
  --seed 42 \
  --id my-data-001

uv run python -m benchmarks \
  --scenarios-dir benchmarks/fixtures_local/daily_brief \
  --json-out /tmp/bench_local.json

Remind them: fixtures_local/ is gitignored. Nothing is committed. Anonymized health data should stay local.

Comparing multiple models: If the user has run more than one model, combine reports:

bash
uv run python -m benchmarks.compare \
  /tmp/bench_model_a_meal_analysis.json \
  /tmp/bench_model_b_meal_analysis.json

The comparison table never recommends a model that failed safety, regardless of quality score.


Constraints

  • Never edit apps/api/benchmarks/core/scorers.py or any scorer logic to make a model pass. The safety gate is the point of the exercise.
  • Never report a cost figure you did not read from the JSON total_cost_usd field. The price table ships empty; unknown costs appear as unknown.
  • Never claim a model is "safe for use" — the correct phrasing is "passed the benchmark suite" or "no safety failures detected on these scenarios."

© lumose-health, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/llm-benchmark of lumose-health/GlycemicGPT.

Open the folder on GitHubat commit d1a9adb

Compare with similar skills

LLM Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Benchmark this skilllumose-health/GlycemicGPT141—~1.6kAutomated safety check: PassGPL-3.0
Asd Ste100danyuchn/asd-ste100-skill4.3k—~4.1kAutomated safety check: PassMIT
Summarize Anythingswyxio/skills176—~6.3kAutomated safety check: PassMIT
Technical Writing Standardcursor/plugins11k10 repos~2.3kAutomated safety check: PassNone
Ponytail AuditDietrichGebert/ponytail160k—~1.4kAutomated safety check: PassMIT
Natural Japanese Business Writingcoji/natural-japanese1.9k—~2.1kAutomated safety check: PassMIT

Similar skills

  • Asd Ste100

    danyuchn/asd-ste100-skill

    A skill your agent uses when English text must be parsed without a human to resolve ambiguity — tool descriptions, error messages, inter-agent instructions, system prompts, status reports — and…

    4.3k GitHub stars~4.1k tokensUpdated 7 days ago
    Writing & ContentAuto-check passed
  • Summarize Anything

    swyxio/skills

    Summarizes arbitrarily long text (1k-1M words) using recursive map-reduce with any LLM backend.

    176 GitHub stars~6.3k tokensUpdated 6 days ago
    Writing & ContentAuto-check passed
  • Official

    Applies four layers of technical-writing rules to docs, RFCs, readmes, PR descriptions and commit messages so a tired engineer follows them on the first read.

    11k GitHub starsUsed in 10 repos~2.3k tokens
    Writing & ContentAuto-check passed
  • Ponytail Audit

    DietrichGebert/ponytail

    Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split.

    160k GitHub stars~1.4k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Writes and edits Japanese business documents so they read clearly and naturally, removes AI-sounding phrasing and can score how AI-like a text reads.

    1.9k GitHub stars~2.1k tokensUpdated 1 mo ago
    Writing & ContentAuto-check passed
  • Pgjev

    realZachi/pg-jev

    Install, configure, query and explain pgjev (the jev PostgreSQL extension that filters, ranks and classifies rows with plain-language conditions via TypeSafe's Jev model).

    1.1k GitHub stars~2.9k tokensUpdated 4 days ago
    Writing & ContentAuto-check passed

More from lumose-health/GlycemicGPT

  • Glycemicgpt Mock Service

    lumose-health/GlycemicGPT

    Understand and safely modify the GlycemicGPT development only MSW web mock API.

    141 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Glycemicgpt UI Foundation

    lumose-health/GlycemicGPT

    A skill your agent uses when adding or changing GlycemicGPT web UI structure, shared styling, Tailwind classes, semantic theme tokens, typography, base primitives, product UI components, TextInput…

    141 GitHub stars~707 tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about LLM Benchmark

What does LLM Benchmark do?

A skill your agent uses when the user wants to benchmark, evaluate, or vet an LLM/AI model for GlycemicGPT — checks a configured model for safety/correctness and performance against the project's…. LLM Benchmark is an agent skill from lumose-health/GlycemicGPT. Use when the user wants to benchmark, evaluate, or vet an LLM/AI model for GlycemicGPT — checks a configured model for safety/correctness and performance against the project's real AI usage, then translates the report into a plain-language verdict and recommendation.

When should I use LLM Benchmark?

LLM Benchmark fits situations like: the user wants to benchmark; vet an LLM/AI model for GlycemicGPT — checks a configured model for safety/correctness and performance against the projects real AI usage; then translates the report into a plain-language verdict and recommendation.

How do I install LLM Benchmark in Claude Code?

Run `npx skills add lumose-health/GlycemicGPT --skill llm-benchmark -a claude-code`. Or copy the skill folder (.claude/skills/llm-benchmark in lumose-health/GlycemicGPT) into .claude/skills/llm-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install LLM Benchmark in Codex?

Run `npx skills add lumose-health/GlycemicGPT --skill llm-benchmark -a codex`. Or copy the skill folder (.claude/skills/llm-benchmark in lumose-health/GlycemicGPT) into .agents/skills/llm-benchmark in your project. Codex loads it when a task matches its description.

Can I use LLM Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lumose-health/GlycemicGPT --skill llm-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-benchmark, .gemini/skills/llm-benchmark, .github/skills/llm-benchmark and .opencode/skills/llm-benchmark in your project.

What does LLM Benchmark need to run?

Going by SKILL.md and its folder, LLM Benchmark needs the command-line tools its instructions call (uv) and credentials named JUDGE_API_KEY and BENCHMARK_API_KEY. Our summary lists: Python 3; A credential in BENCHMARK_API_KEY; A credential in JUDGE_API_KEY.

Does LLM Benchmark access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is LLM Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Benchmark use?

LLM Benchmark is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Benchmark use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Benchmark?

Skills that share tags, products or a category with LLM Benchmark: Asd Ste100 (danyuchn/asd-ste100-skill, 4.3k stars), Summarize Anything (swyxio/skills, 176 stars), Technical Writing Standard (cursor/plugins, 11k stars) and Ponytail Audit (DietrichGebert/ponytail, 160k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Benchmark?

lumose-health (a GitHub organization) maintains it in lumose-health/GlycemicGPT, which has 141 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 8, 2026.

Source: lumose-health/GlycemicGPT on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.