Agent skill

LLM Eval Harness

by daymade in daymade/claude-code-skills

Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.

MITAuto-check passedAI & LLM Engineering

Install LLM Eval Harness

skills CLI
$ npx skills add daymade/claude-code-skills --skill llm-eval-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install daymade/claude-code-skills llm-eval-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/llm-eval-harness .claude/skills/llm-eval-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-eval-harness
GitHub stars
1.4k
Token cost
~4.7k tokens
SKILL.md length
2,120 words
Files
12 (incl. scripts, references, assets)
Skills in repo
103
Repo updated
First seen
Licence
MIT

At a glance

Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.

  • Works in 6 steps: Identify the shape — OpenAI-compatible… → For a new endpoint, availability and… → Run the dimensions the user cares about… → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers Overview, Which dimensions to run, Quick start and Dimension 0 — Availability…, plus 8 more sections
  • Runs Python scripts from its folder; calls uv

What it does

LLM Eval Harness is an agent skill from daymade/claude-code-skills. Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression. Use before hand-rolling a curl loop, when onboarding a provider, or debugging "system prompt 不生效" / a tok/s claim — 测评/压测一个模型或渠道. Not for TTS/voice-clone supplier eval (use the audio skill).

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts, reference files and assets (for example `assets/example_usecases.json`, `references/evaluation_disciplines.md` and `references/production_testing_patterns.md`).

It sits in AI & LLM Engineering, covering LLM evaluation, Prompt engineering and Text to speech and voice. It works with OpenAI. The repository describes itself as: Professional Claude Code skills marketplace featuring production-ready skills for enhanced development workflows. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve Prompt engineering
  • Tasks that involve Text to speech and voice

Example prompts

  • “system prompt 不生效”
  • “/llm-eval-harness”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Identify the shape — OpenAI-compatible (/v1/chat/completions) or Anthropic-Messages
  2. For a new endpoint, availability and fidelity come first — speed numbers for a model
  3. Run the dimensions the user cares about — speed and concurrency for "is it fast/stable",
  4. Report honestly, separating measured from inferred. Lead with the headline the user
  5. Comparing two models? Run the identical probes against each with the same flags, and put
  6. Found a vendor bug worth reporting? Don't paste raw probe output at their support

What it can do on your machine

Read from SKILL.md and the folder at commit 91bed2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Eval Harness loads about 4.7k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 2,120 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~17k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from daymade/claude-code-skills at commit 91bed2b, republished under its MIT licence (© daymade). 2,120 words, ~4,658 tokens.

Download SKILL.mdSave it as .claude/skills/llm-eval-harness/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
llm-eval-harness
description
Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression. Use before hand-rolling a curl loop, when onboarding a provider, or debugging "system prompt 不生效" / a tok/s claim — 测评/压测一个模型或渠道. Not for TTS/voice-clone supplier eval (use the audio skill).

LLM Eval Harness

Overview

Give this skill an endpoint (base_url + model + an API key in an env var) and it measures whether the endpoint actually works and whether the model is fast, stable, protocol-correct, and good enough — instead of trusting the vendor's headline numbers. Six dimensions, usually scattered across ad-hoc scripts that get rewritten (with the same bugs) every time:

DimensionScriptAnswers
Availabilityscripts/availability_probe.pywhich model IDs work here, with 3-state error classification
Request fidelityscripts/fidelity_probe.pydo system prompt / tools / history actually REACH the model?
Speedscripts/speed_probe.pyTTFT + sustained decode tok/s, thinking-aware
Concurrency / stabilityscripts/concurrency_probe.pysuccess rate, p50/p90 latency, where it breaks
Protocol compliancescripts/protocol_probe.pydoes the Anthropic thinking block actually fire when requested, AND does the endpoint accept one already sitting in history on a later turn (N≥10, both are separate checks — see Dimension 3)?
Quality / use-case regressionscripts/usecase_runner.py + blind judgesdoes it pass your accumulated cases?

Which dimensions to run

Route from what the user is actually doing:

  • "接入新模型/新网关先测一下" / onboarding a provider → Availability → Fidelity, then speed/concurrency if adoption looks likely
  • "这个 ID 能不能用 / 写个支持列表" → Availability (and read disciplines §9 before writing any "unavailable" verdict)
  • "模型不听 system prompt / agent 行为怪但没报错" → Fidelity (system canary)
  • "快不快 / tok/s 是真的吗" → Speed; "扛不扛得住" → Concurrency
  • "这个兼容端点是真兼容吗" → Protocol + Fidelity's tools/auth checks
  • "这个模型/渠道用着用着就 400 了 / session 断了 / continue 不下去了" (extended-thinking client, multi-turn session) → Protocol --check history-replay specifically — this symptom is the generation check passing while history-replay silently doesn't; don't stop at generation
  • "换模型质量会不会掉" → Quality regression
  • "要给厂商报 bug" → run the relevant probe with --output, then follow references/vendor_evidence_protocol.md
  • "部署闸门 / 常驻监控 / 告警老误报" → references/production_testing_patterns.md

Don't run all six ritually — pick what answers the user's question.

Key handling (non-negotiable): every script takes the API key by env-var name (--key-env MY_KEY), never the key value on the command line — so it stays out of ps, shell history, and any saved report. Never hardcode a key into a use-case file or a wrapper. Read references/evaluation_disciplines.md for the full reasoning behind this and the other disciplines.

Your private data lives outside this bundle. Use-case libraries, model rosters, and keys belong in ~/.llm-eval/ (or wherever you keep secrets), NOT in this skill directory — the skill is generic and public; your test suite is yours. See "Use-case library" below.

Quick start

Detect what you have, then run the dimensions that apply. For an OpenAI-compatible model:

bash
export MY_KEY=sk-...                       # the key never appears in a command below

# Speed: real-task throughput + sustained decode ceiling
uv run --with openai python scripts/speed_probe.py \
  --base-url https://api.example.com/v1 --model some-model --key-env MY_KEY --mode both

# Concurrency: ramp until it breaks
uv run --with aiohttp python scripts/concurrency_probe.py \
  --url https://api.example.com/v1/chat/completions --model some-model --key-env MY_KEY \
  --format openai --concurrency 10 20 40 60

If the endpoint is Anthropic-Messages-shaped (/v1/messages), also run the protocol probe (below). Pick dimensions by what the user actually asked — don't run all six if they only asked "is it fast?".

Dimension 0 — Availability (which model IDs actually work)

bash
uv run --with aiohttp python scripts/availability_probe.py \
  --base-url <base> --key-env <ENV> --format both \
  --models model-a vendor/model-b @models.txt --output /tmp/avail.json
  • Verdicts are 3-state per failure, not pass/fail: no-channel (no route for this ID), upstream-error (route exists, upstream failing — retry later), empty-content (usually a max_tokens artifact on reasoning models, NOT a broken model). The probe keeps --max-tokens at 8192 by default precisely so thinking models don't read as dead — do not "optimize" it downward.
  • Before writing any "unavailable" list: verify each failing ID exists in official docs. Current-generation model names postdate your training data — WebSearch the vendor's / model developer's model page instead of enumerating guesses; suffixes (-preview, dated variants, tier names) decide routability. Full traps: disciplines §9.
  • Never probe a client-side marker as if it were a real model ID. Bracket-suffixed strings like some-model[1m] are frequently a CLI client's own local convention (Claude Code parses and strips [1m] before the request is ever built — see claude-switch-models-setup's "Configuring Context Window Size" section for the full mechanism) — they never appear on the wire. The probe warns when it sees one, but the discipline is yours: probe the bare model ID the vendor actually documents, not a string copied out of a Claude Code ANTHROPIC_MODEL env var.
  • The gateway's /v1/models listing is one input, never the verdict — real gateways route models the listing omits.

Dimension 0.5 — Request fidelity (does your payload reach the model?)

bash
uv run --with aiohttp python scripts/fidelity_probe.py \
  --base-url <base> --model <model> --key-env <ENV> \
  --format anthropic --check all --repeat 10 --output /tmp/fidelity.json
  • Four checks: system (canary code planted in the system prompt — can the model echo it back?), tools (full round-trip including returning a tool result and verifying the final answer uses it), multiturn (plant facts across turns, recall them all), auth (x-api-key vs Bearer on both protocol endpoints — gateways commonly accept both on one path and only one on the other, and the wrong-header 401 reads exactly like a dead key).
  • This is the dimension that catches the nastiest gateway failure: everything returns 200 and chats fluently, but the system prompt never reached the model — so any tool whose rules live in the system prompt silently misbehaves with no error anywhere. On one real gateway, one model alias never delivered the system prompt in 100+ samples across every legal way of sending it, while sibling aliases delivered intermittently with rates that changed by the hour.
  • Delivery through routed gateways is probabilistic and time-varying — hence --repeat (default 10) and a three-state verdict (delivered / intermittent k-of-N / not-delivered). A clean single-window result is still just that window: re-sample in another time slice before publishing a number (disciplines §12).
  • The probe reports token-accounting corroboration but never trusts it alone — behavioral evidence rules; see disciplines §10 (calibrate the meter) and §11 (why the canary is a neutral external fact, not an identity or obedience test).

Dimension 1 — Speed (thinking-aware)

bash
uv run --with openai python scripts/speed_probe.py \
  --base-url <…/v1> --model <model> --key-env <ENV> --mode both --output /tmp/speed.json
  • mixed runs representative tasks (what real usage feels like); decode forces one long output to find the sustained ceiling (the number to compare against a vendor's claim); both does both.
  • The trap this script exists to avoid: reasoning models stream thinking in a separate reasoning_content field, but completion_tokens counts it. Collecting only content while dividing by completion_tokens produces wildly inflated numbers — a real ~750 tok/s model once measured as 4700 tok/s this way. The script captures both, takes TTFT as the first token of either kind, and reports completion_tokens / (total − TTFT).
  • Read the output correctly: real-task throughput is lower than the decode ceiling because short outputs never reach steady state — that's expected, not a bug. Report both numbers, and note when the model emits thinking (its end-to-end latency includes reasoning time, not just typing).

Dimension 2 — Concurrency / stability

bash
uv run --with aiohttp python scripts/concurrency_probe.py \
  --url <full endpoint URL> --model <model> --key-env <ENV> \
  --format openai|anthropic --concurrency 10 20 40 60 --output /tmp/conc.json
  • Pass several --concurrency levels to ramp and find the ceiling — the level where success rate drops or latency explodes. A model that's fast single-threaded can still collapse at modest concurrency (real example: one provider held 50 concurrent at 0.4s while another dropped requests at just 5 concurrent).
  • The script isolates from any ambient proxy (trust_env=False) and disables keep-alive pooling (force_close) — otherwise you measure the proxy's limit or one pinned upstream replica, not the model. It prints a "concurrency proof" (overlapping request pairs) so you can confirm requests really ran in parallel.
  • Distinguish failure modes from the output: HTTP 429 (clean throttle, retriable) vs a TCP drop that hangs to timeout (much worse for UX) vs 5xx. They imply very different fixes.
Show full SKILL.md (1,010 more words)Show less

Dimension 3 — Protocol compliance (Anthropic thinking block: generation AND history-replay)

bash
uv run python scripts/protocol_probe.py \
  --url <…/v1/messages> --model <model> --key-env <ENV> --check all --repeat 10 --output /tmp/proto.json
  • Only relevant for endpoints claiming Anthropic /v1/messages compatibility. --check all (default) runs BOTH sub-checks — they test different code paths and a vendor can pass one while hard-failing the other:
    • generation: does thinking: {type: enabled} actually produce thinking_delta / signature_delta SSE events when you request it? (the original check)
    • history-replay: does the endpoint ACCEPT a type: "thinking" block that's already sitting in a prior assistant turn, when that history is replayed back on a later turn — exactly what Claude Code and every other agentic client does on every continuation? (--check history-replay to run just this one)
  • Real incident (2026-07-21) that motivated the history-replay check: a Kimi/Moonshot model via a China reseller passed generation fine (emits thinking correctly when asked) but hard-rejected history-replay — 400 "invalid part type: thinking" the instant a prior thinking block came back as input, killing the session on every subsequent turn. Passing generation told us nothing about this; they're orthogonal failure modes (response generation vs. request validation).
  • Don't conclude "vendor/reseller X is broken" from one cross-axis comparison — and don't stop at the first single-axis comparison that confirms a vendor-documented parameter either. This took three rounds to get right, kept in disciplines §§17-19 as the canonical cautionary tale: round 1 compared a different model AND a different reseller at once and (wrongly) blamed the reseller. Round 2 fixed that — same reseller, only the model varied — and the result matched a documented, named vendor parameter (Moonshot's own preserve_thinking), which looked like confirmation. Round 2 was STILL wrong: the probe never left that one reseller, so it couldn't see that the vendor's own direct/native endpoint handled the "rejected" model fine — real production traffic proved it. Read §19 before treating any single-reseller-confirmed result as final; check the vendor's own endpoint or real traffic for the actual channel in question before writing an "X doesn't support thinking" conclusion into anything.
  • Compliance is often probabilistic, not binary, for BOTH checks. One real vendor honored the thinking block on only ~13% of generation requests (vs 100% for two competitors). That's why --repeat defaults to 10; generation's verdict has three states (fully-implemented, intermittent (k/N), not-implemented), history-replay's has its own three-plus states (accepts-thinking-in-history, rejects-thinking-in-history, inconsistent, or inconclusive when errors look unrelated to thinking at all). Never conclude from a single sample.
  • It forces Connection: close per request so a load balancer can't pin all samples to one replica and hide the real distribution (a real probe saw 0/10 with keep-alive vs 17/90 with close on the same endpoint).

Dimension 4 — Quality / use-case regression (blind judge)

This is two halves on purpose: collect, then judge independently.

Step 1 — collect the model's answers to your use-case library:

bash
uv run --with openai python scripts/usecase_runner.py \
  --base-url <…/v1> --model <model> --key-env <ENV> \
  --usecases ~/.llm-eval/usecases.json --output-dir ~/.llm-eval/runs/<model>

Step 2 — judge with independent blind judges (orchestrate inline — do NOT let the model grade itself). For each answer in the run directory, spawn 3 independent Task agents (or fewer for a quick pass). Each judge gets ONLY: the prompt, the answer, and the case's rubric — and is explicitly told it is judging in isolation, with no knowledge of other judges' scores or any prior evaluation (this prevents anchoring). Then aggregate:

  • A case passes only on majority agreement among judges.
  • Compute precision per category (using each case's tags): a category where judges systematically disagree with the rubric is a real weakness — on one real eval, a whole category scored 12.5% precision and exposed a systematic misclassification that a single grader would have missed.
  • Count only explicit judgments. A judge that didn't return a verdict is not a pass — silence ≠ consent. This guards against automation bias.

For the rubric-scoring mechanics (LLM-as-judge thresholds, llm-rubric), you can also compose with the promptfoo-evaluation skill — point its providers at the same endpoint. This harness's blind-judge method and promptfoo's rubric assertions are complementary: use promptfoo for fast per-case pass/fail gating, blind judges for precision on a category you suspect is weak. Full method: references/quality_blind_judge.md.

Use-case library

Keep it OUTSIDE this bundle (e.g. ~/.llm-eval/usecases.json) so it survives skill updates and never lands in a public repo. It's a plain JSON list — version it in a private repo to accumulate a regression suite over time:

json
[
  {"id": "refund-window", "prompt": "A customer asks for a refund 20 days after purchase. Reply as support.",
   "rubric": "1.0 if it correctly cites the 30-day refund window; 0.0 if it refuses or invents a different window.",
   "tags": ["support", "policy"]},
  {"id": "lru-cache", "prompt": "Implement an LRU cache in Python with O(1) get/put.",
   "rubric": "1.0 if get and put are both O(1) via dict + doubly linked list and the self-test passes.",
   "tags": ["code"]}
]

assets/example_usecases.json is a starter you can copy. Only id and prompt are required; rubric, expected, and tags make judging sharper.

Running a full evaluation

When the user says "evaluate / benchmark this model", the typical flow is:

  1. Identify the shape — OpenAI-compatible (/v1/chat/completions) or Anthropic-Messages (/v1/messages)? Hit GET /v1/models or read the vendor docs; don't assume. This decides which probes apply (protocol probe is Anthropic-only). Remember the listing is incomplete evidence either way (disciplines §9).
  2. For a new endpoint, availability and fidelity come first — speed numbers for a model whose system prompt never arrives are answering the wrong question.
  3. Run the dimensions the user cares about — speed and concurrency for "is it fast/stable", add protocol for an Anthropic vendor, add quality when they have a use-case suite. Write each probe's --output JSON to a run directory.
  4. Report honestly, separating measured from inferred. Lead with the headline the user asked about (e.g. "sustained decode ceiling exceeds the vendor's claimed tok/s, while real-task throughput runs lower"). If a number looks impossible (e.g. throughput far above the vendor claim, or a single-sample protocol verdict), treat it as a measurement artifact to investigate, not a result — that skepticism is the whole point of this harness. Rates on routed gateways are additionally time-varying: re-sample another window before freezing any number into a document (disciplines §12).
  5. Comparing two models? Run the identical probes against each with the same flags, and put the two JSON outputs side by side. Keep the test conditions identical (same concurrency levels, same use cases) or the comparison is meaningless.
  6. Found a vendor bug worth reporting? Don't paste raw probe output at their support channel — build the evidence package per references/vendor_evidence_protocol.md (self-audit first, observation wording, request-id table, pre-registered thresholds, adversarial counter-review).

For tests that will run repeatedly against a live system — deployment gates, resident canaries, fault-injection mocks, and the monitoring statistics that lie — see references/production_testing_patterns.md.

Next step

After a run, offer the natural follow-ups:

Evaluation complete for <model>.

Options:
A) Render an HTML dashboard of the results — compose with a visualization skill (Recommended if sharing)
B) Compare against another model — same probes, side-by-side
C) Add the failing cases to ~/.llm-eval/usecases.json as a permanent regression guard
D) Done — the numbers answer the question

© daymade, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts, references, assets) in llm-eval-harness of daymade/claude-code-skills.

  • SKILL.md
  • assets/example_usecases.json
  • references/evaluation_disciplines.md
  • references/production_testing_patterns.md
  • references/quality_blind_judge.md
  • references/vendor_evidence_protocol.md
  • scripts/availability_probe.py
  • scripts/concurrency_probe.py
  • scripts/fidelity_probe.py
  • scripts/protocol_probe.py
  • scripts/speed_probe.py
  • scripts/usecase_runner.py

Open the folder on GitHubat commit 91bed2b

Compare with similar skills

LLM Eval Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Eval Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Eval Harness this skilldaymade/claude-code-skills1.4k—~4.7kAutomated safety check: PassMIT
Agents Best PracticesDenisSergeevitch/agents-best-practices2.4k—~7.4kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Codex Fable5baskduf/FableCodex437—~1.6kAutomated safety check: PassAGPL-3.0
System Prompt Writing Guidecashew-labs/libretto904—~570Automated safety check: PassMIT

Similar skills

  • Agents Best Practices

    DenisSergeevitch/agents-best-practices

    A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.

    2.4k GitHub stars~7.4k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Codex Fable5

    baskduf/FableCodex

    Apply a Claude Fable 5 inspired operating style inside Codex.

    437 GitHub stars~1.6k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • System Prompt Writing Guide

    cashew-labs/libretto

    Lays out a minimal, iteration-first approach to writing system prompts for LLM agents, with model-specific notes for Claude, GPT, Gemini, and Codex.

    904 GitHub stars~570 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Context Audit

    undefined-ui/second-brain-os

    Audit an agent's context layout against the four places: system prompt, tools, history, tail.

    1k GitHub stars~802 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from daymade/claude-code-skills

All 103 skills in this repo
  • Video Comparer

    daymade/claude-code-skills

    This skill should be used when comparing two videos to analyze compression results or quality differences.

    1.4k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check: notes
  • CLI Demo Generator

    daymade/claude-code-skills

    Generates professional animated CLI demos as GIFs using VHS terminal recordings.

    1.4k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Doc To Markdown

    daymade/claude-code-skills

    Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.

    1.4k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Interaction Design Board

    daymade/claude-code-skills

    Generates several distinct, clickable HTML interaction prototypes for one product surface into a Design Board and collects selection/remix feedback before implementation.

    1.4k GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Auto Repo Setup

    daymade/claude-code-skills

    Diagnoses and repairs repository setup and guarded Git workflows for Claude Code or Codex — environment repair, startup sync, hook auditing, collaborator handoff.

    1.4k GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Bigdata Skill

    daymade/claude-code-skills

    Pulls Bigdata.com (RavenPack) financial and news data via the official bigdata-client SDK and /v1/ REST endpoints — structured financials, prices, analyst estimates, entity-sentiment series…

    1.4k GitHub stars~3.7k tokensUpdated today
    Auto-check passed

Works with

Questions about LLM Eval Harness

What does LLM Eval Harness do?

Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression. LLM Eval Harness is an agent skill from daymade/claude-code-skills. Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.

When should I use LLM Eval Harness?

LLM Eval Harness fits situations like: tasks that involve LLM evaluation; tasks that involve Prompt engineering; tasks that involve Text to speech and voice.

How do I install LLM Eval Harness in Claude Code?

Run `npx skills add daymade/claude-code-skills --skill llm-eval-harness -a claude-code`. Or copy the skill folder (llm-eval-harness in daymade/claude-code-skills) into .claude/skills/llm-eval-harness in your project. Claude Code loads it when a task matches its description.

How do I install LLM Eval Harness in Codex?

Run `npx skills add daymade/claude-code-skills --skill llm-eval-harness -a codex`. Or copy the skill folder (llm-eval-harness in daymade/claude-code-skills) into .agents/skills/llm-eval-harness in your project. Codex loads it when a task matches its description.

Can I use LLM Eval Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add daymade/claude-code-skills --skill llm-eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-eval-harness, .gemini/skills/llm-eval-harness, .github/skills/llm-eval-harness and .opencode/skills/llm-eval-harness in your project.

What does LLM Eval Harness need to run?

Going by SKILL.md and its folder, LLM Eval Harness needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does LLM Eval Harness access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is LLM Eval Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does LLM Eval Harness use?

LLM Eval Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Eval Harness use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.

What are the alternatives to LLM Eval Harness?

Skills that share tags, products or a category with LLM Eval Harness: Agents Best Practices (DenisSergeevitch/agents-best-practices, 2.4k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Codex Fable5 (baskduf/FableCodex, 437 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Eval Harness?

daymade (a GitHub user) maintains it in daymade/claude-code-skills, which has 1,444 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 8, 2026.

Source: daymade/claude-code-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.