Agents Best Practices
DenisSergeevitch/agents-best-practices
A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.
Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.
$ npx skills add daymade/claude-code-skills --skill llm-eval-harness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install daymade/claude-code-skills llm-eval-harness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/llm-eval-harness .claude/skills/llm-eval-harness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "llm-eval-harness" agent skill from https://github.com/daymade/claude-code-skills/tree/main/llm-eval-harness into .claude/skills/llm-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-eval-harness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/daymade/claude-code-skills/tree/main/llm-eval-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add daymade/claude-code-skills --skill llm-eval-harness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install daymade/claude-code-skills llm-eval-harness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/llm-eval-harness .agents/skills/llm-eval-harness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "llm-eval-harness" agent skill from https://github.com/daymade/claude-code-skills/tree/main/llm-eval-harness into .agents/skills/llm-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-eval-harness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add daymade/claude-code-skills --skill llm-eval-harness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install daymade/claude-code-skills llm-eval-harness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/llm-eval-harness .cursor/skills/llm-eval-harness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "llm-eval-harness" agent skill from https://github.com/daymade/claude-code-skills/tree/main/llm-eval-harness into .cursor/skills/llm-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-eval-harness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/daymade/claude-code-skills.git --path llm-eval-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add daymade/claude-code-skills --skill llm-eval-harness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install daymade/claude-code-skills llm-eval-harness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/llm-eval-harness .gemini/skills/llm-eval-harness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "llm-eval-harness" agent skill from https://github.com/daymade/claude-code-skills/tree/main/llm-eval-harness into .gemini/skills/llm-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-eval-harness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install daymade/claude-code-skills llm-eval-harnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add daymade/claude-code-skills --skill llm-eval-harness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/llm-eval-harness .github/skills/llm-eval-harness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "llm-eval-harness" agent skill from https://github.com/daymade/claude-code-skills/tree/main/llm-eval-harness into .github/skills/llm-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-eval-harness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add daymade/claude-code-skills --skill llm-eval-harness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install daymade/claude-code-skills llm-eval-harness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/llm-eval-harness .opencode/skills/llm-eval-harness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "llm-eval-harness" agent skill from https://github.com/daymade/claude-code-skills/tree/main/llm-eval-harness into .opencode/skills/llm-eval-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "llm-eval-harness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
llm-eval-harnessTests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.
LLM Eval Harness is an agent skill from daymade/claude-code-skills. Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression. Use before hand-rolling a curl loop, when onboarding a provider, or debugging "system prompt 不生效" / a tok/s claim — 测评/压测一个模型或渠道. Not for TTS/voice-clone supplier eval (use the audio skill).
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts, reference files and assets (for example `assets/example_usecases.json`, `references/evaluation_disciplines.md` and `references/production_testing_patterns.md`).
It sits in AI & LLM Engineering, covering LLM evaluation, Prompt engineering and Text to speech and voice. It works with OpenAI. The repository describes itself as: Professional Claude Code skills marketplace featuring production-ready skills for enhanced development workflows. The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 91bed2b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
LLM Eval Harness loads about 4.7k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 2,120 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from daymade/claude-code-skills at commit 91bed2b, republished under its MIT licence (© daymade). 2,120 words, ~4,658 tokens.
.claude/skills/llm-eval-harness/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Give this skill an endpoint (base_url + model + an API key in an env var) and it
measures whether the endpoint actually works and whether the model is fast, stable,
protocol-correct, and good enough — instead of trusting the vendor's headline numbers.
Six dimensions, usually scattered across ad-hoc scripts that get rewritten (with the
same bugs) every time:
| Dimension | Script | Answers |
|---|---|---|
| Availability | scripts/availability_probe.py | which model IDs work here, with 3-state error classification |
| Request fidelity | scripts/fidelity_probe.py | do system prompt / tools / history actually REACH the model? |
| Speed | scripts/speed_probe.py | TTFT + sustained decode tok/s, thinking-aware |
| Concurrency / stability | scripts/concurrency_probe.py | success rate, p50/p90 latency, where it breaks |
| Protocol compliance | scripts/protocol_probe.py | does the Anthropic thinking block actually fire when requested, AND does the endpoint accept one already sitting in history on a later turn (N≥10, both are separate checks — see Dimension 3)? |
| Quality / use-case regression | scripts/usecase_runner.py + blind judges | does it pass your accumulated cases? |
Route from what the user is actually doing:
--check history-replay specifically — this symptom is the
generation check passing while history-replay silently doesn't; don't stop at generation--output, then follow
references/vendor_evidence_protocol.mdDon't run all six ritually — pick what answers the user's question.
Key handling (non-negotiable): every script takes the API key by env-var name
(--key-env MY_KEY), never the key value on the command line — so it stays out of ps,
shell history, and any saved report. Never hardcode a key into a use-case file or a
wrapper. Read references/evaluation_disciplines.md
for the full reasoning behind this and the other disciplines.
Your private data lives outside this bundle. Use-case libraries, model rosters, and
keys belong in ~/.llm-eval/ (or wherever you keep secrets), NOT in this skill directory —
the skill is generic and public; your test suite is yours. See "Use-case library" below.
Detect what you have, then run the dimensions that apply. For an OpenAI-compatible model:
export MY_KEY=sk-... # the key never appears in a command below
# Speed: real-task throughput + sustained decode ceiling
uv run --with openai python scripts/speed_probe.py \
--base-url https://api.example.com/v1 --model some-model --key-env MY_KEY --mode both
# Concurrency: ramp until it breaks
uv run --with aiohttp python scripts/concurrency_probe.py \
--url https://api.example.com/v1/chat/completions --model some-model --key-env MY_KEY \
--format openai --concurrency 10 20 40 60If the endpoint is Anthropic-Messages-shaped (/v1/messages), also run the protocol probe
(below). Pick dimensions by what the user actually asked — don't run all six if they only
asked "is it fast?".
uv run --with aiohttp python scripts/availability_probe.py \
--base-url <base> --key-env <ENV> --format both \
--models model-a vendor/model-b @models.txt --output /tmp/avail.jsonno-channel (no route for this
ID), upstream-error (route exists, upstream failing — retry later), empty-content
(usually a max_tokens artifact on reasoning models, NOT a broken model). The probe
keeps --max-tokens at 8192 by default precisely so thinking models don't read as
dead — do not "optimize" it downward.-preview, dated variants, tier names) decide routability. Full traps: disciplines §9.some-model[1m] are frequently a CLI client's own local convention (Claude
Code parses and strips [1m] before the request is ever built — see
claude-switch-models-setup's "Configuring Context Window Size" section for the full
mechanism) — they never appear on the wire. The probe warns when it sees one, but the
discipline is yours: probe the bare model ID the vendor actually documents, not a string
copied out of a Claude Code ANTHROPIC_MODEL env var./v1/models listing is one input, never the verdict — real gateways
route models the listing omits.uv run --with aiohttp python scripts/fidelity_probe.py \
--base-url <base> --model <model> --key-env <ENV> \
--format anthropic --check all --repeat 10 --output /tmp/fidelity.jsonsystem (canary code planted in the system prompt — can the model echo
it back?), tools (full round-trip including returning a tool result and verifying
the final answer uses it), multiturn (plant facts across turns, recall them all),
auth (x-api-key vs Bearer on both protocol endpoints — gateways commonly accept
both on one path and only one on the other, and the wrong-header 401 reads exactly
like a dead key).--repeat (default 10) and a three-state verdict (delivered / intermittent k-of-N /
not-delivered). A clean single-window result is still just that window: re-sample in
another time slice before publishing a number (disciplines §12).uv run --with openai python scripts/speed_probe.py \
--base-url <…/v1> --model <model> --key-env <ENV> --mode both --output /tmp/speed.jsonmixed runs representative tasks (what real usage feels like); decode forces one long
output to find the sustained ceiling (the number to compare against a vendor's claim);
both does both.reasoning_content field, but completion_tokens counts it. Collecting only content
while dividing by completion_tokens produces wildly inflated numbers — a real ~750 tok/s
model once measured as 4700 tok/s this way. The script captures both, takes TTFT as the
first token of either kind, and reports completion_tokens / (total − TTFT).uv run --with aiohttp python scripts/concurrency_probe.py \
--url <full endpoint URL> --model <model> --key-env <ENV> \
--format openai|anthropic --concurrency 10 20 40 60 --output /tmp/conc.json--concurrency levels to ramp and find the ceiling — the level where success
rate drops or latency explodes. A model that's fast single-threaded can still collapse at
modest concurrency (real example: one provider held 50 concurrent at 0.4s while another
dropped requests at just 5 concurrent).trust_env=False) and disables keep-alive
pooling (force_close) — otherwise you measure the proxy's limit or one pinned upstream
replica, not the model. It prints a "concurrency proof" (overlapping request pairs) so you
can confirm requests really ran in parallel.uv run python scripts/protocol_probe.py \
--url <…/v1/messages> --model <model> --key-env <ENV> --check all --repeat 10 --output /tmp/proto.json/v1/messages compatibility. --check all
(default) runs BOTH sub-checks — they test different code paths and a vendor can pass one
while hard-failing the other:thinking: {type: enabled} actually produce thinking_delta /
signature_delta SSE events when you request it? (the original check)type: "thinking" block that's already
sitting in a prior assistant turn, when that history is replayed back on a later turn —
exactly what Claude Code and every other agentic client does on every continuation?
(--check history-replay to run just this one)preserve_thinking), which
looked like confirmation. Round 2 was STILL wrong: the probe never left that one reseller, so
it couldn't see that the vendor's own direct/native endpoint handled the "rejected" model fine
— real production traffic proved it. Read §19 before treating any single-reseller-confirmed
result as final; check the vendor's own endpoint or real traffic for the actual channel in
question before writing an "X doesn't support thinking" conclusion into anything.--repeat defaults to 10; generation's verdict has three states (fully-implemented,
intermittent (k/N), not-implemented), history-replay's has its own three-plus states
(accepts-thinking-in-history, rejects-thinking-in-history, inconsistent, or
inconclusive when errors look unrelated to thinking at all). Never conclude from a single
sample.Connection: close per request so a load balancer can't pin all samples to one
replica and hide the real distribution (a real probe saw 0/10 with keep-alive vs 17/90
with close on the same endpoint).This is two halves on purpose: collect, then judge independently.
Step 1 — collect the model's answers to your use-case library:
uv run --with openai python scripts/usecase_runner.py \
--base-url <…/v1> --model <model> --key-env <ENV> \
--usecases ~/.llm-eval/usecases.json --output-dir ~/.llm-eval/runs/<model>Step 2 — judge with independent blind judges (orchestrate inline — do NOT let the model
grade itself). For each answer in the run directory, spawn 3 independent Task agents (or
fewer for a quick pass). Each judge gets ONLY: the prompt, the answer, and the case's
rubric — and is explicitly told it is judging in isolation, with no knowledge of other
judges' scores or any prior evaluation (this prevents anchoring). Then aggregate:
tags): a category where judges
systematically disagree with the rubric is a real weakness — on one real eval, a whole
category scored 12.5% precision and exposed a systematic misclassification that a single
grader would have missed.For the rubric-scoring mechanics (LLM-as-judge thresholds, llm-rubric), you can also
compose with the promptfoo-evaluation skill — point its providers at the same endpoint.
This harness's blind-judge method and promptfoo's rubric assertions are complementary: use
promptfoo for fast per-case pass/fail gating, blind judges for precision on a category you
suspect is weak. Full method: references/quality_blind_judge.md.
Keep it OUTSIDE this bundle (e.g. ~/.llm-eval/usecases.json) so it survives skill updates
and never lands in a public repo. It's a plain JSON list — version it in a private repo to
accumulate a regression suite over time:
[
{"id": "refund-window", "prompt": "A customer asks for a refund 20 days after purchase. Reply as support.",
"rubric": "1.0 if it correctly cites the 30-day refund window; 0.0 if it refuses or invents a different window.",
"tags": ["support", "policy"]},
{"id": "lru-cache", "prompt": "Implement an LRU cache in Python with O(1) get/put.",
"rubric": "1.0 if get and put are both O(1) via dict + doubly linked list and the self-test passes.",
"tags": ["code"]}
]assets/example_usecases.json is a starter you can copy. Only id and prompt are required;
rubric, expected, and tags make judging sharper.
When the user says "evaluate / benchmark this model", the typical flow is:
/v1/chat/completions) or Anthropic-Messages
(/v1/messages)? Hit GET /v1/models or read the vendor docs; don't assume. This decides
which probes apply (protocol probe is Anthropic-only). Remember the listing is incomplete
evidence either way (disciplines §9).--output JSON to a run directory.For tests that will run repeatedly against a live system — deployment gates, resident canaries, fault-injection mocks, and the monitoring statistics that lie — see references/production_testing_patterns.md.
After a run, offer the natural follow-ups:
Evaluation complete for <model>.
Options:
A) Render an HTML dashboard of the results — compose with a visualization skill (Recommended if sharing)
B) Compare against another model — same probes, side-by-side
C) Add the failing cases to ~/.llm-eval/usecases.json as a permanent regression guard
D) Done — the numbers answer the question© daymade, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (scripts, references, assets) in llm-eval-harness of daymade/claude-code-skills.
Open the folder on GitHubat commit 91bed2b
LLM Eval Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| LLM Eval Harness this skilldaymade/claude-code-skills | 1.4k | — | ~4.7k | Automated safety check: Pass | MIT | |
| Agents Best PracticesDenisSergeevitch/agents-best-practices | 2.4k | — | ~7.4k | Automated safety check: Pass | MIT | |
| Azure AI Projects Python SDKmicrosoft/skills | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Fine-Tuning ExpertJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Codex Fable5baskduf/FableCodex | 437 | — | ~1.6k | Automated safety check: Pass | AGPL-3.0 | |
| System Prompt Writing Guidecashew-labs/libretto | 904 | — | ~570 | Automated safety check: Pass | MIT |
DenisSergeevitch/agents-best-practices
A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.
microsoft/skills
Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
baskduf/FableCodex
Apply a Claude Fable 5 inspired operating style inside Codex.
cashew-labs/libretto
Lays out a minimal, iteration-first approach to writing system prompts for LLM agents, with model-specific notes for Claude, GPT, Gemini, and Codex.
undefined-ui/second-brain-os
Audit an agent's context layout against the four places: system prompt, tools, history, tail.
daymade/claude-code-skills
This skill should be used when comparing two videos to analyze compression results or quality differences.
daymade/claude-code-skills
Generates professional animated CLI demos as GIFs using VHS terminal recordings.
daymade/claude-code-skills
Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.
daymade/claude-code-skills
Generates several distinct, clickable HTML interaction prototypes for one product surface into a Design Board and collects selection/remix feedback before implementation.
daymade/claude-code-skills
Diagnoses and repairs repository setup and guarded Git workflows for Claude Code or Codex — environment repair, startup sync, hook auditing, collaborator handoff.
daymade/claude-code-skills
Pulls Bigdata.com (RavenPack) financial and news data via the official bigdata-client SDK and /v1/ REST endpoints — structured financials, prices, analyst estimates, entity-sentiment series…
Works with
Categories
Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression. LLM Eval Harness is an agent skill from daymade/claude-code-skills. Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.
LLM Eval Harness fits situations like: tasks that involve LLM evaluation; tasks that involve Prompt engineering; tasks that involve Text to speech and voice.
Run `npx skills add daymade/claude-code-skills --skill llm-eval-harness -a claude-code`. Or copy the skill folder (llm-eval-harness in daymade/claude-code-skills) into .claude/skills/llm-eval-harness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add daymade/claude-code-skills --skill llm-eval-harness -a codex`. Or copy the skill folder (llm-eval-harness in daymade/claude-code-skills) into .agents/skills/llm-eval-harness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add daymade/claude-code-skills --skill llm-eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-eval-harness, .gemini/skills/llm-eval-harness, .github/skills/llm-eval-harness and .opencode/skills/llm-eval-harness in your project.
Going by SKILL.md and its folder, LLM Eval Harness needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
LLM Eval Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with LLM Eval Harness: Agents Best Practices (DenisSergeevitch/agents-best-practices, 2.4k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Codex Fable5 (baskduf/FableCodex, 437 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
daymade (a GitHub user) maintains it in daymade/claude-code-skills, which has 1,444 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 8, 2026.
Source: daymade/claude-code-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.