Evaluate RAG
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks.
$ npx skills add ancoleman/ai-design-components --skill evaluating-llms -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ancoleman/ai-design-components evaluating-llms --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluating-llms .claude/skills/evaluating-llms && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluating-llms" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/evaluating-llms into .claude/skills/evaluating-llms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ancoleman/ai-design-components/tree/main/skills/evaluating-llmsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ancoleman/ai-design-components --skill evaluating-llms -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ancoleman/ai-design-components evaluating-llms --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/evaluating-llms .agents/skills/evaluating-llms && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluating-llms" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/evaluating-llms into .agents/skills/evaluating-llms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill evaluating-llms -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ancoleman/ai-design-components evaluating-llms --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/evaluating-llms .cursor/skills/evaluating-llms && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluating-llms" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/evaluating-llms into .cursor/skills/evaluating-llms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ancoleman/ai-design-components.git --path skills/evaluating-llms--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ancoleman/ai-design-components --skill evaluating-llms -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ancoleman/ai-design-components evaluating-llms --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/evaluating-llms .gemini/skills/evaluating-llms && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluating-llms" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/evaluating-llms into .gemini/skills/evaluating-llms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ancoleman/ai-design-components evaluating-llmsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ancoleman/ai-design-components --skill evaluating-llms -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/evaluating-llms .github/skills/evaluating-llms && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluating-llms" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/evaluating-llms into .github/skills/evaluating-llms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill evaluating-llms -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ancoleman/ai-design-components evaluating-llms --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/evaluating-llms .opencode/skills/evaluating-llms && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluating-llms" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/evaluating-llms into .opencode/skills/evaluating-llms/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluating-llms", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluating-llmsEvaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks.
Evaluating LLMs is an agent skill from ancoleman/ai-design-components. Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating RAG pipelines, measuring safety (hallucinations, bias), or comparing models for production deployment.
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files, including reference files (for example `examples/python/benchmark_testing.py`, `examples/python/classification_metrics.py` and `examples/python/deepeval_example.py`).
It sits in AI & LLM Engineering, covering LLM evaluation and Retrieval-augmented generation. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python and TypeScript, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pippythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluating LLMs loads about 4.7k tokens when it runs, and up to ~33k if it reads all its reference files. Until then it costs about 60 tokens; SKILL.md has 1,622 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 1,622 words, ~4,666 tokens.
.claude/skills/evaluating-llms/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.Evaluate Large Language Model (LLM) systems using automated metrics, LLM-as-judge patterns, and standardized benchmarks to ensure production quality and safety.
Apply this skill when:
Common triggers:
By Task Type:
| Task Type | Primary Approach | Metrics | Tools |
|---|---|---|---|
| Classification (sentiment, intent) | Automated metrics | Accuracy, Precision, Recall, F1 | scikit-learn |
| Generation (summaries, creative text) | LLM-as-judge + automated | BLEU, ROUGE, BERTScore, Quality rubric | GPT-4/Claude for judging |
| Question Answering | Exact match + semantic similarity | EM, F1, Cosine similarity | Custom evaluators |
| RAG Systems | RAGAS framework | Faithfulness, Answer/Context relevance | RAGAS library |
| Code Generation | Unit tests + execution | Pass@K, Test pass rate | HumanEval, pytest |
| Multi-step Agents | Task completion + tool accuracy | Success rate, Efficiency | Custom evaluators |
By Volume and Cost:
| Samples | Speed | Cost | Recommended Approach |
|---|---|---|---|
| 1,000+ | Immediate | $0 | Automated metrics (regex, JSON validation) |
| 100-1,000 | Minutes | $0.01-0.10 each | LLM-as-judge (GPT-4, Claude) |
| < 100 | Hours | $1-10 each | Human evaluation (pairwise comparison) |
Layered Approach (Recommended for Production):
Test single prompt-response pairs for correctness.
Methods:
Example Use Cases:
Quick Start (Python):
import pytest
from openai import OpenAI
client = OpenAI()
def classify_sentiment(text: str) -> str:
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=[
{"role": "system", "content": "Classify sentiment as positive, negative, or neutral. Return only the label."},
{"role": "user", "content": text}
],
temperature=0
)
return response.choices[0].message.content.strip().lower()
def test_positive_sentiment():
result = classify_sentiment("I love this product!")
assert result == "positive"For complete unit evaluation examples, see examples/python/unit_evaluation.py and examples/typescript/unit-evaluation.ts.
Evaluate RAG systems using RAGAS framework metrics.
Critical Metrics (Priority Order):
Faithfulness (Target: > 0.8) - MOST CRITICAL
Answer Relevance (Target: > 0.7)
Context Relevance (Target: > 0.7)
Context Precision (Target: > 0.5)
Context Recall (Target: > 0.8)
Quick Start (Python with RAGAS):
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_relevancy
from datasets import Dataset
data = {
"question": ["What is the capital of France?"],
"answer": ["The capital of France is Paris."],
"contexts": [["Paris is the capital of France."]],
"ground_truth": ["Paris"]
}
dataset = Dataset.from_dict(data)
results = evaluate(dataset, metrics=[faithfulness, answer_relevancy, context_relevancy])
print(f"Faithfulness: {results['faithfulness']:.2f}")For comprehensive RAG evaluation patterns, see references/rag-evaluation.md and examples/python/ragas_example.py.
Use powerful LLMs (GPT-4, Claude Opus) to evaluate other LLM outputs.
When to Use:
Correlation with Human Judgment: 0.75-0.85 for well-designed rubrics
Best Practices:
Quick Start (Python):
from openai import OpenAI
client = OpenAI()
def evaluate_quality(prompt: str, response: str) -> tuple[int, str]:
"""Returns (score 1-5, reasoning)"""
eval_prompt = f"""
Rate the following LLM response on relevance and helpfulness.
USER PROMPT: {prompt}
LLM RESPONSE: {response}
Provide:
Score: [1-5, where 5 is best]
Reasoning: [1-2 sentences]
"""
result = client.chat.completions.create(
model="gpt-4",
messages=[{"role": "user", "content": eval_prompt}],
temperature=0.3
)
content = result.choices[0].message.content
lines = content.strip().split('\n')
score = int(lines[0].split(':')[1].strip())
reasoning = lines[1].split(':', 1)[1].strip()
return score, reasoningFor detailed LLM-as-judge patterns and prompt templates, see references/llm-as-judge.md and examples/python/llm_as_judge.py.
Measure hallucinations, bias, and toxicity in LLM outputs.
Methods:
Faithfulness to Context (RAG):
Factual Accuracy (Closed-Book):
Self-Consistency:
Types of Bias:
Evaluation Methods:
Stereotype Tests:
Counterfactual Evaluation:
Tools:
For comprehensive safety evaluation patterns, see references/safety-evaluation.md.
Assess model capabilities using standardized benchmarks.
Standard Benchmarks:
| Benchmark | Coverage | Format | Difficulty | Use Case |
|---|---|---|---|---|
| MMLU | 57 subjects (STEM, humanities) | Multiple choice | High school - professional | General intelligence |
| HellaSwag | Sentence completion | Multiple choice | Common sense | Reasoning validation |
| GPQA | PhD-level science | Multiple choice | Very high (expert-level) | Frontier model testing |
| HumanEval | 164 Python problems | Code generation | Medium | Code capability |
| MATH | 12,500 competition problems | Math solving | High school competitions | Math reasoning |
Domain-Specific Benchmarks:
When to Use Benchmarks:
Quick Start (lm-evaluation-harness):
pip install lm-eval
# Evaluate GPT-4 on MMLU
lm_eval --model openai-chat --model_args model=gpt-4 --tasks mmlu --num_fewshot 5For detailed benchmark testing patterns, see references/benchmarks.md and scripts/benchmark_runner.py.
Monitor and optimize LLM quality in production environments.
Compare two LLM configurations:
Metrics:
Real-time quality monitoring:
Sample-based human evaluation:
For production evaluation patterns and monitoring strategies, see references/production-evaluation.md.
For tasks with discrete outputs (sentiment, intent, category).
Metrics:
Quick Start (Python):
from sklearn.metrics import accuracy_score, precision_recall_fscore_support
y_true = ["positive", "negative", "neutral", "positive", "negative"]
y_pred = ["positive", "negative", "neutral", "neutral", "negative"]
accuracy = accuracy_score(y_true, y_pred)
precision, recall, f1, _ = precision_recall_fscore_support(y_true, y_pred, average='weighted')
print(f"Accuracy: {accuracy:.2f}")
print(f"Precision: {precision:.2f}")
print(f"Recall: {recall:.2f}")
print(f"F1 Score: {f1:.2f}")For complete classification evaluation examples, see examples/python/classification_metrics.py.
For open-ended text generation (summaries, creative writing, responses).
Automated Metrics (Use with Caution):
Limitation: Automated metrics correlate weakly with human judgment for creative/subjective generation.
Recommended Approach:
For detailed generation evaluation patterns, see references/evaluation-types.md.
| If Task Is... | Use This Framework | Primary Metric |
|---|---|---|
| RAG system | RAGAS | Faithfulness > 0.8 |
| Classification | scikit-learn metrics | Accuracy, F1 |
| Generation quality | LLM-as-judge | Quality rubric (1-5) |
| Code generation | HumanEval | Pass@1, Test pass rate |
| Model comparison | Benchmark testing | MMLU, HellaSwag scores |
| Safety validation | Hallucination detection | Faithfulness, Fact-check |
| Production monitoring | Online evaluation | User feedback, Business KPIs |
| Library | Use Case | Installation |
|---|---|---|
| RAGAS | RAG evaluation | pip install ragas |
| DeepEval | General LLM evaluation, pytest integration | pip install deepeval |
| LangSmith | Production monitoring, A/B testing | pip install langsmith |
| lm-eval | Benchmark testing (MMLU, HumanEval) | pip install lm-eval |
| scikit-learn | Classification metrics | pip install scikit-learn |
| Application | Hallucination Risk | Bias Risk | Toxicity Risk | Evaluation Priority |
|---|---|---|---|---|
| Customer Support | High | Medium | High | 1. Faithfulness, 2. Toxicity, 3. Bias |
| Medical Diagnosis | Critical | High | Low | 1. Factual Accuracy, 2. Hallucination, 3. Bias |
| Creative Writing | Low | Medium | Medium | 1. Quality/Fluency, 2. Content Policy |
| Code Generation | Medium | Low | Low | 1. Functional Correctness, 2. Security |
| Content Moderation | Low | Critical | Critical | 1. Bias, 2. False Positives/Negatives |
For comprehensive documentation on specific topics:
references/evaluation-types.mdreferences/rag-evaluation.mdreferences/safety-evaluation.mdreferences/benchmarks.mdreferences/llm-as-judge.mdreferences/production-evaluation.mdreferences/metrics-reference.mdPython Examples:
examples/python/unit_evaluation.py - Basic prompt testing with pytestexamples/python/ragas_example.py - RAGAS RAG evaluationexamples/python/deepeval_example.py - DeepEval framework usageexamples/python/llm_as_judge.py - GPT-4 as evaluatorexamples/python/classification_metrics.py - Accuracy, precision, recallexamples/python/benchmark_testing.py - HumanEval exampleTypeScript Examples:
examples/typescript/unit-evaluation.ts - Vitest + OpenAIexamples/typescript/llm-as-judge.ts - GPT-4 evaluationexamples/typescript/langsmith-integration.ts - Production monitoringRun evaluations without loading code into context (token-free):
scripts/run_ragas_eval.py - Run RAGAS evaluation on datasetscripts/compare_models.py - A/B test two modelsscripts/benchmark_runner.py - Run MMLU/HumanEval benchmarksscripts/hallucination_checker.py - Detect hallucinations in outputsExample usage:
# Run RAGAS evaluation on custom dataset
python scripts/run_ragas_eval.py --dataset data/qa_dataset.json --output results.json
# Compare GPT-4 vs Claude on benchmark
python scripts/compare_models.py --model-a gpt-4 --model-b claude-3-opus --tasks mmlu,humanevalRelated Skills:
building-ai-chat: Evaluate AI chat applications (this skill tests what that skill builds)prompt-engineering: Test prompt quality and effectivenesstesting-strategies: Apply testing pyramid to LLM evaluation (unit → integration → E2E)observability: Production monitoring and alerting for LLM qualitybuilding-ci-pipelines: Integrate LLM evaluation into CI/CDWorkflow Integration:
prompt-engineering skill)llm-evaluation skill)building-ai-chat skill)llm-evaluation skill)deploying-applications skill)llm-evaluation + observability skills)1. Over-reliance on Automated Metrics for Generation
2. Ignoring Faithfulness in RAG Systems
3. No Production Monitoring
4. Biased LLM-as-Judge Evaluation
5. Insufficient Benchmark Coverage
6. Missing Safety Evaluation
© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 17 other files (references) in skills/evaluating-llms of ancoleman/ai-design-components.
Open the folder on GitHubat commit 76551b7
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ancoleman/ai-design-components, which our catalogue first saw on October 7, 2026.
Evaluating LLMs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluating LLMs this skillancoleman/ai-design-components | 526 | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Evaluate RAGai-evals-course/evals-skills | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| RAG ArchitectJeffallan/claude-skills | 12k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Jd Gap Analysisstarkyru/learn-ai | 105 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Agent Evalericrisco/rsc-harness | 156 | — | ~3.2k | Automated safety check: Pass | MIT | |
| RAG Observability Evalssickn33/agentic-awesome-skills | 47k | 2 repos | ~3.1k | Automated safety check: Pass | MIT |
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
Jeffallan/claude-skills
Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.
starkyru/learn-ai
Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
sickn33/agentic-awesome-skills
Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.
davepoon/buildwithclaude
Evaluate retrieval and citation behavior for RAG pipelines from deterministic JSONL fixtures.
ancoleman/ai-design-components
Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.
ancoleman/ai-design-components
Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.
ancoleman/ai-design-components
Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.
ancoleman/ai-design-components
Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.
ancoleman/ai-design-components
Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.
ancoleman/ai-design-components
Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.
Categories
Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Evaluating LLMs is an agent skill from ancoleman/ai-design-components. Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks.
Evaluating LLMs fits situations like: testing prompt quality; validating RAG pipelines; measuring safety (hallucinations; comparing models for production deployment.
Run `npx skills add ancoleman/ai-design-components --skill evaluating-llms -a claude-code`. Or copy the skill folder (skills/evaluating-llms in ancoleman/ai-design-components) into .claude/skills/evaluating-llms in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ancoleman/ai-design-components --skill evaluating-llms -a codex`. Or copy the skill folder (skills/evaluating-llms in ancoleman/ai-design-components) into .agents/skills/evaluating-llms in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill evaluating-llms -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluating-llms, .gemini/skills/evaluating-llms, .github/skills/evaluating-llms and .opencode/skills/evaluating-llms in your project.
Going by SKILL.md and its folder, Evaluating LLMs needs Python and TypeScript for the scripts in its folder and the command-line tools its instructions call (pip and python). Our summary lists: Python 3; Node.js.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evaluating LLMs is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 29k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Evaluating LLMs: Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars), RAG Architect (Jeffallan/claude-skills, 12k stars), Jd Gap Analysis (starkyru/learn-ai, 105 stars) and Agent Eval (ericrisco/rsc-harness, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.
Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.