Agent skill

Evaluating LLMs

by ancoleman in ancoleman/ai-design-components

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks.

MITAuto-check passedAI & LLM Engineering

Install Evaluating LLMs

skills CLI
$ npx skills add ancoleman/ai-design-components --skill evaluating-llms -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ancoleman/ai-design-components evaluating-llms --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluating-llms .claude/skills/evaluating-llms && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluating-llms
GitHub stars
526
Used in
1 other repo
Token cost
~4.7k tokens
SKILL.md length
1,622 words
Files
18 (incl. references)
Skills in repo
75
Repo updated
First seen
Licence
MIT

At a glance

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks.

  • Works in 3 steps: Layer 1: Automated metrics for all… → Layer 2: LLM-as-judge for 10% sample… → Layer 3: Human review for 1% edge cases…
  • Testing prompt quality
  • SKILL.md covers When to Use This Skill, Evaluation Strategy Selection, Core Evaluation Patterns and Classification Task Evaluation, plus 7 more sections
  • Runs Python and TypeScript scripts from its folder; calls pip and python

What it does

Evaluating LLMs is an agent skill from ancoleman/ai-design-components. Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating RAG pipelines, measuring safety (hallucinations, bias), or comparing models for production deployment.

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files, including reference files (for example `examples/python/benchmark_testing.py`, `examples/python/classification_metrics.py` and `examples/python/deepeval_example.py`).

It sits in AI & LLM Engineering, covering LLM evaluation and Retrieval-augmented generation. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.

When your agent uses it

  • Testing prompt quality
  • Validating RAG pipelines
  • Measuring safety (hallucinations
  • Comparing models for production deployment

Example prompts

  • “/evaluating-llms”

Requirements

  • Python 3
  • Node.js

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Layer 1: Automated metrics for all outputs (fast, cheap)
  2. Layer 2: LLM-as-judge for 10% sample (nuanced quality)
  3. Layer 3: Human review for 1% edge cases (validation)

What it can do on your machine

Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and TypeScript, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluating LLMs loads about 4.7k tokens when it runs, and up to ~33k if it reads all its reference files. Until then it costs about 60 tokens; SKILL.md has 1,622 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~33k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 1,622 words, ~4,666 tokens.

Download SKILL.mdSave it as .claude/skills/evaluating-llms/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
evaluating-llms
description
Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Use when testing prompt quality, validating RAG pipelines, measuring safety (hallucinations, bias), or comparing models for production deployment.

LLM Evaluation

Evaluate Large Language Model (LLM) systems using automated metrics, LLM-as-judge patterns, and standardized benchmarks to ensure production quality and safety.

When to Use This Skill

Apply this skill when:

  • Testing individual prompts for correctness and formatting
  • Validating RAG (Retrieval-Augmented Generation) pipeline quality
  • Measuring hallucinations, bias, or toxicity in LLM outputs
  • Comparing different models or prompt configurations (A/B testing)
  • Running benchmark tests (MMLU, HumanEval) to assess model capabilities
  • Setting up production monitoring for LLM applications
  • Integrating LLM quality checks into CI/CD pipelines

Common triggers:

  • "How do I test if my RAG system is working correctly?"
  • "How can I measure hallucinations in LLM outputs?"
  • "What metrics should I use to evaluate generation quality?"
  • "How do I compare GPT-4 vs Claude for my use case?"
  • "How do I detect bias in LLM responses?"

Evaluation Strategy Selection

Decision Framework: Which Evaluation Approach?

By Task Type:

Task TypePrimary ApproachMetricsTools
Classification (sentiment, intent)Automated metricsAccuracy, Precision, Recall, F1scikit-learn
Generation (summaries, creative text)LLM-as-judge + automatedBLEU, ROUGE, BERTScore, Quality rubricGPT-4/Claude for judging
Question AnsweringExact match + semantic similarityEM, F1, Cosine similarityCustom evaluators
RAG SystemsRAGAS frameworkFaithfulness, Answer/Context relevanceRAGAS library
Code GenerationUnit tests + executionPass@K, Test pass rateHumanEval, pytest
Multi-step AgentsTask completion + tool accuracySuccess rate, EfficiencyCustom evaluators

By Volume and Cost:

SamplesSpeedCostRecommended Approach
1,000+Immediate$0Automated metrics (regex, JSON validation)
100-1,000Minutes$0.01-0.10 eachLLM-as-judge (GPT-4, Claude)
< 100Hours$1-10 eachHuman evaluation (pairwise comparison)

Layered Approach (Recommended for Production):

  1. Layer 1: Automated metrics for all outputs (fast, cheap)
  2. Layer 2: LLM-as-judge for 10% sample (nuanced quality)
  3. Layer 3: Human review for 1% edge cases (validation)

Core Evaluation Patterns

Unit Evaluation (Individual Prompts)

Test single prompt-response pairs for correctness.

Methods:

  • Exact Match: Response exactly matches expected output
  • Regex Matching: Response follows expected pattern
  • JSON Schema Validation: Structured output validation
  • Keyword Presence: Required terms appear in response
  • LLM-as-Judge: Binary pass/fail using evaluation prompt

Example Use Cases:

  • Email classification (spam/not spam)
  • Entity extraction (dates, names, locations)
  • JSON output formatting validation
  • Sentiment analysis (positive/negative/neutral)

Quick Start (Python):

python
import pytest
from openai import OpenAI

client = OpenAI()

def classify_sentiment(text: str) -> str:
    response = client.chat.completions.create(
        model="gpt-3.5-turbo",
        messages=[
            {"role": "system", "content": "Classify sentiment as positive, negative, or neutral. Return only the label."},
            {"role": "user", "content": text}
        ],
        temperature=0
    )
    return response.choices[0].message.content.strip().lower()

def test_positive_sentiment():
    result = classify_sentiment("I love this product!")
    assert result == "positive"

For complete unit evaluation examples, see examples/python/unit_evaluation.py and examples/typescript/unit-evaluation.ts.

RAG (Retrieval-Augmented Generation) Evaluation

Evaluate RAG systems using RAGAS framework metrics.

Critical Metrics (Priority Order):

  1. Faithfulness (Target: > 0.8) - MOST CRITICAL

    • Measures: Is the answer grounded in retrieved context?
    • Prevents hallucinations
    • If failing: Adjust prompt to emphasize grounding, require citations
  2. Answer Relevance (Target: > 0.7)

    • Measures: How well does the answer address the query?
    • If failing: Improve prompt instructions, add few-shot examples
  3. Context Relevance (Target: > 0.7)

    • Measures: Are retrieved chunks relevant to the query?
    • If failing: Improve retrieval (better embeddings, hybrid search)
  4. Context Precision (Target: > 0.5)

    • Measures: Are relevant chunks ranked higher than irrelevant?
    • If failing: Add re-ranking step to retrieval pipeline
  5. Context Recall (Target: > 0.8)

    • Measures: Are all relevant chunks retrieved?
    • If failing: Increase retrieval count, improve chunking strategy

Quick Start (Python with RAGAS):

python
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy, context_relevancy
from datasets import Dataset

data = {
    "question": ["What is the capital of France?"],
    "answer": ["The capital of France is Paris."],
    "contexts": [["Paris is the capital of France."]],
    "ground_truth": ["Paris"]
}

dataset = Dataset.from_dict(data)
results = evaluate(dataset, metrics=[faithfulness, answer_relevancy, context_relevancy])
print(f"Faithfulness: {results['faithfulness']:.2f}")

For comprehensive RAG evaluation patterns, see references/rag-evaluation.md and examples/python/ragas_example.py.

LLM-as-Judge Evaluation

Use powerful LLMs (GPT-4, Claude Opus) to evaluate other LLM outputs.

When to Use:

  • Generation quality assessment (summaries, creative writing)
  • Nuanced evaluation criteria (tone, clarity, helpfulness)
  • Custom rubrics for domain-specific tasks
  • Medium-volume evaluation (100-1,000 samples)

Correlation with Human Judgment: 0.75-0.85 for well-designed rubrics

Best Practices:

  • Use clear, specific rubrics (1-5 scale with detailed criteria)
  • Include few-shot examples in evaluation prompt
  • Average multiple evaluations to reduce variance
  • Be aware of biases (position bias, verbosity bias, self-preference)

Quick Start (Python):

python
from openai import OpenAI

client = OpenAI()

def evaluate_quality(prompt: str, response: str) -> tuple[int, str]:
    """Returns (score 1-5, reasoning)"""
    eval_prompt = f"""
Rate the following LLM response on relevance and helpfulness.

USER PROMPT: {prompt}
LLM RESPONSE: {response}

Provide:
Score: [1-5, where 5 is best]
Reasoning: [1-2 sentences]
"""
    result = client.chat.completions.create(
        model="gpt-4",
        messages=[{"role": "user", "content": eval_prompt}],
        temperature=0.3
    )
    content = result.choices[0].message.content
    lines = content.strip().split('\n')
    score = int(lines[0].split(':')[1].strip())
    reasoning = lines[1].split(':', 1)[1].strip()
    return score, reasoning

For detailed LLM-as-judge patterns and prompt templates, see references/llm-as-judge.md and examples/python/llm_as_judge.py.

Safety and Alignment Evaluation

Measure hallucinations, bias, and toxicity in LLM outputs.

Hallucination Detection

Methods:

  1. Faithfulness to Context (RAG):

    • Use RAGAS faithfulness metric
    • LLM checks if claims are supported by context
    • Score: Supported claims / Total claims
  2. Factual Accuracy (Closed-Book):

    • LLM-as-judge with access to reliable sources
    • Fact-checking APIs (Google Fact Check)
    • Entity-level verification (dates, names, statistics)
  3. Self-Consistency:

    • Generate multiple responses to same question
    • Measure agreement between responses
    • Low consistency suggests hallucination
Bias Evaluation

Types of Bias:

  • Gender bias (stereotypical associations)
  • Racial/ethnic bias (discriminatory outputs)
  • Cultural bias (Western-centric assumptions)
  • Age/disability bias (ableist or ageist language)

Evaluation Methods:

  1. Stereotype Tests:

    • BBQ (Bias Benchmark for QA): 58,000 question-answer pairs
    • BOLD (Bias in Open-Ended Language Generation)
  2. Counterfactual Evaluation:

    • Generate responses with demographic swaps
    • Example: "Dr. Smith (he/she) recommended..." → compare outputs
    • Measure consistency across variations
Toxicity Detection

Tools:

  • Perspective API (Google): Toxicity, threat, insult scores
  • Detoxify (HuggingFace): Open-source toxicity classifier
  • OpenAI Moderation API: Hate, harassment, violence detection

For comprehensive safety evaluation patterns, see references/safety-evaluation.md.

Benchmark Testing

Assess model capabilities using standardized benchmarks.

Standard Benchmarks:

BenchmarkCoverageFormatDifficultyUse Case
MMLU57 subjects (STEM, humanities)Multiple choiceHigh school - professionalGeneral intelligence
HellaSwagSentence completionMultiple choiceCommon senseReasoning validation
GPQAPhD-level scienceMultiple choiceVery high (expert-level)Frontier model testing
HumanEval164 Python problemsCode generationMediumCode capability
MATH12,500 competition problemsMath solvingHigh school competitionsMath reasoning

Domain-Specific Benchmarks:

  • Medical: MedQA (USMLE), PubMedQA
  • Legal: LegalBench
  • Finance: FinQA, ConvFinQA

When to Use Benchmarks:

  • Comparing multiple models (GPT-4 vs Claude vs Llama)
  • Model selection for specific domains
  • Baseline capability assessment
  • Academic research and publication

Quick Start (lm-evaluation-harness):

bash
pip install lm-eval

# Evaluate GPT-4 on MMLU
lm_eval --model openai-chat --model_args model=gpt-4 --tasks mmlu --num_fewshot 5

For detailed benchmark testing patterns, see references/benchmarks.md and scripts/benchmark_runner.py.

Production Evaluation

Monitor and optimize LLM quality in production environments.

A/B Testing

Compare two LLM configurations:

  • Variant A: GPT-4 (expensive, high quality)
  • Variant B: Claude Sonnet (cheaper, fast)

Metrics:

  • User satisfaction scores (thumbs up/down)
  • Task completion rates
  • Response time and latency
  • Cost per successful interaction
Online Evaluation

Real-time quality monitoring:

  • Response Quality: LLM-as-judge scoring every Nth response
  • User Feedback: Explicit ratings, thumbs up/down
  • Business Metrics: Conversion rates, support ticket resolution
  • Cost Tracking: Tokens used, inference costs
Human-in-the-Loop

Sample-based human evaluation:

  • Random Sampling: Evaluate 10% of responses
  • Confidence-Based: Evaluate low-confidence outputs
  • Error-Triggered: Flag suspicious responses for review

For production evaluation patterns and monitoring strategies, see references/production-evaluation.md.

Show full SKILL.md (652 more words)Show less

Classification Task Evaluation

For tasks with discrete outputs (sentiment, intent, category).

Metrics:

  • Accuracy: Correct predictions / Total predictions
  • Precision: True positives / (True positives + False positives)
  • Recall: True positives / (True positives + False negatives)
  • F1 Score: Harmonic mean of precision and recall
  • Confusion Matrix: Detailed breakdown of prediction errors

Quick Start (Python):

python
from sklearn.metrics import accuracy_score, precision_recall_fscore_support

y_true = ["positive", "negative", "neutral", "positive", "negative"]
y_pred = ["positive", "negative", "neutral", "neutral", "negative"]

accuracy = accuracy_score(y_true, y_pred)
precision, recall, f1, _ = precision_recall_fscore_support(y_true, y_pred, average='weighted')

print(f"Accuracy: {accuracy:.2f}")
print(f"Precision: {precision:.2f}")
print(f"Recall: {recall:.2f}")
print(f"F1 Score: {f1:.2f}")

For complete classification evaluation examples, see examples/python/classification_metrics.py.

Generation Task Evaluation

For open-ended text generation (summaries, creative writing, responses).

Automated Metrics (Use with Caution):

  • BLEU: N-gram overlap with reference text (0-1 score)
  • ROUGE: Recall-oriented overlap (ROUGE-1, ROUGE-L)
  • METEOR: Semantic similarity with stemming
  • BERTScore: Contextual embedding similarity (0-1 score)

Limitation: Automated metrics correlate weakly with human judgment for creative/subjective generation.

Recommended Approach:

  1. Automated metrics: Fast feedback for objective aspects (length, format)
  2. LLM-as-judge: Nuanced quality assessment (relevance, coherence, helpfulness)
  3. Human evaluation: Final validation for subjective criteria (preference, creativity)

For detailed generation evaluation patterns, see references/evaluation-types.md.

Quick Reference Tables

Evaluation Framework Selection
If Task Is...Use This FrameworkPrimary Metric
RAG systemRAGASFaithfulness > 0.8
Classificationscikit-learn metricsAccuracy, F1
Generation qualityLLM-as-judgeQuality rubric (1-5)
Code generationHumanEvalPass@1, Test pass rate
Model comparisonBenchmark testingMMLU, HellaSwag scores
Safety validationHallucination detectionFaithfulness, Fact-check
Production monitoringOnline evaluationUser feedback, Business KPIs
Python Library Recommendations
LibraryUse CaseInstallation
RAGASRAG evaluationpip install ragas
DeepEvalGeneral LLM evaluation, pytest integrationpip install deepeval
LangSmithProduction monitoring, A/B testingpip install langsmith
lm-evalBenchmark testing (MMLU, HumanEval)pip install lm-eval
scikit-learnClassification metricspip install scikit-learn
Safety Evaluation Priority Matrix
ApplicationHallucination RiskBias RiskToxicity RiskEvaluation Priority
Customer SupportHighMediumHigh1. Faithfulness, 2. Toxicity, 3. Bias
Medical DiagnosisCriticalHighLow1. Factual Accuracy, 2. Hallucination, 3. Bias
Creative WritingLowMediumMedium1. Quality/Fluency, 2. Content Policy
Code GenerationMediumLowLow1. Functional Correctness, 2. Security
Content ModerationLowCriticalCritical1. Bias, 2. False Positives/Negatives

Detailed References

For comprehensive documentation on specific topics:

  • Evaluation types (classification, generation, QA, code): references/evaluation-types.md
  • RAG evaluation deep dive (RAGAS framework): references/rag-evaluation.md
  • Safety evaluation (hallucination, bias, toxicity): references/safety-evaluation.md
  • Benchmark testing (MMLU, HumanEval, domain benchmarks): references/benchmarks.md
  • LLM-as-judge best practices and prompts: references/llm-as-judge.md
  • Production evaluation (A/B testing, monitoring): references/production-evaluation.md
  • All metrics definitions and formulas: references/metrics-reference.md

Working Examples

Python Examples:

  • examples/python/unit_evaluation.py - Basic prompt testing with pytest
  • examples/python/ragas_example.py - RAGAS RAG evaluation
  • examples/python/deepeval_example.py - DeepEval framework usage
  • examples/python/llm_as_judge.py - GPT-4 as evaluator
  • examples/python/classification_metrics.py - Accuracy, precision, recall
  • examples/python/benchmark_testing.py - HumanEval example

TypeScript Examples:

  • examples/typescript/unit-evaluation.ts - Vitest + OpenAI
  • examples/typescript/llm-as-judge.ts - GPT-4 evaluation
  • examples/typescript/langsmith-integration.ts - Production monitoring

Executable Scripts

Run evaluations without loading code into context (token-free):

  • scripts/run_ragas_eval.py - Run RAGAS evaluation on dataset
  • scripts/compare_models.py - A/B test two models
  • scripts/benchmark_runner.py - Run MMLU/HumanEval benchmarks
  • scripts/hallucination_checker.py - Detect hallucinations in outputs

Example usage:

bash
# Run RAGAS evaluation on custom dataset
python scripts/run_ragas_eval.py --dataset data/qa_dataset.json --output results.json

# Compare GPT-4 vs Claude on benchmark
python scripts/compare_models.py --model-a gpt-4 --model-b claude-3-opus --tasks mmlu,humaneval

Integration with Other Skills

Related Skills:

  • building-ai-chat: Evaluate AI chat applications (this skill tests what that skill builds)
  • prompt-engineering: Test prompt quality and effectiveness
  • testing-strategies: Apply testing pyramid to LLM evaluation (unit → integration → E2E)
  • observability: Production monitoring and alerting for LLM quality
  • building-ci-pipelines: Integrate LLM evaluation into CI/CD

Workflow Integration:

  1. Write prompt (use prompt-engineering skill)
  2. Unit test prompt (use llm-evaluation skill)
  3. Build AI feature (use building-ai-chat skill)
  4. Integration test RAG pipeline (use llm-evaluation skill)
  5. Deploy to production (use deploying-applications skill)
  6. Monitor quality (use llm-evaluation + observability skills)

Common Pitfalls

1. Over-reliance on Automated Metrics for Generation

  • BLEU/ROUGE correlate weakly with human judgment for creative text
  • Solution: Layer LLM-as-judge or human evaluation

2. Ignoring Faithfulness in RAG Systems

  • Hallucinations are the #1 RAG failure mode
  • Solution: Prioritize faithfulness metric (target > 0.8)

3. No Production Monitoring

  • Models can degrade over time, prompts can break with updates
  • Solution: Set up continuous evaluation (LangSmith, custom monitoring)

4. Biased LLM-as-Judge Evaluation

  • Evaluator LLMs have biases (position bias, verbosity bias)
  • Solution: Average multiple evaluations, use diverse evaluation prompts

5. Insufficient Benchmark Coverage

  • Single benchmark doesn't capture full model capability
  • Solution: Use 3-5 benchmarks across different domains

6. Missing Safety Evaluation

  • Production LLMs can generate harmful content
  • Solution: Add toxicity, bias, and hallucination checks to evaluation pipeline

© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (references) in skills/evaluating-llms of ancoleman/ai-design-components.

  • SKILL.md
  • examples/python/benchmark_testing.py
  • examples/python/classification_metrics.py
  • examples/python/deepeval_example.py
  • examples/python/llm_as_judge.py
  • examples/python/ragas_example.py
  • examples/python/unit_evaluation.py
  • examples/typescript/langsmith-integration.ts
  • examples/typescript/llm-as-judge.ts
  • examples/typescript/unit-evaluation.ts
  • outputs.yaml
  • references/benchmarks.md
  • references/evaluation-types.md
  • references/llm-as-judge.md
  • references/metrics-reference.md
  • references/production-evaluation.md
  • references/rag-evaluation.md
  • … and 1 more

Open the folder on GitHubat commit 76551b7

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ancoleman/ai-design-components, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Evaluating LLMs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluating LLMs compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluating LLMs this skillancoleman/ai-design-components5261 repos~4.7kAutomated safety check: PassMIT
Evaluate RAGai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0
RAG ArchitectJeffallan/claude-skills12k1 repos~2kAutomated safety check: PassMIT
Jd Gap Analysisstarkyru/learn-ai105—~1.9kAutomated safety check: PassMIT
Agent Evalericrisco/rsc-harness156—~3.2kAutomated safety check: PassMIT
RAG Observability Evalssickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMIT

Similar skills

  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • RAG Architect

    Jeffallan/claude-skills

    Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

    12k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Jd Gap Analysis

    starkyru/learn-ai

    Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.

    105 GitHub stars~1.9k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • RAG Observability Evals

    sickn33/agentic-awesome-skills

    Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Evaluation Harness

    davepoon/buildwithclaude

    Evaluate retrieval and citation behavior for RAG pipelines from deterministic JSONL fixtures.

    3.6k GitHub stars~813 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from ancoleman/ai-design-components

All 75 skills in this repo
  • Building AI Chat

    ancoleman/ai-design-components

    Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.

    526 GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed
  • Building Forms

    ancoleman/ai-design-components

    Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.

    526 GitHub stars~3.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Building Tables

    ancoleman/ai-design-components

    Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.

    526 GitHub stars~1.8k tokensUpdated 10 mo ago
    Auto-check passed
  • Creating Dashboards

    ancoleman/ai-design-components

    Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.

    526 GitHub stars~3.5k tokensUpdated 10 mo ago
    Auto-check passed
  • Designing Layouts

    ancoleman/ai-design-components

    Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.

    526 GitHub stars~1.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Displaying Timelines

    ancoleman/ai-design-components

    Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.

    526 GitHub stars~2.7k tokensUpdated 10 mo ago
    Auto-check passed

Questions about Evaluating LLMs

What does Evaluating LLMs do?

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks. Evaluating LLMs is an agent skill from ancoleman/ai-design-components. Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks.

When should I use Evaluating LLMs?

Evaluating LLMs fits situations like: testing prompt quality; validating RAG pipelines; measuring safety (hallucinations; comparing models for production deployment.

How do I install Evaluating LLMs in Claude Code?

Run `npx skills add ancoleman/ai-design-components --skill evaluating-llms -a claude-code`. Or copy the skill folder (skills/evaluating-llms in ancoleman/ai-design-components) into .claude/skills/evaluating-llms in your project. Claude Code loads it when a task matches its description.

How do I install Evaluating LLMs in Codex?

Run `npx skills add ancoleman/ai-design-components --skill evaluating-llms -a codex`. Or copy the skill folder (skills/evaluating-llms in ancoleman/ai-design-components) into .agents/skills/evaluating-llms in your project. Codex loads it when a task matches its description.

Can I use Evaluating LLMs in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill evaluating-llms -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluating-llms, .gemini/skills/evaluating-llms, .github/skills/evaluating-llms and .opencode/skills/evaluating-llms in your project.

What does Evaluating LLMs need to run?

Going by SKILL.md and its folder, Evaluating LLMs needs Python and TypeScript for the scripts in its folder and the command-line tools its instructions call (pip and python). Our summary lists: Python 3; Node.js.

Does Evaluating LLMs access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evaluating LLMs safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluating LLMs use?

Evaluating LLMs is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluating LLMs use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 29k tokens, read only when the agent opens those files.

What are the alternatives to Evaluating LLMs?

Skills that share tags, products or a category with Evaluating LLMs: Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars), RAG Architect (Jeffallan/claude-skills, 12k stars), Jd Gap Analysis (starkyru/learn-ai, 105 stars) and Agent Eval (ericrisco/rsc-harness, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluating LLMs?

ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.

Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.