Agent skill

Testing LLM

by yonatangross in yonatangross/orchestkit

LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner).

MITAuto-check passedAI & LLM Engineering

Install Testing LLM

skills CLI
$ npx skills add yonatangross/orchestkit --skill testing-llm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit testing-llm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/testing-llm .claude/skills/testing-llm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-llm
GitHub stars
289
Token cost
~2.6k tokens
SKILL.md length
809 words
Files
9 (incl. references)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner).

  • Works in 3 steps: Planner: Explores your app and produces… → Generator: Converts Markdown specs into… → Healer (references/healer-agent.md):…
  • Testing AI features
  • SKILL.md covers Quick Reference, Upstream coverage (do not…, When to Use This Skill and LLM Mock Quick Start, plus 9 more sections
  • Calls npx

What it does

Testing LLM is an agent skill from yonatangross/orchestkit. LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner). Use when testing AI features, validating LLM outputs, or building evaluation pipelines.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `checklists/llm-test-checklist.md`, `references/healer-agent.md` and `references/langfuse-v4.md`). Compatibility notes: Claude Code 2.1.277+.

It sits in AI & LLM Engineering, covering Test strategy and Structured output and tool calling. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • Testing AI features
  • Validating LLM outputs
  • Building evaluation pipelines

Example prompts

  • “/testing-llm”

Requirements

  • Python 3
  • Node.js
  • Compatibility (from SKILL.md): Claude Code 2.1.277+.
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, WebFetch, WebSearch

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Planner: Explores your app and produces Markdown test plans. Owned by the
  2. Generator: Converts Markdown specs into Playwright tests, validating selectors
  3. Healer (references/healer-agent.md): Automatically fixes failing tests by replaying failures, inspecting the DOM, and patching…

What it can do on your machine

Read from SKILL.md and the folder at commit 0ef71d2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • WebFetch
    • WebSearch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • deepeval.com
    • playwright.dev
    • docs.ragas.io
    • vcrpy.readthedocs.io
    • docs.scipy.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+.

    From compatibility in the SKILL.md frontmatter.

Context cost

Testing LLM loads about 2.6k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 66 tokens; SKILL.md has 809 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit 0ef71d2, republished under its MIT licence (© yonatangross). 809 words, ~2,572 tokens.

Download SKILL.mdSave it as .claude/skills/testing-llm/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
testing-llm
description
LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner). Use when testing AI features, validating LLM outputs, or building evaluation pipelines.
allowed-tools
Read, Glob, Grep, WebFetch, WebSearch
compatibility
Claude Code 2.1.277+.
license
MIT
user-invocable
false
disable-model-invocation
false
metadata.owner-agent
test-generator
metadata.category
document-asset-creation
metadata.version
2.1.0
metadata.author
OrchestKit
metadata.complexity
medium
metadata.tags
testing, llm, ai, deepeval, ragas, evaluation, mocking

LLM & AI Testing Patterns

Patterns and tools for testing LLM integrations, evaluating AI output quality, mocking responses for deterministic CI, and applying agentic test workflows (planner, generator, healer). Of that trio only the healer keeps a local reference here; the planner and generator stages belong to the testing-e2e skill.

Quick Reference

AreaFilePurpose
Rulesrules/llm-evaluation.mdDeepEval quality metrics, Pydantic schema validation, timeout testing
Rulesrules/llm-mocking.mdMock LLM responses, VCR.py recording, custom request matchers
Referencereferences/ork-delta.mdHouse rules the vendor docs do not carry: GEval and RAGAS API corrections, threshold direction, cassette path, golden-dataset and latency budgets
Referencereferences/healer-agent.mdAuto-fixes failing tests (selectors, waits, dynamic content)
Referencereferences/langfuse-v4.mdLangfuse Python SDK v4 tracing and dataset runs, plus the non-zero throughput assertion
Checklistchecklists/llm-test-checklist.mdComplete LLM testing checklist (setup, coverage, CI/CD)

Upstream coverage (do not restate)

DeepEval, RAGAS, VCR.py and Playwright document themselves. This skill carries only the OrchestKit delta (references/ork-delta.md) plus the house subsets in rules/ and checklists/. Fetch the source below instead of expecting the material here.

TopicSource
Full DeepEval metric catalog and per-metric constructor arguments (the house threshold table and the two-metric quick start stay in this file, rules/llm-evaluation.md and checklists/llm-test-checklist.md)https://deepeval.com/docs/metrics-introduction
GEval custom criteria: evaluation_params, evaluation_steps, criteria (the house import correction stays in references/ork-delta.md)https://deepeval.com/docs/metrics-llm-evals
HallucinationMetric arguments (the house 0.3 ceiling and the inverted-direction warning stay in references/ork-delta.md)https://deepeval.com/docs/metrics-hallucination
RAGAS metric catalog (Faithfulness, LLMContextRecall, FactualCorrectness)https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/
EvaluationDataset construction (the house note on the post-0.2 field names stays in references/ork-delta.md)https://docs.ragas.io/en/stable/concepts/components/eval_dataset/
VCR.py configuration keys (the house record-mode gate and header filters stay in rules/llm-mocking.md)https://vcrpy.readthedocs.io/en/latest/configuration.html
Playwright Planner and Generator agents, init-agents CLI and generated files (the house healer subset stays in references/healer-agent.md)https://playwright.dev/docs/test-agents
Playwright semantic locator ladder used by generated teststesting-e2e skill (rules/e2e-playwright.md) plus https://playwright.dev/docs/locators
Confidence intervals over metric score sampleshttps://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.t.html

When to Use This Skill

  • Testing code that calls LLM APIs (OpenAI, Anthropic, etc.)
  • Validating RAG pipeline output quality
  • Setting up deterministic LLM tests in CI
  • Building evaluation pipelines with quality gates
  • Applying agentic test patterns (plan -> generate -> heal)

LLM Mock Quick Start

Mock LLM responses for fast, deterministic unit tests:

python
from unittest.mock import AsyncMock, patch
import pytest

@pytest.fixture
def mock_llm():
    mock = AsyncMock()
    mock.return_value = {"content": "Mocked response", "confidence": 0.85}
    return mock

@pytest.mark.asyncio
async def test_with_mocked_llm(mock_llm):
    with patch("app.core.model_factory.get_model", return_value=mock_llm):
        result = await synthesize_findings(sample_findings)
    assert result["summary"] is not None

Key rule: NEVER call live LLM APIs in CI. Use mocks for unit tests, VCR.py for integration tests.

DeepEval Quality Quick Start

Validate LLM output quality with multi-dimensional metrics:

python
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric, FaithfulnessMetric

test_case = LLMTestCase(
    input="What is the capital of France?",
    actual_output="The capital of France is Paris.",
    retrieval_context=["Paris is the capital of France."],
)

assert_test(test_case, [
    AnswerRelevancyMetric(threshold=0.7),
    FaithfulnessMetric(threshold=0.8),
])

Library notes (DeepEval, RAGAS)

DeepEval metrics expose a reason field alongside the numeric score when include_reason=True, so a failing CI build gets a human-readable explanation without a second LLM call:

python
metric = AnswerRelevancyMetric(threshold=0.7, include_reason=True)
metric.measure(test_case)
print(metric.score, metric.reason)
# 0.62  "Response addresses the topic but omits the date asked for."

RAGAS uses a class-based metric API — instantiate metric classes and pass an EvaluationDataset. llm= is optional; omit it to use the configured default grader:

python
from ragas import evaluate
from ragas.metrics import Faithfulness, LLMContextRecall

result = evaluate(
    dataset,
    metrics=[Faithfulness(), LLMContextRecall()],
)

Bump floors: deepeval >= 4.0, ragas >= 0.4.

House rules the vendor docs do not state (the inverted HallucinationMetric threshold, the gpt-5-mini grader default, the 95 percent confidence-interval recipe, and the latency, quality-gate and truncation numbers) are recorded in references/ork-delta.md. Read that before writing either library's setup code.

Show full SKILL.md (333 more words)Show less

Quality Metrics Thresholds

MetricThresholdPurpose
Answer Relevancy>= 0.7Response addresses question
Faithfulness>= 0.8Output matches context
Hallucination<= 0.3No fabricated facts
Context Precision>= 0.7Retrieved contexts relevant
Context Recall>= 0.7All relevant contexts retrieved

Structured Output Validation

Always validate LLM output with Pydantic schemas:

python
from pydantic import BaseModel, Field

class LLMResponse(BaseModel):
    answer: str = Field(min_length=1)
    confidence: float = Field(ge=0.0, le=1.0)
    sources: list[str] = Field(default_factory=list)

async def test_structured_output():
    result = await get_llm_response("test query")
    parsed = LLMResponse.model_validate(result)
    assert 0 <= parsed.confidence <= 1.0

VCR.py for Integration Tests

Record and replay LLM API calls for deterministic integration tests:

python
@pytest.fixture(scope="module")
def vcr_config():
    import os
    return {
        "record_mode": "none" if os.environ.get("CI") else "new_episodes",
        "filter_headers": ["authorization", "x-api-key"],
    }

@pytest.mark.vcr()
async def test_llm_integration():
    response = await llm_client.complete("Say hello")
    assert "hello" in response.content.lower()

Agentic Test Workflow

The three-agent pattern for end-to-end test automation:

Planner -> specs/*.md -> Generator -> tests/*.spec.ts -> Healer (auto-fix)
  1. Planner: Explores your app and produces Markdown test plans. Owned by the testing-e2e skill (rules/e2e-ai-agents.md); the CLI and its generated files are documented at https://playwright.dev/docs/test-agents.

  2. Generator: Converts Markdown specs into Playwright tests, validating selectors against the running app. Also owned by testing-e2e (rules/e2e-ai-agents.md); the locator ladder it follows lives in testing-e2e rules/e2e-playwright.md.

  3. Healer (references/healer-agent.md): Automatically fixes failing tests by replaying failures, inspecting the DOM, and patching locators/waits. Max 3 healing attempts per test.

Agent initialization is CLI-only (npx playwright init-agents); there is no config key for it. Only the healing stage keeps a local reference, because its 3-attempt ceiling and its refusal to touch test logic are house limits rather than vendor defaults.

Edge Cases to Always Test

For every LLM integration, cover these paths:

  • Empty/null inputs -- empty strings, None values
  • Long inputs -- truncation behavior near token limits
  • Timeouts -- fail-open vs fail-closed behavior
  • Schema violations -- invalid structured output
  • Prompt injection -- adversarial input resistance
  • Unicode -- non-ASCII characters in prompts and responses

See checklists/llm-test-checklist.md for the complete checklist.

Anti-Patterns

Anti-PatternCorrect Approach
Live LLM calls in CIMock for unit, VCR for integration
Random seedsFixed seeds or mocked responses
Single metric evaluation3-5 quality dimensions
No timeout handlingAlways set < 1s timeout in tests
Hardcoded API keysEnvironment variables, filtered in VCR
Asserting only is not NoneSchema validation + quality metrics
  • ork:testing-unit — Unit testing fundamentals, AAA pattern
  • ork:testing-integration — Integration testing for AI pipelines
  • ork:golden-dataset — Evaluation dataset management
  • ork:testing-e2e owns the Planner and Generator agent workflow and the Playwright locator ladder
  • ork:testing-perf owns the latency and load budgets referenced in references/ork-delta.md

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in src/skills/testing-llm of yonatangross/orchestkit.

  • SKILL.md
  • checklists/llm-test-checklist.md
  • references/healer-agent.md
  • references/langfuse-v4.md
  • references/ork-delta.md
  • rules/_sections.md
  • rules/llm-evaluation.md
  • rules/llm-mocking.md
  • test-cases.json

Open the folder on GitHubat commit 0ef71d2

Compare with similar skills

Testing LLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing LLM compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing LLM this skillyonatangross/orchestkit289—~2.6kAutomated safety check: PassMIT
Cc Skill Project Guidelines Exampledavila7/claude-code-templates32k6 repos~2.2kAutomated safety check: NotesMIT
Project Guidelines Examplevibeeval/vibecosystem5313 repos~2.2kAutomated safety check: NotesMIT
Planning With Filesjarrodwatts/claude-code-config1.1k5 repos~967Automated safety check: PassNone
Tool Use Data Synthesissunny-glow/Auto-BenchMax1.3k—~3.3kAutomated safety check: PassNone
Agent Harness ConstructionKartikLabhshetwar/mind-mentor1477 repos~500Automated safety check: PassApache-2.0

Similar skills

  • Cc Skill Project Guidelines Example

    davila7/claude-code-templates

    Project Guidelines Skill (Example)

    32k GitHub starsUsed in 6 repos~2.2k tokens
    AI & LLM EngineeringAuto-check: notes
  • Project Guidelines Example

    vibeeval/vibecosystem

    Example template for project-specific skill files covering architecture, patterns, testing, and deployment.

    531 GitHub starsUsed in 3 repos~2.2k tokens
    AI & LLM EngineeringAuto-check: notes
  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed
  • Tool Use Data Synthesis

    sunny-glow/Auto-BenchMax

    Synthesize training data for ANY tool-use / agentic benchmark, in ANY repo.

    1.3k GitHub stars~3.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agent Harness Construction

    KartikLabhshetwar/mind-mentor

    Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates.

    147 GitHub starsUsed in 7 repos~500 tokens
    AI & LLM EngineeringAuto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from yonatangross/orchestkit

All 108 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    289 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    289 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    289 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    289 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    289 GitHub stars~4.5k tokensUpdated today
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    289 GitHub stars~3.9k tokensUpdated today
    Auto-check: notes

Questions about Testing LLM

What does Testing LLM do?

LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner). Testing LLM is an agent skill from yonatangross/orchestkit. LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner).

When should I use Testing LLM?

Testing LLM fits situations like: testing AI features; validating LLM outputs; building evaluation pipelines.

How do I install Testing LLM in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill testing-llm -a claude-code`. Or copy the skill folder (src/skills/testing-llm in yonatangross/orchestkit) into .claude/skills/testing-llm in your project. Claude Code loads it when a task matches its description.

How do I install Testing LLM in Codex?

Run `npx skills add yonatangross/orchestkit --skill testing-llm -a codex`. Or copy the skill folder (src/skills/testing-llm in yonatangross/orchestkit) into .agents/skills/testing-llm in your project. Codex loads it when a task matches its description.

Can I use Testing LLM in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill testing-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-llm, .gemini/skills/testing-llm, .github/skills/testing-llm and .opencode/skills/testing-llm in your project.

What does Testing LLM need to run?

Going by SKILL.md and its folder, Testing LLM needs the command-line tools its instructions call (npx). Our summary lists: Python 3; Node.js. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+..

Does Testing LLM access the network?

SKILL.md names 5 domains. As links in the text: deepeval.com, playwright.dev, docs.ragas.io, vcrpy.readthedocs.io and docs.scipy.org. This is read from the text; nothing was executed.

Is Testing LLM safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing LLM use?

Testing LLM is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing LLM use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Testing LLM?

Skills that share tags, products or a category with Testing LLM: Cc Skill Project Guidelines Example (davila7/claude-code-templates, 32k stars), Project Guidelines Example (vibeeval/vibecosystem, 531 stars), Planning With Files (jarrodwatts/claude-code-config, 1.1k stars) and Tool Use Data Synthesis (sunny-glow/Auto-BenchMax, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing LLM?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 289 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.