Cc Skill Project Guidelines Example
davila7/claude-code-templates
Project Guidelines Skill (Example)
LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner).
$ npx skills add yonatangross/orchestkit --skill testing-llm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install yonatangross/orchestkit testing-llm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/testing-llm .claude/skills/testing-llm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "testing-llm" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/testing-llm into .claude/skills/testing-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-llm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/yonatangross/orchestkit/tree/main/src/skills/testing-llmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add yonatangross/orchestkit --skill testing-llm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install yonatangross/orchestkit testing-llm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/skills/testing-llm .agents/skills/testing-llm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "testing-llm" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/testing-llm into .agents/skills/testing-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-llm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill testing-llm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install yonatangross/orchestkit testing-llm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/skills/testing-llm .cursor/skills/testing-llm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "testing-llm" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/testing-llm into .cursor/skills/testing-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-llm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/yonatangross/orchestkit.git --path src/skills/testing-llm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add yonatangross/orchestkit --skill testing-llm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install yonatangross/orchestkit testing-llm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/skills/testing-llm .gemini/skills/testing-llm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "testing-llm" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/testing-llm into .gemini/skills/testing-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-llm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install yonatangross/orchestkit testing-llmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add yonatangross/orchestkit --skill testing-llm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/skills/testing-llm .github/skills/testing-llm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "testing-llm" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/testing-llm into .github/skills/testing-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-llm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill testing-llm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install yonatangross/orchestkit testing-llm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/skills/testing-llm .opencode/skills/testing-llm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "testing-llm" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/testing-llm into .opencode/skills/testing-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "testing-llm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
testing-llmLLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner).
Testing LLM is an agent skill from yonatangross/orchestkit. LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner). Use when testing AI features, validating LLM outputs, or building evaluation pipelines.
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `checklists/llm-test-checklist.md`, `references/healer-agent.md` and `references/langfuse-v4.md`). Compatibility notes: Claude Code 2.1.277+.
It sits in AI & LLM Engineering, covering Test strategy and Structured output and tool calling. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 0ef71d2. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGlobGrepWebFetchWebSearchFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
deepeval.complaywright.devdocs.ragas.iovcrpy.readthedocs.iodocs.scipy.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Claude Code 2.1.277+.
From compatibility in the SKILL.md frontmatter.
Testing LLM loads about 2.6k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 66 tokens; SKILL.md has 809 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from yonatangross/orchestkit at commit 0ef71d2, republished under its MIT licence (© yonatangross). 809 words, ~2,572 tokens.
.claude/skills/testing-llm/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Patterns and tools for testing LLM integrations, evaluating AI output quality, mocking responses for deterministic CI, and applying agentic test workflows (planner, generator, healer). Of that trio only the healer keeps a local reference here; the planner and generator stages belong to the testing-e2e skill.
| Area | File | Purpose |
|---|---|---|
| Rules | rules/llm-evaluation.md | DeepEval quality metrics, Pydantic schema validation, timeout testing |
| Rules | rules/llm-mocking.md | Mock LLM responses, VCR.py recording, custom request matchers |
| Reference | references/ork-delta.md | House rules the vendor docs do not carry: GEval and RAGAS API corrections, threshold direction, cassette path, golden-dataset and latency budgets |
| Reference | references/healer-agent.md | Auto-fixes failing tests (selectors, waits, dynamic content) |
| Reference | references/langfuse-v4.md | Langfuse Python SDK v4 tracing and dataset runs, plus the non-zero throughput assertion |
| Checklist | checklists/llm-test-checklist.md | Complete LLM testing checklist (setup, coverage, CI/CD) |
DeepEval, RAGAS, VCR.py and Playwright document themselves. This skill carries only the
OrchestKit delta (references/ork-delta.md) plus the house subsets in rules/ and
checklists/. Fetch the source below instead of expecting the material here.
| Topic | Source |
|---|---|
Full DeepEval metric catalog and per-metric constructor arguments (the house threshold table and the two-metric quick start stay in this file, rules/llm-evaluation.md and checklists/llm-test-checklist.md) | https://deepeval.com/docs/metrics-introduction |
GEval custom criteria: evaluation_params, evaluation_steps, criteria (the house import correction stays in references/ork-delta.md) | https://deepeval.com/docs/metrics-llm-evals |
HallucinationMetric arguments (the house 0.3 ceiling and the inverted-direction warning stay in references/ork-delta.md) | https://deepeval.com/docs/metrics-hallucination |
RAGAS metric catalog (Faithfulness, LLMContextRecall, FactualCorrectness) | https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/ |
EvaluationDataset construction (the house note on the post-0.2 field names stays in references/ork-delta.md) | https://docs.ragas.io/en/stable/concepts/components/eval_dataset/ |
VCR.py configuration keys (the house record-mode gate and header filters stay in rules/llm-mocking.md) | https://vcrpy.readthedocs.io/en/latest/configuration.html |
Playwright Planner and Generator agents, init-agents CLI and generated files (the house healer subset stays in references/healer-agent.md) | https://playwright.dev/docs/test-agents |
| Playwright semantic locator ladder used by generated tests | testing-e2e skill (rules/e2e-playwright.md) plus https://playwright.dev/docs/locators |
| Confidence intervals over metric score samples | https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.t.html |
Mock LLM responses for fast, deterministic unit tests:
from unittest.mock import AsyncMock, patch
import pytest
@pytest.fixture
def mock_llm():
mock = AsyncMock()
mock.return_value = {"content": "Mocked response", "confidence": 0.85}
return mock
@pytest.mark.asyncio
async def test_with_mocked_llm(mock_llm):
with patch("app.core.model_factory.get_model", return_value=mock_llm):
result = await synthesize_findings(sample_findings)
assert result["summary"] is not NoneKey rule: NEVER call live LLM APIs in CI. Use mocks for unit tests, VCR.py for integration tests.
Validate LLM output quality with multi-dimensional metrics:
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric, FaithfulnessMetric
test_case = LLMTestCase(
input="What is the capital of France?",
actual_output="The capital of France is Paris.",
retrieval_context=["Paris is the capital of France."],
)
assert_test(test_case, [
AnswerRelevancyMetric(threshold=0.7),
FaithfulnessMetric(threshold=0.8),
])DeepEval metrics expose a reason field alongside the numeric score when include_reason=True, so a failing CI build gets a human-readable explanation without a second LLM call:
metric = AnswerRelevancyMetric(threshold=0.7, include_reason=True)
metric.measure(test_case)
print(metric.score, metric.reason)
# 0.62 "Response addresses the topic but omits the date asked for."RAGAS uses a class-based metric API — instantiate metric classes and pass an EvaluationDataset. llm= is optional; omit it to use the configured default grader:
from ragas import evaluate
from ragas.metrics import Faithfulness, LLMContextRecall
result = evaluate(
dataset,
metrics=[Faithfulness(), LLMContextRecall()],
)Bump floors:
deepeval >= 4.0,ragas >= 0.4.
House rules the vendor docs do not state (the inverted HallucinationMetric threshold,
the gpt-5-mini grader default, the 95 percent confidence-interval recipe, and the
latency, quality-gate and truncation numbers) are recorded in references/ork-delta.md.
Read that before writing either library's setup code.
| Metric | Threshold | Purpose |
|---|---|---|
| Answer Relevancy | >= 0.7 | Response addresses question |
| Faithfulness | >= 0.8 | Output matches context |
| Hallucination | <= 0.3 | No fabricated facts |
| Context Precision | >= 0.7 | Retrieved contexts relevant |
| Context Recall | >= 0.7 | All relevant contexts retrieved |
Always validate LLM output with Pydantic schemas:
from pydantic import BaseModel, Field
class LLMResponse(BaseModel):
answer: str = Field(min_length=1)
confidence: float = Field(ge=0.0, le=1.0)
sources: list[str] = Field(default_factory=list)
async def test_structured_output():
result = await get_llm_response("test query")
parsed = LLMResponse.model_validate(result)
assert 0 <= parsed.confidence <= 1.0Record and replay LLM API calls for deterministic integration tests:
@pytest.fixture(scope="module")
def vcr_config():
import os
return {
"record_mode": "none" if os.environ.get("CI") else "new_episodes",
"filter_headers": ["authorization", "x-api-key"],
}
@pytest.mark.vcr()
async def test_llm_integration():
response = await llm_client.complete("Say hello")
assert "hello" in response.content.lower()The three-agent pattern for end-to-end test automation:
Planner -> specs/*.md -> Generator -> tests/*.spec.ts -> Healer (auto-fix)Planner: Explores your app and produces Markdown test plans. Owned by the
testing-e2e skill (rules/e2e-ai-agents.md); the CLI and its generated files are
documented at https://playwright.dev/docs/test-agents.
Generator: Converts Markdown specs into Playwright tests, validating selectors
against the running app. Also owned by testing-e2e (rules/e2e-ai-agents.md); the
locator ladder it follows lives in testing-e2e rules/e2e-playwright.md.
Healer (references/healer-agent.md): Automatically fixes failing tests by replaying failures, inspecting the DOM, and patching locators/waits. Max 3 healing attempts per test.
Agent initialization is CLI-only (npx playwright init-agents); there is no config key
for it. Only the healing stage keeps a local reference, because its 3-attempt ceiling and
its refusal to touch test logic are house limits rather than vendor defaults.
For every LLM integration, cover these paths:
See checklists/llm-test-checklist.md for the complete checklist.
| Anti-Pattern | Correct Approach |
|---|---|
| Live LLM calls in CI | Mock for unit, VCR for integration |
| Random seeds | Fixed seeds or mocked responses |
| Single metric evaluation | 3-5 quality dimensions |
| No timeout handling | Always set < 1s timeout in tests |
| Hardcoded API keys | Environment variables, filtered in VCR |
Asserting only is not None | Schema validation + quality metrics |
ork:testing-unit — Unit testing fundamentals, AAA patternork:testing-integration — Integration testing for AI pipelinesork:golden-dataset — Evaluation dataset managementork:testing-e2e owns the Planner and Generator agent workflow and the Playwright locator ladderork:testing-perf owns the latency and load budgets referenced in references/ork-delta.md© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (references) in src/skills/testing-llm of yonatangross/orchestkit.
Open the folder on GitHubat commit 0ef71d2
Testing LLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Testing LLM this skillyonatangross/orchestkit | 289 | — | ~2.6k | Automated safety check: Pass | MIT | |
| Cc Skill Project Guidelines Exampledavila7/claude-code-templates | 32k | 6 repos | ~2.2k | Automated safety check: Notes | MIT | |
| Project Guidelines Examplevibeeval/vibecosystem | 531 | 3 repos | ~2.2k | Automated safety check: Notes | MIT | |
| Planning With Filesjarrodwatts/claude-code-config | 1.1k | 5 repos | ~967 | Automated safety check: Pass | None | |
| Tool Use Data Synthesissunny-glow/Auto-BenchMax | 1.3k | — | ~3.3k | Automated safety check: Pass | None | |
| Agent Harness ConstructionKartikLabhshetwar/mind-mentor | 147 | 7 repos | ~500 | Automated safety check: Pass | Apache-2.0 |
davila7/claude-code-templates
Project Guidelines Skill (Example)
vibeeval/vibecosystem
Example template for project-specific skill files covering architecture, patterns, testing, and deployment.
jarrodwatts/claude-code-config
Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.
sunny-glow/Auto-BenchMax
Synthesize training data for ANY tool-use / agentic benchmark, in ANY repo.
KartikLabhshetwar/mind-mentor
Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates.
wshobson/agents
Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.
yonatangross/orchestkit
API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.
yonatangross/orchestkit
ADR templates in the Nygard format with context, decision, consequences, and alternatives.
yonatangross/orchestkit
Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.
yonatangross/orchestkit
Structured review processes, conventional comments, language-specific checklists, and feedback templates.
yonatangross/orchestkit
Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.
yonatangross/orchestkit
Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.
Categories
LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner). Testing LLM is an agent skill from yonatangross/orchestkit. LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner).
Testing LLM fits situations like: testing AI features; validating LLM outputs; building evaluation pipelines.
Run `npx skills add yonatangross/orchestkit --skill testing-llm -a claude-code`. Or copy the skill folder (src/skills/testing-llm in yonatangross/orchestkit) into .claude/skills/testing-llm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add yonatangross/orchestkit --skill testing-llm -a codex`. Or copy the skill folder (src/skills/testing-llm in yonatangross/orchestkit) into .agents/skills/testing-llm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill testing-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-llm, .gemini/skills/testing-llm, .github/skills/testing-llm and .opencode/skills/testing-llm in your project.
Going by SKILL.md and its folder, Testing LLM needs the command-line tools its instructions call (npx). Our summary lists: Python 3; Node.js. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+..
SKILL.md names 5 domains. As links in the text: deepeval.com, playwright.dev, docs.ragas.io, vcrpy.readthedocs.io and docs.scipy.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Testing LLM is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Testing LLM: Cc Skill Project Guidelines Example (davila7/claude-code-templates, 32k stars), Project Guidelines Example (vibeeval/vibecosystem, 531 stars), Planning With Files (jarrodwatts/claude-code-config, 1.1k stars) and Tool Use Data Synthesis (sunny-glow/Auto-BenchMax, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 289 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.
Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.