Managed Deep Agents
langchain-ai/langchain-skills
INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.
Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.
$ npx skills add ai-evals-course/evals-skills --skill eval-audit -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ai-evals-course/evals-skills eval-audit --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-audit .claude/skills/eval-audit && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-audit" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/eval-audit into .claude/skills/eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-audit", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ai-evals-course/evals-skills/tree/main/skills/eval-auditType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ai-evals-course/evals-skills --skill eval-audit -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ai-evals-course/evals-skills eval-audit --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-audit .agents/skills/eval-audit && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-audit" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/eval-audit into .agents/skills/eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-audit", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-evals-course/evals-skills --skill eval-audit -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ai-evals-course/evals-skills eval-audit --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-audit .cursor/skills/eval-audit && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-audit" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/eval-audit into .cursor/skills/eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-audit", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ai-evals-course/evals-skills.git --path skills/eval-audit--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ai-evals-course/evals-skills --skill eval-audit -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ai-evals-course/evals-skills eval-audit --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-audit .gemini/skills/eval-audit && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-audit" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/eval-audit into .gemini/skills/eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-audit", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ai-evals-course/evals-skills eval-auditInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ai-evals-course/evals-skills --skill eval-audit -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-audit .github/skills/eval-audit && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-audit" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/eval-audit into .github/skills/eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-audit", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-evals-course/evals-skills --skill eval-audit -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ai-evals-course/evals-skills eval-audit --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-audit .opencode/skills/eval-audit && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-audit" agent skill from https://github.com/ai-evals-course/evals-skills/tree/main/skills/eval-audit into .opencode/skills/eval-audit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-audit", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-auditInspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.
This skill inspects an LLM evaluation setup and returns a prioritized list of problems with concrete next steps. The agent gathers eval artifacts such as traces, evaluator configs, judge prompts, labeled data and metrics dashboards, runs diagnostic checks in six areas, and writes findings ordered by impact, each linked to a fix.
Artifacts come from an observability MCP server (Phoenix, Braintrust, LangSmith, Truesight or similar) when one is connected, otherwise from local CSVs, JSON trace exports, notebooks or evaluation scripts. One check asks whether systematic error analysis was done and whether failure categories were observed in traces or merely brainstormed from generic labels. Another looks at evaluator design, such as whether judges are binary pass or fail.
A separate path covers teams with no eval infrastructure. The skill is not for building a new evaluator from scratch; for that it points to the error-discovery, write-judge-prompt and validate-evaluator skills.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 80d5f7b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
hamel.devarxiv.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
LLM Eval Pipeline Audit loads about 2.5k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 1,222 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ai-evals-course/evals-skills at commit 80d5f7b, republished under its Apache-2.0 licence (© ai-evals-course). 1,222 words, ~2,495 tokens.
.claude/skills/eval-audit/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Inspect an LLM eval pipeline and produce a prioritized list of problems with concrete next steps.
Access to eval artifacts (traces, evaluator configs, judge prompts, labeled data) via an observability MCP server or local files. If none exist, skip to "No Eval Infrastructure."
Check whether the user has an observability MCP server connected (Phoenix, Braintrust, LangSmith, Truesight or similar). If available, use it to pull traces, evaluator definitions, and experiment results. If not, ask for local files: CSVs, JSON trace exports, notebooks, or evaluation scripts.
Work through each area below. Inspect available artifacts, determine whether the problem exists, and record a finding if it does.
Prioritize findings by impact on the user's product. Present the most impactful findings first.
Check: Has the user done systematic error analysis on real or synthetic traces?
Look for: labeled trace datasets, failure category definitions, notes from trace review. If evaluators exist but no documented failure categories, error analysis was likely skipped.
Finding if missing: Evaluators built without error analysis measure generic qualities ("helpfulness", "coherence") instead of actual failure modes. Start with error-discovery, or generate-synthetic-data first if no traces exist.
See: Your AI Product Needs Evals, LLM Evals FAQ
Check: Were failure categories brainstormed or observed?
Generic labels borrowed from research ("hallucination score", "toxicity", "coherence") suggest brainstorming. Application-grounded categories ("missing query constraints", "wrong client tone", "fabricated property features") suggest observation.
Finding if brainstormed: Generic categories miss application-specific failures and produce evaluators that score well on paper but miss real problems. Re-do with error-discovery, starting from traces.
See: Who Validates the Validators?
Check: Are evaluators binary pass/fail?
Flag any that use Likert scales (1-5), letter grades (A-F), or numeric scores without a clear pass/fail threshold.
Finding if not binary: Likert scales are difficult to calibrate. Annotators disagree on the difference between a 3 and a 4, and judges inherit that noise. Consider converting to binary pass/fail with explicit definitions using write-judge-prompt.
See: Creating an LLM Judge That Drives Business Results
Check: Do LLM judge prompts target specific failure modes?
Flag any that evaluate holistically ("Is this response helpful?", "Rate the quality of this output").
Finding if vague: Holistic judges produce unactionable verdicts. Each judge should check exactly one failure mode with explicit pass/fail definitions and few-shot examples. Use write-judge-prompt.
Check: Are code-based checks used where possible?
Flag LLM judges used for objectively checkable criteria: format validation, constraint satisfaction, keyword presence, schema conformance.
Finding if over-relying on judges: Replace objective checks with code (regex, parsing, schema validation, execution tests). Reserve LLM judges for criteria requiring interpretation. Use write-code-eval.
Check: Are similarity metrics used as primary evaluation?
Flag ROUGE, BERTScore, cosine similarity, or embedding distance used as the main evaluator for generation quality.
Finding if present: These metrics measure surface-level overlap, not correctness. They suit retrieval ranking but not generation evaluation. Replace with binary evaluators grounded in specific failure modes.
See: LLM Evals FAQ
Check: Are LLM judges validated against human labels?
Look for: confusion matrices, TPR/TNR measurements, alignment scores. Judges in production with no validation data is a critical finding.
Finding if unvalidated: An unvalidated judge may consistently miss failures or flag passing traces. Measure alignment using TPR and TNR on a held-out test set. Use validate-evaluator.
See: Creating an LLM Judge That Drives Business Results
Check: Is alignment measured with TPR/TNR or with raw accuracy?
Flag "accuracy", "percent agreement", or Cohen's Kappa as the primary alignment metric.
Finding if using accuracy: With class imbalance, raw accuracy is misleading: a judge that always says "Pass" gets 90% accuracy when 90% of traces pass but catches zero failures. Use TPR and TNR, which map directly to bias correction. Use validate-evaluator.
Check: Is there a proper train/dev/test split?
Check whether few-shot examples in judge prompts come from the same data used to measure judge performance.
Finding if leaking: Using evaluation data as few-shot examples inflates alignment scores and hides real judge failures. Split into train (few-shot source), dev (iteration), and test (final measurement). Use validate-evaluator.
Check: Who is reviewing traces?
Determine whether domain experts or outsourced annotators are labeling data.
Finding if outsourced without domain expertise: General annotators catch formatting errors but miss domain-specific failures (wrong medical dosage, incorrect legal citation, mismatched property features). Involve a domain expert.
See: A Field Guide to Improving AI Products
Check: Are reviewers seeing full traces or just final outputs?
Finding if output-only: Reviewing only the final output hides where the pipeline broke. Show the full trace: input, intermediate steps, tool calls, retrieved context, and final output.
Check: How is data displayed to reviewers?
Flag raw JSON, unformatted text, or spreadsheets with trace data in cells.
Finding if raw format: Reviewers spend effort parsing data instead of judging quality. Format in natural representation: render markdown, syntax-highlight code, display tables as tables. Use build-review-interface.
See: LLM Evals FAQ
Check: Is there enough labeled data?
For error analysis, ~100 traces is the rough target for saturation. For judge validation, ~50 Pass and ~50 Fail examples are needed for reliable TPR/TNR. If labeled data is sparse, collect more by sampling traces more effectively:
Finding if insufficient: Small datasets produce unreliable failure rates and wide confidence intervals. Use the sampling strategies above to collect more labeled data, or supplement with generate-synthetic-data.
Check: Is error analysis re-run after significant changes?
Check when error analysis was last performed relative to model switches, prompt rewrites, new features, or production incidents.
Finding if stale: Failure modes shift after pipeline changes, and evaluators built for the old pipeline miss new failure types. Re-run error analysis after every significant change.
Check: Are evaluators maintained?
Look for periodic re-validation of judges or refreshed evaluation datasets.
Finding if set-and-forget: Evaluators degrade as the pipeline evolves. Re-validate judges against fresh human labels and update eval datasets to reflect current usage.
If the user has no eval artifacts (no traces, no evaluators, no labeled data):
error-discovery on a sample of real traces.generate-synthetic-data to create test inputs, run them through the pipeline, then apply error-discovery to the resulting traces.Present findings ordered by impact. For each:
### [Problem Title]
**Status:** [Problem exists / OK / Cannot determine]
[1-2 sentence explanation of the specific problem found]
**Fix:** [Concrete action, referencing a skill or article]Group under the six diagnostic areas. Omit areas where no problems were found.
© ai-evals-course, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/eval-audit of ai-evals-course/evals-skills.
Open the folder on GitHubat commit 80d5f7b
LLM Eval Pipeline Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| LLM Eval Pipeline Audit this skillai-evals-course/evals-skills | 1.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Managed Deep Agentslangchain-ai/langchain-skills | 1.3k | — | ~8.7k | Automated safety check: Notes | MIT | |
| Opik Evaluatecomet-ml/opik-mcp | 219 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Braintrust Agent Evalsgithits-com/githits-cli | 114 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Compliance Drift Evalsucsandman/DashClaw | 310 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Caveman Experiment ManagerJuliusBrussee/caveman | 110k | 1 repos | ~975 | Automated safety check: Pass | Apache-2.0 |
langchain-ai/langchain-skills
INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.
comet-ml/opik-mcp
Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.
githits-com/githits-cli
Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.
ucsandman/DashClaw
Set up compliance exports, drift detection, evaluations, scoring, and learning analytics
JuliusBrussee/caveman
Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.
DenisSergeevitch/agents-best-practices
A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.
ai-evals-course/evals-skills
Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.
ai-evals-course/evals-skills
Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.
ai-evals-course/evals-skills
Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.
ai-evals-course/evals-skills
Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.
ai-evals-course/evals-skills
Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.
ai-evals-course/evals-skills
Write code evaluators for known failure modes with objective rules.
Works with
Categories
Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes. This skill inspects an LLM evaluation setup and returns a prioritized list of problems with concrete next steps. The agent gathers eval artifacts such as traces, evaluator configs, judge prompts, labeled data and metrics dashboards, runs diagnostic checks in six areas, and writes findings ordered by impact, each linked to a fix.
LLM Eval Pipeline Audit fits situations like: inheriting an eval system and unsure whether it can be trusted; checking whether judges have been validated against human labels; finding vanity metrics or skipped error analysis in an eval pipeline; starting point for a team that has no eval infrastructure yet.
Run `npx skills add ai-evals-course/evals-skills --skill eval-audit -a claude-code`. Or copy the skill folder (skills/eval-audit in ai-evals-course/evals-skills) into .claude/skills/eval-audit in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ai-evals-course/evals-skills --skill eval-audit -a codex`. Or copy the skill folder (skills/eval-audit in ai-evals-course/evals-skills) into .agents/skills/eval-audit in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-evals-course/evals-skills --skill eval-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-audit, .gemini/skills/eval-audit, .github/skills/eval-audit and .opencode/skills/eval-audit in your project.
SKILL.md names no scripts, command-line tools or credentials: LLM Eval Pipeline Audit is instructions for the agent only. Our summary lists: Eval artifacts as local files or through an observability MCP server.
SKILL.md names 2 domains. As links in the text: hamel.dev and arxiv.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
LLM Eval Pipeline Audit is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with LLM Eval Pipeline Audit: Managed Deep Agents (langchain-ai/langchain-skills, 1.3k stars), Opik Evaluate (comet-ml/opik-mcp, 219 stars), Braintrust Agent Evals (githits-com/githits-cli, 114 stars) and Compliance Drift Evals (ucsandman/DashClaw, 310 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ai-evals-course (a GitHub organization) maintains it in ai-evals-course/evals-skills, which has 1,468 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 24, 2026.
Source: ai-evals-course/evals-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.