Agent skill

Evals Validate

by tikalk in tikalk/adlc-team-skills

A skill your agent uses when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy).

MITAuto-check passedAI & LLM Engineering

Install Evals Validate

skills CLI
$ npx skills add tikalk/adlc-team-skills --skill evals-validate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tikalk/adlc-team-skills evals-validate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evals/evals-validate .claude/skills/evals-validate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evals-validate
GitHub stars
141
Token cost
~850 tokens
SKILL.md length
374 words
Files
3 (incl. scripts)
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy).

  • Works in 5 steps: Execute Evaluations → Compute Statistical Validation → SLA Compliance Check → …
  • A goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance
  • SKILL.md covers What this skill does, When to use, When NOT to use and Process, plus 1 more section
  • Runs Shell and PowerShell scripts from its folder; calls npx, pytest and python

What it does

Evals Validate is an agent skill from tikalk/adlc-team-skills. Use when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy).

Its SKILL.md is about 850 tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/bash/setup-evals-validate.sh`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Agent skills for the Agentic SDLC: team lifecycle (team-boot, team-learn, team-init, team-repair), software factory, evals, CDR lifecycle with confidence scoring, and… The licence is MIT.

When your agent uses it

  • A goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance
  • Statistical accuracy)

Example prompts

  • “/evals-validate”

Requirements

  • Python 3
  • Node.js
  • A Bash shell
  • PowerShell

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Execute Evaluations
  2. Compute Statistical Validation
  3. SLA Compliance Check
  4. Write Validation Report
  5. Auto-Handoff

What it can do on your machine

Read from SKILL.md and the folder at commit 2dbed36. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and PowerShell), which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • pytest
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evals Validate loads about 850 tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 374 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~850

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tikalk/adlc-team-skills at commit 2dbed36, republished under its MIT licence (© tikalk). 374 words, ~850 tokens.

Download SKILL.mdSave it as .claude/skills/evals-validate/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evals-validate
description
Use when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy).
disable-model-invocation
true

evals-validate

What this skill does

Conducts comprehensive validation of the implemented evaluation system following EDD principles to ensure production readiness through statistical analysis, performance verification, and quality assurance.

Output:

  1. Statistical Validation - TPR/TNR analysis, accuracy metrics, confidence intervals
  2. Performance Validation - SLA compliance verification for evaluation pyramid tiers
  3. Quality Assurance - Goldset integrity, example balance, coverage analysis
  4. Holdout Dataset Validation - Unbiased accuracy assessment on reserved test set
  5. Auto-handoff to /evals-analyze for closed loop trajectory analysis

Key EDD Principles Applied:

  • Principle IV: Evaluation Pyramid - Tier performance SLA validation (Tier 1 <30s, Tier 2 <5min)
  • Principle II: Binary Pass/Fail - Statistical compliance verification
  • Principle IX: Test Data as Code - Holdout dataset validation integrity
  • Principle III: Error Analysis - Pattern stability validation

When to use

  • After /evals-implement: Execute the evaluation suite and measure quality
  • CI/CD Pipeline gate: Run evaluations before release to ensure no regressions
  • Periodic audit: Verify evaluator accuracy on holdout data to check for model drift

When NOT to use

  • Evaluator not generated: Run /evals-implement to build grader files first
  • Analysing failure traces: Use /evals-analyze to extract deep insights from run results

Process

User Input
text
$ARGUMENTS
  • --holdout-only — Validate only on holdout dataset (unbiased validation)
  • --performance-only — Skip statistical analysis, focus on SLA compliance
  • --metrics METRICS — Specific metrics to validate (tpr, tnr, accuracy, performance)
Execution Steps
Phase 1: Execute Evaluations

Runs the underlying framework CLI directly:

  • PromptFoo: npx promptfoo eval --config evals/promptfoo/config.js
  • DeepEval: pytest evals/deepeval/ -v or python evals/deepeval/config.py
Show full SKILL.md (139 more words)Show less
Phase 2: Compute Statistical Validation
  • Parse generated results JSON from evals/results/.
  • Calculate True Positive Rate (TPR) and True Negative Rate (TNR).
  • Calculate overall accuracy with 95% confidence intervals.
  • Ensure no Likert scales or numerical scores leak into results.
Phase 3: SLA Compliance Check
  • Measure execution times for Tier 1 and Tier 2.
  • Verify Tier 1 completes under 30 seconds.
  • Verify Tier 2 completes under 5 minutes.
  • Check headroom analysis (SLA budget consumed).
Phase 4: Write Validation Report
  • Write validation results to evals/results/validation_report.md.
  • Include pass/fail counts, TPR/TNR table, SLA timings, and holdout set results.
Phase 5: Auto-Handoff

Trigger /evals-analyze to close the loop.

Verification

  • Evaluation execution successfully completed with results JSON written to evals/results/
  • evals/results/validation_report.md created with TPR/TNR and SLA metrics
  • Statistical metrics calculated with confidence intervals
  • Headroom and SLA compliance verified
  • Handover summary lists results and validation report path

© tikalk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/evals/evals-validate of tikalk/adlc-team-skills.

  • SKILL.md
  • scripts/bash/setup-evals-validate.sh
  • scripts/powershell/setup-evals-validate.ps1

Open the folder on GitHubat commit 2dbed36

Compare with similar skills

Evals Validate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evals Validate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evals Validate this skilltikalk/adlc-team-skills141—~850Automated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from tikalk/adlc-team-skills

All 44 skills in this repo
  • Workspace

    tikalk/adlc-team-skills

    A skill your agent uses when coordinating a multi-repo workspace — init the .adlc/ structure, discover and link child repos as submodules, or audit workspace health (branch, dirty, unpushed, SHA…

    141 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Team Boot

    tikalk/adlc-team-skills

    A skill your agent uses when a session starts or resumes after compaction (auto via the sessionstart and sessioncompact event hooks) and the team AI directives context — constitution, CDR index…

    141 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Architect Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when ADRs need review, gaps need filling, or ADR status must be approved as Accepted before architecture generation.

    141 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Change Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when reviewing, accepting, rejecting, or deferring ChDRs mined by change-init, validating inferred decisions against their git and issue evidence before promotion to project…

    141 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Change Init

    tikalk/adlc-team-skills

    A skill your agent uses when you want guided mining of git history, structured change-story clustering, or comprehensive rationale recovery before documenting.

    141 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Change Publish

    tikalk/adlc-team-skills

    A skill your agent uses when accepted ChDRs are ready for promotion from drafts to project memory at docs/adlc/memory/chdr/ and the boot-facing chdr.md index needs regenerating.

    141 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Questions about Evals Validate

What does Evals Validate do?

A skill your agent uses when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). Evals Validate is an agent skill from tikalk/adlc-team-skills. Use when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy).

When should I use Evals Validate?

Evals Validate fits situations like: A goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance; statistical accuracy).

How do I install Evals Validate in Claude Code?

Run `npx skills add tikalk/adlc-team-skills --skill evals-validate -a claude-code`. Or copy the skill folder (skills/evals/evals-validate in tikalk/adlc-team-skills) into .claude/skills/evals-validate in your project. Claude Code loads it when a task matches its description.

How do I install Evals Validate in Codex?

Run `npx skills add tikalk/adlc-team-skills --skill evals-validate -a codex`. Or copy the skill folder (skills/evals/evals-validate in tikalk/adlc-team-skills) into .agents/skills/evals-validate in your project. Codex loads it when a task matches its description.

Can I use Evals Validate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tikalk/adlc-team-skills --skill evals-validate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-validate, .gemini/skills/evals-validate, .github/skills/evals-validate and .opencode/skills/evals-validate in your project.

What does Evals Validate need to run?

Going by SKILL.md and its folder, Evals Validate needs a shell and PowerShell for the scripts in its folder and the command-line tools its instructions call (npx, pytest and python). Our summary lists: Python 3; Node.js; A Bash shell; PowerShell.

Does Evals Validate access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evals Validate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evals Validate use?

Evals Validate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evals Validate use?

About 850 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evals Validate?

Skills that share tags, products or a category with Evals Validate: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evals Validate?

tikalk (a GitHub organization) maintains it in tikalk/adlc-team-skills, which has 141 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: tikalk/adlc-team-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.