Agent skill

Evals Implement

by tikalk in tikalk/adlc-team-skills

A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.

MITAuto-check passedTesting & QA

Install Evals Implement

skills CLI
$ npx skills add tikalk/adlc-team-skills --skill evals-implement -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tikalk/adlc-team-skills evals-implement --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evals/evals-implement .claude/skills/evals-implement && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evals-implement
GitHub stars
141
Token cost
~1.4k tokens
SKILL.md length
628 words
Files
3 (incl. scripts)
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.

  • Works in 4 steps: Trace-to-Grader Synthesis (Automated… → Unit Test Generation → Config Generation → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers What this skill does, When to use, When NOT to use and Process, plus 1 more section
  • Runs Shell and PowerShell scripts from its folder; calls pytest

What it does

Evals Implement is an agent skill from tikalk/adlc-team-skills. Use when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/bash/setup-evals-implement.sh`).

It sits in Testing & QA, covering LLM evaluation and Unit testing. It works with Python. The repository describes itself as: Agent skills for the Agentic SDLC: team lifecycle (team-boot, team-learn, team-init, team-repair), software factory, evals, CDR lifecycle with confidence scoring, and… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve Unit testing

Example prompts

  • “/evals-implement”

Requirements

  • Python 3
  • A Bash shell
  • PowerShell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Trace-to-Grader Synthesis (Automated Eval Engineering)
  2. Unit Test Generation
  3. Config Generation
  4. Auto-Handoff

What it can do on your machine

Read from SKILL.md and the folder at commit 2dbed36. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and PowerShell), which the agent can run.

    Shell commands in SKILL.md call:

    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evals Implement loads about 1.4k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 628 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tikalk/adlc-team-skills at commit 2dbed36, republished under its MIT licence (© tikalk). 628 words, ~1,351 tokens.

Download SKILL.mdSave it as .claude/skills/evals-implement/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evals-implement
description
Use when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.
disable-model-invocation
true

evals-implement

What this skill does

Generates the complete executable evaluation implementation following EDD Principle VIII (Close Production Loop) from the published goldset, with automated unit testing to verify evaluator correctness.

Output:

  1. Grader/Metric Implementation - Python evaluators for each goldset criterion with binary pass/fail
    • PromptFoo: Python grader functions with JSON output in evals/{system}/graders/
    • DeepEval: Custom metric classes inheriting from BaseMetric
  2. Evaluator Unit Tests - Automated tests (evals/{system}/tests/test_check_*.py) that run the goldset pass/fail examples against the generated graders to ensure the evaluator itself is accurate
  3. Evaluation Configuration - Complete config file (config.js or config.py) with Tier 1 + Tier 2 evaluation structure
  4. Auto-handoff to /evals-validate to run validation

Key EDD Principles Applied:

  • Principle VIII: Close Production Loop - Failure type gates route to appropriate actions
  • Principle II: Binary Pass/Fail - Ensure graders return strictly 1.0 (pass) or 0.0 (fail)
  • Principle IX: Test Data as Code - Unit test generated code against dataset examples

When to use

  • After /evals-clarify: Convert accepted goldset criteria into executable code
  • Regenerating configs: Re-build evaluator suite after adding new goldset criteria
  • Adding unit tests: Hardening the evaluator itself against regression or bugs

When NOT to use

  • Goldset not published: Run /evals-clarify to generate goldset.json first
  • Running evaluations: Use /evals-validate to run the suite against application outputs

Process

User Input
text
$ARGUMENTS
  • --system SYSTEM — Override active evaluation framework (promptfoo or deepeval)
  • --no-tests — Skip automated unit test generation for graders (not recommended)
Execution Steps
Phase 1: Trace-to-Grader Synthesis (Automated Eval Engineering)
  • Reads evals/{system}/goldset.json.
  • Maps rich evidence fields from the goldset criteria into grader logic (Trace-to-Grader Synthesis):
    • Uses pass_condition and fail_condition as the grader's core rubric.
    • Extracts pass/fail examples to act as raw data anchors and few-shot classification anchors inside the grader logic.
    • Injects Root Cause Analysis and axial_coding notes as contextual prompt guidelines or regex patterns to catch exact failure manifestations.
  • For PromptFoo: Generates Python grader functions (evals/{system}/graders/check_*.py) containing specialized, dynamic LLM-judge templates or regex checks compiled from these goldset inputs.
  • For DeepEval: Generates Custom Metric classes inheriting from BaseMetric compiled from these goldset inputs.
  • All graders conform strictly to the binary pass/fail standard (returning only 1.0 or 0.0, with zero Likert scale leakage).
Phase 2: Unit Test Generation
  • Generates matching unit tests (evals/{system}/tests/test_check_*.py) for each grader.
  • Unit tests verify the grader correctly identifies the goldset's training pass and fail examples.
Show full SKILL.md (252 more words)Show less
Phase 2b: Closed-Loop Grader Self-Tuning
  • Executes generated unit tests (pytest evals/{system}/tests/) to verify evaluator accuracy.
  • Grader Calibration Loop:
    1. Inspects test results to detect any misclassifications (false positives/negatives) on the training cases.
    2. If any test fails, triggers a feedback edit step that parses the failure reasons and automatically adjusts the grader's internal prompt rubric, regex stubs, or score thresholds.
    3. Re-runs pytest to check accuracy.
    4. Repeats for up to 3 iterations (the hard circuit-breaker limit).
  • Holdout Locking: Ensure the holdout validation set (holdout.json) remains completely isolated and is never loaded or exposed to the self-tuning loop (to prevent overfitting).
  • Failure Escalation: If the grader does not converge to 100% training accuracy within 3 iterations, the loop halts, surfaces the failing test case details, and raises an error rather than passing silently.
Phase 3: Config Generation
  • Generates the unified framework configuration file (config.js or config.py).
  • Configures separate Tier 1 (fast checks, <30s, deterministic) and Tier 2 (semantic checks, <5min, LLM-judge) pipelines.
Phase 4: Auto-Handoff

Trigger /evals-validate to run validation.

Verification

  • evals/{system}/graders/ contains Python grader scripts for each criterion compiled dynamically from goldset pass/fail examples and root-cause analyses
  • evals/{system}/tests/ contains matching unit test files
  • Framework config (config.js or config.py) successfully generated
  • Grader calibration self-tuning loop ran and converged to 100% training accuracy within the 3-iteration cap (or raised explicit non-convergence errors)
  • Holdout dataset protection confirmed (validation holdout.json remained completely isolated and untouched during tuning)
  • All grader unit tests pass locally (pytest evals/{system}/tests/)
  • Handover summary lists generated graders, self-tuning iterations, and test results

© tikalk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/evals/evals-implement of tikalk/adlc-team-skills.

  • SKILL.md
  • scripts/bash/setup-evals-implement.sh
  • scripts/powershell/setup-evals-implement.ps1

Open the folder on GitHubat commit 2dbed36

Compare with similar skills

Evals Implement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evals Implement compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evals Implement this skilltikalk/adlc-team-skills141—~1.4kAutomated safety check: PassMIT
Running Testsbrendanhasz/probflow175—~657Automated safety check: PassMIT
Eval Loopjacob-dietle/context-os111—~5.2kAutomated safety check: PassMIT
Eval Driven Devgithub/awesome-copilot40k1 repos~4.4kAutomated safety check: WarnMIT
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence

Similar skills

  • Running Tests

    brendanhasz/probflow

    Run Python unit test suites strictly using the uv package manager and pytest.

    175 GitHub stars~657 tokensUpdated 11 days ago
    Testing & QAAuto-check passed
  • Eval Loop

    jacob-dietle/context-os

    This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target.

    111 GitHub stars~5.2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Eval Driven Dev

    github/awesome-copilot

    Official

    Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~4.4k tokens
    Testing & QAAuto-check: warnings
  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • ONNX Runtime Test Runner

    microsoft/onnxruntime

    Official

    Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.

    22k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed

More from tikalk/adlc-team-skills

All 44 skills in this repo
  • Workspace

    tikalk/adlc-team-skills

    A skill your agent uses when coordinating a multi-repo workspace — init the .adlc/ structure, discover and link child repos as submodules, or audit workspace health (branch, dirty, unpushed, SHA…

    141 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Team Boot

    tikalk/adlc-team-skills

    A skill your agent uses when a session starts or resumes after compaction (auto via the sessionstart and sessioncompact event hooks) and the team AI directives context — constitution, CDR index…

    141 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Architect Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when ADRs need review, gaps need filling, or ADR status must be approved as Accepted before architecture generation.

    141 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Change Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when reviewing, accepting, rejecting, or deferring ChDRs mined by change-init, validating inferred decisions against their git and issue evidence before promotion to project…

    141 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Change Init

    tikalk/adlc-team-skills

    A skill your agent uses when you want guided mining of git history, structured change-story clustering, or comprehensive rationale recovery before documenting.

    141 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Change Publish

    tikalk/adlc-team-skills

    A skill your agent uses when accepted ChDRs are ready for promotion from drafts to project memory at docs/adlc/memory/chdr/ and the boot-facing chdr.md index needs regenerating.

    141 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Evals Implement

What does Evals Implement do?

A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness. Evals Implement is an agent skill from tikalk/adlc-team-skills. Use when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.

When should I use Evals Implement?

Evals Implement fits situations like: tasks that involve LLM evaluation; tasks that involve Unit testing.

How do I install Evals Implement in Claude Code?

Run `npx skills add tikalk/adlc-team-skills --skill evals-implement -a claude-code`. Or copy the skill folder (skills/evals/evals-implement in tikalk/adlc-team-skills) into .claude/skills/evals-implement in your project. Claude Code loads it when a task matches its description.

How do I install Evals Implement in Codex?

Run `npx skills add tikalk/adlc-team-skills --skill evals-implement -a codex`. Or copy the skill folder (skills/evals/evals-implement in tikalk/adlc-team-skills) into .agents/skills/evals-implement in your project. Codex loads it when a task matches its description.

Can I use Evals Implement in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tikalk/adlc-team-skills --skill evals-implement -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-implement, .gemini/skills/evals-implement, .github/skills/evals-implement and .opencode/skills/evals-implement in your project.

What does Evals Implement need to run?

Going by SKILL.md and its folder, Evals Implement needs a shell and PowerShell for the scripts in its folder and the command-line tools its instructions call (pytest). Our summary lists: Python 3; A Bash shell; PowerShell.

Does Evals Implement access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evals Implement safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evals Implement use?

Evals Implement is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evals Implement use?

About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evals Implement?

Skills that share tags, products or a category with Evals Implement: Running Tests (brendanhasz/probflow, 175 stars), Eval Loop (jacob-dietle/context-os, 111 stars), Eval Driven Dev (github/awesome-copilot, 40k stars) and Adk Verify Snippets (google/adk-python, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evals Implement?

tikalk (a GitHub organization) maintains it in tikalk/adlc-team-skills, which has 141 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: tikalk/adlc-team-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.