Agent skill

Evals Clarify

by tikalk in tikalk/adlc-team-skills

A skill your agent uses when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json).

MITAuto-check passedAI & LLM Engineering

Install Evals Clarify

skills CLI
$ npx skills add tikalk/adlc-team-skills --skill evals-clarify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tikalk/adlc-team-skills evals-clarify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evals/evals-clarify .claude/skills/evals-clarify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evals-clarify
GitHub stars
141
Token cost
~1.1k tokens
SKILL.md length
514 words
Files
3 (incl. scripts)
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json).

  • Works in 6 steps: Detect Lightweight Draft Format → Axial Coding & Clustering → Refinement & Adversarial Generation → …
  • Draft eval criteria need refining
  • SKILL.md covers What this skill does, When to use, When NOT to use and Process, plus 1 more section
  • Runs Shell and PowerShell scripts from its folder

What it does

Evals Clarify is an agent skill from tikalk/adlc-team-skills. Use when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json).

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/bash/setup-evals-clarify.sh`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Agent skills for the Agentic SDLC: team lifecycle (team-boot, team-learn, team-init, team-repair), software factory, evals, CDR lifecycle with confidence scoring, and… The licence is MIT.

When your agent uses it

  • Draft eval criteria need refining
  • Acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json)

Example prompts

  • “/evals-clarify”

Requirements

  • A Bash shell
  • PowerShell

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Detect Lightweight Draft Format
  2. Axial Coding & Clustering
  3. Refinement & Adversarial Generation
  4. Holdout Isolation
  5. Publish Goldset
  6. Auto-Handoff

What it can do on your machine

Read from SKILL.md and the folder at commit 2dbed36. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and PowerShell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evals Clarify loads about 1.1k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 514 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tikalk/adlc-team-skills at commit 2dbed36, republished under its MIT licence (© tikalk). 514 words, ~1,149 tokens.

Download SKILL.mdSave it as .claude/skills/evals-clarify/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evals-clarify
description
Use when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json).
disable-model-invocation
true

evals-clarify

What this skill does

Conducts axial coding following EDD Principles III & IX to cluster related failure patterns, refine evaluation criteria, generate adversarial examples, and accept validated drafts into the published goldset.

Output:

  1. Clustered Criteria - Related patterns grouped into coherent evaluation themes
  2. Adversarial Examples - Generated attack scenarios and edge cases for robustness
  3. Published Goldset - Accepted criteria in evals/{system}/goldset.md with full documentation
  4. Holdout Dataset - Reserved test set (20%) for unbiased evaluation validation
  5. JSON Configuration - Auto-generated goldset.json for system consumption
  6. Auto-handoff to /evals-implement for grader generation

Key EDD Principles Applied:

  • Principle III: Error Analysis & Pattern Discovery - Axial coding → theoretical relationships
  • Principle IX: Test Data as Code - Adversarial generation, holdout splits, version control
  • Principle II: Binary Pass/Fail - Maintain strict binary evaluation throughout
  • Principle I: Spec-Driven Contracts - Criteria validate spec compliance

When to use

  • After /evals-specify: Refine and accept draft criteria into goldset
  • Dataset maintenance: Balance pass/fail examples or add adversarial cases
  • Adding holdout split: Isolate validation data from training data

When NOT to use

  • No draft criteria exist: Run /evals-specify to discover patterns first
  • Grader generation: Use /evals-implement to convert accepted goldset into code

Process

User Input
text
$ARGUMENTS
  • --accept IDS — Accept specific draft IDs (e.g., "EVAL-001,EVAL-003")
  • --merge IDS — Merge related criteria (e.g., "EVAL-001+EVAL-002")
  • --split ID — Split complex criterion into multiple focused criteria
  • --holdout-ratio RATIO — Holdout percentage (default: 0.2, range: 0.1-0.3)
Execution Steps
Step 0: Detect Lightweight Draft Format

Check if the draft being reviewed uses the lightweight draft template (indicated by presence of type, evidence, source, revisit-when fields in frontmatter and ## Rejected Alternatives / ## Reason body sections without the full formal template sections).

If lightweight:

  1. Read the draft's captured fields (Context, Decision, Rejected Alternatives, Reason)
  2. Transform to full formal template:
    • ADR: full MADR format with Decision Drivers, Considered Options, Pros/Cons, Constitution Alignment, Related ADRs
    • PDR: full PDR format with Market Forces, Consequences, Alternatives Considered, Links
    • ChDR: full ChDR format with Issue Links, Commits, Consequences, Evidence
    • CDR: full CDR format with Context Type, Target Module, Descriptor, Evidence
    • EVAL: full eval format with Error Analysis, Pass/Fail Examples, Implementation Notes
  3. Enrich from session context (add details the lightweight draft may have omitted)
  4. Present the enriched draft for review

If already full format, proceed with normal review.

Show full SKILL.md (146 more words)Show less
Phase 1: Axial Coding & Clustering
  • Group related draft patterns into coherent themes.
  • Resolve any overlaps or duplicate criteria.
Phase 2: Refinement & Adversarial Generation
  • Generate 3-5 adversarial (attack) examples per criterion to test robustness.
  • Balance pass/fail examples (~50/50 ratio).
Phase 3: Holdout Isolation
  • Isolate exactly 20% of examples as a reserved holdout set (saved to .adlc/memory/evals/holdout.json — machine data stays under .adlc/ per ADR-401).
  • Ensure holdout set is never used in implementation or training.
Phase 4: Publish Goldset
  • Copy accepted drafts to docs/adlc/memory/evals/ (ADR-401 memory root; holdout.json alone stays at .adlc/memory/evals/holdout.json) and update status to accepted.
  • Compile published goldset to evals/{system}/goldset.md (human-readable) and evals/{system}/goldset.json (machine-readable).
Phase 5: Auto-Handoff

Trigger /evals-implement to generate code.

Verification

  • Accepted drafts stored in docs/adlc/memory/evals/EVAL-*.md
  • evals/{system}/goldset.md and goldset.json exist
  • Holdout set .adlc/memory/evals/holdout.json isolated and populated
  • All criteria are strictly binary (no confidence scores or Likert scales)
  • Handover summary lists accepted criteria and adversarial counts

© tikalk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/evals/evals-clarify of tikalk/adlc-team-skills.

  • SKILL.md
  • scripts/bash/setup-evals-clarify.sh
  • scripts/powershell/setup-evals-clarify.ps1

Open the folder on GitHubat commit 2dbed36

Compare with similar skills

Evals Clarify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evals Clarify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evals Clarify this skilltikalk/adlc-team-skills141—~1.1kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from tikalk/adlc-team-skills

All 44 skills in this repo
  • Team Boot

    tikalk/adlc-team-skills

    A skill your agent uses when a session starts or resumes after compaction (auto via the sessionstart and sessioncompact event hooks) and the team AI directives context — constitution, CDR index…

    141 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Architect Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when ADRs need review, gaps need filling, or ADR status must be approved as Accepted before architecture generation.

    141 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Change Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when reviewing, accepting, rejecting, or deferring ChDRs mined by change-init, validating inferred decisions against their git and issue evidence before promotion to project…

    141 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Change Init

    tikalk/adlc-team-skills

    A skill your agent uses when you want guided mining of git history, structured change-story clustering, or comprehensive rationale recovery before documenting.

    141 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Change Publish

    tikalk/adlc-team-skills

    A skill your agent uses when accepted ChDRs are ready for promotion from drafts to project memory at docs/adlc/memory/chdr/ and the boot-facing chdr.md index needs regenerating.

    141 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Evals Analyze

    tikalk/adlc-team-skills

    A skill your agent uses when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.

    141 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about Evals Clarify

What does Evals Clarify do?

A skill your agent uses when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json). Evals Clarify is an agent skill from tikalk/adlc-team-skills.json).

When should I use Evals Clarify?

Evals Clarify fits situations like: draft eval criteria need refining; acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json).

How do I install Evals Clarify in Claude Code?

Run `npx skills add tikalk/adlc-team-skills --skill evals-clarify -a claude-code`. Or copy the skill folder (skills/evals/evals-clarify in tikalk/adlc-team-skills) into .claude/skills/evals-clarify in your project. Claude Code loads it when a task matches its description.

How do I install Evals Clarify in Codex?

Run `npx skills add tikalk/adlc-team-skills --skill evals-clarify -a codex`. Or copy the skill folder (skills/evals/evals-clarify in tikalk/adlc-team-skills) into .agents/skills/evals-clarify in your project. Codex loads it when a task matches its description.

Can I use Evals Clarify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tikalk/adlc-team-skills --skill evals-clarify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-clarify, .gemini/skills/evals-clarify, .github/skills/evals-clarify and .opencode/skills/evals-clarify in your project.

What does Evals Clarify need to run?

Going by SKILL.md and its folder, Evals Clarify needs a shell and PowerShell for the scripts in its folder. Our summary lists: A Bash shell; PowerShell.

Does Evals Clarify access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evals Clarify safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evals Clarify use?

Evals Clarify is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evals Clarify use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evals Clarify?

Skills that share tags, products or a category with Evals Clarify: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evals Clarify?

tikalk (a GitHub organization) maintains it in tikalk/adlc-team-skills, which has 141 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: tikalk/adlc-team-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.