Agent skill

Evals Specify

by tikalk in tikalk/adlc-team-skills

A skill your agent uses when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria.

MITAuto-check passedAI & LLM Engineering

Install Evals Specify

skills CLI
$ npx skills add tikalk/adlc-team-skills --skill evals-specify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tikalk/adlc-team-skills evals-specify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evals/evals-specify .claude/skills/evals-specify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evals-specify
GitHub stars
141
Token cost
~915 tokens
SKILL.md length
394 words
Files
3 (incl. scripts)
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria.

  • Works in 4 steps: Open Coding Analysis → Create Draft Criteria → Create Draft Files → …
  • You want guided bottom-up error analysis
  • SKILL.md covers What this skill does, When to use, When NOT to use and Process, plus 1 more section
  • Runs Shell and PowerShell scripts from its folder

What it does

Evals Specify is an agent skill from tikalk/adlc-team-skills. Use when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria. Optional for routine capture — team-boot writes lightweight eval drafts directly.

Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/bash/setup-evals-specify.sh`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Agent skills for the Agentic SDLC: team lifecycle (team-boot, team-learn, team-init, team-repair), software factory, evals, CDR lifecycle with confidence scoring, and… The licence is MIT.

When your agent uses it

  • You want guided bottom-up error analysis
  • Structured failure taxonomy discovery
  • Comprehensive trace coding before documenting eval criteria

Example prompts

  • “/evals-specify”

Requirements

  • A Bash shell
  • PowerShell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Open Coding Analysis
  2. Create Draft Criteria
  3. Create Draft Files
  4. Auto-Handoff

What it can do on your machine

Read from SKILL.md and the folder at commit 2dbed36. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and PowerShell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evals Specify loads about 915 tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 394 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~915

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tikalk/adlc-team-skills at commit 2dbed36, republished under its MIT licence (© tikalk). 394 words, ~915 tokens.

Download SKILL.mdSave it as .claude/skills/evals-specify/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evals-specify
description
Use when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria. Optional for routine capture — team-boot writes lightweight eval drafts directly.
disable-model-invocation
true

evals-specify

What this skill does

Conducts bottom-up error analysis following EDD Principles III & IX (Error Analysis & Test Data as Code) to discover and document draft evaluation criteria from human observation of system failures.

Output:

  1. Draft Eval Records - Individual EVAL-*.md files in .adlc/drafts/evals/ with open coding notes
  2. Error Pattern Documentation - Bottom-up failure taxonomy from actual traces
  3. Pass/Fail Examples - Real examples that should pass/fail each criterion
  4. Auto-handoff to /evals-clarify for axial coding and clustering

Key EDD Principles Applied:

  • Principle III: Error Analysis & Pattern Discovery - Open coding → failure taxonomy
  • Principle IX: Test Data as Code - Dataset planning and coverage analysis
  • Principle II: Binary Pass/Fail - Maintain strict binary pass/fail conditions
  • Principle V: Trajectory Observability - Track full multi-turn conversation traces

When to use

Note: Routine decision capture is handled by team-boot's continuous capture mechanism, which writes lightweight drafts directly to .adlc/drafts/. This skill is for interactive deep-dive exploration — when you want guided trade-off analysis, multi-option comparison, or structured decision facilitation before documenting.

  • Starting evaluation development: No existing criteria, need discovery from failure logs
  • Production incident analysis: Recent failures require systematic analysis
  • Quality assessment: Discovering and codifying boundary conditions from failures

When NOT to use

  • No failure traces/specs: Generate synthetic traces first, or use /evals-init to set up security baselines
  • Known criteria already exist: Use /evals-clarify to refine or /evals-implement to generate code

Process

Show full SKILL.md (172 more words)Show less
User Input
text
$ARGUMENTS

Treat user input as specific failure areas or error patterns to analyze (e.g., "authentication bypass", "RAG irrelevant results").

  • --traces N — Number of traces to analyze (default: 20, min for theoretical saturation)
  • --source SOURCE — Trace source location (e.g., logs, support tickets)
Execution Steps
Phase 1: Open Coding Analysis
  • Reviews the user-provided failure logs or spec requirements.
  • Conducts open coding of traces to discover recurring failure patterns (EDD Principle III).
  • Identifies: core problem, causal conditions, and consequences.
Phase 2: Create Draft Criteria

Group patterns into draft criteria. For each:

  • Define strict Pass Condition (observable, binary yes/no)
  • Define strict Fail Condition (observable, binary yes/no)
  • Document real pass/fail examples directly from traces
Phase 3: Create Draft Files
  • Copy skills/evals/evals-templates/eval-criterion-template.md to .adlc/drafts/evals/EVAL-{NNN}.md.
  • Populate metadata and error analysis notes.
  • Regenerate index at .adlc/drafts/evals/evals.md.
Phase 4: Auto-Handoff

Trigger /evals-clarify for axial coding and clustering.

Verification

  • Draft files created at .adlc/drafts/evals/EVAL-*.md
  • Index file .adlc/drafts/evals/evals.md updated with draft summaries
  • Each draft contains: status "draft", pass/fail conditions, trace sources, and concrete examples
  • Auto-handoff context produced with list of created drafts

© tikalk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/evals/evals-specify of tikalk/adlc-team-skills.

  • SKILL.md
  • scripts/bash/setup-evals-specify.sh
  • scripts/powershell/setup-evals-specify.ps1

Open the folder on GitHubat commit 2dbed36

Compare with similar skills

Evals Specify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evals Specify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evals Specify this skilltikalk/adlc-team-skills141—~915Automated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from tikalk/adlc-team-skills

All 44 skills in this repo
  • Workspace

    tikalk/adlc-team-skills

    A skill your agent uses when coordinating a multi-repo workspace — init the .adlc/ structure, discover and link child repos as submodules, or audit workspace health (branch, dirty, unpushed, SHA…

    141 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Team Boot

    tikalk/adlc-team-skills

    A skill your agent uses when a session starts or resumes after compaction (auto via the sessionstart and sessioncompact event hooks) and the team AI directives context — constitution, CDR index…

    141 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Architect Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when ADRs need review, gaps need filling, or ADR status must be approved as Accepted before architecture generation.

    141 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Change Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when reviewing, accepting, rejecting, or deferring ChDRs mined by change-init, validating inferred decisions against their git and issue evidence before promotion to project…

    141 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Change Init

    tikalk/adlc-team-skills

    A skill your agent uses when you want guided mining of git history, structured change-story clustering, or comprehensive rationale recovery before documenting.

    141 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Change Publish

    tikalk/adlc-team-skills

    A skill your agent uses when accepted ChDRs are ready for promotion from drafts to project memory at docs/adlc/memory/chdr/ and the boot-facing chdr.md index needs regenerating.

    141 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Questions about Evals Specify

What does Evals Specify do?

A skill your agent uses when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria. Evals Specify is an agent skill from tikalk/adlc-team-skills. Use when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria.

When should I use Evals Specify?

Evals Specify fits situations like: you want guided bottom-up error analysis; structured failure taxonomy discovery; comprehensive trace coding before documenting eval criteria.

How do I install Evals Specify in Claude Code?

Run `npx skills add tikalk/adlc-team-skills --skill evals-specify -a claude-code`. Or copy the skill folder (skills/evals/evals-specify in tikalk/adlc-team-skills) into .claude/skills/evals-specify in your project. Claude Code loads it when a task matches its description.

How do I install Evals Specify in Codex?

Run `npx skills add tikalk/adlc-team-skills --skill evals-specify -a codex`. Or copy the skill folder (skills/evals/evals-specify in tikalk/adlc-team-skills) into .agents/skills/evals-specify in your project. Codex loads it when a task matches its description.

Can I use Evals Specify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tikalk/adlc-team-skills --skill evals-specify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-specify, .gemini/skills/evals-specify, .github/skills/evals-specify and .opencode/skills/evals-specify in your project.

What does Evals Specify need to run?

Going by SKILL.md and its folder, Evals Specify needs a shell and PowerShell for the scripts in its folder. Our summary lists: A Bash shell; PowerShell.

Does Evals Specify access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evals Specify safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evals Specify use?

Evals Specify is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evals Specify use?

About 915 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evals Specify?

Skills that share tags, products or a category with Evals Specify: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evals Specify?

tikalk (a GitHub organization) maintains it in tikalk/adlc-team-skills, which has 141 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: tikalk/adlc-team-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.