Agent skill

Evals Analyze

by tikalk in tikalk/adlc-team-skills

A skill your agent uses when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.

MITAuto-check passedAI & LLM Engineering

Install Evals Analyze

skills CLI
$ npx skills add tikalk/adlc-team-skills --skill evals-analyze -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tikalk/adlc-team-skills evals-analyze --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evals/evals-analyze .claude/skills/evals-analyze && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evals-analyze
GitHub stars
141
Token cost
~1.2k tokens
SKILL.md length
554 words
Files
3 (incl. scripts)
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.

  • Works in 4 steps: Load Evaluation Results → Failure Classification → Action Routing (Close the Loop) → …
  • Evaluation results need triage and loop-closing — spec failures route to deterministic checks
  • SKILL.md covers What this skill does, When to use, When NOT to use and Process, plus 1 more section
  • Runs Shell and PowerShell scripts from its folder

What it does

Evals Analyze is an agent skill from tikalk/adlc-team-skills. Use when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/bash/setup-evals-analyze.sh`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Agent skills for the Agentic SDLC: team lifecycle (team-boot, team-learn, team-init, team-repair), software factory, evals, CDR lifecycle with confidence scoring, and… The licence is MIT.

When your agent uses it

  • Evaluation results need triage and loop-closing — spec failures route to deterministic checks
  • Generalization failures to the evaluator backlog

Example prompts

  • “/evals-analyze”

Requirements

  • A Bash shell
  • PowerShell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Load Evaluation Results
  2. Failure Classification
  3. Action Routing (Close the Loop)
  4. Cross-Functional Insights & PR

What it can do on your machine

Read from SKILL.md and the folder at commit 2dbed36. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and PowerShell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evals Analyze loads about 1.2k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 554 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tikalk/adlc-team-skills at commit 2dbed36, republished under its MIT licence (© tikalk). 554 words, ~1,194 tokens.

Download SKILL.mdSave it as .claude/skills/evals-analyze/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evals-analyze
description
Use when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.
disable-model-invocation
true

evals-analyze

What this skill does

Provides cross-functional team elevation and closed-loop feedback following EDD Principle VIII (Close the Production Loop) by deep-analyzing trajectory failure traces and routing them to correct resolution pathways.

Output:

  1. Trajectory Analysis - Full multi-turn trace analysis with tool calls and context preservation (EDD Principle V)
  2. Failure Routing:
    • Specification Failures (agent logic missing/ambiguous) → subclassified (EVAL-010):
      • Mechanical (fixed, checkable pattern) → fix/extend the existing grader or add a unit test — a deterministic check, not a CDR
      • Judgment gap (missing intent, ambiguity, needs context) → automatically triggers a local call to team-levelup to propose new context rules in adlc branch drafts/cdr/ to fix agent behavior
    • Generalization Failures (grader flawed or lacks edge-case coverage) → Appends evaluator backlog items to the project backlog for ongoing monitoring.
  3. Cross-Functional PR - Creates a team-ai-directives PR with insights and rule updates (EDD Principle X)

Key EDD Principles Applied:

  • Principle VIII: Close Production Loop - Spec failures → fix directives; Gen failures → evaluator backlog
  • Principle V: Trajectory Observability - Full multi-turn traces, not just outputs
  • Principle X: Cross-Functional Observability - PMs, domain experts, and AI engineers collaborate

When to use

  • After /evals-validate: Analyze failures and resolve them
  • Closing a development loop: Translate evaluation failure insights into rule or evaluator fixes
  • Reporting to stakeholders: Generate readable summaries for PMs and domain experts

When NOT to use

  • Evals not yet executed: Run /evals-validate first to generate results in evals/results/
  • Trivial tasks: Closed-loop analysis is overhead for simple features

Process

User Input
text
$ARGUMENTS
  • --focus AREA — Focus analysis on specific areas (e.g., security, quality, performance)
  • --dry-run — Analyze results and print report, but skip PR creation and local skill triggers
Execution Steps
Phase 1: Load Evaluation Results
  • Reads results JSON from evals/results/.
  • Extracts failure cases and full multi-turn conversation traces (including tool calls).
Phase 2: Failure Classification

Categorizes each failure trace:

  • Specification Failure: The agent was correct relative to its context, but the rule/directive was missing, ambiguous, or incorrect. Subclassify (EVAL-010):
    • Mechanical: the gap is a fixed, checkable pattern (banned API, import shape, file-location, syntactic shape) → deterministic check territory
    • Judgment gap: the gap needs intent, context, or cross-file judgement → context rule territory
  • Generalization Failure: The rule was correct, but the agent made a mistake anyway (hallucinated, missed a constraint, or grader lacked edge-case coverage).
Show full SKILL.md (186 more words)Show less
Phase 3: Action Routing (Close the Loop)
  • For Mechanical Specification Failures: Fix/extend the existing binary grader or add a unit test that enforces the pattern. Do NOT propose a CDR for a mechanically-checkable gap — pay once for the check instead of re-deriving it per session.
  • For Judgment-gap Specification Failures: Automatically triggers local skill /team-levelup with the failure trace as input. This creates new rule/persona/example CDRs in adlc branch drafts/cdr/ to fix the agent's behavior.
  • For Generalization Failures: Appends an evaluator backlog item to evals/results/evaluator_backlog.md detailing the needed grader edge-case updates.
Phase 4: Cross-Functional Insights & PR
  • Generates a stakeholder-specific report in evals/results/team_insights.md (tailored for PMs, domain experts, and AI engineers).
  • If git remote and gh CLI are available, commits rule/eval changes in team-ai-directives and opens a draft PR (uses team-levelup logic under the hood).

Verification

  • Trajectory failure traces analyzed and classified
  • Mechanical specification failures routed to grader/unit-test fixes (deterministic checks); judgment-gap failures routed to /team-levelup (proposes CDRs in adlc branch drafts/cdr/)
  • Generalization failures written to evals/results/evaluator_backlog.md
  • Stakeholder report evals/results/team_insights.md generated
  • Draft PR created in team-ai-directives (if applicable)
  • Final report summary presented with PR link and backlog details

© tikalk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/evals/evals-analyze of tikalk/adlc-team-skills.

  • SKILL.md
  • scripts/bash/setup-evals-analyze.sh
  • scripts/powershell/setup-evals-analyze.ps1

Open the folder on GitHubat commit 2dbed36

Compare with similar skills

Evals Analyze next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evals Analyze compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evals Analyze this skilltikalk/adlc-team-skills141—~1.2kAutomated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from tikalk/adlc-team-skills

All 44 skills in this repo
  • Team Boot

    tikalk/adlc-team-skills

    A skill your agent uses when a session starts or resumes after compaction (auto via the sessionstart and sessioncompact event hooks) and the team AI directives context — constitution, CDR index…

    141 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Architect Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when ADRs need review, gaps need filling, or ADR status must be approved as Accepted before architecture generation.

    141 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Change Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when reviewing, accepting, rejecting, or deferring ChDRs mined by change-init, validating inferred decisions against their git and issue evidence before promotion to project…

    141 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Change Init

    tikalk/adlc-team-skills

    A skill your agent uses when you want guided mining of git history, structured change-story clustering, or comprehensive rationale recovery before documenting.

    141 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Change Publish

    tikalk/adlc-team-skills

    A skill your agent uses when accepted ChDRs are ready for promotion from drafts to project memory at docs/adlc/memory/chdr/ and the boot-facing chdr.md index needs regenerating.

    141 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Evals Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json).

    141 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed

Questions about Evals Analyze

What does Evals Analyze do?

A skill your agent uses when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog. Evals Analyze is an agent skill from tikalk/adlc-team-skills. Use when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.

When should I use Evals Analyze?

Evals Analyze fits situations like: evaluation results need triage and loop-closing — spec failures route to deterministic checks; generalization failures to the evaluator backlog.

How do I install Evals Analyze in Claude Code?

Run `npx skills add tikalk/adlc-team-skills --skill evals-analyze -a claude-code`. Or copy the skill folder (skills/evals/evals-analyze in tikalk/adlc-team-skills) into .claude/skills/evals-analyze in your project. Claude Code loads it when a task matches its description.

How do I install Evals Analyze in Codex?

Run `npx skills add tikalk/adlc-team-skills --skill evals-analyze -a codex`. Or copy the skill folder (skills/evals/evals-analyze in tikalk/adlc-team-skills) into .agents/skills/evals-analyze in your project. Codex loads it when a task matches its description.

Can I use Evals Analyze in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tikalk/adlc-team-skills --skill evals-analyze -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-analyze, .gemini/skills/evals-analyze, .github/skills/evals-analyze and .opencode/skills/evals-analyze in your project.

What does Evals Analyze need to run?

Going by SKILL.md and its folder, Evals Analyze needs a shell and PowerShell for the scripts in its folder. Our summary lists: A Bash shell; PowerShell.

Does Evals Analyze access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evals Analyze safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evals Analyze use?

Evals Analyze is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evals Analyze use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evals Analyze?

Skills that share tags, products or a category with Evals Analyze: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evals Analyze?

tikalk (a GitHub organization) maintains it in tikalk/adlc-team-skills, which has 141 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: tikalk/adlc-team-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.