Agent skill

Evals Init

by tikalk in tikalk/adlc-team-skills

A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline.

MITAuto-check passedAI & LLM Engineering

Install Evals Init

skills CLI
$ npx skills add tikalk/adlc-team-skills --skill evals-init -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tikalk/adlc-team-skills evals-init --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tikalk/adlc-team-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evals/evals-init .claude/skills/evals-init && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evals-init
GitHub stars
141
Token cost
~934 tokens
SKILL.md length
300 words
Files
3 (incl. scripts)
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline.

  • Works in 4 steps: Tech Stack Detection → Create Directory Structure → Configuration Copy → …
  • Standing up evals/{system}/ for the first time — scaffolds the EDD directory structure
  • SKILL.md covers What this skill does, When to use, When NOT to use and Process, plus 1 more section
  • Runs Shell and PowerShell scripts from its folder

What it does

Evals Init is an agent skill from tikalk/adlc-team-skills. Use when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline.

Its SKILL.md is about 930 tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/bash/setup-evals-init.sh`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Agent skills for the Agentic SDLC: team lifecycle (team-boot, team-learn, team-init, team-repair), software factory, evals, CDR lifecycle with confidence scoring, and… The licence is MIT.

When your agent uses it

  • Standing up evals/{system}/ for the first time — scaffolds the EDD directory structure
  • Picks PromptFoo
  • DeepEval by tech stack
  • Generates a security baseline

Example prompts

  • “/evals-init”

Requirements

  • Python 3
  • A Bash shell
  • PowerShell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Tech Stack Detection
  2. Create Directory Structure
  3. Configuration Copy
  4. Auto-Handoff

What it can do on your machine

Read from SKILL.md and the folder at commit 2dbed36. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and PowerShell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evals Init loads about 934 tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 300 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~934

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tikalk/adlc-team-skills at commit 2dbed36, republished under its MIT licence (© tikalk). 300 words, ~934 tokens.

Download SKILL.mdSave it as .claude/skills/evals-init/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evals-init
description
Use when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline.
disable-model-invocation
true

evals-init

What this skill does

Initialize the project-level evaluation directory structure following EDD (Eval-Driven Development) principles to prepare for systematic evaluation development. This is completely standalone with zero spec-kit dependencies.

Output:

  1. Directory Structure - evals/{system}/ with proper organization (promptfoo | deepeval)
  2. Security Baseline - Auto-created graders for PII leakage, prompt injection, hallucination detection, misinformation detection
  3. Configuration Files - Standalone config.yml and goldset templates under .adlc/evals/
  4. Auto-handoff to /evals-specify to begin error analysis

Key EDD Principles Applied:

  • Principle I: Spec-Driven Contracts - Evals validate spec compliance
  • Principle II: Binary Pass/Fail - No Likert scales in grader templates
  • Principle IV: Evaluation Pyramid - Tier 1 (fast) + Tier 2 (goldset) structure
  • Principle IX: Test Data as Code - Version control setup for datasets

When to use

  • Starting systematic evaluation: Set up the initial evaluation harness for your application
  • EDD Adoption: Converting from traditional testing to evaluation-driven development
  • Security-first evaluation: Auto-generate baseline security checks from the start

When NOT to use

  • Evals directory already exists: Use /evals-validate to run tests, or /evals-specify to add criteria
  • Evaluating team directives: This is for project-level application behavior testing, not directives compliance

Process

User Input
text
$ARGUMENTS

Parse flags from the arguments first, then treat remaining text as focus areas:

  • --system SYSTEM — Choose promptfoo or deepeval. If omitted, choose interactively based on tech stack.
  • Remaining text — System description (focus setup)
Execution Steps
Phase 1: Tech Stack Detection
  • Scan project manifests (package.json, requirements.txt, Cargo.toml, go.mod, etc.)
  • Recommends PromptFoo for mixed/JS stacks; DeepEval for Python-native stacks
Phase 2: Create Directory Structure

Creates:

evals/
├── {system}/                    # promptfoo | deepeval
│   ├── goldset.md              # Published goldset
│   ├── goldset.json            # Auto-generated for system consumption
│   ├── config.yml              # System-specific configuration
│   ├── config.{js,py}          # Generated system config (.js for promptfoo, .py for deepeval)
│   └── graders/                # Binary pass/fail graders
│       ├── check_pii_leakage.py           # Security baseline
│       ├── check_prompt_injection.py     # Security baseline
│       ├── check_hallucination.py        # Security baseline
│       └── check_misinformation.py       # Security baseline
├── results/                    # Git-ignored run outputs
└── .adlc/
    └── drafts/evals/           # Draft eval records (Markdown + YAML)
Phase 3: Configuration Copy
  • Create .adlc/evals/ if missing.
  • Copy skills/evals/evals-templates/evals-config-template.yml to .adlc/evals/evals-config.yml.
Phase 4: Auto-Handoff

Trigger /evals-specify to begin error analysis.

Verification

  • evals/{system}/goldset.md exists (initially empty)
  • .adlc/evals/evals-config.yml exists
  • Graders directory populated with 4 security baseline python scripts
  • Results directory contains .gitignore to prevent versioning traces
  • Handover report generated with recommended framework and next steps

© tikalk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/evals/evals-init of tikalk/adlc-team-skills.

  • SKILL.md
  • scripts/bash/setup-evals-init.sh
  • scripts/powershell/setup-evals-init.ps1

Open the folder on GitHubat commit 2dbed36

Compare with similar skills

Evals Init next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evals Init compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evals Init this skilltikalk/adlc-team-skills141—~934Automated safety check: PassMIT
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from tikalk/adlc-team-skills

All 44 skills in this repo
  • Workspace

    tikalk/adlc-team-skills

    A skill your agent uses when coordinating a multi-repo workspace — init the .adlc/ structure, discover and link child repos as submodules, or audit workspace health (branch, dirty, unpushed, SHA…

    141 GitHub starsUsed in 1 repo~3.7k tokens
    Auto-check passed
  • Team Boot

    tikalk/adlc-team-skills

    A skill your agent uses when a session starts or resumes after compaction (auto via the sessionstart and sessioncompact event hooks) and the team AI directives context — constitution, CDR index…

    141 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Architect Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when ADRs need review, gaps need filling, or ADR status must be approved as Accepted before architecture generation.

    141 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Change Clarify

    tikalk/adlc-team-skills

    A skill your agent uses when reviewing, accepting, rejecting, or deferring ChDRs mined by change-init, validating inferred decisions against their git and issue evidence before promotion to project…

    141 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Change Init

    tikalk/adlc-team-skills

    A skill your agent uses when you want guided mining of git history, structured change-story clustering, or comprehensive rationale recovery before documenting.

    141 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check passed
  • Change Publish

    tikalk/adlc-team-skills

    A skill your agent uses when accepted ChDRs are ready for promotion from drafts to project memory at docs/adlc/memory/chdr/ and the boot-facing chdr.md index needs regenerating.

    141 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Questions about Evals Init

What does Evals Init do?

A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline. Evals Init is an agent skill from tikalk/adlc-team-skills. Use when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline.

When should I use Evals Init?

Evals Init fits situations like: standing up evals/{system}/ for the first time — scaffolds the EDD directory structure; picks PromptFoo; deepEval by tech stack; generates a security baseline.

How do I install Evals Init in Claude Code?

Run `npx skills add tikalk/adlc-team-skills --skill evals-init -a claude-code`. Or copy the skill folder (skills/evals/evals-init in tikalk/adlc-team-skills) into .claude/skills/evals-init in your project. Claude Code loads it when a task matches its description.

How do I install Evals Init in Codex?

Run `npx skills add tikalk/adlc-team-skills --skill evals-init -a codex`. Or copy the skill folder (skills/evals/evals-init in tikalk/adlc-team-skills) into .agents/skills/evals-init in your project. Codex loads it when a task matches its description.

Can I use Evals Init in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tikalk/adlc-team-skills --skill evals-init -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-init, .gemini/skills/evals-init, .github/skills/evals-init and .opencode/skills/evals-init in your project.

What does Evals Init need to run?

Going by SKILL.md and its folder, Evals Init needs a shell and PowerShell for the scripts in its folder. Our summary lists: Python 3; A Bash shell; PowerShell.

Does Evals Init access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evals Init safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Evals Init use?

Evals Init is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evals Init use?

About 934 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evals Init?

Skills that share tags, products or a category with Evals Init: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evals Init?

tikalk (a GitHub organization) maintains it in tikalk/adlc-team-skills, which has 141 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: tikalk/adlc-team-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.