Agent skill

AI Evaluation Engineering

by devcodex-labs in devcodex-labs/devcodex

AI 评测工程专家 Owner — 当任务涉及模型/Prompt 评测、模型选择、黄金集、评分量表、LLM-as-judge、Judge 校准、重复采样、方差、质量-成本-延迟权衡、提示词回归或模型升级回归时使用;要求把概率性结果转化为可复现、可比较且防污染的评测证据。

AGPL-3.0Auto-check passedAI & LLM Engineering

Install AI Evaluation Engineering

skills CLI
$ npx skills add devcodex-labs/devcodex --skill ai-evaluation-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install devcodex-labs/devcodex ai-evaluation-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/devcodex-labs/devcodex.git skills-src && mkdir -p .claude/skills && cp -r skills-src/content/skills/ai-evaluation-engineering .claude/skills/ai-evaluation-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-evaluation-engineering
GitHub stars
439
Token cost
~434 tokens
SKILL.md length
94 words
Files
3
Skills in repo
70
Repo updated
First seen
Licence
AGPL-3.0

At a glance

AI 评测工程专家 Owner — 当任务涉及模型/Prompt 评测、模型选择、黄金集、评分量表、LLM-as-judge、Judge 校准、重复采样、方差、质量-成本-延迟权衡、提示词回归或模型升级回归时使用;要求把概率性结果转化为可复现、可比较且防污染的评测证据。

  • Works in 7 steps: 冻结 use case、风险等级、失败类型和评测决策用途。 → 建立 train/dev/test 或等价隔离,检查… → 将确定性断言、语义评分、人工评审和业务 outcome 分层,禁止只用单一总分。 → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers 职责, AiEvaluationEngineeringGate, 执行流程 and 输出字段, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

AI Evaluation Engineering is an agent skill from devcodex-labs/devcodex. AI 评测工程专家 Owner — 当任务涉及模型/Prompt 评测、模型选择、黄金集、评分量表、LLM-as-judge、Judge 校准、重复采样、方差、质量-成本-延迟权衡、提示词回归或模型升级回归时使用;要求把概率性结果转化为可复现、可比较且防污染的评测证据。

Its SKILL.md is about 430 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `agents/openai.yaml` and `intent.json`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Intent-driven AI coding workflow runtime for consistent context, skills, approvals, validation, and handoffs across six AI coding hosts. The licence is AGPL-3.0.

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “/ai-evaluation-engineering”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. 冻结 use case、风险等级、失败类型和评测决策用途。
  2. 建立 train/dev/test 或等价隔离,检查 benchmark、Prompt 和检索语料污染。
  3. 将确定性断言、语义评分、人工评审和业务 outcome 分层,禁止只用单一总分。
  4. 校准 Judge:随机顺序、隐藏候选身份、人工样本对照、位置/长度/风格偏差。
  5. 对概率性路径重复采样,报告方差和尾部失败,不用单次成功代表稳定。
  6. 同时测质量、成本、延迟和工具/结构化输出正确性。
  7. 用固定版本 manifest 对比 baseline/candidate,达到阈值才进入发布或模型切换。

What it can do on your machine

Read from SKILL.md and the folder at commit 1dd4525. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Evaluation Engineering loads about 434 tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 94 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~434

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from devcodex-labs/devcodex at commit 1dd4525, republished under its AGPL-3.0 licence (© devcodex-labs). 94 words, ~434 tokens.

Download SKILL.mdSave it as .claude/skills/ai-evaluation-engineering/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
ai-evaluation-engineering
description
AI 评测工程专家 Owner — 当任务涉及模型/Prompt 评测、模型选择、黄金集、评分量表、LLM-as-judge、Judge 校准、重复采样、方差、质量-成本-延迟权衡、提示词回归或模型升级回归时使用;要求把概率性结果转化为可复现、可比较且防污染的评测证据。

AI Evaluation Engineering

职责

设计概率性 AI 系统的评测数据、指标、Judge、重复采样、方差和回归决策。AI Agent Skill 负责系统行为,quality-strategy 负责整体测试组合;本 Skill 负责模型/Prompt 质量证据。

AiEvaluationEngineeringGate

字段要求
evaluationDatasetManifest来源、版本、许可/隐私、任务分层、难例、污染风险和 split
goldenCaseSet输入、期望属性/答案、允许变体、失败标签和维护 owner
metricRubricdeterministic/semantic/human 指标、权重、阈值和不可聚合项
judgeCalibrationJudge 模型/Prompt/版本、盲测、与人工一致性、偏差和漂移
samplingProtocoltemperature/seed、重复次数、置信区间、停止规则和失败重试
varianceReport均值、分布、尾部失败、跨 run/provider 差异和不确定性
costLatencyQualityFrontiertoken/费用/延迟/成功率/质量的 Pareto 权衡
regressionDecisionbaseline/candidate、显著性、阻断阈值、例外和 rollback

执行流程

  1. 冻结 use case、风险等级、失败类型和评测决策用途。
  2. 建立 train/dev/test 或等价隔离,检查 benchmark、Prompt 和检索语料污染。
  3. 将确定性断言、语义评分、人工评审和业务 outcome 分层,禁止只用单一总分。
  4. 校准 Judge:随机顺序、隐藏候选身份、人工样本对照、位置/长度/风格偏差。
  5. 对概率性路径重复采样,报告方差和尾部失败,不用单次成功代表稳定。
  6. 同时测质量、成本、延迟和工具/结构化输出正确性。
  7. 用固定版本 manifest 对比 baseline/candidate,达到阈值才进入发布或模型切换。

输出字段

evaluationDatasetManifest、goldenCaseSet、metricRubric、judgeCalibration、samplingProtocol、varianceReport、costLatencyQualityFrontier、regressionDecision、contaminationCheck、evidenceMatrix。

反模式

  • 用少量“看起来不错”的示例替代评测集。
  • Judge 未校准、知道候选身份或与被评模型同一偏差源。
  • 只跑一次、不报方差,却给出稳定性结论。
  • 把模型名或主观偏好当质量证据。
  • 只看质量分,不披露成本、延迟、结构化输出和工具失败。
  • 在调 Prompt 时反复看 test set,造成隐性污染。
  • 复制历史分数而不锁定 dataset/model/prompt/tool versions。

验证

至少覆盖确定性与概率性双轨、重复采样、Judge 与人工校准、position/verbosity bias、数据污染、provider fallback、工具调用/JSON 合法性、成本延迟预算、版本回归和 inconclusive 路径。

© devcodex-labs, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in content/skills/ai-evaluation-engineering of devcodex-labs/devcodex.

  • SKILL.md
  • agents/openai.yaml
  • intent.json

Open the folder on GitHubat commit 1dd4525

Compare with similar skills

AI Evaluation Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Evaluation Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Evaluation Engineering this skilldevcodex-labs/devcodex439—~434Automated safety check: PassAGPL-3.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from devcodex-labs/devcodex

All 70 skills in this repo
  • Accessibility I18n

    devcodex-labs/devcodex

    无障碍与国际化专家 Owner — 当任务涉及可访问性、键盘操作、焦点、屏幕阅读器、ARIA、语言地区、本地化、RTL、翻译资源、用户可见文案或多语言文档时使用;要求把包容性体验和本地化验证绑定到真实用户路径。

    439 GitHub stars~718 tokensUpdated 21 days ago
    Auto-check passed
  • AI Agent System Architecture

    devcodex-labs/devcodex

    AI Agent 系统架构专家 Owner — 当任务涉及 Agent 路由、工具调用、上下文管理、记忆、状态机、权限、人机协作、可观测性、回放验证或模型辅助治理时使用;要求把 Agent 行为设计成可解释、可恢复、可审计。

    439 GitHub stars~2.4k tokensUpdated 21 days ago
    Auto-check passed
  • API Contract Architecture

    devcodex-labs/devcodex

    API 契约架构专家 Owner — 当任务涉及 public API、HTTP/SDK/CLI 契约、版本兼容、错误模型、分页过滤、幂等、Schema、类型、迁移或消费者影响时使用;要求先冻结消费者契约,再设计实现与验证。

    439 GitHub stars~865 tokensUpdated 21 days ago
    Auto-check passed
  • Architecture Design

    devcodex-labs/devcodex

    架构设计文档编排 Owner — 当用户要求架构设计、系统设计、技术架构或可指导开发、Review 与任务拆分的完整方案时使用;要求从业务流程反推节点、状态、数据、一致性、异常补偿、ADR 与实施任务。

    439 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check passed
  • Audit Common

    devcodex-labs/devcodex

    审查公共维度 G0~G5 + Profile Freshness Check — 所有 audit 子类型必先执行的基础维度层

    439 GitHub stars~4.1k tokensUpdated 21 days ago
    Auto-check passed
  • Audit Session

    devcodex-labs/devcodex

    审计工作流的跨会话状态机 — 在 <audit-root/.audit-state/<session-id.json 持久化轮次/发现项/收敛状态,支持 Token 中断后精准恢复

    439 GitHub stars~1.8k tokensUpdated 21 days ago
    Auto-check passed

Questions about AI Evaluation Engineering

What does AI Evaluation Engineering do?

AI 评测工程专家 Owner — 当任务涉及模型/Prompt 评测、模型选择、黄金集、评分量表、LLM-as-judge、Judge 校准、重复采样、方差、质量-成本-延迟权衡、提示词回归或模型升级回归时使用;要求把概率性结果转化为可复现、可比较且防污染的评测证据。. AI Evaluation Engineering is an agent skill from devcodex-labs/devcodex.

When should I use AI Evaluation Engineering?

AI Evaluation Engineering fits situations like: tasks that involve LLM evaluation.

How do I install AI Evaluation Engineering in Claude Code?

Run `npx skills add devcodex-labs/devcodex --skill ai-evaluation-engineering -a claude-code`. Or copy the skill folder (content/skills/ai-evaluation-engineering in devcodex-labs/devcodex) into .claude/skills/ai-evaluation-engineering in your project. Claude Code loads it when a task matches its description.

How do I install AI Evaluation Engineering in Codex?

Run `npx skills add devcodex-labs/devcodex --skill ai-evaluation-engineering -a codex`. Or copy the skill folder (content/skills/ai-evaluation-engineering in devcodex-labs/devcodex) into .agents/skills/ai-evaluation-engineering in your project. Codex loads it when a task matches its description.

Can I use AI Evaluation Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add devcodex-labs/devcodex --skill ai-evaluation-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-evaluation-engineering, .gemini/skills/ai-evaluation-engineering, .github/skills/ai-evaluation-engineering and .opencode/skills/ai-evaluation-engineering in your project.

What does AI Evaluation Engineering need to run?

SKILL.md names no scripts, command-line tools or credentials: AI Evaluation Engineering is instructions for the agent only.

Does AI Evaluation Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Evaluation Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Evaluation Engineering use?

AI Evaluation Engineering is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Evaluation Engineering use?

About 434 tokens (SKILL.md is roughly 1.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Evaluation Engineering?

Skills that share tags, products or a category with AI Evaluation Engineering: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Evaluation Engineering?

devcodex-labs (a GitHub organization) maintains it in devcodex-labs/devcodex, which has 439 GitHub stars. The repository holds 70 skills in this directory. The repository was last updated on September 17, 2026.

Source: devcodex-labs/devcodex on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.