Official agent skill

Agent Eval Engineering

by langchain-ai in langchain-ai/langchain-skills

Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

OfficialMITAuto-check passedAI & LLM Engineering

Install Agent Eval Engineering

skills CLI
$ npx skills add langchain-ai/langchain-skills --skill eval-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install langchain-ai/langchain-skills eval-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/langchain-ai/langchain-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/config/skills/eval-engineering .claude/skills/eval-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-engineering
GitHub stars
1.3k
Token cost
~4k tokens
SKILL.md length
1,960 words
Files
22 (incl. scripts, references, assets)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

  • Works in 7 steps: Inspect inputs and existing World… → Propose and select a Task → Write and review the Task Spec and World… → …
  • Designing a benchmark or eval suite for a new agent
  • SKILL.md covers Flow, Terms, Reference routing and 1. Inspect inputs and existing…, plus 8 more sections
  • Runs Python scripts from its folder

What it does

Evaluation work starts with inspection of everything that bears on the agent under test: its repository and harness, optional traces, existing tasks and runs, and your goal. From those facts the agent creates a small project-specific World Knowledge Skill and uses it to propose one grounded task. The task spec (`Task.md`) and the world skill are drafted together, shown to you and refined until you approve both.

An approved task is then implemented with its environment (data, services, permissions, state and reset) and a verifier that scores independently, run through the real harness, and the full evidence is inspected so that only failures not caused by the agent get fixed. The world skill is updated with what the run proved before the next task. References cover discovery, environment building, synthetic data, verifier design, calibration, Harbor and multi-turn user simulation, with a service-desk example and templates.

When your agent uses it

  • Designing a benchmark or eval suite for a new agent
  • Turning production traces into reviewed, runnable evaluation tasks
  • Building controlled environments and synthetic data for an agent eval
  • Calibrating verifiers and maintaining a benchmark as the agent changes

Example prompts

  • “Look at this agent repo and its traces, then propose the first eval task and write its Task.md for my review.”
  • “Build the environment and verifier for the approved refund-request task and run it with Harbor.”
  • “Add a multi-turn simulated user to the service desk eval.”

Requirements

  • Harbor, to build and run tasks
  • The agent repository, plus traces if available

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Inspect inputs and existing World knowledge
  2. Propose and select a Task
  3. Write and review the Task Spec and World Skill
  4. Apply Spec2Task
  5. Run and audit
  6. Reconcile project World knowledge
  7. Repeat

What it can do on your machine

Read from SKILL.md and the folder at commit 16a992f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Eval Engineering loads about 4k tokens when it runs, and up to ~36k if it reads all its reference files. Until then it costs about 92 tokens; SKILL.md has 1,960 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~36k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from langchain-ai/langchain-skills at commit 16a992f, republished under its MIT licence (© langchain-ai). 1,960 words, ~3,966 tokens.

Download SKILL.mdSave it as .claude/skills/eval-engineering/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.
name
eval-engineering
description
Inspect an agent repository and optional traces, interview the user, write reviewed Task Specs, build and audit Harbor tasks, and bootstrap reusable project World Knowledge Skills. Use for agent evals, benchmark design, Task generation, controlled Environments, synthetic data, Verifiers, Harbor runs, calibration, or continuous benchmark maintenance.

Eval Engineering

Flow

  • Inspect all inputs first: the repository, Harness, optional traces, existing Tasks and runs, existing World knowledge, and the human goal. Identify the source files and skill references that apply before proposing work.
  • Create or update the small project World Knowledge Skill from reusable facts in those inputs. Use it to propose one grounded Task.
  • Draft the Task Spec and World Skill together. Show both to the user, keep exact Task truth only in Task.md, and refine both until the user approves them.
  • Implement the approved Task, validate its Environment and Verifier, run the real Harness, inspect the full evidence, and fix only non-agent failures.
  • Reconcile the World Skill with what the run proved, then repeat this flow for the next Task.

Terms

  • Task Spec: Task.md, which describes the input, relevant agent conditions, Environment, scoring, fairness, and open decisions for one Task.
  • Task: the runnable instruction, Environment, and Verifier.
  • Harness: the complete agent Harbor runs, including prompts, model loop, tools, hooks, memory, sessions, and adapter.
  • Environment: the files, data, services, identity, permissions, network, clock, and mutable state around the Harness.
  • Verifier: independent checks that score the result or mark a run invalid.
  • World Knowledge Skill: a repository-local skill with reusable project-specific knowledge, references, scripts, assets, and tests that help generate Task Specs and build future Tasks.
  • Spec2Task: the full loop that turns a reviewed Task Spec into an audited runnable Task. Follow Task implementation for its concrete build order and reference routing.

Reference routing

Read each reference when its decision appears:

NeedRead
Inspect source, traces, the Harness, dependencies, access, and existing evalsDiscovery
Bootstrap or update reusable project knowledgeWorld knowledge
Propose Tasks and write the single Task SpecTask design
Build data, services, access, state, and resetEnvironment building
Create structured or natural-language dataSynthetic data
Define independent evidence and scoringVerifier design
Apply Spec2Task to turn a reviewed Spec into an audited TaskTask implementation
Compare model runs and classify failuresCalibration
Package and run Harbor tasksHarbor
Adapt a known benchmark designBenchmark patterns
Build multi-turn conversationsMulti-turn simulation
See World knowledge learned across two TasksService-desk example

Reusable implementation resources:

1. Inspect inputs and existing World knowledge

Review every input the user provides before proposing a Task. Use the guidance that matches each available input:

  • For a repository, Harness, traces, or dependencies, read Discovery.
  • For existing Tasks and runs, inspect their instructions, Environments, Verifiers, rewards, trajectories, and final state. Read Calibration when run quality or failure causes affect the new design.
  • For an existing project World Skill, read World knowledge, then check the sources and reusable methods that affect the new Task.
  • For human goals and constraints, read Task design.
  • For relevant benchmark examples, use the domain index and source callouts in Benchmark patterns.

Inspect the repository before asking questions that source and tests can answer. Follow the active Harness through prompts, models, tools, services, state, effects, and focused tests. Inspect existing Task instructions, parsers, Verifiers, reward paths, and run evidence.

If the user supplies traces, review complete runs or threads. Use traces to learn real requests, dependency behavior, state shapes, errors, and failure conditions. Do not treat a trace answer as independent truth.

If .agents/skills/<project>-world/SKILL.md exists, read it. Follow its routing only for knowledge relevant to the current Task. Check cited repository paths, commands, and scripts when their accuracy affects the design.

2. Propose and select a Task

Read Task design and use the index in Benchmark patterns to find the relevant domain and source callouts. Focus on that domain unless the Task crosses another one. In the first user-facing design response after inspection, propose one Task grounded in repository evidence, supplied traces, existing coverage, or a human priority. State:

  • the real work and capability;
  • the condition that makes the case non-trivial;
  • the Environment and independent evidence it needs;
  • the important failure it can detect;
  • how it differs from existing Tasks; and
  • the main open decision.

In the same response, show the relevant current World Skill content and the specific additions or corrections this Task suggests. If no World Skill exists, show the small initial contents that will help create this Task and future Tasks. Keep the Task's exact request, focal records, expected result, hidden truth, and exact scoring rules out of the World Skill.

Let the user revise the Task proposal and World knowledge together before implementation. Offer alternatives only when a real user choice changes the design.

3. Write and review the Task Spec and World Skill

Copy the Task template to evals/<suite>/tasks/<task-id>/Task.md. Put all Task-specific design in this one file. At the same time, create or update the project World Skill by following World knowledge. Determine the project skill location supported by the active agent and repository. .agents/skills/<project>-world/SKILL.md and .claude/skills/<project>-world/SKILL.md are common landing spots. Follow an established project convention when one exists. Otherwise, explain the proposed location and get user confirmation before creating the skill. Start from the World Skill template when needed.

Keep each Task.md beside the Harbor task it describes:

text
evals/<suite>/tasks/<task-id>/
├── Task.md              # human-reviewed control-plane spec
├── task.toml             # required Harbor configuration
├── instruction.md        # required agent input
├── environment/          # required Environment definition and visible state
│   ├── Dockerfile        # use this or docker-compose.yaml
│   └── docker-compose.yaml # optional; primary service must be main
├── tests/
│   ├── test.sh           # required Harbor Verifier entry point
│   ├── test_*.py         # optional Verifier helpers
│   └── fixtures/         # optional hidden Verifier data
└── solution/
    └── solve.sh          # optional reference path

Never copy or mount Task.md into the evaluated agent's workspace or image. The agent receives instruction.md and only the Environment state intended for the run.

Include:

  • purpose and source evidence;
  • exact input and later turns;
  • only the Harness conditions relevant to this Task;
  • initial state, services, access, visibility, reset, and production differences;
  • required results, prohibited effects, accepted alternatives, and independent Verifier evidence;
  • fairness, leakage risks, and invalid-run conditions; and
  • open decisions and assumptions.

Show the full Task Spec and the World Skill changes to the user. Explain what is already in the World Skill, what this Task adds or corrects, and what stays only in Task.md. Revise both through the same back-and-forth. Mark the Task Spec approved only after explicit approval. Treat World Skill changes as accepted only after the user reviews them. If the user requests an end-to-end build without an approval pause, continue with an agent-reviewed Status: Draft and label the World Skill changes as unreviewed.

If implementation changes the request, visible information, material Environment behavior, or scoring boundary, update Task.md and show the change. Set its status back to Draft. Show the diff and require explicit reapproval before setting it to Approved again.

Show full SKILL.md (905 more words)Show less

4. Apply Spec2Task

Follow Task implementation. It gives the build order and routes each decision to the Environment, synthetic-data, Verifier, Harbor, and calibration references.

For an existing project, use its pinned or supported Harbor version. Otherwise, use the installed supported version and record it. Upgrade only with user approval and a stated compatibility reason. Use the installed CLI help as the command contract.

Before a scored model run:

  1. Confirm the model, trial count, judge, timeout, and maximum expected cost with the user unless the user already authorized that run plan.
  2. Complete the package audit in Harbor. Confirm every required file, entry point, path, permission, configuration value, mount, service, and reward output needed for this exact Task is present and works through Harbor. Confirm hidden Task, Verifier, solution, and secret material is absent from the agent-visible image and workspace.
  3. Check setup and trial isolation in the way that fits the Environment. A fresh container or worktree can provide isolation by replacement. A reused mutable service needs a checked reset. Immutable frozen data needs only a checked load. See Task implementation.
  4. Exercise every operation the Task depends on.
  5. Run the reference path when one exists.
  6. Test the Verifier with a clear valid result, a valid alternative, a realistic wrong result, a shortcut, a prohibited collateral change, and missing or corrupt evidence.
  7. Confirm every completed Verifier path writes a valid reward and useful evidence without exposing hidden truth or secrets.

5. Run and audit

Run the actual Harness through Harbor. Read the complete trajectory, not only the reward. Inspect:

  • messages, model calls, tool calls, results, retries, and errors;
  • initial and final Environment state and external effects;
  • service, setup, readiness, reset, and cleanup evidence;
  • each Verifier criterion, its evidence, decision, and error; and
  • the resolved Harness, model, Environment, and judge configuration.

Classify each unsuccessful run as an agent capability failure, missing information, Harness defect, Environment defect, Verifier false rejection, Verifier false acceptance, leakage, or infrastructure failure. Fix non-agent failures before using the score.

Model comparison is an optional calibration strategy, not a completion rule. When it would answer a real uncertainty, compare a weaker model, the target model, or a stronger model and repeat trials when behavior is variable. Read every selected trace. Contrast can expose unclear inputs, brittle setup, leakage, shortcuts, or reward hacks. Pass rates and model ordering do not prove Task quality.

Read Calibration for the complete audit method.

6. Reconcile project World knowledge

Use World knowledge throughout Task design, implementation, and audit. Add or correct project-specific knowledge when the work supplies evidence that would help another Task. This can include Task patterns, Environment methods, data creation, Verifier evidence, run procedures, scripts, assets, and examples.

After the audit, reconcile the World Skill with what the completed Task proved. Show the user:

  • the proposed reusable knowledge;
  • the evidence supporting it;
  • how another Task would use it;
  • where it should live; and
  • what remains specific to the completed Task.

Remove or narrow ideas that the Task disproved. If the user asked for autonomous end-to-end updates without a pause, make the smallest supported update, show it in the final review, and do not imply that the human approved the generalization.

Create only SKILL.md at first. Add references/, scripts/, assets/, or tests/ only when their real contents justify them.

Keep the completed Task's request, focal state, expected result, and exact criteria in its collocated Task.md. Do not copy broad guidance that is already clear in this skill. Record the project-specific adaptation of that guidance.

7. Repeat

Use Tasks two and three to test the World Skill. Check whether it reduces rediscovery, improves Task Specs, preserves important relationships, reuses a proven operation, or prevents a known Verifier defect. Correct rules that are missing, stale, or too broad.

When several materially different Tasks have exercised the shared knowledge and the construction and verification methods are clear, the next cycle can propose several independent Task Specs:

  1. Mine new repository, trace, and human evidence.
  2. Use the World Skill to generate distinct Task Specs.
  3. Have the human review the specs.
  4. Build independent approved Tasks in parallel.
  5. Audit every Task individually.
  6. Update World knowledge only with reusable corrections.

Continue this loop as production behavior, user priorities, agents, and models change.

Dependencies, access, and safety

Map required systems, data, roles, network needs, and safe setup methods. Never read, print, copy, store, or ask the human to paste secret values. Tell the human what dependency is needed, why it is needed, and how the project expects access to be provided. Default to controlled local, frozen, or simulated dependencies. Never write to production during an eval. Treat access, startup, reset, timeout, judge, and Verifier failures as invalid runs, not failed agent work.

Complete only when

  • Task.md matches the built instruction, Environment, and Verifier.
  • The Task is solvable from agent-visible or normally discoverable information.
  • The Environment starts reliably and isolates trials by replacement, reset, or immutable state as appropriate, without leaking hidden truth.
  • Valid and invalid Verifier cases behave as intended.
  • At least one real Harness run was read in full.
  • Non-agent failures were repaired or reported as unresolved limits.
  • Model comparison, when used, includes trace review rather than pass rates alone.
  • The project World Skill was created or updated with the Task Spec, reviewed during the work, and reconciled with the final evidence.
  • The user receives the Task path, run command, results, evidence, and remaining limits.

© langchain-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 21 other files (scripts, references, assets) in config/skills/eval-engineering of langchain-ai/langchain-skills.

  • SKILL.md
  • agents/openai.yaml
  • assets/task/Task.md.template
  • assets/world-skill/SKILL.md.template
  • references/calibration.md
  • references/discovery.md
  • references/environment-building.md
  • references/examples/service-desk.md
  • references/harbor.md
  • references/multi-turn-simulation/guide.md
  • references/multi-turn-simulation/harbor_example.py
  • references/multi-turn-simulation/model_user.py
  • references/multi-turn-simulation/runner.py
  • references/patterns.md
  • … and 8 more

Open the folder on GitHubat commit 16a992f

Compare with similar skills

Agent Eval Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Eval Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Eval Engineering this skilllangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Synthetic Eval Data Generatorai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.0
GAIA Agent Benchmarkingamd/gaia1.6k—~1.8kAutomated safety check: PassMIT
Chatbox Session RAG Evalchatboxai/chatbox42k—~758Automated safety check: PassGPL-3.0
Windmill AI Evalswindmill-labs/windmill18k—~969Automated safety check: NotesCustom licence
Octocode Benchmark Runnerbgauryy/octocode946—~2.1kAutomated safety check: PassMIT

Similar skills

  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Chatbox Session RAG Eval

    chatboxai/chatbox

    Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.

    42k GitHub stars~758 tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Windmill AI Evals

    windmill-labs/windmill

    Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

    18k GitHub stars~969 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    946 GitHub stars~2.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

    305 GitHub stars~497 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from langchain-ai/langchain-skills

All 22 skills in this repo
  • Langgraph Human In The Loop

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when implementing human-in-the-loop patterns, pausing for approval, or handling errors in LangGraph.

    1.3k GitHub starsUsed in 2 repos~4.1k tokens
    Auto-check passed
  • Swarm Parallel Dispatch

    langchain-ai/langchain-skills

    Official

    Fans a list of independent items out to subagents in parallel, merges the results back into a table and supports retrying only the rows that failed.

    1.3k GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed
  • LangGraph Decision Models

    langchain-ai/langchain-skills

    Official

    Routes LangGraph agents with typed decision models that return probabilities, and finds LLM calls that only exist to produce a routing decision.

    1.3k GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Deep Agents Core

    langchain-ai/langchain-skills

    Official

    Explains how to build agents with the Deep Agents framework: create_deep_agent, the built-in middleware, the harness, SKILL.md format and configuration options.

    1.3k GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Langchain Dependencies

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when setting up a new project or when asked about package versions, installation, or dependency management for LangChain, LangGraph, LangSmith, or Deep Agents.

    1.3k GitHub starsUsed in 1 repo~3.6k tokens
    Auto-check passed
  • Langchain Middleware

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when you need human-in-the-loop approval, custom middleware, or structured output.

    1.3k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Questions about Agent Eval Engineering

What does Agent Eval Engineering do?

Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier. Evaluation work starts with inspection of everything that bears on the agent under test: its repository and harness, optional traces, existing tasks and runs, and your goal. From those facts the agent creates a small project-specific World Knowledge Skill and uses it to propose one grounded task.

When should I use Agent Eval Engineering?

Agent Eval Engineering fits situations like: designing a benchmark or eval suite for a new agent; turning production traces into reviewed, runnable evaluation tasks; building controlled environments and synthetic data for an agent eval; calibrating verifiers and maintaining a benchmark as the agent changes.

How do I install Agent Eval Engineering in Claude Code?

Run `npx skills add langchain-ai/langchain-skills --skill eval-engineering -a claude-code`. Or copy the skill folder (config/skills/eval-engineering in langchain-ai/langchain-skills) into .claude/skills/eval-engineering in your project. Claude Code loads it when a task matches its description.

How do I install Agent Eval Engineering in Codex?

Run `npx skills add langchain-ai/langchain-skills --skill eval-engineering -a codex`. Or copy the skill folder (config/skills/eval-engineering in langchain-ai/langchain-skills) into .agents/skills/eval-engineering in your project. Codex loads it when a task matches its description.

Can I use Agent Eval Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add langchain-ai/langchain-skills --skill eval-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-engineering, .gemini/skills/eval-engineering, .github/skills/eval-engineering and .opencode/skills/eval-engineering in your project.

What does Agent Eval Engineering need to run?

Going by SKILL.md and its folder, Agent Eval Engineering needs Python for the scripts in its folder. Our summary lists: Harbor, to build and run tasks; The agent repository, plus traces if available.

Does Agent Eval Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Eval Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Eval Engineering use?

Agent Eval Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Eval Engineering use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 32k tokens, read only when the agent opens those files.

What are the alternatives to Agent Eval Engineering?

Skills that share tags, products or a category with Agent Eval Engineering: Synthetic Eval Data Generator (ai-evals-course/evals-skills, 1.5k stars), GAIA Agent Benchmarking (amd/gaia, 1.6k stars), Chatbox Session RAG Eval (chatboxai/chatbox, 42k stars) and Windmill AI Evals (windmill-labs/windmill, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Eval Engineering?

langchain-ai (a GitHub organization, an official publisher) maintains it in langchain-ai/langchain-skills, which has 1,270 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 5, 2026.

Source: langchain-ai/langchain-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.