Evaluation
PrimeIntellect-ai/prime-envs
Install and run a verifiers environment — smoke testing during development and full benchmark evals.
Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add github/awesome-copilot --skill eval-driven-dev -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install github/awesome-copilot eval-driven-dev --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-driven-dev .claude/skills/eval-driven-dev && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-driven-dev" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/eval-driven-dev into .claude/skills/eval-driven-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-driven-dev", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/github/awesome-copilot/tree/main/skills/eval-driven-devType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add github/awesome-copilot --skill eval-driven-dev -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install github/awesome-copilot eval-driven-dev --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval-driven-dev .agents/skills/eval-driven-dev && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-driven-dev" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/eval-driven-dev into .agents/skills/eval-driven-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-driven-dev", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill eval-driven-dev -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install github/awesome-copilot eval-driven-dev --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval-driven-dev .cursor/skills/eval-driven-dev && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-driven-dev" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/eval-driven-dev into .cursor/skills/eval-driven-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-driven-dev", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/github/awesome-copilot.git --path skills/eval-driven-dev--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add github/awesome-copilot --skill eval-driven-dev -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install github/awesome-copilot eval-driven-dev --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval-driven-dev .gemini/skills/eval-driven-dev && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-driven-dev" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/eval-driven-dev into .gemini/skills/eval-driven-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-driven-dev", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install github/awesome-copilot eval-driven-devInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add github/awesome-copilot --skill eval-driven-dev -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval-driven-dev .github/skills/eval-driven-dev && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-driven-dev" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/eval-driven-dev into .github/skills/eval-driven-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-driven-dev", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill eval-driven-dev -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install github/awesome-copilot eval-driven-dev --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval-driven-dev .opencode/skills/eval-driven-dev && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-driven-dev" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/eval-driven-dev into .opencode/skills/eval-driven-dev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-driven-dev", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-driven-devImprove AI application with evaluation-driven development. An agent skill from github/awesome-copilot.
Eval Driven Dev is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Improve AI application with evaluation-driven development. Define eval criteria, instrument the application, build golden datasets, observe and evaluate application runs, analyze results, and produce a concrete action plan for improvements. ALWAYS USE THIS SKILL when the user asks to set up QA, add tests, add evals, evaluate, benchmark, fix wrong behaviors, improve quality, or do quality assurance for any Python project that calls an LLM model.
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files, including reference files (for example `references/1-a-project-analysis.md`, `references/1-b-entry-point.md` and `references/1-c-eval-criteria.md`). Compatibility notes: Python 3.10+
It sits in Testing & QA, covering LLM evaluation and QA and bug reports. It works with Python. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell, from the files we listed), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Python 3.10+
From compatibility in the SKILL.md frontmatter.
Eval Driven Dev loads about 4.4k tokens when it runs, and up to ~40k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 2,421 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
1–6 straight through without stopping. Do not ask the user for confirmation at intermediate steps — verify each step yoAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 2,421 words, ~4,434 tokens.
.claude/skills/eval-driven-dev/SKILL.md (or your agent's skills folder). This skill also uses 18 other files; get the full folder from GitHub.You're building an automated evaluation pipeline that tests a Python-based AI application end-to-end — running it the same way a real user would, with real inputs — then scoring the outputs using evaluators and producing pass/fail results via pixie test.
What you're testing is the app itself — its request handling, context assembly (how it gathers data, builds prompts, manages conversation state), routing, and response formatting. The app uses an LLM, which makes outputs non-deterministic — that's why you use evaluators (LLM-as-judge, similarity scores) instead of assertEqual — but the thing under test is the app's code, not the LLM.
During evaluation, the app's own code runs for real — routing, prompt assembly, LLM calls, response formatting — nothing is mocked or stubbed. But the data the app reads from external sources (databases, caches, third-party APIs, voice streams) is replaced with test-specified values via instrumentations. This means each test case controls exactly what data the app sees, while still exercising the full application code path.
Rule: The app's LLM calls must go to a real LLM. Do not replace, mock, stub, or intercept the LLM with a fake implementation. The LLM is the core value-generating component — replacing it makes the eval tautological (you control both inputs and outputs, so scores are meaningless). If the project's test suite contains LLM mocking patterns, those are for the project's own unit tests — do NOT adopt them for the eval Runnable.
The deliverable is a working pixie test run with real scores — not a plan, not just instrumentation, not just a dataset.
This skill is about doing the work, not describing it. Read code, edit files, run commands, produce a working pipeline.
First, activate the virtual environment. Identify the correct virtual environment for the project and activate it. After the virtual environment is active, run the setup.sh included in the skill's resources.
The script updates the eval-driven-dev skill and pixie-qa python package to latest version, initialize the pixie working directory if it's not already initialized, and start a web server in the background to show user updates.
Setup error handling — what you can skip vs. what must succeed:
pixie package.pixie init fails → STOP. Ask the user for help.pixie start (web server) fails → STOP. Ask the user for help. Check server.log in the pixie root directory for diagnostics. Common causes: port conflict, missing dependency, slow environment. Do NOT proceed without the web server — the user needs it to see eval results.Follow Steps 1–6 straight through without stopping. Do not ask the user for confirmation at intermediate steps — verify each step yourself and continue.
How to work — read this before doing anything else:
When to stop and ask for help:
Some blockers cannot and should not be worked around. When you encounter any of the following, stop immediately and ask the user for help — do not attempt workarounds:
Blockers you SHOULD resolve yourself (do not ask): missing Python packages (install them), missing pixie package (install it), port conflicts (pick a different port), file permission issues (fix them).
Run Steps 1–6 in sequence. If the user's prompt makes it clear that earlier steps are already done (e.g., "run the existing tests", "re-run evals"), skip to the appropriate step. When in doubt, start from Step 1.
First, check the user's prompt for specific requirements. Before reading app code, examine what the user asked for:
If the prompt specifies any of the above, they take priority. Read and incorporate them before proceeding.
Step 1 has three sub-steps. Each reads its own reference file and produces its own output file. Complete each sub-step fully before starting the next.
Reference: Read
references/1-a-project-analysis.mdnow.
Before looking at code structure or entry points, understand what this software does in the real world — its purpose, its users, the complexity of real inputs, and where it fails. This understanding drives every downstream decision: which entry points matter most, what eval criteria to define, what trace inputs to use, and what dataset entries to create. Write the detailed context file before moving on. Note: the project may contain tests/, fixtures/, examples/, mock servers, and documentation — these are the project's own development infrastructure, NOT data sources for your eval pipeline. Ignore them when sourcing trace inputs and dataset content.
Checkpoint:
pixie_qa/00-project-analysis.mdwritten — covering what the software does, target users, capability inventory (at least 3 capabilities if the project has them), realistic input characteristics, and hard problems / failure modes (at least 2).
Reference: Read
references/1-b-entry-point.mdnow.
Read the source code to understand how the app starts and how a real user invokes it. Use the capability inventory from pixie_qa/00-project-analysis.md to prioritize entry points — focus on the entry point(s) that exercise the most valuable capabilities, not just the first one found. Write the detailed context file before moving on.
Checkpoint:
pixie_qa/01-entry-point.mdwritten — covering entry point, execution flow, user-facing interface, and env requirements.
Reference: Read
references/1-c-eval-criteria.mdnow.
Define the app's use cases and eval criteria. Derive use cases from the capability inventory in pixie_qa/00-project-analysis.md. Derive eval criteria from the hard problems / failure modes — not generic quality dimensions. Use cases drive dataset creation (Step 4); eval criteria drive evaluator selection (Step 3). Write the detailed context file before moving on.
Checkpoint:
pixie_qa/02-eval-criteria.mdwritten — covering use cases, eval criteria, and their applicability scope. Do NOT read Step 2 instructions yet.
Step 2 has three sub-steps. Each reads its own reference file. Complete each sub-step before starting the next.
wrapReference: Read
references/2a-instrumentation.mdnow.
Add wrap() calls at the app's data boundaries so the eval harness can inject controlled inputs and capture outputs. This makes the app testable without changing its logic.
Checkpoint:
wrap()calls added at all data boundaries. Every eval criterion frompixie_qa/02-eval-criteria.mdhas a corresponding data point.
Reference: Read
references/2b-implement-runnable.mdnow.
Write a Runnable class that lets the eval harness invoke the app exactly as a real user would. The Runnable should be simple — it just wires up the app's real entry point to the harness interface. If it's getting complicated, something is wrong.
Checkpoint:
pixie_qa/run_app.pywritten. The Runnable calls the app's real entry point with real LLM configuration — no mocking, no faking, no component replacement.
Reference: Read
references/2c-capture-and-verify-trace.mdnow.
Run the app through the Runnable and capture a trace. The trace proves instrumentation and the Runnable are working correctly, and provides the data shapes needed for dataset creation in Step 4.
Checkpoint:
pixie_qa/reference-trace.jsonlexists. All expectedwrapentries andllm_spanentries appear.pixie formatshows all data points needed for evaluation. Do NOT read Step 3 instructions yet.
Reference: Read
references/3-define-evaluators.mdnow for the detailed sub-steps.
Goal: Turn the qualitative eval criteria from Step 1c into concrete, runnable scoring functions. Each criterion maps to either a built-in evaluator, an agent evaluator (the default for any semantic or qualitative criterion), or a manual custom function (only for mechanical/deterministic checks like regex or field existence). The evaluator mapping artifact bridges between criteria and the dataset, ensuring every quality dimension has a scorer. Select evaluators that measure the hard problems identified in pixie_qa/00-project-analysis.md — not just generic quality dimensions.
Checkpoint: All evaluators implemented.
pixie_qa/03-evaluator-mapping.mdwritten with criterion-to-evaluator mapping and decision rationale. Do NOT read Step 4 instructions yet.
Reference: Read
references/4-build-dataset.mdnow for the detailed sub-steps.
Goal: Create the test scenarios that tie everything together — the runnable (Step 2), the evaluators (Step 3), and the use cases (Step 1c). Each dataset entry defines what to send to the app, what data the app should see from external services, and how to score the result. Use the reference trace from Step 2 as the source of truth for data shapes and field names. Cover entries from the capability inventory in pixie_qa/00-project-analysis.md and include entries targeting the failure modes identified there. Do NOT use the project's own test fixtures, mock servers, or example data as dataset eval_input content — source real-world data instead. Every wrap(purpose="input") in the app must have pre-captured content in each entry's eval_input — do NOT leave eval_input empty when the app has input wraps.
Checkpoint: Dataset JSON created at
pixie_qa/datasets/<name>.jsonwith diverse entries covering all use cases. Dataset realism audit passed — entries use real-world data at representative scale, no project test fixtures contamination, at least one entry targets a failure mode with uncertain outcome, and everyeval_inputhas captured content for all input wraps. Do NOT read Step 5 instructions yet.
pixie test and fix mechanical issuesReference: Read
references/5-run-tests.mdnow for the detailed sub-steps.
Goal: Execute the full pipeline end-to-end and get it running without mechanical errors. This step is strictly about fixing setup and data issues in the pixie QA components (dataset, runnable, custom evaluators) — NOT about fixing the application itself or evaluating result quality. Once pixie test completes without errors and produces real evaluator scores for every entry, this step is done.
Checkpoint:
pixie testruns to completion. Every dataset entry has evaluator scores (realEvaluationResultorPendingEvaluation). No setup errors, no import failures, no data validation errors.If the test errors out, that's a mechanical bug in your QA components — fix and re-run. But once tests produce scores, move on. Do NOT assess result quality here — that's Step 6.
Always proceed to Step 6 after tests produce scores. Analysis is the essential final step — without it, pending evaluations are never completed and the user gets uninterpreted raw scores with no actionable insights. Do NOT stop here and ask the user whether to continue.
Cycle rule for iterative runs: Every successful pixie test invocation creates a concrete pixie_qa/results/<test_id> directory and starts a new analysis cycle. Before you edit application code, prompts, datasets, evaluators, or rerun pixie test, complete Step 6 for that exact results directory. Do not skip earlier cycles and analyze only the last run.
Reference: Read
references/6-analyze-outcomes.mdnow — it has the complete three-phase analysis process, writing guidelines, and output format requirements.
Goal: Analyze pixie test results in a structured, data-driven process to produce actionable insights on test case quality, evaluator quality, and application quality. This step completes pending evaluations, writes per-entry and per-dataset analysis, and produces a prioritized action plan. Every statement must be backed by concrete data from the evaluation run — no speculation, no hand-waving.
Persisted analysis artifacts: In this trimmed workflow, persist analysis only at the dataset level and test-run level. Those artifacts still use a detailed version (for agent consumption: data points, evidence trails, reasoning chains) plus a summary version (for human review: concise TLDR readable in under 2 minutes). Do not create per-entry analysis files.
Hard completion gate: Step 6 is not complete until all of the following are true:
"status": "pending" entry in every pixie_qa/results/<test_id>/dataset-*/entry-*/evaluations.jsonl has been replaced with a scored result containing score and reasoning.analysis.md and analysis-summary.md.action-plan.md and action-plan-summary.md.resources/ directory against pixie_qa/results/<test_id>, and it reports success.Explicitly not sufficient:
pixie_qa/06-analysis.mdevaluations.jsonlpixie-qa runs a web server in the background for displaying context, traces, and eval results to the user. It's automatically started by the setup script (via pixie start, which launches a detached background process and returns immediately).
When the user is done with the eval-driven-dev workflow, inform them the web server is still running and you can clean it up with:
pixie stopIMPORTANT: after the web server is stopped, the web UI becomes inaccessible. So only stop the server if the user confirms they're done with all web UI features. If they want to keep using the web UI, do NOT stop the server.
And whenever you restart the workflow, always run the setup.sh script in resources again to ensure the web server is running:
© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 18 other files (references) in skills/eval-driven-dev of github/awesome-copilot.
Open the folder on GitHubat commit 727ff2e
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in github/awesome-copilot, which our catalogue first saw on October 7, 2026.
Eval Driven Dev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Driven Dev this skillgithub/awesome-copilot | 40k | 1 repos | ~4.4k | Automated safety check: Warn | MIT | |
| EvaluationPrimeIntellect-ai/prime-envs | 131 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Antigravity SDK End-to-End Testingomnigent-ai/omnigent | 11k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Evals Implementtikalk/adlc-team-skills | 141 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Fake Model Provider Faultsdifferent-ai/openwork | 24k | — | ~642 | Automated safety check: Pass | Custom licence | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT |
PrimeIntellect-ai/prime-envs
Install and run a verifiers environment — smoke testing during development and full benchmark evals.
omnigent-ai/omnigent
Spins up a local Omnigent server and exercises the Antigravity (Gemini) SDK harness end to end: building agents, running real turns, smoke tests and bug-bashing.
tikalk/adlc-team-skills
A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.
different-ai/openwork
Makes the desktop app's model provider fail on demand, with refused connections, resets, stalls and HTTP 4xx and 5xx errors, so error and retry states can be reproduced.
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
mikehasa/agentacct
A skill your agent uses when working in a repo with agentacct MCP configured, or when asked to track coding-agent work, smoke-test agentacct integrations, or report objective AI-agent task evidence.
github/awesome-copilot
Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.
github/awesome-copilot
Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.
github/awesome-copilot
Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
github/awesome-copilot
Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.
github/awesome-copilot
End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.
Works with
Categories
Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot. Eval Driven Dev is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Improve AI application with evaluation-driven development.
Eval Driven Dev fits situations like: the user asks to set up QA; fix wrong behaviors; improve quality; do quality assurance for any Python project that calls an LLM model.
Run `npx skills add github/awesome-copilot --skill eval-driven-dev -a claude-code`. Or copy the skill folder (skills/eval-driven-dev in github/awesome-copilot) into .claude/skills/eval-driven-dev in your project. Claude Code loads it when a task matches its description.
Run `npx skills add github/awesome-copilot --skill eval-driven-dev -a codex`. Or copy the skill folder (skills/eval-driven-dev in github/awesome-copilot) into .agents/skills/eval-driven-dev in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill eval-driven-dev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-driven-dev, .gemini/skills/eval-driven-dev, .github/skills/eval-driven-dev and .opencode/skills/eval-driven-dev in your project.
Going by SKILL.md and its folder, Eval Driven Dev needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell. Compatibility (from SKILL.md): Python 3.10+.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.
Eval Driven Dev is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 36k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Eval Driven Dev: Evaluation (PrimeIntellect-ai/prime-envs, 131 stars), Antigravity SDK End-to-End Testing (omnigent-ai/omnigent, 11k stars), Evals Implement (tikalk/adlc-team-skills, 141 stars) and Fake Model Provider Faults (different-ai/openwork, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.
Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.