Eval Driven Dev
github/awesome-copilot
Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.
Install and run a verifiers environment — smoke testing during development and full benchmark evals.
$ npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install PrimeIntellect-ai/prime-envs evaluation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/PrimeIntellect-ai/prime-envs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluation .claude/skills/evaluation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluation" agent skill from https://github.com/PrimeIntellect-ai/prime-envs/tree/main/skills/evaluation into .claude/skills/evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/PrimeIntellect-ai/prime-envs/tree/main/skills/evaluationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install PrimeIntellect-ai/prime-envs evaluation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrimeIntellect-ai/prime-envs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/evaluation .agents/skills/evaluation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluation" agent skill from https://github.com/PrimeIntellect-ai/prime-envs/tree/main/skills/evaluation into .agents/skills/evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install PrimeIntellect-ai/prime-envs evaluation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrimeIntellect-ai/prime-envs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/evaluation .cursor/skills/evaluation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluation" agent skill from https://github.com/PrimeIntellect-ai/prime-envs/tree/main/skills/evaluation into .cursor/skills/evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/PrimeIntellect-ai/prime-envs.git --path skills/evaluation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install PrimeIntellect-ai/prime-envs evaluation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrimeIntellect-ai/prime-envs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/evaluation .gemini/skills/evaluation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluation" agent skill from https://github.com/PrimeIntellect-ai/prime-envs/tree/main/skills/evaluation into .gemini/skills/evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install PrimeIntellect-ai/prime-envs evaluationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/PrimeIntellect-ai/prime-envs.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/evaluation .github/skills/evaluation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluation" agent skill from https://github.com/PrimeIntellect-ai/prime-envs/tree/main/skills/evaluation into .github/skills/evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install PrimeIntellect-ai/prime-envs evaluation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PrimeIntellect-ai/prime-envs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/evaluation .opencode/skills/evaluation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluation" agent skill from https://github.com/PrimeIntellect-ai/prime-envs/tree/main/skills/evaluation into .opencode/skills/evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluationInstall and run a verifiers environment — smoke testing during development and full benchmark evals.
Evaluation is an agent skill from PrimeIntellect-ai/prime-envs. Install and run a verifiers environment — smoke testing during development and full benchmark evals. Covers the eval CLI for the tasksets. Use while developing/iterating on an environment, when running/evaluating/benchmarking one, or when the user mentions eval, tasksets, eval runs, or smoke tests.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering LLM evaluation and QA and bug reports. The repository describes itself as: Environments by the Prime Intellect Research Team. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 89dcc7f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvjqFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENSERPER_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluation loads about 4.6k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 2,251 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from PrimeIntellect-ai/prime-envs at commit 89dcc7f, republished under its Apache-2.0 licence (© PrimeIntellect-ai). 2,251 words, ~4,640 tokens.
.claude/skills/evaluation/SKILL.md (or your agent's skills folder).Running an environment in this repo — both while developing/iterating on it (smoke-testing changes) and when benchmarking a finished env. Run everything from the repo root, always via uv.
Environments are v1 tasksets: a package that exports a Taskset. Run one with the eval CLI, by taskset id: uv run eval <taskset-id>.
Editable, local install from the repo root (not from inside the env dir). Envs live under
environments/, grouped into semantic group folders (e.g. environments/math/, environments/swe/):
uv pip install -e path/to/envEditable means code edits are picked up without reinstalling — so during development you can change the env and immediately re-run it. Only re-run this install after editing pyproject.toml (e.g. new deps). v1 tasksets pin a verifiers pre-release; if the install complains, add --prerelease=if-necessary (never the blanket --prerelease=allow, which bumps unrelated deps to betas — e.g. a sympy rc + mpmath alpha that can't import together).
Harbor-based envs (SWE/terminal tasksets) must depend on verifiers[harbor], not plain verifiers — harbor is an optional extra of verifiers, and without it the eval crashes at taskset load with ModuleNotFoundError: No module named 'harbor.models'.
In a dirty worktree, use uv run --no-sync ... after the editable install so machine-local uv settings cannot rewrite the repository's uv.lock. Do not export UV_FROZEN for an eval: subprocess runtimes inherit it, and their unlocked inline uv scripts cannot start in frozen mode.
Run an eval from the repository root, not from inside its virtualenv. NLTK's import protection rejects dependencies installed beneath the current working directory. For a separate temporary venv, activate it by absolute path and run uv run --active --no-project eval ... from outside that venv.
After dependency edits, sync the environment project before testing imports:
uv sync --project path/to/env --all-extrasFor gated Hugging Face datasets, verify access by downloading a small required file;
repository metadata can remain public even when the authenticated account cannot read the
artifacts. Keep HF_TOKEN on the evaluator. If the agent needs gated corpus files, stage only
those files into its runtime and keep answer keys, rubrics, and source provenance evaluator-side.
If a fresh environment import fails because resolved package versions are incompatible, fix the environment's pyproject.toml bounds and sync again. Do not patch generated .venv files.
For OfficeQA Pro V2, authenticate the evaluator for the gated databricks/officeqa-pro-v2 dataset at revision 65a2b315780417bc50d7bfe6e5bdb904e63fda65. Use the registered Prime VM image reference prime/primeintellect/officeqa-pro-v2:65a2b3157804-v1, containing all 1,435 parsed documents at /workspace/documents. Prime VM lookup rejects the OCI digest form with VM image not found. Before evaluating, use prime images list --search officeqa --output json --plain and verify the container artifact’s fullImagePath digest is sha256:f99fb779f7f473fd7b596da7690992d0bdb20cccf89a0a4083cbeb2852c01b42 so the corpus stays aligned with the pinned question CSV.
PI-OfficeQa-Pro-v2 (pi_officeqa_pro_v2) loads 82 tasks by default with 15 repairs and 8 exclusions. It blocks Hugging Face, GitHub, their content hosts, and common mirrors through its native task network policy during sandbox execution. Other web sources remain available; evaluator-side dataset downloads and provider-hosted search services outside the sandbox are unaffected.
OfficeQA-Pro-v2 (officeqa_pro_v2) loads all 90 unmodified tasks with the original regex grader at zero numerical tolerance and unrestricted networking. Both variants reuse the same public Prime image; select the environment explicitly when reporting scores.
PI-OfficeQa-Pro-v2 uses a ReferenceJudge subclass with a packaged prompt covering numeric values, signs, units, precision, and list order, defaulting to openai/gpt-5.6-luna with medium reasoning. Configure it under --env.taskset.task.judges.0.*, independently of the solver's -m setting. Judge calls use evaluator-side Prime credentials and are recorded in trace.info["judge"] and trace.extra_usage; these scores use LLM judgment rather than the upstream deterministic scorer. The prompt grades the last <FINAL_ANSWER> block when present, otherwise the stated final answer; XML tags and harmless prose remain optional.
When a judge's id names an environment package, export its judge class in that package's __all__ so Verifiers can discover it. For judge-only checks, load saved traces with vf.Trace[YourTaskData] or vf.WireTrace so custom reference-answer fields survive deserialization, then call await task.score(trace). A bare vf.Trace can discard those fields. An in-process smoke with --no-serve can validate the task and judge when the worker pool fails; record that mode because it does not validate the served path.
For reference-judge prompts that request an explanation, end with a boxed verdict, such as \boxed{yes}. Verifiers 0.3.1 prioritizes boxed text; an earlier explanation labeled "Final answer" can otherwise make its verdict parser pick up incidental yes/no words before the actual verdict.
Run a 3x1 (3 tasks, 1 rollout each) in plain-log mode to confirm the env loads and scores end-to-end. Spanning a few tasks surfaces weird things (bad rows, edge-case prompts, inconsistent scoring) that a single task hides:
uv run --no-sync eval <taskset-id> -n 3 -r 1 --no-rich -v --no-push--no-rich turns off the live dashboard (which is on by default) for plain logs, -v prints prompts/completions, and --no-push keeps development runs local. This is the inner loop while developing — re-run it after each change to verify the dataset loads, the rollout runs, and the rubric scores as expected. Fix any errors here before scaling up.
Run all tasks (omit -n) and pick -r so the total sample is representative — usually >500 total rollouts (num_tasks × r) is a good target. For large datasets -r 1 may already clear that; for small ones, bump -r.
Recommended full-eval invocation:
uv run eval <taskset-id> -r <r> -c <N> --no-richThe run is always saved to disk (see Inspect output) — there's no save flag.
eval)| Flag | Meaning | When |
|---|---|---|
-m <slug> | model id (default deepseek/deepseek-v4-flash) | to override the default |
-n <N> | number of tasks; omit for all | smoke (-n 3) vs full (omit) |
-r <N> | rollouts per task (>=2 if the taskset has @group_rewards) | almost always set |
-c <N> | max episodes in flight (default 128) | raise for cheap envs, lower for sandboxed |
-s | shuffle tasks before taking the first -n | sampling a subset of a big dataset |
-v | debug logs (prompts/completions) | developing/debugging |
--no-rich | plain logs instead of the live dashboard | non-interactive or captured runs |
-o <dir> | directory that groups runs; each run writes below its run.dir | pinning a known output root |
--resume | re-run a saved config's missing or errored episodes in place | use with @ <run-dir>/configs/resolved/eval.json |
Most tasksets yield records as they are consumed. In Verifiers 0.3.1, -n N uses a lazy head
before the runner collects those N tasks, but automationbench_env, general_agent,
tau2_bench, and tau3_bench still construct their complete task lists. -s calls
Taskset.shuffle() first, which must materialize the complete finite taskset, so avoid -s
when a smoke test should construct only a prefix.
Pass typed, dotted flags or a TOML file — there is no -a JSON blob:
uv run eval wikispeedia --env.taskset.min-path-length 5 --env.taskset.max-path-length 8
uv run eval <taskset-id> @ eval.toml # a launch TOML is re-runnableCommon knobs: --env.agent.max-turns, --env.agent.max-total-tokens, --sampling.max-tokens, --sampling.temperature, --env.agent.harness.id, --env.agent.runtime.type. uv run eval <taskset-id> -h prints the full typed help, narrowed to the chosen taskset/harness.
TasksetConfig.system_prompt is reserved for --env.taskset.system-prompt, a path to a prompt file that overrides every task during iteration. Environment configs must not redefine it as inline text; use a distinct field such as task_system_prompt for an inline/default task prompt.
List-valued flags on union configs (e.g. --env.agent.runtime.labels on the prime runtime) must be passed as a space-separated JSON value — --env.agent.runtime.labels '["ci"]'. The bare form (... labels ci) and the =-form with JSON both parse as a string and fail validation with Input should be a valid list.
A dotted override below a nested RuntimeConfig can replace that whole runtime with the selected provider's defaults. After overriding any nested runtime field, inspect configs/resolved/eval.json and confirm its image, workdir, runtime type, and resources before treating the smoke test as representative.
A harness belongs to an agent seat: on the default single-agent env the seat is agent (--env.agent.harness.*); a multi-agent env names its own seats (e.g. --env.solver.harness.* and --env.judge.harness.* on agentic-judge). A taskset that bundles its own harness runs with it by default; otherwise pass --env.agent.harness.id. Select where rollouts execute with --env.agent.runtime.type:
In Verifiers 0.3.1, the execution runtime belongs to the environment's agent role, not HarnessConfig. If a bundled harness must run with the interpreter that installed its package, export a custom Env/EnvConfig and set its agent default to AgentConfig(runtime=SubprocessConfig()).
Toolsets running on Docker or Prime runtimes require Verifiers 0.3.1 to be installed from its source checkout, whether or not they are colocated with the harness. The framework builds and uploads its own source distribution with the environment package; a wheel-only Verifiers install has no pyproject.toml to build and fails before the tool server starts.
Prime port exposure can return before its public route is reachable from another Prime sandbox. A taskset that supports a separate Prime tool-server runtime must check the exposed route from Prime before handing it to the agent; a localhost probe inside the server only proves that the process is listening.
TaskData.network_allow=[] is the Verifiers 0.3.1 default-deny policy for Docker and Prime solver runtimes; Prime requires vm=true. Verifiers applies it after trusted task setup and retains its own interception and MCP routes. Subprocess and Modal do not implement task-data network policies, although a task-owned Modal runtime can set network_access=false. A secondary runtime created directly with make_runtime() must use a restricted config and call prepare_execution(...) after setup, before it runs agent-controlled code.
uv run eval <taskset-id> --env.agent.runtime.type subprocess # local process
uv run eval <taskset-id> --env.agent.runtime.type docker # local container (needs local docker)
uv run eval <taskset-id> --env.agent.runtime.type prime # remote prime VM (default; needs auth)
uv run eval <taskset-id> --env.agent.runtime.type modal # remote modal sandbox (needs auth)Harbor tasksets should reuse HarborTask for environment upload and packaged test.sh
scoring. When gated verifier inputs must not enter the agent box, export a HarborEnv
subclass, set HarborData.verifier, and collect only the solution artifacts that the
separate verifier needs.
For packages whose task Dockerfile only selects a base image and copies task assets, resolve
the FROM image, set upload_environment=True, and let HarborTask.setup stage the assets;
Verifiers deliberately does not build task Dockerfiles.
Harbor can also load an unpublished dataset directly from Git. Set HarborConfig.dataset
to "" and repo to a pinned GitHub URL of the form
https://github.com/<org>/<repo>/tree/<commit>/<tasks-subdir>; the implicit repository
loader discovers the task directories, downloads them together, and caches the export.
Reuse HarborTaskset.load() instead of implementing another archive downloader.
RewardKit judges routed through an OpenAI-compatible endpoint need LiteLLM's provider prefix in
front of the endpoint model id: a pinference model id such as openai/gpt-5.5 is passed as
openai/openai/gpt-5.5. Pass that model, the endpoint, and a credential template through
HarborData.verifier_env; the packaged RewardKit verifier does not consume Verifiers judge
sampling settings. Configure task-owned judge endpoints under --env.taskset.task.judge.*.
Prime containers accept explicit registry refs such as GHCR, but Prime VM image builds require Docker Hub or Prime registry refs; point VM smoke tests at a published Prime mirror.
When creating a data archive for an image on macOS, run tar with COPYFILE_DISABLE=1 and
--no-mac-metadata; otherwise BSD tar can add ._* AppleDouble files. Verify the file count
inside a disposable sandbox before running the image at scale.
For publicly supported tasksets, use canonical platform image references such as
recursive-task-synthesis/<task-id>:latest. Platform publishing supports namespaced image
names via prime images push --platform-image; do not prepend prime/primeintellect/ to
platform references.
Images altered or fixed relative to upstream can be uploaded to the team's userspace. Use the exact image reference returned by Prime and point the affected task definitions explicitly to those images.
Preserve existing upstream or platform references for unchanged images. Do not introduce automatic prefix rewriting or fallback between platform and team images.
When polling Prime status from zsh, do not assign to status: it is a read-only shell parameter.
Use a task-specific variable such as image_state.
For Harbor packages whose Dockerfile lives in environment/, pass that directory as both the
Docker build context and the Dockerfile parent. Dockerfile COPY paths are relative to the build
context; passing the task-package root makes valid paths such as COPY cli.c ... fail when the
file is at environment/cli.c.
Prime image-build initiation can return a temporary HTTP 502. A batch publisher should retry that specific response with a bounded delay, while allowing other failures to surface for diagnosis. For concurrent image publishing, catch per-image submission failures at the worker boundary and record the task ID; one failed image must not terminate the remaining independent submissions.
The built-in bash harness (a bash + edit agent) is the fallback when no harness id is given; null is a tool-less chat loop (MCP tools only). For "bring your own search" tasksets that ship no search tool of their own (e.g. openseeker, redsearcher, s1-deepresearch, arxivmath), give the agent a Serper-backed search tool with --env.agent.harness.search true (needs SERPER_API_KEY in the eval environment):
uv run eval openseeker --env.agent.harness.id bash --env.agent.harness.search true --env.agent.runtime.type prime-c caps episodes in flight (default 128). An episode may contain multiple agent runs; those run one at a time by default. Set --env.max-concurrent-agents <N> only when an environment should play several agents from the same episode concurrently. The upper bound on simultaneous agent runs is therefore roughly -c * --env.max-concurrent-agents.
Raise -c for cheap, non-sandboxed tasksets; keep it lower for sandboxed ones (containers/remote runtimes) and tune up from there. Eval uses the env-server pool by default; its settings live below --serve.pool.* (for example, --serve.pool.type static --serve.pool.num-workers 4), while --serve.max-concurrent optionally caps episodes per worker. Use --no-serve only when in-process execution is intentional.
The taskset's default model is usually fine (deepseek/deepseek-v4-flash). Otherwise pick by tier (-m <slug>):
deepseek/deepseek-v4-flash or z-ai/glm-5.1.openai/gpt-5.4 or openai/gpt-5.5 at medium reasoning.openai/gpt-5.5 is the strongest but expensive. Avoid it by default; use it only as a last resort to confirm a task is solvable when the cheaper models get zero reward.The Rich dashboard is on by default and shows a live reward <mean> · err <share> headline, but it doesn't render in a plain captured shell. To watch it live, run inside tmux and drive it with send-keys:
tmux new-session -d -s eval
tmux send-keys -t eval 'uv run eval <taskset-id> -r <r>' Enter
# watch it:
tmux attach -t eval # or: tmux capture-pane -t eval -pFor headless/automated runs, prefer --no-rich and follow the logs directly.
Each run is saved to a fresh per-run dir (so runs never overwrite each other):
outputs/<run-dir>/
configs/eval.toml # launch config when the run starts from CLI flags
configs/resolved/eval.json # complete config; use this to resume
traces.jsonl # one full episode per line
logs/attempt_<n>/eval.log # logs for one run or resume attempt
logs/latest -> attempt_<n> # current attempt# find the newest run
ls -dt outputs/*/ | head
# avg reward across rollouts (recomputed — aggregates aren't stored)
jq -s '[.[].traces[] | [.rewards[]? | .score * .weight] | add // 0] | if length > 0 then add / length else 0 end' outputs/<...>/traces.jsonlThe dashboard shows the avg reward live; with --no-rich, recompute it from traces.jsonl (each line is an episode whose traces[] carry the per-reward breakdown). Skim a few episodes for sanity before publishing.
© PrimeIntellect-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/evaluation of PrimeIntellect-ai/prime-envs.
Open the folder on GitHubat commit 89dcc7f
Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluation this skillPrimeIntellect-ai/prime-envs | 130 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Eval Driven Devgithub/awesome-copilot | 40k | 1 repos | ~4.4k | Automated safety check: Warn | MIT | |
| QA Methodologymagnus919/agent-skills | 113 | — | ~3.2k | Automated safety check: Pass | MIT | |
| Early Experience DataOSU-NLP-Group/EarlyExperience | 102 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Eval Guidemicrosoft/eval-guide | 138 | — | ~22k | Automated safety check: Warn | MIT | |
| Write A Specdifferent-ai/openwork | 24k | — | ~3.3k | Automated safety check: Pass | Custom licence |
github/awesome-copilot
Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.
magnus919/agent-skills
Design and apply QA methodology for software teams: test strategy, regression testing, CI failure triage, test automation, quality gates and metrics, risk-based testing, exploratory testing, test…
OSU-NLP-Group/EarlyExperience
A skill your agent uses whenever the user asks to generate, collect, inspect, or prepare early-experience training data (Implicit World Modeling or Self-Reflection, in the sense of arXiv:2510.08558)…
microsoft/eval-guide
Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.
different-ai/openwork
Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer.
CALLE-AI/awesome-phone-call-agents
Monitor business phone lines and deployed voice agents with LineCanary — scheduled CALL-E test calls that walk the caller journey, assert structured results, diff against baselines and alert on…
PrimeIntellect-ai/prime-envs
Where task data and taskset code live across prime-tasks and prime-envs, and how to fix a broken task.
Categories
Install and run a verifiers environment — smoke testing during development and full benchmark evals. Evaluation is an agent skill from PrimeIntellect-ai/prime-envs. Install and run a verifiers environment — smoke testing during development and full benchmark evals.
Evaluation fits situations like: tasks that involve LLM evaluation; tasks that involve QA and bug reports.
Run `npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a claude-code`. Or copy the skill folder (skills/evaluation in PrimeIntellect-ai/prime-envs) into .claude/skills/evaluation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a codex`. Or copy the skill folder (skills/evaluation in PrimeIntellect-ai/prime-envs) into .agents/skills/evaluation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation, .gemini/skills/evaluation, .github/skills/evaluation and .opencode/skills/evaluation in your project.
Going by SKILL.md and its folder, Evaluation needs the command-line tools its instructions call (uv and jq) and credentials named HF_TOKEN and SERPER_API_KEY. Our summary lists: Python 3; Docker.
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evaluation: Eval Driven Dev (github/awesome-copilot, 40k stars), QA Methodology (magnus919/agent-skills, 113 stars), Early Experience Data (OSU-NLP-Group/EarlyExperience, 102 stars) and Eval Guide (microsoft/eval-guide, 138 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
PrimeIntellect-ai (a GitHub organization) maintains it in PrimeIntellect-ai/prime-envs, which has 130 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 8, 2026.
Source: PrimeIntellect-ai/prime-envs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.