Install and run a verifiers environment — smoke testing during development and full benchmark evals.

Apache-2.0Auto-check passedTesting & QA

Install Evaluation

skills CLI
$ npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PrimeIntellect-ai/prime-envs evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PrimeIntellect-ai/prime-envs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluation .claude/skills/evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluation
GitHub stars
130
Token cost
~4.6k tokens
SKILL.md length
2,251 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Install and run a verifiers environment — smoke testing during development and full benchmark evals.

  • Tasks that involve LLM evaluation
  • SKILL.md covers Setup, Smoke test, Full eval and Live dashboard in tmux, plus 1 more section
  • Calls uv and jq; reaches github.com; needs HF_TOKEN and SERPER_API_KEY
  • Tasks that involve QA and bug reports

What it does

Evaluation is an agent skill from PrimeIntellect-ai/prime-envs. Install and run a verifiers environment — smoke testing during development and full benchmark evals. Covers the eval CLI for the tasksets. Use while developing/iterating on an environment, when running/evaluating/benchmarking one, or when the user mentions eval, tasksets, eval runs, or smoke tests.

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering LLM evaluation and QA and bug reports. The repository describes itself as: Environments by the Prime Intellect Research Team. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve QA and bug reports

Example prompts

  • “/evaluation”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 89dcc7f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • SERPER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluation loads about 4.6k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 2,251 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PrimeIntellect-ai/prime-envs at commit 89dcc7f, republished under its Apache-2.0 licence (© PrimeIntellect-ai). 2,251 words, ~4,640 tokens.

Download SKILL.mdSave it as .claude/skills/evaluation/SKILL.md (or your agent's skills folder).
name
evaluation
description
Install and run a verifiers environment — smoke testing during development and full benchmark evals. Covers the `eval` CLI for the tasksets. Use while developing/iterating on an environment, when running/evaluating/benchmarking one, or when the user mentions eval, tasksets, eval runs, or smoke tests.

Evaluation

Running an environment in this repo — both while developing/iterating on it (smoke-testing changes) and when benchmarking a finished env. Run everything from the repo root, always via uv.

Environments are v1 tasksets: a package that exports a Taskset. Run one with the eval CLI, by taskset id: uv run eval <taskset-id>.

Setup

Editable, local install from the repo root (not from inside the env dir). Envs live under environments/, grouped into semantic group folders (e.g. environments/math/, environments/swe/):

bash
uv pip install -e path/to/env

Editable means code edits are picked up without reinstalling — so during development you can change the env and immediately re-run it. Only re-run this install after editing pyproject.toml (e.g. new deps). v1 tasksets pin a verifiers pre-release; if the install complains, add --prerelease=if-necessary (never the blanket --prerelease=allow, which bumps unrelated deps to betas — e.g. a sympy rc + mpmath alpha that can't import together).

Harbor-based envs (SWE/terminal tasksets) must depend on verifiers[harbor], not plain verifiers — harbor is an optional extra of verifiers, and without it the eval crashes at taskset load with ModuleNotFoundError: No module named 'harbor.models'.

In a dirty worktree, use uv run --no-sync ... after the editable install so machine-local uv settings cannot rewrite the repository's uv.lock. Do not export UV_FROZEN for an eval: subprocess runtimes inherit it, and their unlocked inline uv scripts cannot start in frozen mode.

Run an eval from the repository root, not from inside its virtualenv. NLTK's import protection rejects dependencies installed beneath the current working directory. For a separate temporary venv, activate it by absolute path and run uv run --active --no-project eval ... from outside that venv.

After dependency edits, sync the environment project before testing imports:

bash
uv sync --project path/to/env --all-extras

For gated Hugging Face datasets, verify access by downloading a small required file; repository metadata can remain public even when the authenticated account cannot read the artifacts. Keep HF_TOKEN on the evaluator. If the agent needs gated corpus files, stage only those files into its runtime and keep answer keys, rubrics, and source provenance evaluator-side.

If a fresh environment import fails because resolved package versions are incompatible, fix the environment's pyproject.toml bounds and sync again. Do not patch generated .venv files.

For OfficeQA Pro V2, authenticate the evaluator for the gated databricks/officeqa-pro-v2 dataset at revision 65a2b315780417bc50d7bfe6e5bdb904e63fda65. Use the registered Prime VM image reference prime/primeintellect/officeqa-pro-v2:65a2b3157804-v1, containing all 1,435 parsed documents at /workspace/documents. Prime VM lookup rejects the OCI digest form with VM image not found. Before evaluating, use prime images list --search officeqa --output json --plain and verify the container artifact’s fullImagePath digest is sha256:f99fb779f7f473fd7b596da7690992d0bdb20cccf89a0a4083cbeb2852c01b42 so the corpus stays aligned with the pinned question CSV.

PI-OfficeQa-Pro-v2 (pi_officeqa_pro_v2) loads 82 tasks by default with 15 repairs and 8 exclusions. It blocks Hugging Face, GitHub, their content hosts, and common mirrors through its native task network policy during sandbox execution. Other web sources remain available; evaluator-side dataset downloads and provider-hosted search services outside the sandbox are unaffected.

OfficeQA-Pro-v2 (officeqa_pro_v2) loads all 90 unmodified tasks with the original regex grader at zero numerical tolerance and unrestricted networking. Both variants reuse the same public Prime image; select the environment explicitly when reporting scores.

PI-OfficeQa-Pro-v2 uses a ReferenceJudge subclass with a packaged prompt covering numeric values, signs, units, precision, and list order, defaulting to openai/gpt-5.6-luna with medium reasoning. Configure it under --env.taskset.task.judges.0.*, independently of the solver's -m setting. Judge calls use evaluator-side Prime credentials and are recorded in trace.info["judge"] and trace.extra_usage; these scores use LLM judgment rather than the upstream deterministic scorer. The prompt grades the last <FINAL_ANSWER> block when present, otherwise the stated final answer; XML tags and harmless prose remain optional.

When a judge's id names an environment package, export its judge class in that package's __all__ so Verifiers can discover it. For judge-only checks, load saved traces with vf.Trace[YourTaskData] or vf.WireTrace so custom reference-answer fields survive deserialization, then call await task.score(trace). A bare vf.Trace can discard those fields. An in-process smoke with --no-serve can validate the task and judge when the worker pool fails; record that mode because it does not validate the served path.

For reference-judge prompts that request an explanation, end with a boxed verdict, such as \boxed{yes}. Verifiers 0.3.1 prioritizes boxed text; an earlier explanation labeled "Final answer" can otherwise make its verdict parser pick up incidental yes/no words before the actual verdict.

Smoke test

Run a 3x1 (3 tasks, 1 rollout each) in plain-log mode to confirm the env loads and scores end-to-end. Spanning a few tasks surfaces weird things (bad rows, edge-case prompts, inconsistent scoring) that a single task hides:

bash
uv run --no-sync eval <taskset-id> -n 3 -r 1 --no-rich -v --no-push

--no-rich turns off the live dashboard (which is on by default) for plain logs, -v prints prompts/completions, and --no-push keeps development runs local. This is the inner loop while developing — re-run it after each change to verify the dataset loads, the rollout runs, and the rubric scores as expected. Fix any errors here before scaling up.

Full eval

Run all tasks (omit -n) and pick -r so the total sample is representative — usually >500 total rollouts (num_tasks × r) is a good target. For large datasets -r 1 may already clear that; for small ones, bump -r.

Recommended full-eval invocation:

bash
uv run eval <taskset-id> -r <r> -c <N> --no-rich

The run is always saved to disk (see Inspect output) — there's no save flag.

Key flags (eval)
FlagMeaningWhen
-m <slug>model id (default deepseek/deepseek-v4-flash)to override the default
-n <N>number of tasks; omit for allsmoke (-n 3) vs full (omit)
-r <N>rollouts per task (>=2 if the taskset has @group_rewards)almost always set
-c <N>max episodes in flight (default 128)raise for cheap envs, lower for sandboxed
-sshuffle tasks before taking the first -nsampling a subset of a big dataset
-vdebug logs (prompts/completions)developing/debugging
--no-richplain logs instead of the live dashboardnon-interactive or captured runs
-o <dir>directory that groups runs; each run writes below its run.dirpinning a known output root
--resumere-run a saved config's missing or errored episodes in placeuse with @ <run-dir>/configs/resolved/eval.json

Most tasksets yield records as they are consumed. In Verifiers 0.3.1, -n N uses a lazy head before the runner collects those N tasks, but automationbench_env, general_agent, tau2_bench, and tau3_bench still construct their complete task lists. -s calls Taskset.shuffle() first, which must materialize the complete finite taskset, so avoid -s when a smoke test should construct only a prefix.

Configuring the env

Pass typed, dotted flags or a TOML file — there is no -a JSON blob:

bash
uv run eval wikispeedia --env.taskset.min-path-length 5 --env.taskset.max-path-length 8
uv run eval <taskset-id> @ eval.toml          # a launch TOML is re-runnable

Common knobs: --env.agent.max-turns, --env.agent.max-total-tokens, --sampling.max-tokens, --sampling.temperature, --env.agent.harness.id, --env.agent.runtime.type. uv run eval <taskset-id> -h prints the full typed help, narrowed to the chosen taskset/harness.

TasksetConfig.system_prompt is reserved for --env.taskset.system-prompt, a path to a prompt file that overrides every task during iteration. Environment configs must not redefine it as inline text; use a distinct field such as task_system_prompt for an inline/default task prompt.

List-valued flags on union configs (e.g. --env.agent.runtime.labels on the prime runtime) must be passed as a space-separated JSON value — --env.agent.runtime.labels '["ci"]'. The bare form (... labels ci) and the =-form with JSON both parse as a string and fail validation with Input should be a valid list.

A dotted override below a nested RuntimeConfig can replace that whole runtime with the selected provider's defaults. After overriding any nested runtime field, inspect configs/resolved/eval.json and confirm its image, workdir, runtime type, and resources before treating the smoke test as representative.

Show full SKILL.md (1,049 more words)Show less
Harness & runtime (sandboxed / agentic tasksets)

A harness belongs to an agent seat: on the default single-agent env the seat is agent (--env.agent.harness.*); a multi-agent env names its own seats (e.g. --env.solver.harness.* and --env.judge.harness.* on agentic-judge). A taskset that bundles its own harness runs with it by default; otherwise pass --env.agent.harness.id. Select where rollouts execute with --env.agent.runtime.type:

In Verifiers 0.3.1, the execution runtime belongs to the environment's agent role, not HarnessConfig. If a bundled harness must run with the interpreter that installed its package, export a custom Env/EnvConfig and set its agent default to AgentConfig(runtime=SubprocessConfig()).

Toolsets running on Docker or Prime runtimes require Verifiers 0.3.1 to be installed from its source checkout, whether or not they are colocated with the harness. The framework builds and uploads its own source distribution with the environment package; a wheel-only Verifiers install has no pyproject.toml to build and fails before the tool server starts.

Prime port exposure can return before its public route is reachable from another Prime sandbox. A taskset that supports a separate Prime tool-server runtime must check the exposed route from Prime before handing it to the agent; a localhost probe inside the server only proves that the process is listening.

TaskData.network_allow=[] is the Verifiers 0.3.1 default-deny policy for Docker and Prime solver runtimes; Prime requires vm=true. Verifiers applies it after trusted task setup and retains its own interception and MCP routes. Subprocess and Modal do not implement task-data network policies, although a task-owned Modal runtime can set network_access=false. A secondary runtime created directly with make_runtime() must use a restricted config and call prepare_execution(...) after setup, before it runs agent-controlled code.

bash
uv run eval <taskset-id> --env.agent.runtime.type subprocess  # local process
uv run eval <taskset-id> --env.agent.runtime.type docker      # local container (needs local docker)
uv run eval <taskset-id> --env.agent.runtime.type prime       # remote prime VM (default; needs auth)
uv run eval <taskset-id> --env.agent.runtime.type modal       # remote modal sandbox (needs auth)

Harbor tasksets should reuse HarborTask for environment upload and packaged test.sh scoring. When gated verifier inputs must not enter the agent box, export a HarborEnv subclass, set HarborData.verifier, and collect only the solution artifacts that the separate verifier needs.

For packages whose task Dockerfile only selects a base image and copies task assets, resolve the FROM image, set upload_environment=True, and let HarborTask.setup stage the assets; Verifiers deliberately does not build task Dockerfiles.

Harbor can also load an unpublished dataset directly from Git. Set HarborConfig.dataset to "" and repo to a pinned GitHub URL of the form https://github.com/<org>/<repo>/tree/<commit>/<tasks-subdir>; the implicit repository loader discovers the task directories, downloads them together, and caches the export. Reuse HarborTaskset.load() instead of implementing another archive downloader.

RewardKit judges routed through an OpenAI-compatible endpoint need LiteLLM's provider prefix in front of the endpoint model id: a pinference model id such as openai/gpt-5.5 is passed as openai/openai/gpt-5.5. Pass that model, the endpoint, and a credential template through HarborData.verifier_env; the packaged RewardKit verifier does not consume Verifiers judge sampling settings. Configure task-owned judge endpoints under --env.taskset.task.judge.*.

Prime containers accept explicit registry refs such as GHCR, but Prime VM image builds require Docker Hub or Prime registry refs; point VM smoke tests at a published Prime mirror.

When creating a data archive for an image on macOS, run tar with COPYFILE_DISABLE=1 and --no-mac-metadata; otherwise BSD tar can add ._* AppleDouble files. Verify the file count inside a disposable sandbox before running the image at scale.

For publicly supported tasksets, use canonical platform image references such as recursive-task-synthesis/<task-id>:latest. Platform publishing supports namespaced image names via prime images push --platform-image; do not prepend prime/primeintellect/ to platform references.

Images altered or fixed relative to upstream can be uploaded to the team's userspace. Use the exact image reference returned by Prime and point the affected task definitions explicitly to those images.

Preserve existing upstream or platform references for unchanged images. Do not introduce automatic prefix rewriting or fallback between platform and team images.

When polling Prime status from zsh, do not assign to status: it is a read-only shell parameter. Use a task-specific variable such as image_state.

For Harbor packages whose Dockerfile lives in environment/, pass that directory as both the Docker build context and the Dockerfile parent. Dockerfile COPY paths are relative to the build context; passing the task-package root makes valid paths such as COPY cli.c ... fail when the file is at environment/cli.c.

Prime image-build initiation can return a temporary HTTP 502. A batch publisher should retry that specific response with a bounded delay, while allowing other failures to surface for diagnosis. For concurrent image publishing, catch per-image submission failures at the worker boundary and record the task ID; one failed image must not terminate the remaining independent submissions.

The built-in bash harness (a bash + edit agent) is the fallback when no harness id is given; null is a tool-less chat loop (MCP tools only). For "bring your own search" tasksets that ship no search tool of their own (e.g. openseeker, redsearcher, s1-deepresearch, arxivmath), give the agent a Serper-backed search tool with --env.agent.harness.search true (needs SERPER_API_KEY in the eval environment):

bash
uv run eval openseeker --env.agent.harness.id bash --env.agent.harness.search true --env.agent.runtime.type prime
Concurrency

-c caps episodes in flight (default 128). An episode may contain multiple agent runs; those run one at a time by default. Set --env.max-concurrent-agents <N> only when an environment should play several agents from the same episode concurrently. The upper bound on simultaneous agent runs is therefore roughly -c * --env.max-concurrent-agents.

Raise -c for cheap, non-sandboxed tasksets; keep it lower for sandboxed ones (containers/remote runtimes) and tune up from there. Eval uses the env-server pool by default; its settings live below --serve.pool.* (for example, --serve.pool.type static --serve.pool.num-workers 4), while --serve.max-concurrent optionally caps episodes per worker. Use --no-serve only when in-process execution is intentional.

Models

The taskset's default model is usually fine (deepseek/deepseek-v4-flash). Otherwise pick by tier (-m <slug>):

  • Cheap — deepseek/deepseek-v4-flash or z-ai/glm-5.1.
  • Good (more capable) — openai/gpt-5.4 or openai/gpt-5.5 at medium reasoning.
  • Very hard only — openai/gpt-5.5 is the strongest but expensive. Avoid it by default; use it only as a last resort to confirm a task is solvable when the cheaper models get zero reward.

Live dashboard in tmux

The Rich dashboard is on by default and shows a live reward <mean> · err <share> headline, but it doesn't render in a plain captured shell. To watch it live, run inside tmux and drive it with send-keys:

bash
tmux new-session -d -s eval
tmux send-keys -t eval 'uv run eval <taskset-id> -r <r>' Enter
# watch it:
tmux attach -t eval        # or: tmux capture-pane -t eval -p

For headless/automated runs, prefer --no-rich and follow the logs directly.

Inspect output

Each run is saved to a fresh per-run dir (so runs never overwrite each other):

text
outputs/<run-dir>/
  configs/eval.toml           # launch config when the run starts from CLI flags
  configs/resolved/eval.json  # complete config; use this to resume
  traces.jsonl                 # one full episode per line
  logs/attempt_<n>/eval.log    # logs for one run or resume attempt
  logs/latest -> attempt_<n>   # current attempt
bash
# find the newest run
ls -dt outputs/*/ | head

# avg reward across rollouts (recomputed — aggregates aren't stored)
jq -s '[.[].traces[] | [.rewards[]? | .score * .weight] | add // 0] | if length > 0 then add / length else 0 end' outputs/<...>/traces.jsonl

The dashboard shows the avg reward live; with --no-rich, recompute it from traces.jsonl (each line is an episode whose traces[] carry the per-reward breakdown). Skim a few episodes for sanity before publishing.

© PrimeIntellect-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/evaluation of PrimeIntellect-ai/prime-envs.

Open the folder on GitHubat commit 89dcc7f

Compare with similar skills

Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluation this skillPrimeIntellect-ai/prime-envs130—~4.6kAutomated safety check: PassApache-2.0
Eval Driven Devgithub/awesome-copilot40k1 repos~4.4kAutomated safety check: WarnMIT
QA Methodologymagnus919/agent-skills113—~3.2kAutomated safety check: PassMIT
Early Experience DataOSU-NLP-Group/EarlyExperience102—~4.1kAutomated safety check: PassMIT
Eval Guidemicrosoft/eval-guide138—~22kAutomated safety check: WarnMIT
Write A Specdifferent-ai/openwork24k—~3.3kAutomated safety check: PassCustom licence

Similar skills

  • Eval Driven Dev

    github/awesome-copilot

    Official

    Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~4.4k tokens
    Testing & QAAuto-check: warnings
  • QA Methodology

    magnus919/agent-skills

    Design and apply QA methodology for software teams: test strategy, regression testing, CI failure triage, test automation, quality gates and metrics, risk-based testing, exploratory testing, test…

    113 GitHub stars~3.2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Early Experience Data

    OSU-NLP-Group/EarlyExperience

    A skill your agent uses whenever the user asks to generate, collect, inspect, or prepare early-experience training data (Implicit World Modeling or Self-Reflection, in the sense of arXiv:2510.08558)…

    102 GitHub stars~4.1k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Eval Guide

    microsoft/eval-guide

    Official

    Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

    138 GitHub stars~22k tokensUpdated 3 mo ago
    Testing & QAAuto-check: warnings
  • Write A Spec

    different-ai/openwork

    Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer.

    24k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Linecanary Monitor

    CALLE-AI/awesome-phone-call-agents

    Monitor business phone lines and deployed voice agents with LineCanary — scheduled CALL-E test calls that walk the caller journey, assert structured results, diff against baselines and alert on…

    106 GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed

More from PrimeIntellect-ai/prime-envs

  • Prime Tasks

    PrimeIntellect-ai/prime-envs

    Where task data and taskset code live across prime-tasks and prime-envs, and how to fix a broken task.

    130 GitHub stars~504 tokensUpdated today
    Auto-check passed

Questions about Evaluation

What does Evaluation do?

Install and run a verifiers environment — smoke testing during development and full benchmark evals. Evaluation is an agent skill from PrimeIntellect-ai/prime-envs. Install and run a verifiers environment — smoke testing during development and full benchmark evals.

When should I use Evaluation?

Evaluation fits situations like: tasks that involve LLM evaluation; tasks that involve QA and bug reports.

How do I install Evaluation in Claude Code?

Run `npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a claude-code`. Or copy the skill folder (skills/evaluation in PrimeIntellect-ai/prime-envs) into .claude/skills/evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Evaluation in Codex?

Run `npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a codex`. Or copy the skill folder (skills/evaluation in PrimeIntellect-ai/prime-envs) into .agents/skills/evaluation in your project. Codex loads it when a task matches its description.

Can I use Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PrimeIntellect-ai/prime-envs --skill evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation, .gemini/skills/evaluation, .github/skills/evaluation and .opencode/skills/evaluation in your project.

What does Evaluation need to run?

Going by SKILL.md and its folder, Evaluation needs the command-line tools its instructions call (uv and jq) and credentials named HF_TOKEN and SERPER_API_KEY. Our summary lists: Python 3; Docker.

Does Evaluation access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluation use?

Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluation use?

About 4.6k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluation?

Skills that share tags, products or a category with Evaluation: Eval Driven Dev (github/awesome-copilot, 40k stars), QA Methodology (magnus919/agent-skills, 113 stars), Early Experience Data (OSU-NLP-Group/EarlyExperience, 102 stars) and Eval Guide (microsoft/eval-guide, 138 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluation?

PrimeIntellect-ai (a GitHub organization) maintains it in PrimeIntellect-ai/prime-envs, which has 130 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 8, 2026.

Source: PrimeIntellect-ai/prime-envs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.