Agent skill

Generate Verifiers Env

by adithya-s-k in adithya-s-k/FineEnvs

Builds a Verifiers (PrimeIntellect) variant of an RL environment.

Apache-2.0Auto-check passedEducation

Install Generate Verifiers Env

skills CLI
$ npx skills add adithya-s-k/FineEnvs --skill generate-verifiers-env -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install adithya-s-k/FineEnvs generate-verifiers-env --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/adithya-s-k/FineEnvs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/generate-verifiers-env .claude/skills/generate-verifiers-env && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
generate-verifiers-env
GitHub stars
421
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
629 words
Files
2 (incl. references)
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Builds a Verifiers (PrimeIntellect) variant of an RL environment.

  • Works in 4 steps: The toolkit class → Standalone tool functions for vf.ToolEnv → The rubric → …
  • Someone asks to scaffold a Verifiers env
  • SKILL.md covers Concept, Archetypes, Two consumption paths (always… and Recommended file layout, plus 5 more sections
  • Calls uv

What it does

Generate Verifiers Env is an agent skill from adithya-s-k/FineEnvs. Builds a Verifiers (PrimeIntellect) variant of an RL environment. Use whenever someone asks to scaffold a Verifiers env, port to Verifiers, build an in-process toolkit, set up a vf.ToolEnv with a Rubric, or wire up a TRL GRPOTrainer rollout. Verifiers is the right framework when the user wants in-process tools (no HTTP server), structured tool calling driven by plain Python functions, composable reward rubrics with multiple grader functions, fast iteration with no Docker, or the cleanest path from prototype to…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/architecture.md`).

It sits in Education, covering Quizzes and assessments, Reinforcement learning and Structured output and tool calling. It works with Python and Docker. The repository describes itself as: FineEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs. The licence is Apache-2.0.

When your agent uses it

  • Someone asks to scaffold a Verifiers env
  • Port to Verifiers
  • Build an in-process toolkit
  • Set up a vf.ToolEnv with a Rubric

Example prompts

  • “make a verifiers env for X”
  • “wrap my game in verifiers”
  • “set up a vf.ToolEnv”
  • “/generate-verifiers-env”

Requirements

  • Python 3
  • Docker

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. The toolkit class
  2. Standalone tool functions for vf.ToolEnv
  3. The rubric
  4. Rollout — rollout.py

What it can do on your machine

Read from SKILL.md and the folder at commit b0f4c2f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • docs.primeintellect.ai
    • pypi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Generate Verifiers Env loads about 2.3k tokens when it runs, and up to ~3.6k if it reads all its reference files. Until then it costs about 207 tokens; SKILL.md has 629 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~207
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from adithya-s-k/FineEnvs at commit b0f4c2f, republished under its Apache-2.0 licence (© adithya-s-k). 629 words, ~2,322 tokens.

Download SKILL.mdSave it as .claude/skills/generate-verifiers-env/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
generate-verifiers-env
description
Builds a Verifiers (PrimeIntellect) variant of an RL environment. Use whenever someone asks to scaffold a Verifiers env, port to Verifiers, build an in-process toolkit, set up a `vf.ToolEnv` with a Rubric, or wire up a TRL `GRPOTrainer` rollout. Verifiers is the right framework when the user wants in-process tools (no HTTP server), structured tool calling driven by plain Python functions, composable reward rubrics with multiple grader functions, fast iteration with no Docker, or the cleanest path from prototype to TRL training. Output is a runnable `<env_dir>/verifiers/` folder with `env.py` (toolkit + standalone tool functions + `create_verifiers_env`), `rollout.py`, and `pyproject.toml`. Use for prompts like "make a verifiers env for X", "wrap my game in verifiers", or "set up a vf.ToolEnv".

generate-verifiers-env

Build the Verifiers variant of an env. Verifiers is in-process — no HTTP server, no Docker, no HF Space. The trainer (or a manual rollout) imports tool functions directly from env.py.

Concept

PrimeIntellect Verifiers is a Python library — not a server framework. It provides vf.ToolEnv (multi-turn rollout), vf.Rubric (composable async graders), and adapters into TRL GRPOTrainer. The trainer or rollout owns the LLM client; the env owns the tools and the grader.

When the user has a shared domain module (<domain>.py) and wants a Verifiers variant, wrap it as a toolkit class plus standalone tool functions. Don't duplicate domain logic.

Archetypes

ArchetypeHallmarks
Pure-Python gameOne @tool-style function, terminal reward via rubric checking the trajectory.
Stateful sandbox in-processToolkit owns the sandbox (E2B, browser); initialize() is lazy; cleanup() is mandatory in finally.
Vision envDrive the toolkit manually (skip vf.ToolEnv since vision content blocks aren't first-class in verifiers' rollout). Send the screenshot in the user message each turn.

Two consumption paths (always provide both)

Path A — DesktopToolkit-style class (used by TRL adapter + manual rollout)
python
class WordleToolkit:
    def __init__(self): ...
    def initialize(self): ...     # lazy E2B / state init
    def cleanup(self): ...        # kill sandbox
    def reset(self): ...          # new episode
    def guess(self, word: str) -> str:
        """Submit a 5-letter word guess. Returns colored feedback."""
        ...

Public methods are introspected as tools by the TRL adapter. Docstrings become tool descriptions.

Path B — vf.ToolEnv for native verifiers env.evaluate(client, model)
python
def create_verifiers_env():
    import verifiers as vf
    from datasets import Dataset
    dataset = Dataset.from_list([{"question": t["task"], "answer": t["expected_output"]} for t in TASKS])
    async def correctness(completion, answer, **kwargs) -> float:
        # read from the completion trajectory; return 0.0–1.0
        ...
    rubric = vf.Rubric(funcs=[correctness])
    return vf.ToolEnv(tools=TOOL_FUNCTIONS, max_turns=8, dataset=dataset, rubric=rubric, system_prompt="...")

TOOL_FUNCTIONS is a list of plain Python functions (not bound methods). They can share state via a module-level toolkit instance.

The user picks the actual paths. The canonical shape:

<env_dir>/verifiers/
├── pyproject.toml      # verifiers + e2b-* + datasets + python-dotenv + openai
├── __init__.py
├── env.py              # Toolkit class + standalone tool fns + create_verifiers_env()
├── rollout.py          # Drives the toolkit manually with the openai client
└── README.md

Implementation order

1. The toolkit class
  • __init__ takes config (api_key="", app="firefox", etc.). Don't create the sandbox here — too eager.
  • initialize() is the lazy creation hook. Always call it from each tool method.
  • cleanup() kills the sandbox. Always call it from finally in the rollout.
  • reset() calls cleanup() + reinitializes. Used between episodes by the TRL adapter.
  • Each tool method:
    • takes typed args (used for OpenAI tool-schema generation via inspect)
    • has a docstring (becomes the tool description — first paragraph only)
    • calls self.initialize() first, mutates state, returns a string
2. Standalone tool functions for vf.ToolEnv

Module-level shared toolkit, plus thin wrappers:

python
_shared: Optional[WordleToolkit] = None
def _kit():
    global _shared
    if _shared is None:
        _shared = WordleToolkit()
    return _shared

def guess(word: str) -> str:
    """Submit a 5-letter word guess."""
    return _kit().guess(word)

TOOL_FUNCTIONS = [guess]

Why both? The TRL adapter wants the toolkit class (per-rollout instance, isolated state). vf.ToolEnv wants free functions. Don't pick one — provide both.

3. The rubric

Rubrics are composable graders. Each grader is async def func(completion, answer, **kwargs) -> float. Combine multiple in a vf.Rubric(funcs=[...]) and they're averaged (or weighted, see verifiers docs).

For a single-criterion env, one grader suffices:

python
async def correctness(completion, answer, **kwargs) -> float:
    if not completion: return 0.0
    last = completion[-1].get("content", "") if isinstance(completion[-1], dict) else str(completion[-1])
    return 1.0 if answer.strip() in last.strip() else 0.0

For multi-criterion (e.g. computer-use envs that need both terminate(success) AND a state check):

python
async def correctness(completion, answer, **kwargs) -> float:
    seen_success = any("terminated: success" in str(m) for m in completion)
    seen_expected = any(answer in str(m) for m in completion)
    return 1.0 if (seen_success and seen_expected) else (0.5 if seen_success else 0.0)
Show full SKILL.md (232 more words)Show less
4. Rollout — rollout.py

Build OpenAI tool schemas from the function signatures + docstrings via inspect:

python
def func_to_openai_tool(fn):
    sig = inspect.signature(fn)
    hints = get_type_hints(fn)
    doc = (fn.__doc__ or "").strip().split("\n\n")[0]
    properties, required = {}, []
    for name, p in sig.parameters.items():
        ann = hints.get(name, str)
        origin = get_origin(ann)
        if origin in (list, "list"):
            inner = get_args(ann)
            properties[name] = {"type": "array", "items": {"type": "integer" if (inner and inner[0] is int) else "string"}}
        elif ann is int:    properties[name] = {"type": "integer"}
        elif ann is float:  properties[name] = {"type": "number"}
        elif ann is bool:   properties[name] = {"type": "boolean"}
        else:               properties[name] = {"type": "string"}
        if p.default is inspect.Parameter.empty:
            required.append(name)
    return {"type": "function", "function": {
        "name": fn.__name__, "description": doc,
        "parameters": {"type": "object", "properties": properties, "required": required},
    }}

This pattern works for any toolkit. Use it as the standard adapter from Python signatures to OpenAI tool schemas.

For multimodal envs, drive the toolkit manually (don't use vf.ToolEnv since vision-content blocks aren't first-class in verifiers' rollout). Send the latest screenshot in the user message every turn:

python
text, b64 = kit._ctrl.screenshot()      # if you exposed _ctrl
messages.append({"role": "user", "content": [
    {"type": "text", "text": "Latest screenshot:"},
    {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}},
]})

Validation gates

  1. Toolkit imports cleanly — uv run python -c "from env import DesktopToolkit, TOOL_FUNCTIONS"
  2. vf.ToolEnv builds — uv run python -c "from env import create_verifiers_env; env = create_verifiers_env(); print(env)"
  3. Manual rollout — MAX_TURNS=3 uv run python rollout.py runs end-to-end. Hits a real backend (E2B or whatever the env uses).

Gotchas

  • ModuleNotFoundError: attrs — e2b-desktop transitively needs attrs but doesn't pin it. Add attrs>=23.0 to dependencies.
  • TypedDict vs dataclass for verifiers data structures — most are TypedDicts. Access by key, not attribute. (Same trap exists in skyrl-gym; we hit it during the desktop_env port.)
  • Tool-schema **kwargs is forbidden — vLLM (used by some trainers) can't introspect **kwargs for JSON schema generation. Define explicit params, even if empty.
  • Don't return huge strings — verifiers passes the result through to the model verbatim. A 100KB log dump will blow your context. Truncate / summarize in the tool method.

Reference

  • references/architecture.md — vf.ToolEnv internals + Rubric composition + TRL adapter shape

Official documentation

© adithya-s-k, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .claude/skills/generate-verifiers-env of adithya-s-k/FineEnvs.

  • SKILL.md
  • references/architecture.md

Open the folder on GitHubat commit b0f4c2f

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in adithya-s-k/FineEnvs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Generate Verifiers Env next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Generate Verifiers Env compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Generate Verifiers Env this skilladithya-s-k/FineEnvs4211 repos~2.3kAutomated safety check: PassApache-2.0
DeepTutor CLIHKUDS/DeepTutor41k—~2.3kAutomated safety check: PassApache-2.0
Agent Prompt Quality Barmastra-ai/mastra29k—~2kAutomated safety check: PassCustom licence
Auto Improvecrimeacs/auto-improve135—~651Automated safety check: PassMIT
Pre Pushartcc/freelingo158—~732Automated safety check: PassAGPL-3.0
Sciatlas Idea Evaluatezjunlp/SciAtlas160—~2.2kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.3k tokensUpdated 3 days ago
    EducationAuto-check passed
  • Agent Prompt Quality Bar

    mastra-ai/mastra

    Universal quality bar and final audit rubric for any agent system prompt.

    29k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Auto Improve

    crimeacs/auto-improve

    GAN-style iterative improvement loop for any text artifact. An agent skill from crimeacs/auto-improve.

    135 GitHub stars~651 tokensUpdated 2 mo ago
    EducationAuto-check passed
  • Pre Push

    artcc/freelingo

    A skill your agent uses when the user asks to check before pushing, pre-push, verificar antes de pushear, run all checks, or quiere validar que todo pasa antes de hacer push.

    158 GitHub stars~732 tokensUpdated today
    EducationAuto-check passed
  • Sciatlas Idea Evaluate

    zjunlp/SciAtlas

    Use only the current SciAtlas automated review workflow (reviewpipeline) to take a novice user from zero setup to a final automated review of a research idea or paper, including setup, registration…

    160 GitHub stars~2.2k tokensUpdated 7 days ago
    EducationAuto-check: notes
  • Clawpathy Autoresearch

    ClawBio/ClawBio

    Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from adithya-s-k/FineEnvs

All 8 skills in this repo
  • Generate Nemo Gym Env

    adithya-s-k/FineEnvs

    Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs.

    421 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Generate Openenv Env

    adithya-s-k/FineEnvs

    Builds an OpenEnv (Hugging Face) variant of an RL environment.

    421 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Generate Ors Env

    adithya-s-k/FineEnvs

    Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package.

    421 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check: notes
  • Rl Env From Description

    adithya-s-k/FineEnvs

    Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym.

    421 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check: notes
  • Article Frontmatter

    adithya-s-k/FineEnvs

    Configure article metadata via MDX frontmatter. An agent skill from adithya-s-k/FineEnvs.

    421 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Create HTML Embed

    adithya-s-k/FineEnvs

    Create self-contained D3 HTML embed charts for the research article template.

    421 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Generate Verifiers Env

What does Generate Verifiers Env do?

Builds a Verifiers (PrimeIntellect) variant of an RL environment. Generate Verifiers Env is an agent skill from adithya-s-k/FineEnvs. Builds a Verifiers (PrimeIntellect) variant of an RL environment.

When should I use Generate Verifiers Env?

Generate Verifiers Env fits situations like: someone asks to scaffold a Verifiers env; port to Verifiers; build an in-process toolkit; set up a vf.ToolEnv with a Rubric.

How do I install Generate Verifiers Env in Claude Code?

Run `npx skills add adithya-s-k/FineEnvs --skill generate-verifiers-env -a claude-code`. Or copy the skill folder (.claude/skills/generate-verifiers-env in adithya-s-k/FineEnvs) into .claude/skills/generate-verifiers-env in your project. Claude Code loads it when a task matches its description.

How do I install Generate Verifiers Env in Codex?

Run `npx skills add adithya-s-k/FineEnvs --skill generate-verifiers-env -a codex`. Or copy the skill folder (.claude/skills/generate-verifiers-env in adithya-s-k/FineEnvs) into .agents/skills/generate-verifiers-env in your project. Codex loads it when a task matches its description.

Can I use Generate Verifiers Env in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adithya-s-k/FineEnvs --skill generate-verifiers-env -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generate-verifiers-env, .gemini/skills/generate-verifiers-env, .github/skills/generate-verifiers-env and .opencode/skills/generate-verifiers-env in your project.

What does Generate Verifiers Env need to run?

Going by SKILL.md and its folder, Generate Verifiers Env needs the command-line tools its instructions call (uv). Our summary lists: Python 3; Docker.

Does Generate Verifiers Env access the network?

SKILL.md names 3 domains. As links in the text: github.com, docs.primeintellect.ai and pypi.org. This is read from the text; nothing was executed.

Is Generate Verifiers Env safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Generate Verifiers Env use?

Generate Verifiers Env is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Generate Verifiers Env use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Generate Verifiers Env?

Skills that share tags, products or a category with Generate Verifiers Env: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), Agent Prompt Quality Bar (mastra-ai/mastra, 29k stars), Auto Improve (crimeacs/auto-improve, 135 stars) and Pre Push (artcc/freelingo, 158 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Generate Verifiers Env?

adithya-s-k (a GitHub user) maintains it in adithya-s-k/FineEnvs, which has 421 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 6, 2026.

Source: adithya-s-k/FineEnvs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.