Agent skill

Rl Env From Description

by adithya-s-k in adithya-s-k/FineEnvs

Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Rl Env From Description

skills CLI
$ npx skills add adithya-s-k/FineEnvs --skill rl-env-from-description -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install adithya-s-k/FineEnvs rl-env-from-description --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/adithya-s-k/FineEnvs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/rl-env-from-description .claude/skills/rl-env-from-description && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rl-env-from-description
GitHub stars
461
Token cost
~2.3k tokens
SKILL.md length
978 words
Files
6 (incl. references)
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym.

  • Works in 4 steps: Interview the user (focused, not… → Pick the closest archetype → Implement in dependency order → …
  • Someone describes an environment they want to build (I want to train an agent that does X
  • SKILL.md covers When to use, Recommended layout (suggest,…, The four-step flow and What success looks like, plus 3 more sections
  • Needs OPENAI_API_KEY

What it does

Rl Env From Description is an agent skill from adithya-s-k/FineEnvs. Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym. Use whenever someone describes an environment they want to build ("I want to train an agent that does X", "make an env where the model has to Y"), asks to scaffold a new env, asks to port an existing env to one of these frameworks, or asks how to design tools/rewards/state for a new env. Use even when the user does not explicitly say "RL…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/interview.md`, `references/nemo_gym.md` and `references/openenv.md`).

It sits in AI & LLM Engineering, covering Reinforcement learning, Plain language and style rules and Structured output and tool calling. It works with SQL. The repository describes itself as: FineEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs. The licence is Apache-2.0.

When your agent uses it

  • Someone describes an environment they want to build (I want to train an agent that does X
  • Make an env where the model has to Y)
  • Asks to scaffold a new env
  • Asks to port an existing env to one of these frameworks

Example prompts

  • “I want to train an agent that does X”
  • “make an env where the model has to Y”
  • “RL environment”
  • “/rl-env-from-description”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Interview the user (focused, not exhaustive)
  2. Pick the closest archetype
  3. Implement in dependency order
  4. Validate end-to-end before declaring done

What it can do on your machine

Read from SKILL.md and the folder at commit 5e0b46e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • huggingface.co
    • openrewardstandard.io
    • docs.openreward.ai
    • pypi.org
    • docs.primeintellect.ai
    • docs.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rl Env From Description loads about 2.3k tokens when it runs, and up to ~5.5k if it reads all its reference files. Until then it costs about 206 tokens; SKILL.md has 978 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~206
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:89
    etc.), check for the relevant secret in `.env` and stop with a clear error if it's missing.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from adithya-s-k/FineEnvs at commit 5e0b46e, republished under its Apache-2.0 licence (© adithya-s-k). 978 words, ~2,326 tokens.

Download SKILL.mdSave it as .claude/skills/rl-env-from-description/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
rl-env-from-description
description
Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym. Use whenever someone describes an environment they want to build ("I want to train an agent that does X", "make an env where the model has to Y"), asks to scaffold a new env, asks to port an existing env to one of these frameworks, or asks how to design tools/rewards/state for a new env. Use even when the user does not explicitly say "RL environment" — descriptions like "agent that browses the web", "tool-calling agent for SQL", or "game-playing agent" all qualify. Drives the full flow — clarifying interview, env-name selection, shared-domain extraction, per-framework implementation, and rollout-based smoke tests.

RL Env From Description

Convert a plain-English description of an RL training environment into runnable code across OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym. Two other framework variants (SkyRL Gym, GEM) are secondary and only relevant for text-action-with-tag-parsing envs — produce them only if the user asks.

When to use

  • A user describes an env in plain English (a goal, an action surface, or a reward shape) and wants code.
  • A user asks to "build an env for X", "scaffold an RL env", "port this env to OpenEnv/ORS/Verifiers/NeMo Gym".
  • A user already has a runnable env in one framework and wants the same env in others.
  • A user asks "what's the right way to design my reward / state / tool surface for this task" — start with the interview below, then implement.

Do not use for: training runs (TRL/GRPO config), evaluation harness work, or general agent-design questions that don't end with new env code.

A clean shape that scales well — but the user gets to pick the actual paths:

<env_dir>/                          # whatever the user names it
├── <domain>.py                     # SHARED pure logic (e.g. game.py)
├── tasks.py                        # SHARED list of task dicts (optional)
├── openenv/                        # OpenEnv variant (HTTP, MCP)
├── ors/                            # ORS variant (HTTP, REST + SSE)
├── verifiers/                      # Verifiers variant (in-process)
└── nemo_gym/                       # NeMo Gym variant (HTTP, REST + cookies)

Inside each framework folder, the public contract is:

  • pyproject.toml — framework-specific deps
  • __init__.py
  • One implementation file (server.py for ORS, server/<env>_environment.py for OpenEnv, env.py for Verifiers, server.py for NeMo Gym)
  • rollout.py — runs an LLM against the env end-to-end
  • README.md — one-page consumption guide

Always ask the user where they want files written. If they don't have a preference, propose the layout above. Don't force it.

The four-step flow

Step 1 — Interview the user (focused, not exhaustive)

Ask only the questions that determine architecture. The full bank lives in references/interview.md; the must-cover set is:

  1. What does the agent DO? One sentence describing the goal and the loop.
  2. Action surface — structured tool calls (most cases) or free text with tag parsing (rare; only when the model has no tool-calling support).
  3. State — does anything persist across turns? Per-session sandbox? In-memory dict? Nothing?
  4. Reward — when does it fire (per-step, on terminate, post-episode)? What's the success criterion?
  5. External backends — sandbox (E2B), web service, none?
  6. Termination — fixed turn cap, model-emits-terminate, or a derived condition.
  7. Where should the files live? Project-relative path; never assume.

If the user has already given enough signal in their description (e.g. they cited an existing env they want to mirror), skip questions whose answers are obvious. Don't make people repeat themselves.

When in doubt about an architectural choice, propose a default with a one-line rationale and let the user veto.

Step 2 — Pick the closest archetype

Match the user's description to one of these archetypes; tell them which archetype you're using and why:

ArchetypeHallmarksTypical reward shape
Pure-Python gameDeterministic, single tool, no external services, multi-turnTerminal reward (1.0/0.0) or per-step from game state
Stateful sandboxReal backend (E2B Code Interpreter, browser, DB), structured tool calls, state persists across callsExternal grader (string match, unit tests, LLM judge)
Vision / computer-useScreenshots + mouse/keyboard, 19-tool action surface modelled on Anthropic's computer_20251124Terminal reward via terminate(status) tool
Text-action with parsingModel emits free text containing tags; env parses (use only if the model has no tool-calling support)Per-step from parsed action results
Step 3 — Implement in dependency order

The shared module first, then per-framework variants. Order doesn't matter between frameworks.

  1. <env_dir>/<domain>.py + <env_dir>/tasks.py — the only file that contains domain logic. Frameworks just wrap it. Keep it pure-Python; no framework imports.
  2. OpenEnv — read references/openenv.md (planner-level) or trigger generate-openenv-env skill (full workflow). Use MCPEnvironment + @mcp.tool + create_app(...) in server/app.py.
  3. ORS — read references/ors.md or trigger generate-ors-env. Use Environment + @tool methods + ToolOutput(blocks=[...], reward=..., finished=...). Per-tool-call reward is the framework's defining feature.
  4. Verifiers — read references/verifiers.md or trigger generate-verifiers-env. Plain Python tool functions on a toolkit class; vf.ToolEnv + vf.Rubric for native consumption.
  5. NeMo Gym — read references/nemo_gym.md or trigger generate-nemo-gym-env. SimpleResourcesServer with one app.post("/<tool>") per tool; cookie sessions; post-episode /verify reward.
Show full SKILL.md (350 more words)Show less
Step 4 — Validate end-to-end before declaring done

Each framework folder gets ONE smoke rollout against a small LLM (Qwen via HF Router by default, or OpenAI if OPENAI_API_KEY is set). The rollout must:

  • Discover the tools the env exposes (don't hardcode names — except for NeMo Gym, which has no list_tools()).
  • Drive a 3–5 turn loop and print every tool call + result.
  • Fail loudly if a tool call errors. (MAX_TURNS=3 for the smoke check.)

If the env needs an external backend (E2B, etc.), check for the relevant secret in .env and stop with a clear error if it's missing.

What success looks like

A user typing "make me an env where the agent plays connect-four at path/to/connect_four/" should end with:

  • path/to/connect_four/game.py (the shared engine), tasks.py (a few starting positions)
  • path/to/connect_four/openenv/, .../ors/, .../verifiers/, .../nemo_gym/ all runnable
  • 4 green rollout smoke tests (NeMo Gym tested via deployed Space if local Ray init fails on shared nodes)
  • Whatever README convention the project uses, updated

…in one continuous flow, with the user only answering 5–7 questions along the way.

Reference docs

  • references/interview.md — full question bank with example answers
  • references/openenv.md — OpenEnv-specific implementation notes (planner-level; defers to generate-openenv-env)
  • references/ors.md — ORS planner-level (defers to generate-ors-env)
  • references/verifiers.md — Verifiers planner-level (defers to generate-verifiers-env)
  • references/nemo_gym.md — NeMo Gym planner-level (defers to generate-nemo-gym-env)

When the user wants only one framework variant, trigger the framework-specific skill directly: generate-openenv-env, generate-ors-env, generate-verifiers-env, or generate-nemo-gym-env.

Hard guardrails

  • Don't impose a folder layout. Suggest the recommended one once; respect the user's choice if they want different paths.
  • Don't skip the shared domain module. Cross-framework consistency is impossible without it. Every framework variant must wrap the same <domain>.py — never duplicate logic.
  • Don't run training. This skill ends with rollouts. Training/eval is a separate concern.
  • Don't invent APIs. When unsure about a framework's actual call shape, read its references/architecture.md (in the framework-specific skill) before writing code.
  • Coordinate spaces matter for vision envs. Declare the convention (pixel vs normalized 0–1000) in the prompt. Qwen2.5-VL emits 0–1000 normalized; the rollout adapter must rescale.

Official documentation

© adithya-s-k, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in .claude/skills/rl-env-from-description of adithya-s-k/FineEnvs.

  • SKILL.md
  • references/interview.md
  • references/nemo_gym.md
  • references/openenv.md
  • references/ors.md
  • references/verifiers.md

Open the folder on GitHubat commit 5e0b46e

Compare with similar skills

Rl Env From Description next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rl Env From Description compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rl Env From Description this skilladithya-s-k/FineEnvs461—~2.3kAutomated safety check: NotesApache-2.0
AI ToolsDbxstudio/dbx-studio1021 repos~643Automated safety check: PassApache-2.0
Slime Useryzlnew/infra-skills149—~3.2kAutomated safety check: PassNone
Building Agentsericrisco/rsc-harness180—~5kAutomated safety check: PassMIT
SQL Explainermergisi/awesome-openclaw-agents4k—~300Automated safety check: PassMIT
Planning With Filesjarrodwatts/claude-code-config1.1k5 repos~967Automated safety check: PassNone

Similar skills

  • AI Tools

    Dbxstudio/dbx-studio

    Reference for all AI tools available in DBX Studio's AI chat system.

    102 GitHub starsUsed in 1 repo~643 tokens
    AI & LLM EngineeringAuto-check passed
  • Slime User

    yzlnew/infra-skills

    Guide for using SLIME (LLM post-training framework for RL Scaling).

    149 GitHub stars~3.2k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Building Agents

    ericrisco/rsc-harness

    A skill your agent uses when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server —…

    180 GitHub stars~5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • SQL Explainer

    mergisi/awesome-openclaw-agents

    Paste a SQL query and get a plain English explanation of what it does.

    4k GitHub stars~300 tokensUpdated 13 days ago
    DatabasesAuto-check passed
  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from adithya-s-k/FineEnvs

All 8 skills in this repo
  • Generate Verifiers Env

    adithya-s-k/FineEnvs

    Builds a Verifiers (PrimeIntellect) variant of an RL environment.

    461 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Generate Nemo Gym Env

    adithya-s-k/FineEnvs

    Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs.

    461 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Generate Openenv Env

    adithya-s-k/FineEnvs

    Builds an OpenEnv (Hugging Face) variant of an RL environment.

    461 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Generate Ors Env

    adithya-s-k/FineEnvs

    Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package.

    461 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check: notes
  • Article Frontmatter

    adithya-s-k/FineEnvs

    Configure article metadata via MDX frontmatter. An agent skill from adithya-s-k/FineEnvs.

    461 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Create HTML Embed

    adithya-s-k/FineEnvs

    Create self-contained D3 HTML embed charts for the research article template.

    461 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Rl Env From Description

What does Rl Env From Description do?

Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym. Rl Env From Description is an agent skill from adithya-s-k/FineEnvs. Turns a user's plain-English description of an RL training environment into runnable code across the four target frameworks — OpenEnv, OpenReward (ORS), Verifiers, and NeMo Gym.

When should I use Rl Env From Description?

Rl Env From Description fits situations like: someone describes an environment they want to build (I want to train an agent that does X; make an env where the model has to Y); asks to scaffold a new env; asks to port an existing env to one of these frameworks.

How do I install Rl Env From Description in Claude Code?

Run `npx skills add adithya-s-k/FineEnvs --skill rl-env-from-description -a claude-code`. Or copy the skill folder (.claude/skills/rl-env-from-description in adithya-s-k/FineEnvs) into .claude/skills/rl-env-from-description in your project. Claude Code loads it when a task matches its description.

How do I install Rl Env From Description in Codex?

Run `npx skills add adithya-s-k/FineEnvs --skill rl-env-from-description -a codex`. Or copy the skill folder (.claude/skills/rl-env-from-description in adithya-s-k/FineEnvs) into .agents/skills/rl-env-from-description in your project. Codex loads it when a task matches its description.

Can I use Rl Env From Description in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add adithya-s-k/FineEnvs --skill rl-env-from-description -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rl-env-from-description, .gemini/skills/rl-env-from-description, .github/skills/rl-env-from-description and .opencode/skills/rl-env-from-description in your project.

What does Rl Env From Description need to run?

Going by SKILL.md and its folder, Rl Env From Description needs credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Rl Env From Description access the network?

SKILL.md names 7 domains. As links in the text: github.com, huggingface.co, openrewardstandard.io, docs.openreward.ai, pypi.org, docs.primeintellect.ai and docs.nvidia.com. This is read from the text; nothing was executed.

Is Rl Env From Description safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Rl Env From Description use?

Rl Env From Description is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rl Env From Description use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.2k tokens, read only when the agent opens those files.

What are the alternatives to Rl Env From Description?

Skills that share tags, products or a category with Rl Env From Description: AI Tools (Dbxstudio/dbx-studio, 102 stars), Slime User (yzlnew/infra-skills, 149 stars), Building Agents (ericrisco/rsc-harness, 180 stars) and SQL Explainer (mergisi/awesome-openclaw-agents, 4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rl Env From Description?

adithya-s-k (a GitHub user) maintains it in adithya-s-k/FineEnvs, which has 461 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 8, 2026.

Source: adithya-s-k/FineEnvs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.