Agent skill

Optimize Nethack Policy

by Linzwcs in Linzwcs/EvoPolicyGym

Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark.

MITAuto-check passed

Install Optimize Nethack Policy

skills CLI
$ npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Linzwcs/EvoPolicyGym optimize-nethack-policy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Linzwcs/EvoPolicyGym.git skills-src && mkdir -p .claude/skills && cp -r skills-src/environments/nle/nethack/skills/optimize-nethack-policy .claude/skills/optimize-nethack-policy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
optimize-nethack-policy
GitHub stars
175
Token cost
~1.9k tokens
SKILL.md length
929 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark.

  • Works in 6 steps: Prompt handler: respond to , direction,… → Map memory: combine visible characters,… → Immediate safety: monitor HP, hunger,… → …
  • SKILL.md covers Respect the ABI, Actions, Observation and Build an Episode-local…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Optimize Nethack Policy is an agent skill from Linzwcs/EvoPolicyGym. Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: EvoPolicyGym is infrastructure for evaluating coding agents and generating training experience through Autonomous Policy Evolution. The licence is MIT.

Example prompts

  • “/optimize-nethack-policy”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Prompt handler: respond to , direction, eating, and yes/no
  2. Map memory: combine visible characters, colors, glyphs, current position,
  3. Immediate safety: monitor HP, hunger, conditions, nearby monsters, and
  4. Exploration: prefer reachable low-visit frontiers, search near plausible
  5. Resource state: track public inventory descriptions and observed outcome
  6. Progress plan: descend when prepared, collect useful public items, and

What it can do on your machine

Read from SKILL.md and the folder at commit a3d9669. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Optimize Nethack Policy loads about 1.9k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 929 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Linzwcs/EvoPolicyGym at commit a3d9669, republished under its MIT licence (© Linzwcs). 929 words, ~1,852 tokens.

Download SKILL.mdSave it as .claude/skills/optimize-nethack-policy/SKILL.md (or your agent's skills folder).
name
optimize-nethack-policy
description
Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark.

Optimize an NLE NetHack Policy

Maximize mean shaped return in deterministic but hidden NetHackScore Episodes. The task uses NLE 1.3.0, NetHack 3.6.7, a fixed neutral human male Monk, the standard 23-action task set, and at most 5,000 Policy steps. Each Policy failure receives a large negative return, so validity and termination come first.

Respect the ABI

  • Export make_policy(context) from policy.py and return an object with act(observation).
  • Return an exact integer from 0 through 22. Booleans, floats, containers, and out-of-range integers are invalid; the Benchmark never repairs an Action.
  • A fresh Policy process and instance are created for every Episode. Keep map, inventory, prompt, and plan state only between act() calls in that Episode.
  • Read public static configuration from context.environment_parameters. Never infer, request, or search for Environment seeds or Host paths.

Actions

text
 0 more               8 northwest         16 run_northwest
 1 north              9 run_north         17 up
 2 east              10 run_east          18 down
 3 south             11 run_south         19 wait
 4 west              12 run_west          20 kick
 5 northeast         13 run_northeast     21 eat
 6 southeast         14 run_southeast     22 search
 7 southwest         15 run_southwest

Actions are NLE keyboard inputs. In a prompt, the same raw key can mean an answer rather than movement. For example Actions 1 through 8 emit k l j h u n b y; Action 21 emits e, Action 22 emits s, and Action 0 emits Enter. Inspect message and input_mode before interpreting an Action as an ordinary command.

Observation

The observation is a dictionary:

  • screen.glyphs: int16 TensorValue, shape (21, 79);
  • screen.chars: uint8 TensorValue, shape (21, 79), row-major visible character codes;
  • screen.colors: uint8 TensorValue, shape (21, 79);
  • stats: named public NLE blstats, including x, y, score, HP, depth, gold, experience, turn, hunger, encumbrance, dungeon level, and conditions;
  • message: the current public NetHack message;
  • inventory: only populated entries with letter, description, glyph, and object_class;
  • input_mode: normal, yes_no, get_line, or more.

The Policy does not receive NLE's privileged internal array, raw seeds, ttyrec paths, rewards, or evaluator state. Decode tensors from TensorValue.data; do not import or reach into the trusted NLE Environment.

Build an Episode-local controller

Use explicit layers that can be tested separately:

  1. Prompt handler: respond to <More>, direction, eating, and yes/no questions before normal navigation. A prompt that is mistaken for movement commonly creates frozen-step penalties or loops.
  2. Map memory: combine visible characters, colors, glyphs, current position, and dungeon level. Track visited, blocked, dangerous, door, corridor, staircase, item, and monster cells without assuming a hidden seed route.
  3. Immediate safety: monitor HP, hunger, conditions, nearby monsters, and escape space. Do not start a long run or repeated search while threatened.
  4. Exploration: prefer reachable low-visit frontiers, search near plausible dead ends, handle doors deliberately, and avoid immediate reversal loops.
  5. Resource state: track public inventory descriptions and observed outcome messages. Initiating eat and selecting an inventory letter are separate decisions.
  6. Progress plan: descend when prepared, collect useful public items, and balance score gain against survival and frozen-step penalties.

The standard 23-action profile is deliberately narrower than a full NetHack keyboard. Do not write plans that require unavailable commands such as wield, wear, quaff, read, open, or arbitrary inventory letters.

Optimize the actual metric

The scalar score is mean Episode return from NetHackScore-v0: NetHack score deltas plus a -0.01 penalty when repeated Environment steps do not advance game time. mean_game_score, depth, deaths, ascensions, truncations, and Policy failures are secondary diagnostics. A Policy failure receives -max_episode_steps, not the partial return accumulated before failure. frozen_steps, mean_frozen_steps, frozen_step_fraction, and mean_frozen_penalty expose the exact aggregate contribution of unchanged-turn Actions to this score; they do not select or summarize trajectory content.

Small batches have high procedural variance. First eliminate failures and frozen prompt loops, then compare unchanged candidates on equal larger batches. Use 1--4 Episodes per search submission until memory and runtime are measured.

Show full SKILL.md (347 more words)Show less

Use Feedback safely

Training Feedback includes aggregate scores plus complete raw evidence for all Episodes in that submission. Read artifact-manifest.json, then use its ordinals to align every trajectory-*.jsonl.gz transition with the exact Policy-visible observations in observations-*.npz. Open NPZ files with numpy.load(..., allow_pickle=False). No Environment or Host code has already selected frames, created screenshots, rendered videos, or decided which segments are important.

Choose your own analysis. You may decode terminal chars and colors, inspect glyphs or named stats, render selected observations with Pillow, export a segment with ImageIO, compare messages and inventory, or use another derived representation. Keep non-submitted scripts, selected frames, derived images, and notes under analysis/; only program/ is submitted. Bulk evidence from older submissions can be evicted, while the newest complete submission is protected, so copy only the derived material you decide is useful before a later submission.

Validation and Assessment are Host-only aggregate phases and publish no detailed Artifacts back to the workspace. Never attempt to recover hidden seeds, Host paths, ttyrec data, or privileged simulator state from Feedback. Treat small-batch score movement as noisy and use equal Episode counts when comparing candidates.

Before finish, compare the candidate Program digests and the actual source changes you intended. Submit behaviorally distinct candidates that test real strategy alternatives. Renaming variables, retaining unused branches, or passing a flag that does not change reachable behavior wastes a Validation slot even though it creates a different source digest. Exact aggregate ties on the same hidden Validation Episodes are a reason to inspect candidate distinctness, not evidence that the Host should inspect or rewrite Programs.

Iterate in stages

  1. Submit the packaged baseline on a small batch and inspect the manifest.
  2. Select and decode the observations or trajectory segments needed to explain failures and repeated behavior.
  3. Make prompt handling total and eliminate invalid Actions/timeouts.
  4. Add persistent dungeon-level map memory and loop detection.
  5. Improve frontier choice, door/search behavior, and staircase handling.
  6. Add HP, hunger, inventory, monster, and condition-aware interrupts.
  7. Remove behaviorally redundant candidates, then compare promising distinct candidates with equal Validation budgets before final handoff.

© Linzwcs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in environments/nle/nethack/skills/optimize-nethack-policy of Linzwcs/EvoPolicyGym.

Open the folder on GitHubat commit a3d9669

Compare with similar skills

Optimize Nethack Policy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Optimize Nethack Policy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Optimize Nethack Policy this skillLinzwcs/EvoPolicyGym175—~1.9kAutomated safety check: PassMIT
Benchmark Optimization Loopaffaan-m/ECC276k1 repos~664Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k3 repos~654Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k—~412Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k—~330Automated safety check: PassMIT
SQL Optimizationgithub/awesome-copilot40k2 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…

    276k GitHub starsUsed in 1 repo~664 tokens
    Auto-check passed
  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    276k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed
  • Benchmark

    affaan-m/ECC

    このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

    276k GitHub stars~412 tokensUpdated 4 days ago
    Auto-check passed
  • Benchmark

    affaan-m/ECC

    使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。

    276k GitHub stars~330 tokensUpdated 4 days ago
    Auto-check passed
  • SQL Optimization

    github/awesome-copilot

    Official

    Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…

    40k GitHub starsUsed in 2 repos~2.3k tokens
    DatabasesAuto-check passed
  • Benchmark

    androidx/androidx

    Benchmarking and improving the performance of Jetpack Compose.

    6.1k GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed

More from Linzwcs/EvoPolicyGym

  • Optimize Balatro Policy

    Linzwcs/EvoPolicyGym

    Build, modularize, improve, test, and select a reward-aligned EvoPolicyGym Bot system for the Jackdaw Balatro Benchmark.

    175 GitHub stars~4.8k tokensUpdated 20 days ago
    Auto-check passed
  • Evopolicygym

    Linzwcs/EvoPolicyGym

    Operate and extend EvoPolicyGym through its public SDK. An agent skill from Linzwcs/EvoPolicyGym.

    175 GitHub stars~890 tokensUpdated 20 days ago
    Auto-check passed
  • Improve an EvoPolicyGym Policy Program for Crafter local-symbolic-v1 observations.

    175 GitHub stars~805 tokensUpdated 20 days ago
    Auto-check passed
  • Optimize Crafter Policy

    Linzwcs/EvoPolicyGym

    Improve an EvoPolicyGym Policy Program for the additive RGB Crafter survival-development Benchmark.

    175 GitHub stars~1.7k tokensUpdated 20 days ago
    Auto-check passed

Questions about Optimize Nethack Policy

What does Optimize Nethack Policy do?

Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark. Optimize Nethack Policy is an agent skill from Linzwcs/EvoPolicyGym. Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark.

How do I install Optimize Nethack Policy in Claude Code?

Run `npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a claude-code`. Or copy the skill folder (environments/nle/nethack/skills/optimize-nethack-policy in Linzwcs/EvoPolicyGym) into .claude/skills/optimize-nethack-policy in your project. Claude Code loads it when a task matches its description.

How do I install Optimize Nethack Policy in Codex?

Run `npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a codex`. Or copy the skill folder (environments/nle/nethack/skills/optimize-nethack-policy in Linzwcs/EvoPolicyGym) into .agents/skills/optimize-nethack-policy in your project. Codex loads it when a task matches its description.

Can I use Optimize Nethack Policy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimize-nethack-policy, .gemini/skills/optimize-nethack-policy, .github/skills/optimize-nethack-policy and .opencode/skills/optimize-nethack-policy in your project.

What does Optimize Nethack Policy need to run?

SKILL.md names no scripts, command-line tools or credentials: Optimize Nethack Policy is instructions for the agent only.

Does Optimize Nethack Policy access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Optimize Nethack Policy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Optimize Nethack Policy use?

Optimize Nethack Policy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Optimize Nethack Policy use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Optimize Nethack Policy?

Skills that share tags, products or a category with Optimize Nethack Policy: Benchmark Optimization Loop (affaan-m/ECC, 276k stars), Benchmark (affaan-m/ECC, 276k stars), Benchmark (affaan-m/ECC, 276k stars) and Benchmark (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Optimize Nethack Policy?

Linzwcs (a GitHub user) maintains it in Linzwcs/EvoPolicyGym, which has 175 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 18, 2026.

Source: Linzwcs/EvoPolicyGym on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.