Benchmark Optimization Loop
affaan-m/ECC
Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…
Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark.
$ npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Linzwcs/EvoPolicyGym optimize-nethack-policy --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Linzwcs/EvoPolicyGym.git skills-src && mkdir -p .claude/skills && cp -r skills-src/environments/nle/nethack/skills/optimize-nethack-policy .claude/skills/optimize-nethack-policy && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "optimize-nethack-policy" agent skill from https://github.com/Linzwcs/EvoPolicyGym/tree/main/environments/nle/nethack/skills/optimize-nethack-policy into .claude/skills/optimize-nethack-policy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-nethack-policy", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Linzwcs/EvoPolicyGym/tree/main/environments/nle/nethack/skills/optimize-nethack-policyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Linzwcs/EvoPolicyGym optimize-nethack-policy --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Linzwcs/EvoPolicyGym.git skills-src && mkdir -p .agents/skills && cp -r skills-src/environments/nle/nethack/skills/optimize-nethack-policy .agents/skills/optimize-nethack-policy && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "optimize-nethack-policy" agent skill from https://github.com/Linzwcs/EvoPolicyGym/tree/main/environments/nle/nethack/skills/optimize-nethack-policy into .agents/skills/optimize-nethack-policy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-nethack-policy", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Linzwcs/EvoPolicyGym optimize-nethack-policy --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Linzwcs/EvoPolicyGym.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/environments/nle/nethack/skills/optimize-nethack-policy .cursor/skills/optimize-nethack-policy && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "optimize-nethack-policy" agent skill from https://github.com/Linzwcs/EvoPolicyGym/tree/main/environments/nle/nethack/skills/optimize-nethack-policy into .cursor/skills/optimize-nethack-policy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-nethack-policy", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Linzwcs/EvoPolicyGym.git --path environments/nle/nethack/skills/optimize-nethack-policy--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Linzwcs/EvoPolicyGym optimize-nethack-policy --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Linzwcs/EvoPolicyGym.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/environments/nle/nethack/skills/optimize-nethack-policy .gemini/skills/optimize-nethack-policy && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "optimize-nethack-policy" agent skill from https://github.com/Linzwcs/EvoPolicyGym/tree/main/environments/nle/nethack/skills/optimize-nethack-policy into .gemini/skills/optimize-nethack-policy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-nethack-policy", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Linzwcs/EvoPolicyGym optimize-nethack-policyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Linzwcs/EvoPolicyGym.git skills-src && mkdir -p .github/skills && cp -r skills-src/environments/nle/nethack/skills/optimize-nethack-policy .github/skills/optimize-nethack-policy && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "optimize-nethack-policy" agent skill from https://github.com/Linzwcs/EvoPolicyGym/tree/main/environments/nle/nethack/skills/optimize-nethack-policy into .github/skills/optimize-nethack-policy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-nethack-policy", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Linzwcs/EvoPolicyGym optimize-nethack-policy --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Linzwcs/EvoPolicyGym.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/environments/nle/nethack/skills/optimize-nethack-policy .opencode/skills/optimize-nethack-policy && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "optimize-nethack-policy" agent skill from https://github.com/Linzwcs/EvoPolicyGym/tree/main/environments/nle/nethack/skills/optimize-nethack-policy into .opencode/skills/optimize-nethack-policy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "optimize-nethack-policy", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
optimize-nethack-policyImprove an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark.
Optimize Nethack Policy is an agent skill from Linzwcs/EvoPolicyGym. Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: EvoPolicyGym is infrastructure for evaluating coding agents and generating training experience through Autonomous Policy Evolution. The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit a3d9669. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Optimize Nethack Policy loads about 1.9k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 929 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Linzwcs/EvoPolicyGym at commit a3d9669, republished under its MIT licence (© Linzwcs). 929 words, ~1,852 tokens.
.claude/skills/optimize-nethack-policy/SKILL.md (or your agent's skills folder).Maximize mean shaped return in deterministic but hidden NetHackScore Episodes. The task uses NLE 1.3.0, NetHack 3.6.7, a fixed neutral human male Monk, the standard 23-action task set, and at most 5,000 Policy steps. Each Policy failure receives a large negative return, so validity and termination come first.
make_policy(context) from policy.py and return an object with
act(observation).act() calls in that Episode.context.environment_parameters.
Never infer, request, or search for Environment seeds or Host paths. 0 more 8 northwest 16 run_northwest
1 north 9 run_north 17 up
2 east 10 run_east 18 down
3 south 11 run_south 19 wait
4 west 12 run_west 20 kick
5 northeast 13 run_northeast 21 eat
6 southeast 14 run_southeast 22 search
7 southwest 15 run_southwestActions are NLE keyboard inputs. In a prompt, the same raw key can mean an
answer rather than movement. For example Actions 1 through 8 emit k l j h u n b y; Action 21 emits e, Action 22 emits s, and Action 0 emits Enter. Inspect
message and input_mode before interpreting an Action as an ordinary command.
The observation is a dictionary:
screen.glyphs: int16 TensorValue, shape (21, 79);screen.chars: uint8 TensorValue, shape (21, 79), row-major visible
character codes;screen.colors: uint8 TensorValue, shape (21, 79);stats: named public NLE blstats, including x, y, score, HP, depth,
gold, experience, turn, hunger, encumbrance, dungeon level, and conditions;message: the current public NetHack message;inventory: only populated entries with letter, description, glyph,
and object_class;input_mode: normal, yes_no, get_line, or more.The Policy does not receive NLE's privileged internal array, raw seeds,
ttyrec paths, rewards, or evaluator state. Decode tensors from
TensorValue.data; do not import or reach into the trusted NLE Environment.
Use explicit layers that can be tested separately:
<More>, direction, eating, and yes/no
questions before normal navigation. A prompt that is mistaken for movement
commonly creates frozen-step penalties or loops.eat and selecting an inventory letter are separate
decisions.The standard 23-action profile is deliberately narrower than a full NetHack keyboard. Do not write plans that require unavailable commands such as wield, wear, quaff, read, open, or arbitrary inventory letters.
The scalar score is mean Episode return from NetHackScore-v0: NetHack score
deltas plus a -0.01 penalty when repeated Environment steps do not advance
game time. mean_game_score, depth, deaths, ascensions, truncations, and Policy
failures are secondary diagnostics. A Policy failure receives
-max_episode_steps, not the partial return accumulated before failure.
frozen_steps, mean_frozen_steps, frozen_step_fraction, and
mean_frozen_penalty expose the exact aggregate contribution of unchanged-turn
Actions to this score; they do not select or summarize trajectory content.
Small batches have high procedural variance. First eliminate failures and frozen prompt loops, then compare unchanged candidates on equal larger batches. Use 1--4 Episodes per search submission until memory and runtime are measured.
Training Feedback includes aggregate scores plus complete raw evidence for all
Episodes in that submission. Read artifact-manifest.json, then use its
ordinals to align every trajectory-*.jsonl.gz transition with the exact
Policy-visible observations in observations-*.npz. Open NPZ files with
numpy.load(..., allow_pickle=False). No Environment or Host code has already
selected frames, created screenshots, rendered videos, or decided which
segments are important.
Choose your own analysis. You may decode terminal chars and colors, inspect
glyphs or named stats, render selected observations with Pillow, export a
segment with ImageIO, compare messages and inventory, or use another derived
representation. Keep non-submitted scripts, selected frames, derived images,
and notes under analysis/; only program/ is submitted. Bulk evidence from
older submissions can be evicted, while the newest complete submission is
protected, so copy only the derived material you decide is useful before a
later submission.
Validation and Assessment are Host-only aggregate phases and publish no detailed Artifacts back to the workspace. Never attempt to recover hidden seeds, Host paths, ttyrec data, or privileged simulator state from Feedback. Treat small-batch score movement as noisy and use equal Episode counts when comparing candidates.
Before finish, compare the candidate Program digests and the actual source
changes you intended. Submit behaviorally distinct candidates that test real
strategy alternatives. Renaming variables, retaining unused branches, or
passing a flag that does not change reachable behavior wastes a Validation
slot even though it creates a different source digest. Exact aggregate ties on
the same hidden Validation Episodes are a reason to inspect candidate
distinctness, not evidence that the Host should inspect or rewrite Programs.
© Linzwcs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in environments/nle/nethack/skills/optimize-nethack-policy of Linzwcs/EvoPolicyGym.
Open the folder on GitHubat commit a3d9669
Optimize Nethack Policy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Optimize Nethack Policy this skillLinzwcs/EvoPolicyGym | 175 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Benchmark Optimization Loopaffaan-m/ECC | 276k | 1 repos | ~664 | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 276k | 3 repos | ~654 | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 276k | — | ~412 | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 276k | — | ~330 | Automated safety check: Pass | MIT | |
| SQL Optimizationgithub/awesome-copilot | 40k | 2 repos | ~2.3k | Automated safety check: Pass | MIT |
affaan-m/ECC
Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…
affaan-m/ECC
Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…
affaan-m/ECC
このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.
affaan-m/ECC
使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。
github/awesome-copilot
Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…
androidx/androidx
Benchmarking and improving the performance of Jetpack Compose.
Linzwcs/EvoPolicyGym
Build, modularize, improve, test, and select a reward-aligned EvoPolicyGym Bot system for the Jackdaw Balatro Benchmark.
Linzwcs/EvoPolicyGym
Operate and extend EvoPolicyGym through its public SDK. An agent skill from Linzwcs/EvoPolicyGym.
Linzwcs/EvoPolicyGym
Improve an EvoPolicyGym Policy Program for Crafter local-symbolic-v1 observations.
Linzwcs/EvoPolicyGym
Improve an EvoPolicyGym Policy Program for the additive RGB Crafter survival-development Benchmark.
Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark. Optimize Nethack Policy is an agent skill from Linzwcs/EvoPolicyGym. Improve an EvoPolicyGym Policy Program for the deterministic NLE NetHackScore Benchmark.
Run `npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a claude-code`. Or copy the skill folder (environments/nle/nethack/skills/optimize-nethack-policy in Linzwcs/EvoPolicyGym) into .claude/skills/optimize-nethack-policy in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a codex`. Or copy the skill folder (environments/nle/nethack/skills/optimize-nethack-policy in Linzwcs/EvoPolicyGym) into .agents/skills/optimize-nethack-policy in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Linzwcs/EvoPolicyGym --skill optimize-nethack-policy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimize-nethack-policy, .gemini/skills/optimize-nethack-policy, .github/skills/optimize-nethack-policy and .opencode/skills/optimize-nethack-policy in your project.
SKILL.md names no scripts, command-line tools or credentials: Optimize Nethack Policy is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Optimize Nethack Policy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Optimize Nethack Policy: Benchmark Optimization Loop (affaan-m/ECC, 276k stars), Benchmark (affaan-m/ECC, 276k stars), Benchmark (affaan-m/ECC, 276k stars) and Benchmark (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Linzwcs (a GitHub user) maintains it in Linzwcs/EvoPolicyGym, which has 175 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 18, 2026.
Source: Linzwcs/EvoPolicyGym on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.