Agent skill

Karpathy Practice Environments

by LearnPrompt in LearnPrompt/andrej-karpathy-skills

Build practice environments and evaluation gyms where agents can try, fail, and learn from feedback.

MITAuto-check passedEducation

Install Karpathy Practice Environments

skills CLI
$ npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-practice-environments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LearnPrompt/andrej-karpathy-skills karpathy-practice-environments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LearnPrompt/andrej-karpathy-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/karpathy-practice-environments .claude/skills/karpathy-practice-environments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
karpathy-practice-environments
GitHub stars
110
Token cost
~2.1k tokens
SKILL.md length
204 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Build practice environments and evaluation gyms where agents can try, fail, and learn from feedback.

  • The user wants to set up automated evaluation for an agent
  • SKILL.md covers Core Principle, The Three Training Data Types, Building a Practice Gym… and Practice Environment Design…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Needs a sandbox for testing LLM capabilities

What it does

Karpathy Practice Environments is an agent skill from LearnPrompt/andrej-karpathy-skills. Build practice environments and evaluation gyms where agents can try, fail, and learn from feedback. Use this skill when the user wants to set up automated evaluation for an agent, needs a sandbox for testing LLM capabilities, wants to build a practice gym for skills, needs an RL-style environment for agent improvement, or says "eval environment", "practice gym", "agent sandbox", "automated testing for agents", "RL environment", "agent eval loop". Based on Karpathy LLM Textbook and environments posts.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Reinforcement learning. The repository describes itself as: Karpathy-inspired Agent Skills collection. The licence is MIT.

When your agent uses it

  • The user wants to set up automated evaluation for an agent
  • Needs a sandbox for testing LLM capabilities
  • Wants to build a practice gym for skills
  • Needs an RL-style environment for agent improvement

Example prompts

  • “eval environment”
  • “practice gym”
  • “agent sandbox”
  • “/karpathy-practice-environments”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9e46dec. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • x.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Karpathy Practice Environments loads about 2.1k tokens when it runs. Until then it costs about 134 tokens; SKILL.md has 204 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LearnPrompt/andrej-karpathy-skills at commit 9e46dec, republished under its MIT licence (© LearnPrompt). 204 words, ~2,146 tokens.

Download SKILL.mdSave it as .claude/skills/karpathy-practice-environments/SKILL.md (or your agent's skills folder).
name
karpathy-practice-environments
description
Build practice environments and evaluation gyms where agents can try, fail, and learn from feedback. Use this skill when the user wants to set up automated evaluation for an agent, needs a sandbox for testing LLM capabilities, wants to build a practice gym for skills, needs an RL-style environment for agent improvement, or says "eval environment", "practice gym", "agent sandbox", "automated testing for agents", "RL environment", "agent eval loop". Based on Karpathy LLM Textbook and environments posts.
disable-model-invocation
false
user-invocable
true
related_skills
karpathy-education-first, karpathy-understanding-first, karpathy-meta-reflection, karpathy-autoresearch

Skill 14: LLM Textbook + Practice Environments(LLM教科书 + 练习环境)

Source: https://x.com/karpathy/status/1885026028428681698 | https://x.com/karpathy/status/1960803117689397543 "Take LLMs to school" | "Environments for RL" posts

Core Principle

LLMs learn like students. Give them textbooks, worked problems, and practice gyms.

Karpathy's insight: the missing ingredient in LLM training isn't more parameters — it's practice environments where agents can try → fail → see feedback → try again. Like a student who reads theory but also needs problem sets.

Apply this to your agents: build environments where they can practice a skill with automatic scoring, not just generate output into the void.

The Three Training Data Types

Every good learning environment needs all three:

1. EXPOSITION (pretrain equivalent)
   → Background knowledge, concepts, context
   → The "textbook chapter" the agent reads before practicing
   
2. WORKED EXAMPLES (SFT equivalent)
   → Complete input-output pairs with reasoning shown
   → "Here's a solved problem — learn the pattern"
   
3. PRACTICE PROBLEMS with feedback (RL equivalent)
   → Problems with verifiable correct answers
   → Automatic scoring so agent knows how it did
   → Enough variety that memorization doesn't work

Building a Practice Gym (General Template)

For any skill you want an agent to get better at:

python
#!/usr/bin/env python3
"""
Practice Gym for: [SKILL_NAME]
Karpathy-style RL environment for agent skill development
"""

import json
import random
from typing import Callable

class PracticeGym:
    """
    A practice environment where an agent can repeatedly attempt
    a task and receive automatic feedback.
    """
    
    def __init__(self, task_generator: Callable, scorer: Callable):
        self.task_generator = task_generator  # generates new practice problems
        self.scorer = scorer                   # returns 0.0-1.0 score for an attempt
        self.history = []
    
    def sample_task(self):
        """Generate a new practice problem."""
        return self.task_generator()
    
    def evaluate(self, task, attempt):
        """Score an agent's attempt. Returns dict with score + feedback."""
        score = self.scorer(task, attempt)
        result = {
            "task": task,
            "attempt": attempt,
            "score": score,
            "passed": score >= 0.8,
        }
        self.history.append(result)
        return result
    
    def summary(self):
        """Summarize performance across all attempts."""
        if not self.history:
            return "No attempts yet."
        scores = [r["score"] for r in self.history]
        return {
            "total_attempts": len(scores),
            "avg_score": sum(scores) / len(scores),
            "pass_rate": sum(1 for s in scores if s >= 0.8) / len(scores),
            "recent_trend": "improving" if scores[-1] > scores[0] else "declining"
        }

Practice Environment Design Prompt

Build a practice gym for any skill:

Design a practice environment for training an LLM agent to [SKILL].

Skill description: [WHAT THE AGENT SHOULD GET BETTER AT]
Current performance: [HOW WELL IT DOES NOW]
Target performance: [WHAT GOOD LOOKS LIKE]

Design the gym with:

1. TASK GENERATOR
   - Input format: [what the task looks like]
   - Variation parameters: [what changes between tasks]
   - Difficulty levels: [easy / medium / hard criteria]
   - Example task: [concrete example]

2. SCORER  
   - What makes an attempt correct? (exact match / rubric / functional test)
   - Score breakdown: [what earns partial credit]
   - Automatic vs human evaluation: [which parts can be automated]
   - Example: good attempt vs bad attempt with scores

3. CURRICULUM
   - Start: [simplest tasks to build confidence]
   - Progress: [how to increase difficulty as performance improves]
   - Mastery criterion: [when is the skill "learned"?]

4. FEEDBACK FORMAT
   What should the agent receive after each attempt?
   - Score (0-1)
   - Specific error: [what went wrong]
   - Hint for next attempt: [one actionable tip]

Quick Eval Loop Prompt

For rapidly testing an agent on a set of practice problems:

Run an evaluation loop on this agent capability.

Agent task: [WHAT THE AGENT IS SUPPOSED TO DO]

Test cases:
1. Input: [test input 1]
   Expected output: [expected 1]
   
2. Input: [test input 2]
   Expected output: [expected 2]
   
3. Input: [test input 3]
   Expected output: [expected 3]

For each test case:
1. Attempt the task
2. Compare to expected output
3. Score: PASS / PARTIAL / FAIL
4. Explain the error (if any) in one sentence

Final report:
- Pass rate: N/3
- Most common failure mode: [pattern in errors]
- Suggested improvement: [one specific fix]

The Curriculum Design Template

Build a structured learning progression:

Design a learning curriculum for an agent to master [SKILL].

Like Karpathy's "take LLMs to school" approach — structure it as:

WEEK 1 — Foundations (exposition):
- Core concepts to internalize
- 3-5 worked examples with full reasoning traces
- Quiz: 5 basic problems with answers

WEEK 2 — Pattern Recognition (practice):
- 20 practice problems, graduated difficulty
- Automatic scoring criteria
- Common error analysis

WEEK 3 — Generalization (RL-style):
- Novel problems the agent hasn't seen
- Real-world variants
- Adversarial examples (edge cases designed to fail)

Mastery test: [describe the final eval that confirms the skill is learned]

Scoring Rubrics for Common Skills

Code Generation
Score 1.0: Correct output, handles edge cases, clean style
Score 0.8: Correct output, misses 1 edge case
Score 0.5: Core logic correct, fails on some inputs
Score 0.2: Wrong approach but partially useful
Score 0.0: Doesn't compile or clearly wrong
Reasoning/Analysis
Score 1.0: Correct conclusion + correct reasoning chain + appropriate confidence
Score 0.8: Correct conclusion, minor reasoning gap
Score 0.5: Correct conclusion, wrong reasoning
Score 0.2: Wrong conclusion but shows relevant knowledge
Score 0.0: Wrong conclusion, no relevant reasoning
Following Instructions
Score 1.0: All instructions followed, output matches spec exactly
Score 0.8: Minor deviation from spec, intent preserved
Score 0.5: Core task done, 1-2 instructions ignored
Score 0.2: Significant departure from instructions
Score 0.0: Instructions ignored entirely

Environment-in-a-Gist

Karpathy's pattern: publish the environment spec (not the implementation) so anyone can build their gym:

Write a Gist-style environment spec for [SKILL_GYM].

Format:
# [GYM_NAME] — Practice Environment Spec
## What this trains
## Task format
## Scoring criteria  
## Sample task + ideal response
## Curriculum (3 stages)
## How to evaluate mastery

This is a spec, not code. Anyone with an LLM should be able to implement it.

Workflow

属于工作流:月度体检(第3步)

位置上游下游
第3步(练习)karpathy-understanding-first(识别盲区后)karpathy-education-first(把练习成果教学化)

完整链路:meta-reflection → understanding-first → practice-environments → education-first

Prompt Contract

text
Build a practice environment for <SKILL_OR_CAPABILITY>. Include: 1) Exposition — what this skill is and why it matters (3-5 sentences), 2) Worked examples — 2-3 complete input→output demonstrations, 3) Practice tasks — 5 exercises with increasing difficulty, 4) Automatic grading — for each task define pass/fail criteria that can be checked without human judgment, 5) Retry loop — if failed, what feedback to give and how to adjust difficulty.

Verification Checklist

  • 练习环境有明确的技能定义(不是模糊的「变强」)
  • Worked examples 覆盖了典型和边界情况
  • 评分标准不需要人工判断(自动化)
  • 失败时有具体反馈(不只是「错了」)
  • 难度有递进(不是一步到位)
  • Agent 可以自主重试直到通过

© LearnPrompt, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in karpathy-practice-environments of LearnPrompt/andrej-karpathy-skills.

Open the folder on GitHubat commit 9e46dec

Compare with similar skills

Karpathy Practice Environments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Karpathy Practice Environments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Karpathy Practice Environments this skillLearnPrompt/andrej-karpathy-skills110—~2.1kAutomated safety check: PassMIT
Generate Verifiers Envadithya-s-k/FineEnvs4561 repos~2.3kAutomated safety check: PassApache-2.0
Pieter AbbeelK-Dense-AI/mimeo282—~1.5kAutomated safety check: PassMIT
Stuart RussellK-Dense-AI/mimeo282—~1.6kAutomated safety check: PassMIT
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Fix Art IssuesOpenPipe/ART11k—~840Automated safety check: NotesApache-2.0

Similar skills

  • Generate Verifiers Env

    adithya-s-k/FineEnvs

    Builds a Verifiers (PrimeIntellect) variant of an RL environment.

    456 GitHub starsUsed in 1 repo~2.3k tokens
    EducationAuto-check passed
  • Pieter Abbeel

    K-Dense-AI/mimeo

    Applies the reasoning of Pieter Abbeel, robotics and reinforcement learning expert, UC Berkeley professor, and co-founder of Covariant.

    282 GitHub stars~1.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Stuart Russell

    K-Dense-AI/mimeo

    Applies the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial Intelligence: A Modern Approach'.

    282 GitHub stars~1.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Fix Art Issues

    OpenPipe/ART

    Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.

    11k GitHub stars~840 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • AI Engineering Project Tutor

    rohitg00/ai-engineering-from-scratch

    Tutors a learner through one stage of a hands-on AI engineering project per session: lesson, prediction, code, grader run and reflection, with hints but never full solutions.

    66k GitHub stars~1.6k tokensUpdated today
    EducationAuto-check passed

More from LearnPrompt/andrej-karpathy-skills

All 15 skills in this repo
  • Karpathy Methodology Index

    LearnPrompt/andrej-karpathy-skills

    Apply Andrej Karpathy AI methodology and principles from his 2023-2026 insights. Use this skill when the user wants to apply Karpathy-style thinking, needs…

    110 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy Agentic Engineering

    LearnPrompt/andrej-karpathy-skills

    Apply Karpathy-style agentic engineering to any coding or building task.

    110 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed
  • AutoResearch Loop

    LearnPrompt/andrej-karpathy-skills

    Sets up an autonomous research loop where an agent runs experiments on git branches, logs results and proposes the next iteration while you approve each hypothesis change.

    110 GitHub stars~1.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy Education First

    LearnPrompt/andrej-karpathy-skills

    Apply the education-first mindset — make everything you build teachable, create nano-project explanations, write for beginners.

    110 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy Idea Files

    LearnPrompt/andrej-karpathy-skills

    Create and share ideas as abstract Gist-style specs instead of code — letting agents or others implement.

    110 GitHub stars~1.5k tokensUpdated 3 mo ago
    Auto-check passed
  • Karpathy LLM Simulator

    LearnPrompt/andrej-karpathy-skills

    Use LLM as a simulator of expert debates and opposing viewpoints instead of getting a single sycophantic answer.

    110 GitHub stars~1.3k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Karpathy Practice Environments

What does Karpathy Practice Environments do?

Build practice environments and evaluation gyms where agents can try, fail, and learn from feedback. Karpathy Practice Environments is an agent skill from LearnPrompt/andrej-karpathy-skills. Build practice environments and evaluation gyms where agents can try, fail, and learn from feedback.

When should I use Karpathy Practice Environments?

Karpathy Practice Environments fits situations like: the user wants to set up automated evaluation for an agent; needs a sandbox for testing LLM capabilities; wants to build a practice gym for skills; needs an RL-style environment for agent improvement.

How do I install Karpathy Practice Environments in Claude Code?

Run `npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-practice-environments -a claude-code`. Or copy the skill folder (karpathy-practice-environments in LearnPrompt/andrej-karpathy-skills) into .claude/skills/karpathy-practice-environments in your project. Claude Code loads it when a task matches its description.

How do I install Karpathy Practice Environments in Codex?

Run `npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-practice-environments -a codex`. Or copy the skill folder (karpathy-practice-environments in LearnPrompt/andrej-karpathy-skills) into .agents/skills/karpathy-practice-environments in your project. Codex loads it when a task matches its description.

Can I use Karpathy Practice Environments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LearnPrompt/andrej-karpathy-skills --skill karpathy-practice-environments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/karpathy-practice-environments, .gemini/skills/karpathy-practice-environments, .github/skills/karpathy-practice-environments and .opencode/skills/karpathy-practice-environments in your project.

What does Karpathy Practice Environments need to run?

SKILL.md names no scripts, command-line tools or credentials: Karpathy Practice Environments is instructions for the agent only. Our summary lists: Python 3.

Does Karpathy Practice Environments access the network?

SKILL.md names 1 domain. As links in the text: x.com. This is read from the text; nothing was executed.

Is Karpathy Practice Environments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Karpathy Practice Environments use?

Karpathy Practice Environments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Karpathy Practice Environments use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Karpathy Practice Environments?

Skills that share tags, products or a category with Karpathy Practice Environments: Generate Verifiers Env (adithya-s-k/FineEnvs, 456 stars), Pieter Abbeel (K-Dense-AI/mimeo, 282 stars), Stuart Russell (K-Dense-AI/mimeo, 282 stars) and AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Karpathy Practice Environments?

LearnPrompt (a GitHub user) maintains it in LearnPrompt/andrej-karpathy-skills, which has 110 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on July 10, 2026.

Source: LearnPrompt/andrej-karpathy-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.