Agent skill

nanoGPT Training Guide

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Walks through nanoGPT, Karpathy's compact GPT implementation: training on Shakespeare, reproducing GPT-2, fine-tuning GPT-2 checkpoints and training on your own text.

MITAuto-check passedAI & LLM Engineering

Install nanoGPT Training Guide

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill nanogpt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs nanogpt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/01-model-architecture/nanogpt .claude/skills/nanogpt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nanogpt
GitHub stars
13k
Used in
2 other repos
Token cost
~1.7k tokens
SKILL.md length
316 words
Files
4 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Walks through nanoGPT, Karpathy's compact GPT implementation: training on Shakespeare, reproducing GPT-2, fine-tuning GPT-2 checkpoints and training on your own text.

  • Learning how a GPT model and its training loop work
  • SKILL.md covers Quick start, Common workflows, When to use vs alternatives and Common issues, plus 3 more sections
  • Calls python and pip
  • Training a small character-level model on a laptop CPU

What it does

nanoGPT is presented as a roughly 300-line model file plus a roughly 300-line training loop in plain PyTorch, meant for learning how GPT works and for experimenting with transformer variants. The quick start trains a character-level model on Shakespeare, which is light enough for a CPU, and the skill gives rough training times: about five minutes on CPU or about one minute on a GPU.

Other workflows reproduce GPT-2 at 124M parameters on OpenWebText with multi-GPU training, fine-tune a pretrained GPT-2 checkpoint by setting the init_from option, and train on a custom dataset through your own prepare script followed by train.py with a dataset flag. Data preparation for OpenWebText is said to take about an hour and the full run about four days on eight A100s. The skill points to Hugging Face Transformers, Megatron-LM and LitGPT as alternatives for production or large-scale work.

When your agent uses it

  • Learning how a GPT model and its training loop work
  • Training a small character-level model on a laptop CPU
  • Fine-tuning a pretrained GPT-2 checkpoint on your own text
  • Reproducing the GPT-2 124M run on OpenWebText

Example prompts

  • “Train a character-level model on the Shakespeare data with nanoGPT and show me a sample.”
  • “Fine-tune gpt2 on my notes in data/notes using nanoGPT and a custom prepare script.”
  • “Explain the structure of model.py in nanoGPT, layer by layer.”

Requirements

  • Python with torch, numpy, transformers, datasets, tiktoken, wandb and tqdm
  • Multiple GPUs for the OpenWebText run

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

nanoGPT Training Guide loads about 1.7k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 316 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 316 words, ~1,686 tokens.

Download SKILL.mdSave it as .claude/skills/nanogpt/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
nanogpt
description
Educational GPT implementation in ~300 lines. Reproduces GPT-2 (124M) on OpenWebText. Clean, hackable code for learning transformers. By Andrej Karpathy. Perfect for understanding GPT architecture from scratch. Train on Shakespeare (CPU) or OpenWebText (multi-GPU).
version
1.0.0
author
Orchestra Research
license
MIT
tags
Model Architecture, NanoGPT, GPT-2, Educational, Andrej Karpathy, Transformer, Minimalist, From Scratch, Training
dependencies
torch, transformers, datasets, tiktoken, wandb

nanoGPT - Minimalist GPT Training

Quick start

nanoGPT is a simplified GPT implementation designed for learning and experimentation.

Installation:

bash
pip install torch numpy transformers datasets tiktoken wandb tqdm

Train on Shakespeare (CPU-friendly):

bash
# Prepare data
python data/shakespeare_char/prepare.py

# Train (5 minutes on CPU)
python train.py config/train_shakespeare_char.py

# Generate text
python sample.py --out_dir=out-shakespeare-char

Output:

ROMEO:
What say'st thou? Shall I speak, and be a man?

JULIET:
I am afeard, and yet I'll speak; for thou art
One that hath been a man, and yet I know not
What thou art.

Common workflows

Workflow 1: Character-level Shakespeare

Complete training pipeline:

bash
# Step 1: Prepare data (creates train.bin, val.bin)
python data/shakespeare_char/prepare.py

# Step 2: Train small model
python train.py config/train_shakespeare_char.py

# Step 3: Generate text
python sample.py --out_dir=out-shakespeare-char

Config (config/train_shakespeare_char.py):

python
# Model config
n_layer = 6          # 6 transformer layers
n_head = 6           # 6 attention heads
n_embd = 384         # 384-dim embeddings
block_size = 256     # 256 char context

# Training config
batch_size = 64
learning_rate = 1e-3
max_iters = 5000
eval_interval = 500

# Hardware
device = 'cpu'  # Or 'cuda'
compile = False # Set True for PyTorch 2.0

Training time: ~5 minutes (CPU), ~1 minute (GPU)

Workflow 2: Reproduce GPT-2 (124M)

Multi-GPU training on OpenWebText:

bash
# Step 1: Prepare OpenWebText (takes ~1 hour)
python data/openwebtext/prepare.py

# Step 2: Train GPT-2 124M with DDP (8 GPUs)
torchrun --standalone --nproc_per_node=8 \
  train.py config/train_gpt2.py

# Step 3: Sample from trained model
python sample.py --out_dir=out

Config (config/train_gpt2.py):

python
# GPT-2 (124M) architecture
n_layer = 12
n_head = 12
n_embd = 768
block_size = 1024
dropout = 0.0

# Training
batch_size = 12
gradient_accumulation_steps = 5 * 8  # Total batch ~0.5M tokens
learning_rate = 6e-4
max_iters = 600000
lr_decay_iters = 600000

# System
compile = True  # PyTorch 2.0

Training time: ~4 days (8× A100)

Workflow 3: Fine-tune pretrained GPT-2

Start from OpenAI checkpoint:

python
# In train.py or config
init_from = 'gpt2'  # Options: gpt2, gpt2-medium, gpt2-large, gpt2-xl

# Model loads OpenAI weights automatically
python train.py config/finetune_shakespeare.py

Example config (config/finetune_shakespeare.py):

python
# Start from GPT-2
init_from = 'gpt2'

# Dataset
dataset = 'shakespeare_char'
batch_size = 1
block_size = 1024

# Fine-tuning
learning_rate = 3e-5  # Lower LR for fine-tuning
max_iters = 2000
warmup_iters = 100

# Regularization
weight_decay = 1e-1
Workflow 4: Custom dataset

Train on your own text:

python
# data/custom/prepare.py
import numpy as np

# Load your data
with open('my_data.txt', 'r') as f:
    text = f.read()

# Create character mappings
chars = sorted(list(set(text)))
stoi = {ch: i for i, ch in enumerate(chars)}
itos = {i: ch for i, ch in enumerate(chars)}

# Tokenize
data = np.array([stoi[ch] for ch in text], dtype=np.uint16)

# Split train/val
n = len(data)
train_data = data[:int(n*0.9)]
val_data = data[int(n*0.9):]

# Save
train_data.tofile('data/custom/train.bin')
val_data.tofile('data/custom/val.bin')

Train:

bash
python data/custom/prepare.py
python train.py --dataset=custom

When to use vs alternatives

Use nanoGPT when:

  • Learning how GPT works
  • Experimenting with transformer variants
  • Teaching/education purposes
  • Quick prototyping
  • Limited compute (can run on CPU)

Simplicity advantages:

  • ~300 lines: Entire model in model.py
  • ~300 lines: Training loop in train.py
  • Hackable: Easy to modify
  • No abstractions: Pure PyTorch

Use alternatives instead:

  • HuggingFace Transformers: Production use, many models
  • Megatron-LM: Large-scale distributed training
  • LitGPT: More architectures, production-ready
  • PyTorch Lightning: Need high-level framework

Common issues

Issue: CUDA out of memory

Reduce batch size or context length:

python
batch_size = 1  # Reduce from 12
block_size = 512  # Reduce from 1024
gradient_accumulation_steps = 40  # Increase to maintain effective batch

Issue: Training too slow

Enable compilation (PyTorch 2.0+):

python
compile = True  # 2× speedup

Use mixed precision:

python
dtype = 'bfloat16'  # Or 'float16'

Issue: Poor generation quality

Train longer:

python
max_iters = 10000  # Increase from 5000

Lower temperature:

python
# In sample.py
temperature = 0.7  # Lower from 1.0
top_k = 200       # Add top-k sampling

Issue: Can't load GPT-2 weights

Install transformers:

bash
pip install transformers

Check model name:

python
init_from = 'gpt2'  # Valid: gpt2, gpt2-medium, gpt2-large, gpt2-xl

Advanced topics

Model architecture: See references/architecture.md for GPT block structure, multi-head attention, and MLP layers explained simply.

Training loop: See references/training.md for learning rate schedule, gradient accumulation, and distributed data parallel setup.

Data preparation: See references/data.md for tokenization strategies (character-level vs BPE) and binary format details.

Hardware requirements

  • Shakespeare (char-level):

    • CPU: 5 minutes
    • GPU (T4): 1 minute
    • VRAM: <1GB
  • GPT-2 (124M):

    • 1× A100: ~1 week
    • 8× A100: ~4 days
    • VRAM: ~16GB per GPU
  • GPT-2 Medium (350M):

    • 8× A100: ~2 weeks
    • VRAM: ~40GB per GPU

Performance:

  • With compile=True: 2× speedup
  • With dtype=bfloat16: 50% memory reduction

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in 01-model-architecture/nanogpt of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/architecture.md
  • references/data.md
  • references/training.md

Open the folder on GitHubat commit 773a529

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

nanoGPT Training Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

nanoGPT Training Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
nanoGPT Training Guide this skillOrchestra-Research/AI-Research-SKILLs13k2 repos~1.7kAutomated safety check: PassMIT
Alphagenome Finetuninggenomicsxai/alphagenome-pytorch162—~1kAutomated safety check: PassApache-2.0
Deep Learningericrisco/rsc-harness180—~3.4kAutomated safety check: PassMIT
Quaxnstarman/quax143—~5.5kAutomated safety check: PassApache-2.0
ML Experiment IterationLeeroo-AI/superml195—~4.8kAutomated safety check: PassApache-2.0
Alphagenome Predictionsgenomicsxai/alphagenome-pytorch162—~868Automated safety check: PassApache-2.0

Similar skills

  • Alphagenome Finetuning

    genomicsxai/alphagenome-pytorch

    Fine-tune or transfer-learn AlphaGenome-PyTorch on custom genomic data — pick a mode (linear probe, LoRA, Locon, full), train on BigWig tracks with agt finetune, use adapters, delta checkpoints…

    162 GitHub stars~1k tokensUpdated 25 days ago
    AI & LLM EngineeringAuto-check passed
  • Deep Learning

    ericrisco/rsc-harness

    A skill your agent uses when training or debugging a neural net in PyTorch — the forward/loss/backward/step loop and its silent bugs, mixed precision (AMP), AdamW/LR schedules, DDP/FSDP/ZeRO…

    180 GitHub stars~3.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Quax

    nstarman/quax

    A skill your agent uses when writing, reviewing, or debugging JAX code that involves quax — custom array-ish objects (physical units, LoRA, sparse, symbolic zero, named axes), quax.quaxify…

    143 GitHub stars~5.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • ML Experiment Iteration

    Leeroo-AI/superml

    Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues.

    195 GitHub stars~4.8k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Alphagenome Predictions

    genomicsxai/alphagenome-pytorch

    Run AlphaGenome-PyTorch to get genomic track predictions — via the agt predict CLI (single locus, BED regions, whole chromosomes, raw FASTA sequences, or per-gene count tables/AnnData), variant…

    162 GitHub stars~868 tokensUpdated 25 days ago
    AI & LLM EngineeringAuto-check passed
  • ExecuTorch Model Export

    pytorch/executorch

    Export a PyTorch model to .pte format for ExecuTorch. Use when converting models, lowering to edge, or generating .pte files.

    5.1k GitHub stars~238 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Questions about nanoGPT Training Guide

What does nanoGPT Training Guide do?

Walks through nanoGPT, Karpathy's compact GPT implementation: training on Shakespeare, reproducing GPT-2, fine-tuning GPT-2 checkpoints and training on your own text. nanoGPT is presented as a roughly 300-line model file plus a roughly 300-line training loop in plain PyTorch, meant for learning how GPT works and for experimenting with transformer variants. The quick start trains a character-level model on Shakespeare, which is light enough for a CPU, and the skill gives rough training times: about five minutes on CPU or about one minute on a GPU.

When should I use nanoGPT Training Guide?

nanoGPT Training Guide fits situations like: learning how a GPT model and its training loop work; training a small character-level model on a laptop CPU; fine-tuning a pretrained GPT-2 checkpoint on your own text; reproducing the GPT-2 124M run on OpenWebText.

How do I install nanoGPT Training Guide in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill nanogpt -a claude-code`. Or copy the skill folder (01-model-architecture/nanogpt in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/nanogpt in your project. Claude Code loads it when a task matches its description.

How do I install nanoGPT Training Guide in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill nanogpt -a codex`. Or copy the skill folder (01-model-architecture/nanogpt in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/nanogpt in your project. Codex loads it when a task matches its description.

Can I use nanoGPT Training Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill nanogpt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nanogpt, .gemini/skills/nanogpt, .github/skills/nanogpt and .opencode/skills/nanogpt in your project.

What does nanoGPT Training Guide need to run?

Going by SKILL.md and its folder, nanoGPT Training Guide needs the command-line tools its instructions call (python and pip). Our summary lists: Python with torch, numpy, transformers, datasets, tiktoken, wandb and tqdm; Multiple GPUs for the OpenWebText run.

Does nanoGPT Training Guide access the network?

SKILL.md names 2 domains. As links in the text: github.com and huggingface.co. This is read from the text; nothing was executed.

Is nanoGPT Training Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does nanoGPT Training Guide use?

nanoGPT Training Guide is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does nanoGPT Training Guide use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.1k tokens, read only when the agent opens those files.

What are the alternatives to nanoGPT Training Guide?

Skills that share tags, products or a category with nanoGPT Training Guide: Alphagenome Finetuning (genomicsxai/alphagenome-pytorch, 162 stars), Deep Learning (ericrisco/rsc-harness, 180 stars), Quax (nstarman/quax, 143 stars) and ML Experiment Iteration (Leeroo-AI/superml, 195 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains nanoGPT Training Guide?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,405 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.