Agent skill

LLM From Scratch Guide

by wentorai in wentorai/research-plugins

Build a ChatGPT-like LLM from scratch using PyTorch step by step

MITAuto-check passedAI & LLM Engineering

Install LLM From Scratch Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill llm-from-scratch-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins llm-from-scratch-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/domains/ai-ml/llm-from-scratch-guide .claude/skills/llm-from-scratch-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-from-scratch-guide
GitHub stars
298
Used in
1 other repo
Token cost
~1.6k tokens
SKILL.md length
657 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Build a ChatGPT-like LLM from scratch using PyTorch step by step

  • Tasks that involve Deep learning
  • SKILL.md covers Overview, Installation and Setup, Core Learning Pipeline and Research Applications, plus 2 more sections
  • Calls git, python and pip; reaches github.com

What it does

LLM From Scratch Guide is an agent skill from wentorai/research-plugins. Build a ChatGPT-like LLM from scratch using PyTorch step by step

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning. It works with PyTorch and OpenAI. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Deep learning

Example prompts

  • “/llm-from-scratch-guide”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • sebastianraschka.com
    • pytorch.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM From Scratch Guide loads about 1.6k tokens when it runs. Until then it costs about 22 tokens; SKILL.md has 657 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~22
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 657 words, ~1,627 tokens.

Download SKILL.mdSave it as .claude/skills/llm-from-scratch-guide/SKILL.md (or your agent's skills folder).
name
llm-from-scratch-guide
description
Build a ChatGPT-like LLM from scratch using PyTorch step by step

LLM From Scratch Guide

Overview

LLMs-from-scratch is a comprehensive educational repository with over 87,000 stars on GitHub that teaches you how to build a ChatGPT-like large language model from the ground up using PyTorch. Created by Sebastian Raschka, a machine learning researcher and author, the project provides a complete pipeline covering data preparation, tokenization, attention mechanisms, pretraining, and instruction finetuning.

Unlike tutorials that treat LLMs as black boxes, this project demystifies every component by walking through the full implementation. Each chapter corresponds to a Jupyter notebook with clear explanations, diagrams, and runnable code. The repository accompanies the book "Build a Large Language Model (From Scratch)" and serves as a standalone learning resource for researchers and engineers who want deep understanding of transformer-based language models.

The project is particularly valuable for academic researchers who need to understand the internals of LLMs for their own research, whether that involves modifying architectures, running ablation studies, or developing domain-specific language models for scientific applications.

Installation and Setup

Clone the repository and set up a Python environment with the required dependencies:

bash
git clone https://github.com/rasbt/LLMs-from-scratch.git
cd LLMs-from-scratch

# Create a virtual environment
python -m venv llm-env
source llm-env/bin/activate

# Install dependencies
pip install -r requirements.txt

The project requires Python 3.10+ and PyTorch 2.0+. For GPU-accelerated training, ensure you have CUDA installed. The notebooks can also run on CPU for smaller model configurations, though training times will be significantly longer.

Key dependencies include:

  • PyTorch >= 2.0 for model implementation and training
  • tiktoken for BPE tokenization compatible with OpenAI models
  • matplotlib for training visualization
  • jupyter for interactive notebook execution

Core Learning Pipeline

The project is organized into sequential chapters that build on each other:

Chapter 1: Understanding Large Language Models

Covers the conceptual foundations of LLMs, including the transformer architecture, the difference between encoder and decoder models, and how pretraining and finetuning work at a high level.

Chapter 2: Working with Text Data

Implements text tokenization from scratch, including byte-pair encoding (BPE). You build a custom tokenizer and learn how text is converted to numerical representations:

python
# Tokenization example from the project
import tiktoken

tokenizer = tiktoken.get_encoding("gpt2")
text = "Large language models are fascinating."
token_ids = tokenizer.encode(text)
decoded = tokenizer.decode(token_ids)
Chapter 3: Coding Attention Mechanisms

Implements self-attention, multi-head attention, and causal (masked) attention from scratch. This is the core computational primitive of transformers:

python
# Simplified multi-head attention
class MultiHeadAttention(nn.Module):
    def __init__(self, d_in, d_out, context_length, num_heads, dropout=0.0):
        super().__init__()
        self.W_query = nn.Linear(d_in, d_out, bias=False)
        self.W_key = nn.Linear(d_in, d_out, bias=False)
        self.W_value = nn.Linear(d_in, d_out, bias=False)
        self.out_proj = nn.Linear(d_out, d_out)
        self.num_heads = num_heads
        self.head_dim = d_out // num_heads
Chapter 4: Implementing a GPT Model

Assembles the full GPT architecture using the attention mechanism, layer normalization, feed-forward networks, and positional embeddings.

Chapter 5: Pretraining on Unlabeled Data

Trains the GPT model on a text corpus using next-token prediction. Covers the training loop, loss computation, learning rate scheduling, and gradient clipping.

Show full SKILL.md (269 more words)Show less
Chapter 6: Finetuning for Text Classification

Adapts the pretrained model for downstream classification tasks, demonstrating how to add a classification head and finetune on labeled data.

Chapter 7: Instruction Finetuning

Converts the pretrained model into an instruction-following assistant using supervised finetuning on instruction-response pairs, similar to how ChatGPT is trained.

Research Applications

This resource is invaluable for several research scenarios:

  • Architecture ablation studies: Modify individual components (attention heads, layer count, embedding dimensions) and measure their impact on performance
  • Domain-specific pretraining: Use the pipeline to pretrain models on scientific corpora (biomedical literature, physics papers, chemical databases)
  • Tokenizer research: Experiment with different tokenization strategies for specialized vocabularies
  • Efficient training methods: Test techniques like gradient accumulation, mixed precision, and learning rate warmup
  • Interpretability research: Inspect attention patterns and intermediate representations at every layer

For researchers working with limited compute, the project includes configurations for small models (124M parameters) that can be trained on a single GPU in reasonable time, making it practical for experimentation and prototyping.

Integration with Research Workflows

Combine this project with other tools in your research stack:

  • Use Weights & Biases or MLflow for experiment tracking during pretraining runs
  • Export trained models to Hugging Face Hub for sharing and reproducibility
  • Integrate with PyTorch Lightning for distributed training across multiple GPUs
  • Apply LoRA or QLoRA adapters from the bonus chapters for parameter-efficient finetuning

The bonus materials in the repository cover additional topics like DPO (Direct Preference Optimization), loading pretrained weights from Hugging Face, and converting models between different formats.

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/domains/ai-ml/llm-from-scratch-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LLM From Scratch Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM From Scratch Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM From Scratch Guide this skillwentorai/research-plugins2981 repos~1.6kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Docstringpytorch/pytorch104k2 repos~2.6kAutomated safety check: PassCustom licence
CI Metricspytorch/pytorch104k—~1.1kAutomated safety check: PassCustom licence
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Benchmark Pyreflyfacebook/pyrefly7.1k—~1.8kAutomated safety check: PassMIT

Similar skills

  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Docstring

    pytorch/pytorch

    Write docstrings for PyTorch functions and methods following PyTorch conventions.

    104k GitHub starsUsed in 2 repos~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Metrics

    pytorch/pytorch

    Query PyTorch CI, GitHub Actions, HUD, Grafana, and infrastructure metrics.

    104k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Benchmark Pyrefly

    facebook/pyrefly

    Official

    Run Pyrefly benchmarks locally via Buck or Cargo, including PyTorch real-world LSP benchmarks.

    7.1k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Document Public APIs

    pytorch/pytorch

    Document undocumented public APIs in PyTorch by removing functions from coverageignorefunctions and coverageignoreclasses in docs/source/conf.py, running Sphinx coverage, and adding the appropriate…

    104k GitHub stars~4.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Works with

Questions about LLM From Scratch Guide

What does LLM From Scratch Guide do?

Build a ChatGPT-like LLM from scratch using PyTorch step by step. LLM From Scratch Guide is an agent skill from wentorai/research-plugins.

When should I use LLM From Scratch Guide?

LLM From Scratch Guide fits situations like: tasks that involve Deep learning.

How do I install LLM From Scratch Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill llm-from-scratch-guide -a claude-code`. Or copy the skill folder (skills/domains/ai-ml/llm-from-scratch-guide in wentorai/research-plugins) into .claude/skills/llm-from-scratch-guide in your project. Claude Code loads it when a task matches its description.

How do I install LLM From Scratch Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill llm-from-scratch-guide -a codex`. Or copy the skill folder (skills/domains/ai-ml/llm-from-scratch-guide in wentorai/research-plugins) into .agents/skills/llm-from-scratch-guide in your project. Codex loads it when a task matches its description.

Can I use LLM From Scratch Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill llm-from-scratch-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-from-scratch-guide, .gemini/skills/llm-from-scratch-guide, .github/skills/llm-from-scratch-guide and .opencode/skills/llm-from-scratch-guide in your project.

What does LLM From Scratch Guide need to run?

Going by SKILL.md and its folder, LLM From Scratch Guide needs the command-line tools its instructions call (git, python and pip). Our summary lists: Python 3.

Does LLM From Scratch Guide access the network?

SKILL.md names 3 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: sebastianraschka.com and pytorch.org. This is read from the text; nothing was executed.

Is LLM From Scratch Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM From Scratch Guide use?

LLM From Scratch Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM From Scratch Guide use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM From Scratch Guide?

Skills that share tags, products or a category with LLM From Scratch Guide: CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Docstring (pytorch/pytorch, 104k stars), CI Metrics (pytorch/pytorch, 104k stars) and Cuda Index Width (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM From Scratch Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.