Agent skill

Deep Learning Book

by alirezarezvani in alirezarezvani/claude-skills

Study companion and working knowledge base for the Deep Learning textbook by Goodfellow, Bengio & Courville (MIT Press, 2016), read free at deeplearningbook.org.

MITAuto-check passedAI & LLM Engineering

Install Deep Learning Book

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill deep-learning-book -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills deep-learning-book --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/deep-learning-book/skills/deep-learning-book .claude/skills/deep-learning-book && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deep-learning-book
GitHub stars
28k
Token cost
~2.9k tokens
SKILL.md length
1,126 words
Files
37 (incl. scripts, references, assets)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

Study companion and working knowledge base for the Deep Learning textbook by Goodfellow, Bengio & Courville (MIT Press, 2016), read free at deeplearningbook.org.

  • Teaching this book
  • SKILL.md covers How to Use This Skill, Core Frameworks & Mental Models, Chapter Index and Topic Index, plus 3 more sections
  • Calls python3
  • Planning a route through it

What it does

Deep Learning Book is an agent skill from alirezarezvani/claude-skills. Study companion and working knowledge base for the Deep Learning textbook by Goodfellow, Bengio & Courville (MIT Press, 2016), read free at deeplearningbook.org. Indexes all 20 chapters, carries a 2016-to-2026 delta layer naming what the book got right, what was superseded (transformers, AdamW, diffusion, double descent) and what still holds, and ships four deterministic tools: a prerequisite-aware reading-path planner, a training-failure diagnostic, a capacity-and-regularization planner, and a…

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 38 other files, including scripts, reference files and assets (for example `assets/chapter_worksheet.md`, `assets/example_layer_spec.json` and `assets/study_log_template.md`).

It sits in AI & LLM Engineering, covering Deep learning and Translation. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • Teaching this book
  • Planning a route through it
  • Deciding whether a chapters advice is still current
  • Translating its math into a training decision

Example prompts

  • “/deep-learning-book”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • deeplearningbook.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deep Learning Book loads about 2.9k tokens when it runs, and up to ~8.3k if it reads all its reference files. Until then it costs about 200 tokens; SKILL.md has 1,126 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~200
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 1,126 words, ~2,863 tokens.

Download SKILL.mdSave it as .claude/skills/deep-learning-book/SKILL.md (or your agent's skills folder). This skill also uses 36 other files; get the full folder from GitHub.
name
deep-learning-book
description
Study companion and working knowledge base for the Deep Learning textbook by Goodfellow, Bengio & Courville (MIT Press, 2016), read free at deeplearningbook.org. Indexes all 20 chapters, carries a 2016-to-2026 delta layer naming what the book got right, what was superseded (transformers, AdamW, diffusion, double descent) and what still holds, and ships four deterministic tools: a prerequisite-aware reading-path planner, a training-failure diagnostic, a capacity-and-regularization planner, and a parameter/FLOP/activation-memory calculator. Use when studying or teaching this book, planning a route through it, deciding whether a chapter's advice is still current, or translating its math into a training decision. It points at the official chapters — it never reproduces them.
license
MIT
metadata.version
1.0.0
metadata.author
Alireza Rezvani
metadata.category
engineering
metadata.updated
2026-08-25

Deep Learning — Study Companion

Source book: Deep Learning, Ian Goodfellow, Yoshua Bengio & Aaron Courville (MIT Press, 2016) · 20 chapters, 3 parts · read free at deeplearningbook.org · companion compiled 2026-08-25.

This is a companion, not a copy. The book is copyrighted, and its site states that the HTML-only format exists to discourage copying under the authors' MIT Press contract. Nothing here reproduces its text. Every chapter file is original synthesis — what the chapter establishes, how to use it, where it has aged — plus a link to the official chapter. Read the book at the link; use this to navigate it, keep it current, and turn it into decisions. See references/rights_and_use.md.

How to Use This Skill

  • No argument — load the core frameworks below.
  • A topic — ask about regularization, saddle points, partition function; resolved through the Topic Index, then that chapter file is read before answering.
  • chNN — load that chapter's file.
  • "is this still true?" — the 2016→2026 delta layer, in every chapter file and in references/book_to_2026_delta.md.
  • "where do I start?" — run scripts/reading_path_planner.py.

When asked about something outside these 20 chapters, say so and route to the delta reference rather than improvising the book's position on material published after it.


Core Frameworks & Mental Models

The (T, P, E) frame — ch05

Name the task, the performance measure, and the experience in one sentence before any model code. Most failed projects failed at P: an unstated metric, or a proxy whose relationship to the real objective was never checked.

Every loss is a negative log-likelihood — ch03, ch06

Choose the output distribution, then take its negative log. Gaussian → MSE, Bernoulli → binary cross-entropy, categorical → cross-entropy, Laplace → MAE. "Which loss?" is always the question "which distribution?" in disguise. Modern contrastive and preference objectives sit outside this frame — a real limit of the book, not a gap in your understanding.

KL asymmetry decides your failure mode — ch03, ch19, ch20

D(p‖q) ≠ D(q‖p). Forward KL is mode-covering (blurry averages); reverse KL is mode-seeking (sharp but partial). This single fact predicts VAE blur, GAN mode collapse, and the characteristic over-confidence of mean-field variational posteriors.

Train-error-first triage — ch11, ch05

High training error → capacity or optimization is the bottleneck; more data will not help. Low training error with a large validation gap → data or regularization. This is the highest-value heuristic in the book. scripts/training_diagnostics.py runs it.

Capacity, the gap, and the U-curve's caveat — ch05, ch07

Regularization trades variance for bias. But the classical U-shaped capacity curve is incomplete: past the interpolation threshold, test error can fall again (double descent, 2019–2020, post-dating the book). Practical consequence: when a large model overfits, try more data, more regularization or longer training before shrinking it.

Architecture is a prior, not a trick — ch09, ch10, ch15

Convolution asserts translation equivariance and locality. Recurrence asserts that the past compresses into a state. A distributed representation asserts that factors combine combinatorially. When the assertion is false, the architecture cannot be rescued by tuning — and when it is true, it beats capacity. This is also why Vision Transformers need more data than ConvNets: they discard the prior and buy it back with examples.

Depth's real cost is gradient flow and activation memory — ch06, ch08, ch10

Backprop is the chain rule scheduled well: one forward-pass-equivalent of compute, and memory proportional to stored activations. Depth fails through vanishing/exploding gradients and ill-conditioning, which is why residual connections, normalization and clipping exist.

The partition function organizes Part III — ch16, ch17, ch18, ch19

For undirected models, the likelihood gradient needs samples from the model itself. Four escape routes: sample it (CD/PCD), sidestep it algebraically (pseudolikelihood, score matching), learn around it (NCE), or estimate it for evaluation (AIS). Score matching's descendants are today's diffusion models — which is why Part III repays reading even though its models did not survive.

Diagnose before you redesign — ch04, ch08, ch11

Gradient norm exploding → clip. Norm large but loss flat → ill-conditioning. Norm near zero with high loss → saturation or dead units. NaN → numerics first. Change one thing per experiment.


Show full SKILL.md (472 more words)Show less

Chapter Index

#TitleKey content
ch01Introductionrepresentation learning, depth as composition, curse of dimensionality
ch02Linear Algebranorms, SVD, eigendecomposition, conditioning, PCA
ch03Probability & Information Theorydistributions, entropy, KL, cross-entropy
ch04Numerical Computationunder/overflow, conditioning, gradient descent, KKT
ch05Machine Learning Basicscapacity, bias–variance, No Free Lunch, MLE, manifolds
ch06Deep Feedforward Networksoutput/hidden units, universal approximation, backprop
ch07Regularizationnorm penalties, augmentation, early stopping, dropout
ch08OptimizationSGD, momentum, init, Adam, batch norm, saddles
ch09Convolutional Networkssparse interactions, sharing, equivariance, pooling
ch10Sequence ModelingBPTT, vanishing gradients, LSTM/GRU, attention
ch11Practical Methodologymetrics, baselines, the data-vs-capacity rule, debugging
ch12Applicationsscaling, compression, vision, speech, NLP (dated)
ch13Linear Factor ModelsPPCA, factor analysis, ICA, sparse coding
ch14Autoencodersundercomplete, sparse, denoising, contractive
ch15Representation Learningtransfer, distributed codes, disentanglement
ch16Structured Probabilistic Modelsdirected/undirected, energy-based, d-separation
ch17Monte Carlo Methodsimportance sampling, MCMC, Gibbs, mixing
ch18Confronting the Partition FunctionCD/PCD, pseudolikelihood, score matching, NCE, AIS
ch19Approximate InferenceELBO, EM, mean field, amortization
ch20Deep Generative ModelsBoltzmann machines, VAE, GAN, autoregressive

Topic Index

  • Activation functions, ReLU, GELU → ch06
  • Adam, AdamW, adaptive optimizers → ch08, ch07
  • Attention, transformers → ch10, ch12
  • Autoencoders, denoising, sparse → ch14, ch13
  • Backpropagation, autodiff → ch06
  • Batch / layer normalization → ch08
  • Bias–variance, double descent → ch05
  • Convolution, pooling, receptive field → ch09
  • Cross-entropy, KL divergence, entropy → ch03
  • Diffusion, score matching → ch18, ch14, ch20
  • Dropout, weight decay, early stopping → ch07
  • ELBO, variational inference, EM → ch19
  • Energy-based models, graphical models → ch16
  • GANs, VAEs, generative taxonomy → ch20
  • Gradient clipping, exploding/vanishing → ch10, ch08
  • Hyperparameter search → ch11
  • Initialization → ch08
  • LSTM, GRU, BPTT, teacher forcing → ch10
  • Maximum likelihood, MAP → ch05, ch03
  • MCMC, Gibbs, importance sampling → ch17
  • Numerical stability, softmax, log-space → ch04
  • Partition function, CD, PCD, NCE → ch18, ch16
  • PCA, ICA, factor analysis → ch13, ch02
  • Representation learning, transfer, probes → ch15, ch01
  • Saddle points, ill-conditioning → ch08, ch04
  • SVD, eigendecomposition, condition number → ch02
  • Training diagnostics, metric choice → ch11
  • Universal approximation → ch06

Supporting Files

Tools

bash
S=engineering/deep-learning-book/skills/deep-learning-book/scripts
python3 $S/reading_path_planner.py --goal "train a transformer" --background applied --hours-per-week 5
python3 $S/training_diagnostics.py --train-loss 0.02 --val-loss 1.9 --grad-norm 0.4 --epochs 30
python3 $S/capacity_planner.py --params 12000000 --train-examples 50000 --train-error 0.01 --val-error 0.22
python3 $S/model_arithmetic.py --spec-sample

Every tool supports --help, --sample and --output json, uses the standard library only, and returns typed exit codes.


Scope & Limits

This companion covers the 2016 edition's 20 chapters and the delta between them and 2026 practice. It does not cover: reinforcement learning beyond passing mention, LLM training infrastructure, RLHF/DPO alignment, agentic systems, MLOps tooling, or fairness and safety evaluation — none of which the book treats. For production ML engineering use engineering-team/senior-ml-engineer; for LLM cost work use engineering/llm-cost-optimizer.

When a question lands outside the book, say the book does not cover it and cite the delta reference for what replaced its position. A companion that quietly extrapolates is worse than one that names its boundary.

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 36 other files (scripts, references, assets) in engineering/deep-learning-book/skills/deep-learning-book of alirezarezvani/claude-skills.

  • SKILL.md
  • assets/chapter_worksheet.md
  • assets/example_layer_spec.json
  • assets/study_log_template.md
  • chapters/ch01-introduction.md
  • chapters/ch02-linear-algebra.md
  • chapters/ch03-probability-information-theory.md
  • chapters/ch04-numerical-computation.md
  • chapters/ch05-machine-learning-basics.md
  • chapters/ch06-deep-feedforward-networks.md
  • chapters/ch07-regularization.md
  • chapters/ch08-optimization.md
  • chapters/ch09-convolutional-networks.md
  • chapters/ch10-sequence-modeling.md
  • chapters/ch11-practical-methodology.md
  • chapters/ch12-applications.md
  • chapters/ch13-linear-factor-models.md
  • chapters/ch14-autoencoders.md
  • chapters/ch15-representation-learning.md
  • … and 18 more

Open the folder on GitHubat commit 19392f7

Compare with similar skills

Deep Learning Book next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deep Learning Book compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deep Learning Book this skillalirezarezvani/claude-skills28k—~2.9kAutomated safety check: PassMIT
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Neuron Nki Writinguw-syfi/vibesys103—~5kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
Add Oponnx/onnx22k—~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Neuron Nki Writing

    uw-syfi/vibesys

    Guide for writing and modifying NKI kernels. An agent skill from uw-syfi/vibesys.

    103 GitHub stars~5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Op

    onnx/onnx

    Add a new ONNX operator or update an existing operator to a new opset version.

    22k GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Questions about Deep Learning Book

What does Deep Learning Book do?

Study companion and working knowledge base for the Deep Learning textbook by Goodfellow, Bengio & Courville (MIT Press, 2016), read free at deeplearningbook.org. Deep Learning Book is an agent skill from alirezarezvani/claude-skills.org.

When should I use Deep Learning Book?

Deep Learning Book fits situations like: teaching this book; planning a route through it; deciding whether a chapters advice is still current; translating its math into a training decision.

How do I install Deep Learning Book in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill deep-learning-book -a claude-code`. Or copy the skill folder (engineering/deep-learning-book/skills/deep-learning-book in alirezarezvani/claude-skills) into .claude/skills/deep-learning-book in your project. Claude Code loads it when a task matches its description.

How do I install Deep Learning Book in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill deep-learning-book -a codex`. Or copy the skill folder (engineering/deep-learning-book/skills/deep-learning-book in alirezarezvani/claude-skills) into .agents/skills/deep-learning-book in your project. Codex loads it when a task matches its description.

Can I use Deep Learning Book in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill deep-learning-book -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deep-learning-book, .gemini/skills/deep-learning-book, .github/skills/deep-learning-book and .opencode/skills/deep-learning-book in your project.

What does Deep Learning Book need to run?

Going by SKILL.md and its folder, Deep Learning Book needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Deep Learning Book access the network?

SKILL.md names 1 domain. As links in the text: deeplearningbook.org. This is read from the text; nothing was executed.

Is Deep Learning Book safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Deep Learning Book use?

Deep Learning Book is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deep Learning Book use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.

What are the alternatives to Deep Learning Book?

Skills that share tags, products or a category with Deep Learning Book: Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars), Neuron Nki Writing (uw-syfi/vibesys, 103 stars), Add Uint Support (pytorch/pytorch, 104k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deep Learning Book?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,788 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.