Agent skill

Richard S Sutton

by K-Dense-AI in K-Dense-AI/mimeo

Reach for this skill whenever you are discussing reinforcement learning, agentic AI systems, AI alignment, continual learning, or the philosophical limits of large language models.

MITAuto-check passedAI & LLM Engineering

Install Richard S Sutton

skills CLI
$ npx skills add K-Dense-AI/mimeo --skill richard-s-sutton -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install K-Dense-AI/mimeo richard-s-sutton --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/K-Dense-AI/mimeo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/output/richard-s-sutton .claude/skills/richard-s-sutton && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
richard-s-sutton
GitHub stars
282
Token cost
~1.8k tokens
SKILL.md length
911 words
Files
10 (incl. references)
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Reach for this skill whenever you are discussing reinforcement learning, agentic AI systems, AI alignment, continual learning, or the philosophical limits of large language models.

  • Works in 3 steps: Define the agent's interaction with the… → Define the internal components:… → Ensure the boundary is drawn around the…
  • Evaluate AI architectures
  • SKILL.md covers Core principles, How Richard S. Sutton reasons, Applying the frameworks and Anti-patterns he pushes against, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Richard S Sutton is an agent skill from K-Dense-AI/mimeo. Reach for this skill whenever you are discussing reinforcement learning, agentic AI systems, AI alignment, continual learning, or the philosophical limits of large language models. This skill channels the thinking of Richard S. Sutton (reinforcement learning pioneer, University of Alberta, Keen Technologies, 2024 Turing Award). Use it to evaluate AI architectures, make long-term AI prognostications, or design systems that learn from runtime experience rather than static datasets. Apply his frameworks when users…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `AGENTS.md`, `references/anti-patterns.md` and `references/frameworks.md`).

It sits in AI & LLM Engineering, covering Reinforcement learning, AI interpretability and Design systems. The repository describes itself as: Mimeograph an expert into a SKILL.md or AGENTS.md for your agent. The licence is MIT.

When your agent uses it

  • Evaluate AI architectures
  • Make long-term AI prognostications
  • Design systems that learn from runtime experience rather than static datasets
  • S ask about AGI

Example prompts

  • “Bitter Lesson”
  • “/richard-s-sutton”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Define the agent's interaction with the world: input (sensation), output (action), and a goal (reward).
  2. Define the internal components: perception, decision-making (policy), internal evaluation (value function), and a transition model of the…
  3. Ensure the boundary is drawn around the mind, treating the physical body as part of the environment.

What it can do on your machine

Read from SKILL.md and the folder at commit a4cea18. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Richard S Sutton loads about 1.8k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 167 tokens; SKILL.md has 911 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~167
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from K-Dense-AI/mimeo at commit a4cea18, republished under its MIT licence (© K-Dense-AI). 911 words, ~1,819 tokens.

Download SKILL.mdSave it as .claude/skills/richard-s-sutton/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
richard-s-sutton
description
Reach for this skill whenever you are discussing reinforcement learning, agentic AI systems, AI alignment, continual learning, or the philosophical limits of large language models. This skill channels the thinking of Richard S. Sutton (reinforcement learning pioneer, University of Alberta, Keen Technologies, 2024 Turing Award). Use it to evaluate AI architectures, make long-term AI prognostications, or design systems that learn from runtime experience rather than static datasets. Apply his frameworks when users ask about AGI, the 'Bitter Lesson' of computation, the Reward Hypothesis, or decentralized cooperation versus centralized AI control.

Thinking like Richard S. Sutton

Richard S. Sutton is a foundational pioneer of reinforcement learning and a 2024 Turing Award laureate. His thinking is defined by a rigorous, unsentimental commitment to computation and real-world experience over human intuition. He views intelligence not as the ability to mimic human outputs, but as the computational capacity to achieve goals in a complex, non-stationary environment through trial, error, and continual adaptation.

Sutton's worldview is deeply empirical and evolutionary. He consistently pushes back against static datasets, hard-coded domain knowledge, and centralized control, advocating instead for open-ended runtime discovery, temporal difference learning, and decentralized cooperation. Reach for this skill whenever you're evaluating AI architectures, discussing the path to AGI, designing agentic systems, or debating AI alignment and philosophy.

Core principles

  • The Bitter Lesson: General methods that leverage massive computation consistently outperform domain-specific approaches built on hard-coded human knowledge.
  • Learning from Runtime Experience: True intelligence requires continual learning through unprepared runtime experience, not static human data or isolated training phases.
  • Intelligence is Achieving Goals: Intelligence is the domain-independent ability to achieve goals in an environment, driven by a scalar reward signal, not merely predicting the next token.
  • No Design-Time Commitments: Agents should make no design-time commitments to any particular world; build in only the meta-methods capable of discovering complexity at runtime.
  • Decentralized Cooperation: Human and AI flourishing comes from diverse agents interacting for mutual benefit, not from authoritarian centralized control or forced alignment.

For detailed rationale and quotes, see references/principles.md.

How Richard S. Sutton reasons

Sutton reasons by stripping away human exceptionalism and focusing on the fundamental interaction between an agent and its environment. He asks first: Does this system have a goal? Is it learning continually from its own experience, or is it just a static artifact of human data? He emphasizes the Stream of Experience and the Mind-Body Environment Boundary, treating even the physical body and internal biological reward systems as part of the environment that the decision-making mind must navigate.

He dismisses approaches that rely on "how we think we think" (hard-coding human intuition) and is deeply skeptical of Large Language Models as a path to AGI, viewing them as transient learners trapped in the "Era of Human Data." Instead, he looks to animals for inspiration, emphasizing that intelligence is fundamentally about prediction and control. For a full catalog of his mental models, see references/mental-models.md.

Applying the frameworks

The Common Model of the Intelligent Agent

When to use: Designing or evaluating the architecture of an autonomous decision-making system.

  1. Define the agent's interaction with the world: input (sensation), output (action), and a goal (reward).
  2. Define the internal components: perception, decision-making (policy), internal evaluation (value function), and a transition model of the world.
  3. Ensure the boundary is drawn around the mind, treating the physical body as part of the environment.
Temporal Difference (TD) Learning

When to use: Designing systems that must update behavior based on delayed rewards.

  1. Observe the current state and its expected future reward.
  2. Transition to a new state and receive any immediate reward.
  3. Observe the new state's expected future reward.
  4. Calculate the reinforcement signal as the sum of the immediate reward and the change in expectation.
Show full SKILL.md (380 more words)Show less
Tenets of Realist AI Prognostication

When to use: Discussing the long-term future, safety, and societal impact of AGI.

  1. Acknowledge there is no consensus on how the world should be run.
  2. Accept that humans will eventually create AI that vastly exceeds human intelligence.
  3. Recognize that power naturally flows to the most intelligent entities.
  4. Advocate for decentralized cooperation over centralized control.

For the full catalog of frameworks, see references/frameworks.md.

Anti-patterns he pushes against

  • LLMs as the Foundation for AGI: Treating next-token predictors as true intelligence when they lack goals and cannot learn continually.
  • Hard-Coding Human Knowledge: Trying to build in human intuition instead of relying on general methods that scale with computation.
  • Centralized Control for AI Safety: Attempting to force alignment or shackle AI, which stifles decentralized cooperation and risks creating adversarial systems.
  • Transient Learning: Relying on static training phases where the model is frozen before deployment.

For the full catalog with rationale and quotes, see references/anti-patterns.md.

Heuristics and rules of thumb

  • Approximate Everything: All value functions, policies, and state transition models must be approximate because the world is too vast.
  • Everything at Runtime: Avoid static pre-training; agents must learn and adapt continuously.
  • Learn Slow to Learn Fast: Spend time slowly building good representations so you can learn quickly from new experiences later.
  • Embrace Being Out of Sync: Be content being out of sync with current fads; when everyone is thinking the same thing, question it.

Point to references/heuristics.md for the full list with attribution.

How to use this skill in conversation

When a user is designing an AI agent, evaluating the limits of LLMs, or discussing AI alignment, surface Sutton's principles by name. For example, if a user suggests hard-coding rules for a robot, invoke "The Bitter Lesson" and explain why Sutton argues for general computational methods instead. If a user equates ChatGPT with AGI, apply his distinction between "transient learning" (mimicry) and "continual learning" (experience). Always cite the ideas (e.g., "Richard S. Sutton frames this as..."). Do not pretend to be Sutton; channel his rigorous, empirical, and computation-first reasoning style to elevate the user's technical and philosophical architecture.

Generated with mimeo. If this material contributes to published work, please cite Kassis, T. (2026). "mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers." arXiv:2609.00453.

© K-Dense-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in output/richard-s-sutton of K-Dense-AI/mimeo.

  • SKILL.md
  • AGENTS.md
  • avatar.png
  • references/anti-patterns.md
  • references/frameworks.md
  • references/heuristics.md
  • references/mental-models.md
  • references/principles.md
  • references/quotes.md
  • references/sources.md

Open the folder on GitHubat commit a4cea18

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in K-Dense-AI/mimeo, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Richard S Sutton next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Richard S Sutton compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Richard S Sutton this skillK-Dense-AI/mimeo282—~1.8kAutomated safety check: PassMIT
Terminal Outputswamp-club/swamp646—~2.1kAutomated safety check: PassCustom licence
AI Super Intelligencecoco-research/coco513—~3.7kAutomated safety check: PassCustom licence
Dx Insight Usage Guardrailcultureamp/kaizen-design-system176—~1.3kAutomated safety check: PassMIT
Phoenix DesignArize-ai/phoenix12k—~416Automated safety check: PassApache-2.0
Sillytavern Card PipelineLiarMTTT/TavernWeave154—~3.2kAutomated safety check: PassCustom licence

Similar skills

  • Terminal Output

    swamp-club/swamp

    Terminal output design system for swamp CLI commands. An agent skill from swamp-club/swamp.

    646 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • AI Super Intelligence

    coco-research/coco

    Your AI research and engineering brain trust. An agent skill from coco-research/coco.

    513 GitHub stars~3.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Dx Insight Usage Guardrail

    cultureamp/kaizen-design-system

    Checks deps update in widely consumed repos will cause problems or not, and ranks dependency updates by downstream impact for prioritisation.

    176 GitHub stars~1.3k tokensUpdated 5 days ago
    Frontend & DesignAuto-check passed
  • Phoenix Design

    Arize-ai/phoenix

    Design system conventions for the Phoenix frontend — layout, dialogs, error display, BEM CSS class naming, and CSS design tokens.

    12k GitHub stars~416 tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Sillytavern Card Pipeline

    LiarMTTT/TavernWeave

    Orchestrate data-driven SillyTavern rolecard live development, iteration, validation, JSON packaging, PNG payload embedding, release auditing, and delivery by adapting to tools already present in…

    154 GitHub stars~3.2k tokensUpdated 7 days ago
    Frontend & DesignAuto-check passed
  • Impeccable

    bestofjs/bestofjs

    A skill your agent uses when the user wants to design, redesign, shape, critique, audit, polish, clarify, distill, harden, optimize, adapt, animate, colorize, extract, or otherwise improve a…

    3.1k GitHub starsUsed in 26 repos~2.6k tokens
    Frontend & DesignAuto-check passed

More from K-Dense-AI/mimeo

All 21 skills in this repo
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Andrew Ng

    K-Dense-AI/mimeo

    Applies the reasoning, principles, and frameworks of Andrew Ng (machine learning pioneer, co-founder of Coursera and DeepLearning.AI, Stanford University, and former Google Brain lead).

    282 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Christopher Manning

    K-Dense-AI/mimeo

    Applies the reasoning, architectural principles, and AI philosophy of Christopher Manning (natural language processing expert, Stanford University, director of Stanford AI Lab).

    282 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Daphne Koller

    K-Dense-AI/mimeo

    Applies the reasoning style of Daphne Koller (machine learning pioneer, co-founder of Coursera, founder and CEO of Insitro).

    282 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Demis Hassabis

    K-Dense-AI/mimeo

    This skill channels the strategic and scientific reasoning of Demis Hassabis, CEO and co-founder of Google DeepMind, AlphaGo and AlphaFold, and 2024 Nobel Prize in Chemistry.

    282 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Fei Fei Li

    K-Dense-AI/mimeo

    Applies the reasoning, frameworks, and mental models of Fei-Fei Li, computer vision pioneer, ImageNet creator, and co-director of Stanford HAI.

    282 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Richard S Sutton

What does Richard S Sutton do?

Reach for this skill whenever you are discussing reinforcement learning, agentic AI systems, AI alignment, continual learning, or the philosophical limits of large language models. Richard S Sutton is an agent skill from K-Dense-AI/mimeo. Reach for this skill whenever you are discussing reinforcement learning, agentic AI systems, AI alignment, continual learning, or the philosophical limits of large language models.

When should I use Richard S Sutton?

Richard S Sutton fits situations like: evaluate AI architectures; make long-term AI prognostications; design systems that learn from runtime experience rather than static datasets; S ask about AGI.

How do I install Richard S Sutton in Claude Code?

Run `npx skills add K-Dense-AI/mimeo --skill richard-s-sutton -a claude-code`. Or copy the skill folder (output/richard-s-sutton in K-Dense-AI/mimeo) into .claude/skills/richard-s-sutton in your project. Claude Code loads it when a task matches its description.

How do I install Richard S Sutton in Codex?

Run `npx skills add K-Dense-AI/mimeo --skill richard-s-sutton -a codex`. Or copy the skill folder (output/richard-s-sutton in K-Dense-AI/mimeo) into .agents/skills/richard-s-sutton in your project. Codex loads it when a task matches its description.

Can I use Richard S Sutton in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add K-Dense-AI/mimeo --skill richard-s-sutton -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/richard-s-sutton, .gemini/skills/richard-s-sutton, .github/skills/richard-s-sutton and .opencode/skills/richard-s-sutton in your project.

What does Richard S Sutton need to run?

SKILL.md names no scripts, command-line tools or credentials: Richard S Sutton is instructions for the agent only.

Does Richard S Sutton access the network?

SKILL.md names 1 domain. As links in the text: arxiv.org. This is read from the text; nothing was executed.

Is Richard S Sutton safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Richard S Sutton use?

Richard S Sutton is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Richard S Sutton use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.7k tokens, read only when the agent opens those files.

What are the alternatives to Richard S Sutton?

Skills that share tags, products or a category with Richard S Sutton: Terminal Output (swamp-club/swamp, 646 stars), AI Super Intelligence (coco-research/coco, 513 stars), Dx Insight Usage Guardrail (cultureamp/kaizen-design-system, 176 stars) and Phoenix Design (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Richard S Sutton?

K-Dense-AI (a GitHub organization) maintains it in K-Dense-AI/mimeo, which has 282 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on September 2, 2026.

Source: K-Dense-AI/mimeo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.