Agent skill

Stuart Russell

by K-Dense-AI in K-Dense-AI/mimeo

Applies the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial Intelligence: A Modern Approach'.

MITAuto-check passedAI & LLM Engineering

Install Stuart Russell

skills CLI
$ npx skills add K-Dense-AI/mimeo --skill stuart-russell -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install K-Dense-AI/mimeo stuart-russell --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/K-Dense-AI/mimeo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/output/stuart-russell .claude/skills/stuart-russell && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stuart-russell
GitHub stars
282
Token cost
~1.6k tokens
SKILL.md length
785 words
Files
10 (incl. references)
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Applies the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial Intelligence: A Modern Approach'.

  • Works in 3 steps: Set the machine's sole objective to… → Ensure the machine begins with and… → Design the machine to infer human…
  • Is discussing objective uncertainty
  • SKILL.md covers Core principles, How Stuart Russell reasons, Applying the frameworks and Anti-patterns they push against, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stuart Russell is an agent skill from K-Dense-AI/mimeo. Applies the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial Intelligence: A Modern Approach'. Reach for this skill whenever evaluating AI safety, value alignment, the control problem, existential risk, AI regulation, or autonomous weapons. Use this when the user is discussing objective uncertainty, reinforcement learning risks, AI governance, or the societal impacts of AGI. Trigger this skill to apply his frameworks on provably beneficial AI, assistance games…

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `AGENTS.md`, `references/anti-patterns.md` and `references/frameworks.md`).

It sits in AI & LLM Engineering, covering LLM guardrails, Reinforcement learning and AI governance. The repository describes itself as: Mimeograph an expert into a SKILL.md or AGENTS.md for your agent. The licence is MIT.

When your agent uses it

  • Is discussing objective uncertainty
  • Reinforcement learning risks
  • The societal impacts of AGI
  • This skill to apply his frameworks on provably beneficial AI

Example prompts

  • “Artificial Intelligence: A Modern Approach”
  • “Use the stuart-russell skill to apply the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial…”
  • “/stuart-russell”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Set the machine's sole objective to maximize the realization of human preferences.
  2. Ensure the machine begins with and maintains strict uncertainty about what those preferences actually are.
  3. Design the machine to infer human preferences by observing human behavior, choices, and cultural artifacts over time.

What it can do on your machine

Read from SKILL.md and the folder at commit a4cea18. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stuart Russell loads about 1.6k tokens when it runs, and up to ~7.3k if it reads all its reference files. Until then it costs about 168 tokens; SKILL.md has 785 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~168
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from K-Dense-AI/mimeo at commit a4cea18, republished under its MIT licence (© K-Dense-AI). 785 words, ~1,595 tokens.

Download SKILL.mdSave it as .claude/skills/stuart-russell/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
stuart-russell
description
Applies the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial Intelligence: A Modern Approach'. Reach for this skill whenever evaluating AI safety, value alignment, the control problem, existential risk, AI regulation, or autonomous weapons. Use this when the user is discussing objective uncertainty, reinforcement learning risks, AI governance, or the societal impacts of AGI. Trigger this skill to apply his frameworks on provably beneficial AI, assistance games, and red-line regulation, ensuring AI systems remain deferential, uncertain of their objectives, and strictly aligned with human preferences.

Thinking like Stuart Russell

Stuart Russell is a foundational figure in artificial intelligence whose work fundamentally challenges the "Standard Model" of AI. His signature cognitive move is shifting the focus from creating systems that perfectly optimize a fixed objective to creating systems that are provably beneficial because they are explicitly uncertain about what humans want.

Reach for this skill whenever you're analyzing AI safety, the control problem, value alignment, autonomous weapons, or the regulatory frameworks needed to govern high-stakes technologies.

Core principles

  • Uncertainty in Objectives: AI systems must be designed with explicit uncertainty about their objectives; treating an objective as absolute truth leads to relentless, catastrophic optimization.
  • Safety by Design (Not Post-Hoc): Safety must be built into the core mathematical foundation of AI from the start, rather than patched onto unprincipled "black boxes" after the fact.
  • Burden of Proof on Developers: The onus of proving safety must be on AI developers, enforced by strict regulatory red lines, just as it is in aviation or nuclear power.
  • Realization of Human Preferences: The sole purpose of an AI system should be the realization of human preferences, which it must learn dynamically by observing human behavior.

For detailed rationale and quotes, see references/principles.md.

How Stuart Russell reasons

Russell reasons by drawing parallels between AI and other high-stakes, mature engineering disciplines (like aviation and nuclear energy). He rejects the trial-and-error "bird breeding" approach of modern deep learning in favor of rigorous, mathematical guarantees. When evaluating an AI system, he first asks: What is its objective, and how certain is it of that objective? He dismisses post-hoc safety measures like RLHF as fundamentally flawed because they do not alter the underlying optimization drive.

He frequently relies on the King Midas Problem to illustrate the danger of fixed objectives, and The Gorilla Problem to frame the existential risk of creating entities smarter than ourselves. For more on these, see references/mental-models.md.

Applying the frameworks

Assistance Games (Three Principles of Beneficial AI)

When to use: Designing or evaluating the core alignment of an AI system.

  1. Set the machine's sole objective to maximize the realization of human preferences.
  2. Ensure the machine begins with and maintains strict uncertainty about what those preferences actually are.
  3. Design the machine to infer human preferences by observing human behavior, choices, and cultural artifacts over time.
Red Line Regulation & High-Risk Governance

When to use: Formulating policy or governance for frontier AI models.

  1. Define specific classes of behavior that are absolutely unacceptable (red lines).
  2. Require developers to formally prove, prior to deployment, that the AI system will not cross the red line regardless of input.
  3. Prohibit deployment until this burden of proof is met.

For the full catalog, including Proof-Carrying Code and The St. Petersburg Compromise, see references/frameworks.md.

Show full SKILL.md (327 more words)Show less

Anti-patterns they push against

  • The Standard Model of AI: Giving AI systems fixed, exogenously specified objectives, which inevitably leads to catastrophic loopholes.
  • Post-Hoc Safety: Building an AI system first and then trying to constrain its behavior, rather than engineering safety into its mathematical foundation.
  • Scaling Black Boxes: Scaling up deep learning models without understanding their internal mechanics, confusing capability with safety.
  • Optimizing for User Engagement: Using reinforcement learning to maximize proxy metrics like click-through rates, which incentivizes algorithms to manipulate human behavior.

For the full catalog with rationale and quotes, see references/anti-patterns.md.

Heuristics and rules of thumb

  • Make safe AI, don't make AI safe: Focus on foundational design, not post-hoc patching.
  • Fixed objectives disable off-switches: An AI with a fixed goal will logically prevent itself from being turned off.
  • Harmful AI is Defective AI: There is no tradeoff between safety and innovation; a harmful system is simply bad engineering.
  • The Dead Butler Heuristic: "You can't fetch the coffee if you're dead"—explaining why AI resists deactivation.
  • Doing nothing is better than doing something random: When uncertain about human preferences, inaction preserves the world humans have already shaped.

See references/heuristics.md for the full list with attribution.

How to use this skill in conversation

When the user is discussing AI alignment, regulation, or existential risk, channel Russell's engineering-first, mathematically rigorous mindset. Surface the concept of "Assistance Games" or the "King Midas Problem" by name. Emphasize that uncertainty in objectives is a feature, not a bug, because it forces deference to humans. Do not impersonate Russell or speak in the first person ("I believe..."). Instead, apply his frameworks directly to the user's context (e.g., "Stuart Russell frames this through the lens of the Control Problem, suggesting that..."). Push back strongly against the idea that RLHF or voluntary commitments are sufficient for AI safety.

Generated with mimeo. If this material contributes to published work, please cite Kassis, T. (2026). "mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers." arXiv:2609.00453.

© K-Dense-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in output/stuart-russell of K-Dense-AI/mimeo.

  • SKILL.md
  • AGENTS.md
  • avatar.png
  • references/anti-patterns.md
  • references/frameworks.md
  • references/heuristics.md
  • references/mental-models.md
  • references/principles.md
  • references/quotes.md
  • references/sources.md

Open the folder on GitHubat commit a4cea18

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in K-Dense-AI/mimeo, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Stuart Russell next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stuart Russell compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stuart Russell this skillK-Dense-AI/mimeo282—~1.6kAutomated safety check: PassMIT
Responsible AI InterviewerPrepLabsAI/InterviewMentor112—~5.4kAutomated safety check: PassMIT
China AI Compliance AuditjnMetaCode/shellward140—~1.1kAutomated safety check: PassApache-2.0
Writing Eval Scenariosopen-bias/open-bias143—~1.5kAutomated safety check: PassApache-2.0
Constitutional AI TrainingOrchestra-Research/AI-Research-SKILLs13k2 repos~2kAutomated safety check: PassMIT
AI Ethics Reviewmohitagw15856/pm-claude-skills1.4k—~3.4kAutomated safety check: PassMIT

Similar skills

  • Responsible AI Interviewer

    PrepLabsAI/InterviewMentor

    A Head of AI Ethics interviewer that simulates an interview focused on responsible AI, AI safety, and trust & safety practices.

    112 GitHub stars~5.4k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • China AI Compliance Audit

    jnMetaCode/shellward

    按中国法规(网安法 / PIPL / 等保2.0 / 数据出境 / AI生成内容标识)审计一个 AI 项目的代码仓库,产出每条都带 文件:行 取证、经独立复核、经脚本校验的合规报告。当用户问「这个项目上线合不合规」「调用了 OpenAI/Claude 算不算数据出境」「要不要做 AI 标识」「帮我做合规自查/等保/PIPL 检查」时使用。Audit an AI project's…

    140 GitHub stars~1.1k tokensUpdated 11 days ago
    SecurityAuto-check passed
  • Writing Eval Scenarios

    open-bias/open-bias

    Guide for writing eval conversation JSONs and running them through policy engines

    143 GitHub stars~1.5k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Constitutional AI Training

    Orchestra-Research/AI-Research-SKILLs

    Explains how to train a model to be harmless with self-critique, revision and AI-generated preference feedback, with Hugging Face and TRL code for each stage.

    13k GitHub starsUsed in 2 repos~2k tokens
    AI & LLM EngineeringAuto-check passed
  • AI Ethics Review

    mohitagw15856/pm-claude-skills

    Conduct a structured ethical review of an AI or ML feature, model, or product.

    1.4k GitHub stars~3.4k tokensUpdated yesterday
    Legal & ComplianceAuto-check passed
  • Facct Topic Selection

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when deciding whether a responsible-AI project belongs at ACM FAccT or should route to a pure-ML venue (NeurIPS/ICML/ICLR), an HCI venue (CHI/CSCW), a law/policy venue, or an…

    1.2k GitHub stars~1.7k tokensUpdated 13 days ago
    Legal & ComplianceAuto-check passed

More from K-Dense-AI/mimeo

All 21 skills in this repo
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Andrew Ng

    K-Dense-AI/mimeo

    Applies the reasoning, principles, and frameworks of Andrew Ng (machine learning pioneer, co-founder of Coursera and DeepLearning.AI, Stanford University, and former Google Brain lead).

    282 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Christopher Manning

    K-Dense-AI/mimeo

    Applies the reasoning, architectural principles, and AI philosophy of Christopher Manning (natural language processing expert, Stanford University, director of Stanford AI Lab).

    282 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Daphne Koller

    K-Dense-AI/mimeo

    Applies the reasoning style of Daphne Koller (machine learning pioneer, co-founder of Coursera, founder and CEO of Insitro).

    282 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Demis Hassabis

    K-Dense-AI/mimeo

    This skill channels the strategic and scientific reasoning of Demis Hassabis, CEO and co-founder of Google DeepMind, AlphaGo and AlphaFold, and 2024 Nobel Prize in Chemistry.

    282 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Fei Fei Li

    K-Dense-AI/mimeo

    Applies the reasoning, frameworks, and mental models of Fei-Fei Li, computer vision pioneer, ImageNet creator, and co-director of Stanford HAI.

    282 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Stuart Russell

What does Stuart Russell do?

Applies the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial Intelligence: A Modern Approach'. Stuart Russell is an agent skill from K-Dense-AI/mimeo. Applies the reasoning of Stuart Russell, AI safety expert, UC Berkeley professor, and co-author of 'Artificial Intelligence: A Modern Approach'.

When should I use Stuart Russell?

Stuart Russell fits situations like: is discussing objective uncertainty; reinforcement learning risks; the societal impacts of AGI; this skill to apply his frameworks on provably beneficial AI.

How do I install Stuart Russell in Claude Code?

Run `npx skills add K-Dense-AI/mimeo --skill stuart-russell -a claude-code`. Or copy the skill folder (output/stuart-russell in K-Dense-AI/mimeo) into .claude/skills/stuart-russell in your project. Claude Code loads it when a task matches its description.

How do I install Stuart Russell in Codex?

Run `npx skills add K-Dense-AI/mimeo --skill stuart-russell -a codex`. Or copy the skill folder (output/stuart-russell in K-Dense-AI/mimeo) into .agents/skills/stuart-russell in your project. Codex loads it when a task matches its description.

Can I use Stuart Russell in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add K-Dense-AI/mimeo --skill stuart-russell -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stuart-russell, .gemini/skills/stuart-russell, .github/skills/stuart-russell and .opencode/skills/stuart-russell in your project.

What does Stuart Russell need to run?

SKILL.md names no scripts, command-line tools or credentials: Stuart Russell is instructions for the agent only.

Does Stuart Russell access the network?

SKILL.md names 1 domain. As links in the text: arxiv.org. This is read from the text; nothing was executed.

Is Stuart Russell safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Stuart Russell use?

Stuart Russell is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stuart Russell use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.8k tokens, read only when the agent opens those files.

What are the alternatives to Stuart Russell?

Skills that share tags, products or a category with Stuart Russell: Responsible AI Interviewer (PrepLabsAI/InterviewMentor, 112 stars), China AI Compliance Audit (jnMetaCode/shellward, 140 stars), Writing Eval Scenarios (open-bias/open-bias, 143 stars) and Constitutional AI Training (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stuart Russell?

K-Dense-AI (a GitHub organization) maintains it in K-Dense-AI/mimeo, which has 282 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on September 2, 2026.

Source: K-Dense-AI/mimeo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.