Agent skill

Thinking Thought Experiment

by tjboudreaux in tjboudreaux/cc-thinking-skills

When a real test is too rare, large, or irreversible, run a controlled counterfactual: isolate one variable, fix conditions, trace the mechanistic chain, and bound what the result implies.

MITAuto-check passed

Install Thinking Thought Experiment

skills CLI
$ npx skills add tjboudreaux/cc-thinking-skills --skill thinking-thought-experiment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tjboudreaux/cc-thinking-skills thinking-thought-experiment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tjboudreaux/cc-thinking-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/thinking-thought-experiment .claude/skills/thinking-thought-experiment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
thinking-thought-experiment
GitHub stars
1.6k
Token cost
~937 tokens
SKILL.md length
484 words
Files
1
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

When a real test is too rare, large, or irreversible, run a controlled counterfactual: isolate one variable, fix conditions, trace the mechanistic chain, and bound what the result implies.

  • Works in 6 steps: State the question and isolation. Name… → Fix initial conditions. Specify system… → Trace the mechanism step by step. From… → …
  • SKILL.md covers When to Use, When NOT to Use, Procedure and Output, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Thinking Thought Experiment is an agent skill from tjboudreaux/cc-thinking-skills. When a real test is too rare, large, or irreversible, run a controlled counterfactual: isolate one variable, fix conditions, trace the mechanistic chain, and bound what the result implies.

Its SKILL.md is about 940 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 28 eval-informed mental models and critical-thinking skills for Claude Code, GitHub Copilot, Codex, Cursor, and other Agent Skills-compatible tools. The licence is MIT.

Example prompts

  • “/thinking-thought-experiment”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. State the question and isolation. Name exactly one primary variable or counterfactual change. Freeze all other conditions as the control…
  2. Fix initial conditions. Specify system state, load, configuration, actors, and what is not changed. Write values concrete enough that…
  3. Trace the mechanism step by step. From t0, record what fails, queues, retries, or adapts next—and why—using known components and policies…
  4. Extract invariants and break points. Note what still holds (invariants) and the first step where the system violates a requirement…
  5. Bound implications. Map insights only to actions or checks justified by the chain (limits, guards, monitoring, redesign). Label…
  6. Name a discriminating real check, then stop. For the weakest link, state the cheapest observation or experiment that would confirm or kill…

What it can do on your machine

Read from SKILL.md and the folder at commit 7b8fece. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Thinking Thought Experiment loads about 937 tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 484 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~937

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tjboudreaux/cc-thinking-skills at commit 7b8fece, republished under its MIT licence (© tjboudreaux). 484 words, ~937 tokens.

Download SKILL.mdSave it as .claude/skills/thinking-thought-experiment/SKILL.md (or your agent's skills folder).
name
thinking-thought-experiment
description
When a real test is too rare, large, or irreversible, run a controlled counterfactual: isolate one variable, fix conditions, trace the mechanistic chain, and bound what the result implies.
disable-model-invocation
true

Thought Experiment

When empiricism is out of reach, run a disciplined counterfactual: one isolated change, fixed conditions, step-by-step mechanism, and a hard bound on implications.

When to Use

  • You need behavior under failure, scale, or policy you cannot cheaply trigger or measure (region outage, 100x load, one-way architecture).
  • A decision is expensive or irreversible and a mental trace can surface break points before commit.
  • Edge cases are too costly to stage, but a mechanistic chain can still expose missing controls.

When NOT to Use

  • A cheap real test exists (load test, flag, query, spike) → run the test; do not substitute imagination.
  • Adversarial security attack-path work → use red-team structure, not free-form scenarios.
  • You already know the mechanism and only need a decision under known facts → decide; do not dramatize.
  • Vague "what if everything" brainstorming without a single isolated variable → tighten or stop.

Procedure

  1. State the question and isolation. Name exactly one primary variable or counterfactual change. Freeze all other conditions as the control world. Reject multi-variable "and also" scenarios.
  2. Fix initial conditions. Specify system state, load, configuration, actors, and what is not changed. Write values concrete enough that another agent could replay the setup.
  3. Trace the mechanism step by step. From t0, record what fails, queues, retries, or adapts next—and why—using known components and policies only. No hand-wavy "then everything collapses"; each step needs a causal link.
  4. Extract invariants and break points. Note what still holds (invariants) and the first step where the system violates a requirement (capacity, correctness, safety, UX). Mark assumptions that, if false, void the chain.
  5. Bound implications. Map insights only to actions or checks justified by the chain (limits, guards, monitoring, redesign). Label speculative leaps beyond the isolation as out of bound.
  6. Name a discriminating real check, then stop. For the weakest link, state the cheapest observation or experiment that would confirm or kill it. Stop after one controlled chain with bounded implications; if a link is cheaply testable now, exit to that test instead of further imagination.
Show full SKILL.md (147 more words)Show less

Output

Emit a thought-experiment record:

  • question: what behavior or decision is under test
  • isolated_variable: single change vs control world
  • initial_conditions: frozen state and non-changes
  • consequence_chain: ordered mechanistic steps
  • invariants: what still holds
  • break_points: first requirement failures and critical assumptions
  • implication_bound: actions/checks justified by the chain only
  • discriminating_check: cheapest real observation to confirm or kill the weak link

Verification

  • Isolation check: more than one free variable without a stated control → invalid; reset.
  • Mechanism check: any step without a causal link to a known component/policy → rewrite or drop.
  • Implication bound: recommendations not entailed by the chain are out of scope.
  • Empiricism override: if a real test became available mid-analysis, stop the thought experiment and test.
  • Over-application guard: do not use this skill for ordinary debugging you can reproduce, or as a substitute for red-team threat modeling.
  • Stop: one isolated counterfactual → full chain → bounded implications + discriminating check; no scenario sprawl.

© tjboudreaux, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/thinking-thought-experiment of tjboudreaux/cc-thinking-skills.

Open the folder on GitHubat commit 7b8fece

Compare with similar skills

Thinking Thought Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Thinking Thought Experiment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Thinking Thought Experiment this skilltjboudreaux/cc-thinking-skills1.6k—~937Automated safety check: PassMIT
Finding ExperimentsPostHog/posthog40k—~826Automated safety check: PassCustom licence
ExperimentsArize-ai/phoenix12k—~1.8kAutomated safety check: PassCustom licence
Scroll Experiencesickn33/agentic-awesome-skills47k2 repos~534Automated safety check: PassMIT
Webgl Experiencenexu-io/open-design100k—~903Automated safety check: PassApache-2.0
Creating ExperimentsPostHog/posthog40k—~2.7kAutomated safety check: PassCustom licence

Similar skills

  • Finding Experiments

    PostHog/posthog

    Official

    Resolves a PostHog experiment reference from natural language to a concrete experiment ID by browsing experiment-list (not feature-flag tools), with disambiguation when multiple experiments match.

    40k GitHub stars~826 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Experiments

    Arize-ai/phoenix

    Run, read, and compare dataset-backed experiments to find evidence that a prompt or pipeline is improving.

    12k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Scroll Experience

    sickn33/agentic-awesome-skills

    Expert in building immersive scroll-driven experiences - parallax storytelling, scroll animations, interactive narratives, and cinematic web experiences.

    47k GitHub starsUsed in 2 repos~534 tokens
    Writing & ContentAuto-check passed
  • Webgl Experience

    nexu-io/open-design

    A full-screen, real-time WebGL/WebGL2 experience — animated shaders, 3D scenes, generative visuals, particle fields — rendered live on the GPU with a typographic overlay.

    100k GitHub stars~903 tokensUpdated today
    Game DevelopmentAuto-check passed
  • Creating Experiments

    PostHog/posthog

    Official

    Guides agents through experiment creation: reading the project's setup with experiment-setup-context, defining the hypothesis, configuring rollout and bucketing, setting up analytics and running…

    40k GitHub stars~2.7k tokensUpdated today
    Marketing & SEOAuto-check passed
  • Scroll Experience

    davila7/claude-code-templates

    Expert in building immersive scroll-driven experiences - parallax storytelling, scroll animations, interactive narratives, and cinematic web experiences.

    33k GitHub starsUsed in 3 repos~1.5k tokens
    Frontend & DesignAuto-check passed

More from tjboudreaux/cc-thinking-skills

All 26 skills in this repo
  • Thinking Bounded Rationality

    tjboudreaux/cc-thinking-skills

    A skill your agent uses when search or investigation could run forever.

    1.6k GitHub stars~703 tokensUpdated 2 mo ago
    Auto-check passed
  • Thinking Circle Of Competence

    tjboudreaux/cc-thinking-skills

    A skill your agent uses when a specific claim may lack grounding.

    1.6k GitHub stars~733 tokensUpdated 2 mo ago
    Auto-check passed
  • Thinking Lindy Effect

    tjboudreaux/cc-thinking-skills

    A skill your agent uses when longevity of a non-perishable option matters.

    1.6k GitHub stars~650 tokensUpdated 2 mo ago
    Auto-check passed
  • Thinking Via Negativa

    tjboudreaux/cc-thinking-skills

    A skill your agent uses when the reflex is to add a feature, layer, or process.

    1.6k GitHub stars~735 tokensUpdated 2 mo ago
    Auto-check passed
  • Thinking Cynefin

    tjboudreaux/cc-thinking-skills

    When the right response mode is unclear, classify the cause-effect domain first; decompose disorder.

    1.6k GitHub stars~651 tokensUpdated 2 mo ago
    Auto-check passed
  • Thinking Five Whys Plus

    tjboudreaux/cc-thinking-skills

    When a fault is localized and the proximate cause is known but the systemic root is not, chain evidence-linked whys with a counterfactual stop and a countermeasure.

    1.6k GitHub stars~881 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Thinking Thought Experiment

What does Thinking Thought Experiment do?

When a real test is too rare, large, or irreversible, run a controlled counterfactual: isolate one variable, fix conditions, trace the mechanistic chain, and bound what the result implies. Thinking Thought Experiment is an agent skill from tjboudreaux/cc-thinking-skills. When a real test is too rare, large, or irreversible, run a controlled counterfactual: isolate one variable, fix conditions, trace the mechanistic chain, and bound what the result implies.

How do I install Thinking Thought Experiment in Claude Code?

Run `npx skills add tjboudreaux/cc-thinking-skills --skill thinking-thought-experiment -a claude-code`. Or copy the skill folder (skills/thinking-thought-experiment in tjboudreaux/cc-thinking-skills) into .claude/skills/thinking-thought-experiment in your project. Claude Code loads it when a task matches its description.

How do I install Thinking Thought Experiment in Codex?

Run `npx skills add tjboudreaux/cc-thinking-skills --skill thinking-thought-experiment -a codex`. Or copy the skill folder (skills/thinking-thought-experiment in tjboudreaux/cc-thinking-skills) into .agents/skills/thinking-thought-experiment in your project. Codex loads it when a task matches its description.

Can I use Thinking Thought Experiment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tjboudreaux/cc-thinking-skills --skill thinking-thought-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/thinking-thought-experiment, .gemini/skills/thinking-thought-experiment, .github/skills/thinking-thought-experiment and .opencode/skills/thinking-thought-experiment in your project.

What does Thinking Thought Experiment need to run?

SKILL.md names no scripts, command-line tools or credentials: Thinking Thought Experiment is instructions for the agent only.

Does Thinking Thought Experiment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Thinking Thought Experiment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Thinking Thought Experiment use?

Thinking Thought Experiment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Thinking Thought Experiment use?

About 937 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Thinking Thought Experiment?

Skills that share tags, products or a category with Thinking Thought Experiment: Finding Experiments (PostHog/posthog, 40k stars), Experiments (Arize-ai/phoenix, 12k stars), Scroll Experience (sickn33/agentic-awesome-skills, 47k stars) and Webgl Experience (nexu-io/open-design, 100k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Thinking Thought Experiment?

tjboudreaux (a GitHub user) maintains it in tjboudreaux/cc-thinking-skills, which has 1,613 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on August 7, 2026.

Source: tjboudreaux/cc-thinking-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.