Agent skill

Deterministic Simulation

by nerdsane in nerdsane/temper

Deterministic simulation testing (DST), TigerBeetle/FoundationDB style, and DST-driven development.

Apache-2.0Auto-check passedTesting & QA

Install Deterministic Simulation

skills CLI
$ npx skills add nerdsane/temper --skill deterministic-simulation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nerdsane/temper deterministic-simulation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nerdsane/temper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/deterministic-simulation .claude/skills/deterministic-simulation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deterministic-simulation
GitHub stars
102
Token cost
~1.4k tokens
SKILL.md length
789 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Deterministic simulation testing (DST), TigerBeetle/FoundationDB style, and DST-driven development.

  • Works in 5 steps: Define the harness first. Before… → Implement. Production code, running… → Run seeds until the invariants hold. Not… → …
  • Testing a stateful system (Temper and anything like it - databases
  • SKILL.md covers What it is, What it is NOT - the mistakes…, Requirements on the code under… and DST-driven development, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Deterministic Simulation is an agent skill from nerdsane/temper. Deterministic simulation testing (DST), TigerBeetle/FoundationDB style, and DST-driven development. Use when building or testing a stateful system (Temper and anything like it - databases, queues, engines, protocol code), when a task mentions simulation, seeds, fault injection, or invariants, or before writing tests for concurrent/distributed logic. Not for frontend apps.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Chaos engineering. The repository describes itself as: A machine tool for agents: a verified runtime for systems agents build. The licence is Apache-2.0.

When your agent uses it

  • Testing a stateful system (Temper and anything like it - databases
  • A task mentions simulation
  • Fault injection
  • Before writing tests for concurrent/distributed logic

Example prompts

  • “/deterministic-simulation”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Define the harness first. Before implementing, extend the simulator with the scenario: the workload, the faults, and the invariants - the…
  2. Implement. Production code, running inside the simulation.
  3. Run seeds until the invariants hold. Not one seed - many. A green run on one seed is one execution, not correctness.
  4. A failing seed found later is committed as a regression case and stays in the suite forever.
  5. Fix by root cause. Never fix by weakening the invariant or narrowing the workload.

What it can do on your machine

Read from SKILL.md and the folder at commit 1be4b7b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deterministic Simulation loads about 1.4k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 789 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nerdsane/temper at commit 1be4b7b, republished under its Apache-2.0 licence (© nerdsane). 789 words, ~1,437 tokens.

Download SKILL.mdSave it as .claude/skills/deterministic-simulation/SKILL.md (or your agent's skills folder).
name
deterministic-simulation
description
Deterministic simulation testing (DST), TigerBeetle/FoundationDB style, and DST-driven development. Use when building or testing a stateful system (Temper and anything like it - databases, queues, engines, protocol code), when a task mentions simulation, seeds, fault injection, or invariants, or before writing tests for concurrent/distributed logic. Not for frontend apps.

Deterministic simulation testing

What it is

The whole system runs inside a simulator that owns every source of nondeterminism: time, randomness, scheduling, network, disk. One seed drives one execution. The same seed replays the exact same execution, byte for byte. The simulator injects faults (crashes, partitions, delayed and dropped messages, disk errors) while invariants are checked continuously. Different seeds explore different executions; a failing seed is a permanent, replayable reproduction of a bug.

Reference implementations and what each proved:

  • FoundationDB: a deterministic simulator running the whole cluster in one process; BUGGIFY markers in production code cooperatively inject faults with some probability when simulating; "swizzle-clogging" (clog a random subset of nodes' networks one by one, unclog in random order) finds the deep interleavings. Their bar: if production hits a bug the simulator could have expressed, that is a simulator gap to fix.
  • TigerBeetle (VOPR): an entire cluster of real code under network, storage, and process faults at ~1000x real-time (a virtual clock means simulated time runs as fast as the CPU allows); runs continuously across many cores and seeds; assumes the disk WILL fail - corruption and misdirected reads/writes are in the fault model, not just crashes.
  • Antithesis: DST as a service over unmodified systems.

In this repo: the simulator lives in temper-runtime (sim module) with temper-store-sim as the simulated store; DST suites are platform_e2e_dst and system_entity_dst in crates/temper-platform/tests/.

What it is NOT - the mistakes agents make

  • Not an integration test. An integration test runs the system against real dependencies on real time and passes or fails once. A simulation runs thousands of seeded executions against simulated dependencies with faults injected.
  • Not a mock-based unit test. Mocks replace the system's parts to isolate one piece. In DST the PRODUCTION CODE runs - all of it, unmodified. Only the environment (clock, network, disk, scheduler, entropy) is simulated.
  • Not a parallel reimplementation. You do not write a second version of the logic and compare. The one real implementation runs in the simulator. A simplified MODEL may exist as an oracle to check results against, but the thing under test is always the production code.
  • Not "tests that use a seed." If any nondeterminism leaks (a real clock read, an unseeded RNG, thread timing, iteration order of an unordered map), replay breaks and the whole method is void. Determinism is the load-bearing property.

Requirements on the code under test

  • All time via an injected clock. Never read the wall clock directly.
  • All randomness from one seeded source the simulator provides.
  • Single logical thread of execution, or scheduling fully controlled by the simulator.
  • All I/O (network, disk, external services) behind interfaces the simulator can implement.
  • No dependency on unordered iteration, real timers, or ambient environment.

If the code cannot meet these, that is an architecture finding to raise, not a reason to fall back to integration tests.

Show full SKILL.md (319 more words)Show less

DST-driven development

For systems like Temper this replaces test-driven development. The loop:

  1. Define the harness first. Before implementing, extend the simulator with the scenario: the workload, the faults, and the invariants - the things that must never happen (lost write, double apply, stuck state machine, divergent replicas). Run it. The invariant must fail now - a harness that cannot catch the missing behavior proves nothing.
  2. Implement. Production code, running inside the simulation.
  3. Run seeds until the invariants hold. Not one seed - many. A green run on one seed is one execution, not correctness.
  4. A failing seed found later is committed as a regression case and stays in the suite forever.
  5. Fix by root cause. Never fix by weakening the invariant or narrowing the workload.

Writing good simulations

  • Coverage lives in the workload and fault schedule, not the framework. It is easy to build a simulator that explores almost nothing. Vary operation mixes, timings, fault frequencies; check that interesting states are actually reached.
  • Invariants are properties, not examples: "no acknowledged write is ever lost", not "this call returns 3".
  • Keep seeds cheap. A virtual clock costs nothing to advance - simulated hours run in wall-clock seconds. Fast executions buy more seeds per CI run, and more seeds are more coverage.
  • Put cooperative fault points in production code (FoundationDB's BUGGIFY pattern): rare branches - a timeout firing early, a message reordered - taken with small probability only under simulation. The code helps the simulator find its own weaknesses.
  • Report failures as: seed, invariant violated, minimal event trace. The seed IS the bug report.

Limits - say them, do not hide them

DST cannot catch: bugs in the simulator's model of the environment, behavior of real external systems, real-clock/performance issues, and nondeterminism the harness failed to capture. Code changes invalidate old seeds' meaning (the seed replays a different execution). DST complements live verification; it does not replace the Definition of Done's live run.

© nerdsane, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/deterministic-simulation of nerdsane/temper.

Open the folder on GitHubat commit 1be4b7b

Compare with similar skills

Deterministic Simulation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deterministic Simulation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deterministic Simulation this skillnerdsane/temper102—~1.4kAutomated safety check: PassApache-2.0
Fake Model Provider Faultsdifferent-ai/openwork24k—~642Automated safety check: PassCustom licence
Testinguw-syfi/vibesys103—~1.5kAutomated safety check: PassMIT
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Audit Reviewtestflows/TestFlows-GitHub-Hetzner-Runners102—~2.1kAutomated safety check: PassCustom licence
Deploy Workflownwiizo/ccswarm153—~582Automated safety check: PassMIT

Similar skills

  • Fake Model Provider Faults

    different-ai/openwork

    Makes the desktop app's model provider fail on demand, with refused connections, resets, stalls and HTTP 4xx and 5xx errors, so error and retry states can be reproduced.

    24k GitHub stars~642 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Testing

    uw-syfi/vibesys

    Write, change, review, or fix tests in any language in this repository.

    103 GitHub stars~1.5k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Audit Review

    testflows/TestFlows-GitHub-Hetzner-Runners

    Perform deep feature audits with transition-matrix and logical fault-injection validation.

    102 GitHub stars~2.1k tokensUpdated 16 days ago
    DevOps & CloudAuto-check passed
  • Deploy Workflow

    nwiizo/ccswarm

    Release deployment process for ccswarm. An agent skill from nwiizo/ccswarm.

    153 GitHub stars~582 tokensUpdated 23 days ago
    Testing & QAAuto-check passed
  • E2E Deploy Rhdh

    redhat-developer/rhdh

    Deploy RHDH to an OpenShift cluster using local-run.sh for E2E test execution, with autonomous error recovery for deployment failures

    172 GitHub stars~2.6k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from nerdsane/temper

  • Verify Temper

    nerdsane/temper

    Launch the Temper kernel locally and verify a change end to end - build, serve, drive the OData surface, run the verification cascade and DST proof.

    102 GitHub stars~726 tokensUpdated 3 days ago
    Auto-check passed

Questions about Deterministic Simulation

What does Deterministic Simulation do?

Deterministic simulation testing (DST), TigerBeetle/FoundationDB style, and DST-driven development. Deterministic Simulation is an agent skill from nerdsane/temper. Deterministic simulation testing (DST), TigerBeetle/FoundationDB style, and DST-driven development.

When should I use Deterministic Simulation?

Deterministic Simulation fits situations like: testing a stateful system (Temper and anything like it - databases; A task mentions simulation; fault injection; before writing tests for concurrent/distributed logic.

How do I install Deterministic Simulation in Claude Code?

Run `npx skills add nerdsane/temper --skill deterministic-simulation -a claude-code`. Or copy the skill folder (.agents/skills/deterministic-simulation in nerdsane/temper) into .claude/skills/deterministic-simulation in your project. Claude Code loads it when a task matches its description.

How do I install Deterministic Simulation in Codex?

Run `npx skills add nerdsane/temper --skill deterministic-simulation -a codex`. Or copy the skill folder (.agents/skills/deterministic-simulation in nerdsane/temper) into .agents/skills/deterministic-simulation in your project. Codex loads it when a task matches its description.

Can I use Deterministic Simulation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nerdsane/temper --skill deterministic-simulation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deterministic-simulation, .gemini/skills/deterministic-simulation, .github/skills/deterministic-simulation and .opencode/skills/deterministic-simulation in your project.

What does Deterministic Simulation need to run?

SKILL.md names no scripts, command-line tools or credentials: Deterministic Simulation is instructions for the agent only.

Does Deterministic Simulation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Deterministic Simulation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deterministic Simulation use?

Deterministic Simulation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deterministic Simulation use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deterministic Simulation?

Skills that share tags, products or a category with Deterministic Simulation: Fake Model Provider Faults (different-ai/openwork, 24k stars), Testing (uw-syfi/vibesys, 103 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars) and Audit Review (testflows/TestFlows-GitHub-Hetzner-Runners, 102 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deterministic Simulation?

nerdsane (a GitHub user) maintains it in nerdsane/temper, which has 102 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 4, 2026.

Source: nerdsane/temper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.