Agent skill

Elodin Monte Carlo

by elodin-sys in elodin-sys/elodin

Develop and calibrate simulations against experimental truth data using elodin monte-carlo.

Apache-2.0Auto-check passed

Install Elodin Monte Carlo

skills CLI
$ npx skills add elodin-sys/elodin --skill elodin-monte-carlo -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install elodin-sys/elodin elodin-monte-carlo --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/elodin-sys/elodin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/elodin-monte-carlo .claude/skills/elodin-monte-carlo && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
elodin-monte-carlo
GitHub stars
547
Token cost
~2.9k tokens
SKILL.md length
1,349 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Develop and calibrate simulations against experimental truth data using elodin monte-carlo.

  • Works in 8 steps: Vendor truth data with provenance → Build a dependency-free reference module → Reconstruct missing channels from physics → …
  • Vendoring real telemetry as a reference profile
  • SKILL.md covers The development loop, 1. Vendor truth data with…, 2. Build a dependency-free… and 3. Reconstruct missing…, plus 6 more sections
  • Calls ruff, cargo and python

What it does

Elodin Monte Carlo is an agent skill from elodin-sys/elodin. Develop and calibrate simulations against experimental truth data using elodin monte-carlo. Use when vendoring real telemetry as a reference profile, adding a truth-replay ghost entity, writing campaign specs/hooks and run scoring, reconstructing missing data channels from physics, or iterating on guidance and physics models with campaign feedback.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Elodin simulation and flight software monorepo. The licence is Apache-2.0.

When your agent uses it

  • Vendoring real telemetry as a reference profile
  • Adding a truth-replay ghost entity
  • Writing campaign specs/hooks and run scoring
  • Reconstructing missing data channels from physics

Example prompts

  • “/elodin-monte-carlo”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Vendor truth data with provenance
  2. Build a dependency-free reference module
  3. Reconstruct missing channels from physics
  4. Truth ghost: replay reality inside the sim
  5. Campaign as the test harness
  6. Diagnose with the data, not by staring at code
  7. Tracking recorded profiles: control lessons
  8. Operational pitfalls

What it can do on your machine

Read from SKILL.md and the folder at commit 729022c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ruff
    • cargo
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Elodin Monte Carlo loads about 2.9k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 1,349 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from elodin-sys/elodin at commit 729022c, republished under its Apache-2.0 licence (© elodin-sys). 1,349 words, ~2,881 tokens.

Download SKILL.mdSave it as .claude/skills/elodin-monte-carlo/SKILL.md (or your agent's skills folder).
name
elodin-monte-carlo
description
Develop and calibrate simulations against experimental truth data using elodin monte-carlo. Use when vendoring real telemetry as a reference profile, adding a truth-replay ghost entity, writing campaign specs/hooks and run scoring, reconstructing missing data channels from physics, or iterating on guidance and physics models with campaign feedback.

Truth-Data-Driven Simulation with Elodin Monte Carlo

The most effective way to build a credible simulation is to anchor it to experimental truth data and use elodin monte-carlo as the test harness: every model change is judged by a 30-run campaign against recorded reality, in under a minute. This skill codifies that workflow. The canonical worked example is examples/apollo-lander/ (with the full methodology in its WHITEPAPER.md).

The development loop

text
vendor truth data -> build reference module -> sim + truth ghost + graphs
      ^                                                    |
      |                                                    v
 narrow spec.toml <- read report <- run campaign <- score runs vs truth
                         |
                         v
              export worst run CSV -> diagnose time series -> fix model

Run the campaign after every meaningful change. Distribution deltas (RMSE, success rate, margins, dispersion) tell you immediately whether a change helped, broke something, or just moved noise.

1. Vendor truth data with provenance

  • Keep the verbatim source file untouched in data/ next to a derived, SI-unit version the sim actually loads. Cite the source URL in README/docs.
  • Write a sanity_check() in the reference module that re-derives the cleaned data from the raw file and asserts agreement, plus checks documented anchor values. Make python reference.py print the profile and the check.
  • Distrust every column. Real-world traps hit in practice: a "meters" column converted at x0.3405 instead of x0.3048 (recompute conversions from the original-unit column); digitized charts with out-of-order and duplicate timestamps (sort, average duplicates); head samples from a different measurement frame (a state-vector update step) that must be trimmed and back-extrapolated; a sensor channel that is not the state you want (radar slant range is not altitude).
  • Cross-validate datasets against each other and against documented events: mission reports, transcripts, and timelines give anchor points (event times, velocity callouts, fuel consumed) that catch unit and frame errors no amount of smoothing will.

2. Build a dependency-free reference module

Put all truth handling in one stdlib-only Python module (see examples/apollo-lander/reference.py) shared by the sim and any external controller:

  • Resample onto a uniform grid, despike with a median filter, smooth with a moving average. Only enforce monotonicity if the physics demands it — real profiles wiggle, and the wiggles are history worth keeping.
  • Expose interpolators (altitude(t), descent_rate(t), ...) plus the raw arrays. In sim.py, convert once to jnp.asarray and close over them in systems: they become JIT-time constants (large baked constants are interned and shared across campaign workers).

3. Reconstruct missing channels from physics

Truth datasets are rarely complete. When a needed channel is missing (e.g. horizontal velocity existed only as a chart image), reconstruct it:

  • Integrate the vehicle dynamics along the channels you do have, using documented schedules (throttle history, event times) for the rest.
  • Allocate consistently: if total thrust is known and the vertical share is implied by the recorded altitude, the horizontal share is the remainder — this makes the reconstructed profile flyable by construction, which matters when a controller later tracks it (see section 7).
  • Calibrate segment-by-segment through documented anchors so the profile passes through known values exactly; integrate the best-documented segment unscaled and let it yield the unknown initial condition.
  • Iterate fixed-point if channels couple (two passes usually converge).
  • Document the uncertainty (~±10% is typical) and assert anchors in sanity_check().

4. Truth ghost: replay reality inside the sim

Render the recorded vehicle next to the simulated one. The robust pattern:

python
# Kinematic ghost: NO el.Body — gravity/integrator/telemetry systems never match it.
world.spawn([
    StaticSceneObject(el.WorldPos(...)),           # world_pos only
    el.C(TruthMarker, jnp.array([1.0])),           # scopes the playback query
    el.C(Altitude, ...), el.C(VerticalSpeed, ...), # graph channels
], name="lander_truth")

@el.system
def truth_playback(
    tick: el.Query[el.SimulationTick], truth: el.Query[TruthMarker]
) -> el.Query[el.WorldPos, Altitude, VerticalSpeed]:
    t_s = tick[0] * SIM_TIME_STEP
    alt = jnp.interp(t_s, ref_time, ref_altitude)
    ...
    return truth.map((el.WorldPos, Altitude, VerticalSpeed), lambda _: (...))

Hard-won rules:

  • Never give the ghost an el.Body. If physics systems match it, gravity integrates its velocity while replay snaps its position — sawtooth motion, runaway velocities in the graphs, and corrupted exports.
  • Do not drive the ghost from pre_step DB writes: explicit writes land on their own clock and interleave with telemetry commits, producing CSV exports with holes. The in-sim el.SimulationTick playback keeps truth on exactly the simulated vehicle's commit clock, so exports line up row-for-row.
  • A marker component is the cheapest way to scope a playback query to one entity.

Put sim-vs-truth pairs on every KDL graph, and include raw measurement curves (e.g. slant range next to true altitude) — visible divergence between a sensor and the state teaches more than hiding it.

5. Campaign as the test harness

  • el.monte_carlo.params_spec(...) defaults drive single runs; spec.toml drives campaigns. Keep ranges mirrored, and anchor ranges in data (a fuel chart pins the propellant load; navigation accuracy pins IC dispersions). When control authority is saturated (e.g. a full-throttle braking burn), keep IC ranges tight — the real system could not recover big errors either, and neither can yours.
  • Score every run against truth in post_step and emit el.monte_carlo.result(traj_rmse=..., pitch_rmse=..., miss_distance=..., soft_landing=...). Fit metrics (RMSE vs truth) are what turn a campaign from a stress test into a calibration engine: report the best-fit run's params, narrow spec.toml around them, repeat (or automate the loop, see examples/apollo-lander/calibrate.py).
  • Latch event metrics in-sim, at the event. A touchdown-speed metric read one tick after contact reads the zeroed post-contact state; capture it in the same system that detects the event, before state is clobbered.
  • Keep the LHS seed fixed while iterating so campaign-to-campaign deltas reflect your changes, not resampling. If you edit spec.toml, re-run elodin monte-carlo sample (the runner warns when a --plan CSV is older than its sibling spec).
  • Control concurrency with --workers N / workers = N in campaign.toml (exactly N runs at once). S10_MAX_INFLIGHT is only the low-level process-count escape hatch.
  • Use [[build]] steps in campaign.toml to compile external FSW (and generate configs) once before workers start; [env] sets campaign-wide process environment. Put per-worker resources under [resources.ports] — numeric bases give a validated static plan, "auto" allocates per run — and read them with el.monte_carlo.port("name", default) or ELODIN_MC_PORT_<NAME>.
  • At scale: scratch_dir = "auto" moves per-run DB IO to tmpfs when the artifact disk can't sustain workers x write IOPS; [retention] prune_on_pass/prune_on_fail globs bound disk without hook-side cleanup; export_db(..., pattern="rocket.*") exports only scored components; and a [quality] max_behind_deadline_frac = 0.05 gate marks load-degraded real-time runs degraded instead of silently ingesting skewed samples. Failed runs carry a one-line failure_reason in results.csv.
Show full SKILL.md (426 more words)Show less

6. Diagnose with the data, not by staring at code

  • Export any run and interrogate it with quick Python:

    bash
    elodin-db export dbs/<campaign>/runs/run_0000012/db \
        --format csv --join --flatten -o /tmp/run12

    Print a time-series table of the suspect channels (state, reference, command, actuator) every N seconds — most control/physics bugs are obvious in 20 rows.

  • Rank runs by the failing metric (post_run_result.json per run) and export the worst one, not the best.

  • Treat "too perfect" as a bug. Metrics that are identically zero or pegged to a clamp value usually mean the measurement is wrong (read after state-clobber, trivially satisfied criterion), not that the system is great.

  • Verify exports are clean: row counts equal across entities, no empty cells, no physically impossible values. Data quality bugs masquerade as physics.

  • Use the editor for visual regressions (attitude behavior, ghost overlap, terrain seating); confirm emergent event times against the historical record.

7. Tracking recorded profiles: control lessons

When a controller tracks the truth profile, the campaign will find every weakness. Patterns that survived:

  • Feed-forward the reconstructed profile; bound the feedback. Clamp each feedback channel's authority (e.g. ±0.8 m/s²) around the flyable feed-forward. Unbounded feedback lets channels fight over a saturated actuator and limit-cycle.
  • If the feed-forward is not flyable (demands more than the actuator), the tracker oscillates no matter the gains — fix the reference (section 3), not the gains.
  • Prefer one unified tracking law with state-dependent limits over discrete phase switches; let phases emerge (a speed-dependent tilt budget reproduces a pitchover; a demand-dependent throttle latch reproduces throttle-down).
  • Add terminal logic: near the end, stop chasing the time-indexed reference and null residual errors — event timing shifts with small tracking errors, and arriving early must not mean arriving with the reference's residual velocity.
  • Capped, slowly-fading position trim erodes IC offsets without destabilizing the velocity loop.

8. Operational pitfalls

  • Stale processes hold UDP ports. Prefer direct world.recipe() sidecars plus s10 readiness probes. Linux campaigns use cgroup teardown to reap reparented daemons; --keep-existing opts out of scoped pre-reaping.
  • Return valid = False from post_run hooks for infrastructure failures so they are reported separately from scored misses.
  • CI hygiene per repo rules: ruff format && ruff check --fix for the Python, cargo fmt/cargo clippy -- -Dwarnings for external controllers.

Reference example

examples/apollo-lander/ exercises every pattern above: vendored NASA telemetry with unit-bug handling and anchor checks, a physics-based horizontal-profile reconstruction, a kinematic truth ghost on el.SimulationTick, an external Rust LGC tracking the profile through a SITL UDP bridge, campaign hooks producing fit metrics and landing-ellipse statistics, and a calibration loop. Its WHITEPAPER.md documents the full methodology and the honest caveats.

© elodin-sys, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/elodin-monte-carlo of elodin-sys/elodin.

Open the folder on GitHubat commit 729022c

Compare with similar skills

Elodin Monte Carlo next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Elodin Monte Carlo compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Elodin Monte Carlo this skillelodin-sys/elodin547—~2.9kAutomated safety check: PassApache-2.0
Monte Carlo Remediationsickn33/agentic-awesome-skills47k1 repos~4kAutomated safety check: PassApache-2.0
Monte Carlo Preventsickn33/agentic-awesome-skills47k1 repos~3.3kAutomated safety check: PassMIT
Monte Carlo Context Detectionsickn33/agentic-awesome-skills47k1 repos~2.6kAutomated safety check: WarnMIT
Monte Carlo Push Ingestionsickn33/agentic-awesome-skills47k1 repos~4.6kAutomated safety check: PassMIT
Monte Carlo Asset Healthsickn33/agentic-awesome-skills47k1 repos~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Monte Carlo Remediation

    sickn33/agentic-awesome-skills

    Investigate and remediate data quality alerts using Monte Carlo MCP tools.

    47k GitHub starsUsed in 1 repo~4k tokens
    DevelopmentAuto-check passed
  • Monte Carlo Prevent

    sickn33/agentic-awesome-skills

    Surfaces Monte Carlo data observability context (table health, alerts, lineage, blast radius) before SQL/dbt edits.

    47k GitHub starsUsed in 1 repo~3.3k tokens
    Data & AnalyticsAuto-check passed
  • Monte Carlo Context Detection

    sickn33/agentic-awesome-skills

    Route data-related requests to the right Monte Carlo skill or workflow.

    47k GitHub starsUsed in 1 repo~2.6k tokens
    Data & AnalyticsAuto-check: warnings
  • Monte Carlo Push Ingestion

    sickn33/agentic-awesome-skills

    Expert guide for pushing metadata, lineage, and query logs to Monte Carlo from any data warehouse.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    DatabasesAuto-check passed
  • Monte Carlo Asset Health

    sickn33/agentic-awesome-skills

    Curated upstream guidance for Monte Carlo Asset Health; use when the workflow matches the user goal.

    47k GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Monte Carlo Monitor Creation

    sickn33/agentic-awesome-skills

    Guides creation of Monte Carlo monitors via MCP tools, producing monitors-as-code YAML for CI/CD deployment.

    47k GitHub starsUsed in 1 repo~2.9k tokens
    DevOps & CloudAuto-check passed

More from elodin-sys/elodin

All 14 skills in this repo
  • Branch Regression

    elodin-sys/elodin

    Compare two git branches (usually the current branch vs main) by running every example on each, capturing exit codes, logs, and editor screenshots, then diffing the results.

    547 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Elodin Cranelift

    elodin-sys/elodin

    Work with the Cranelift JIT MLIR backend. An agent skill from elodin-sys/elodin.

    547 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Elodin DB

    elodin-sys/elodin

    Work with Elodin-DB, the time-series telemetry database. An agent skill from elodin-sys/elodin.

    547 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Elodin Dev

    elodin-sys/elodin

    Develop and contribute to the Elodin codebase. An agent skill from elodin-sys/elodin.

    547 GitHub stars~896 tokensUpdated yesterday
    Auto-check passed
  • Elodin Editor Dev

    elodin-sys/elodin

    Contribute to the Elodin Editor, the 3D viewer and graphing tool.

    547 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Elodin Headless Capture

    elodin-sys/elodin

    Run the Elodin Editor without a physical display in Gamescope, take screenshots, and record video through PipeWire and GStreamer.

    547 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes

Questions about Elodin Monte Carlo

What does Elodin Monte Carlo do?

Develop and calibrate simulations against experimental truth data using elodin monte-carlo. Elodin Monte Carlo is an agent skill from elodin-sys/elodin. Develop and calibrate simulations against experimental truth data using elodin monte-carlo.

When should I use Elodin Monte Carlo?

Elodin Monte Carlo fits situations like: vendoring real telemetry as a reference profile; adding a truth-replay ghost entity; writing campaign specs/hooks and run scoring; reconstructing missing data channels from physics.

How do I install Elodin Monte Carlo in Claude Code?

Run `npx skills add elodin-sys/elodin --skill elodin-monte-carlo -a claude-code`. Or copy the skill folder (.cursor/skills/elodin-monte-carlo in elodin-sys/elodin) into .claude/skills/elodin-monte-carlo in your project. Claude Code loads it when a task matches its description.

How do I install Elodin Monte Carlo in Codex?

Run `npx skills add elodin-sys/elodin --skill elodin-monte-carlo -a codex`. Or copy the skill folder (.cursor/skills/elodin-monte-carlo in elodin-sys/elodin) into .agents/skills/elodin-monte-carlo in your project. Codex loads it when a task matches its description.

Can I use Elodin Monte Carlo in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add elodin-sys/elodin --skill elodin-monte-carlo -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/elodin-monte-carlo, .gemini/skills/elodin-monte-carlo, .github/skills/elodin-monte-carlo and .opencode/skills/elodin-monte-carlo in your project.

What does Elodin Monte Carlo need to run?

Going by SKILL.md and its folder, Elodin Monte Carlo needs the command-line tools its instructions call (ruff, cargo and python). Our summary lists: Python 3.

Does Elodin Monte Carlo access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Elodin Monte Carlo safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Elodin Monte Carlo use?

Elodin Monte Carlo is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Elodin Monte Carlo use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Elodin Monte Carlo?

Skills that share tags, products or a category with Elodin Monte Carlo: Monte Carlo Remediation (sickn33/agentic-awesome-skills, 47k stars), Monte Carlo Prevent (sickn33/agentic-awesome-skills, 47k stars), Monte Carlo Context Detection (sickn33/agentic-awesome-skills, 47k stars) and Monte Carlo Push Ingestion (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Elodin Monte Carlo?

elodin-sys (a GitHub organization) maintains it in elodin-sys/elodin, which has 547 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 9, 2026.

Source: elodin-sys/elodin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.