---
name: author-rhai-tests
description: >
  Author and review LunCoSim behavioral, asset-backed, component, mission,
  visual, and requirements-verification tests. Use when a test observes a
  Twin, USD stage, Modelica participant, runtime asset, or authored policy;
  keep the test in the Twin's Rhai scenario and run it through production.
  Use Rust tests only for generic engine mechanisms that Rhai cannot observe.
---

# Author tests in the right layer

Use this skill whenever a test is intended to prove what a mission Twin does,
what an authored component looks like, whether a USD relationship is wired,
whether a Modelica participant responds, or whether a public runtime command
or query produces the required result.

## Ownership rule

The default decision is simple: if the test includes an authored asset or
runtime behavior, author it in Rhai beside the Twin asset and execute it with
the production `luncosim` scene-test/runtime surface. This includes USD,
SysML/KerML, Modelica, terrain, materials, component files, scene paths, and
visual evidence. A Rust implementation does not make an observable behavior a
Rust-owned test.

Rust tests are reserved for generic mechanisms that Rhai cannot observe
without already depending on the mechanism under test: pure math/lowering,
parser and schema contracts, serialization, generic path/identity resolution,
and lifecycle/resource seams. Keep those fixtures inline or temporary and
generic; do not embed a repository or Twin asset path in a Rust test. A Rust
resolver test may prove the generic TwinRoots contract, but it must not become
a second asset acceptance runner.

The source-of-truth split is:

| Fact or behavior | Authoritative source | Test owner |
| --- | --- | --- |
| Prim identity, topology, transforms, dimensions, materials, physics schemas | USD | Rhai scene gate |
| Requirement intent, traceability, scalar limits, verification names | SysML/KerML | Rhai verifier reading the mounted SysML report |
| Continuous equations and participant state | Modelica | Rhai observer through the public runtime surface |
| Sequence, stimulus, policy, verdict, visual inspection rubric | Rhai | Rhai |
| Generic parser, resolver, serializer, or engine invariant | Rust owner crate | Rust unit/integration test |

Do not copy a threshold, prim path, clock value, or component parameter into a
test when it is already authored in SysML or USD. Load the authoritative
source, then use the generic helpers in
`assets/scripting/tools/sysml_requirements.rhai` and the public query/command
surface to evaluate it.

## Authoring workflow

1. Split the behavior into the smallest independently verifiable component
   (bus, leg, wheel, ramp, panel, tank, joint, controller, or mission phase).
   Give the component its own authored requirement/verification mapping and
   Rhai observer when the contract is independently useful.
2. Inspect the owning USD/SysML/Modelica source and existing tool libraries
   before adding helpers. Put reusable mechanics in a namespaced Rhai library;
   keep the observer short and declarative. If a helper needs engine state that
   the public API cannot expose, add one generic Rust capability at its owner,
   then consume it from Rhai.
3. Make positive conformance evidence the default: prove that the required
   component, topology, relationship, datum, or runtime outcome is present and
   correct. Do not write a negative test merely to assert that an obsolete
   implementation name, old shape, or superseded path is absent; update the
   positive requirement and observe the required result instead.

   Add a negative case only when rejection or safe failure is itself a real
   contract, for example malformed source, non-finite values, missing
   safety-critical relationships, invalid units or clocks, stale-generation
   mutation, unsupported commands, or a required fail-safe response. Such a
   case must be bounded, non-destructive, and reach the public
   diagnostic/verdict boundary without crashing or silently substituting a
   default. A historical regression example alone is not sufficient reason
   to add a negative test.
4. For requirements, read the mounted SysML snapshot and evaluate composed USD
   evidence. Keep the requirement ID/limit in SysML and emit structured check
   evidence (name, measured value, units, criterion, source revision, and
   pass/fail) from Rhai.

For support-gated motion, distinguish authored geometry from runtime state:
`QueryPhysicsState.support_footprint_count` is only the number of declared
probes, while `support_contact_count` is the latest evaluated contact count
and `support_sample_tick` proves freshness. Treat a missing (`null`) contact
count as unavailable evidence and fail the check; never infer contact from a
non-empty footprint.
5. For visual requirements, define the camera, lighting/time contract,
   reference artifact, measurable geometry/placement rubric, and capture
   window. A screenshot is evidence only when the observer records the exact
   camera/source revision and verdict; visual inspection must not be reduced to
   an unbounded “looks good” assertion.
6. Run the narrowest production gate. Use `--validate` only as parse/preflight
   evidence; it is not a behavior or visual verdict. For a live Editor session,
   use `RunRhai`/`run_rhai_test.sh` or an attached `RunScenario` and preserve
   the current process and camera. Do not rebuild Rust for a Rhai-only change.
   Scene discovery is recursive. Give independent Editor fixtures nested
   one-scene directories so each windowed run mounts its own Twin and preview
   state. The production test wrappers give each process a run-scoped
   `LUNCOSIM_CONFIG`, isolating saved workspace/session state as well as
   ephemeral settings and runtime overlays.
7. Repeat with deterministic clocks and explicit seeds. Record the clock
   contract, timestep/substeps, source revision, and finite-state result. A
   repeatability check compares the same sampled evidence, not merely a zero
   exit code. For an owner-only hook, exercise valid context through its real
   production owner and rejection from an authored off-cycle invocation. Use
   `RunRhai`'s `Application/Repl/Evaluation` route for deliberate policy
   inspection; it does not stand in for the live owner context.
   World-bound `RunRhai` requests drain one per application update in FIFO
   order, and excess queue submissions or over-budget invocations return
   terminal errors. Keep public command behavior assertions in authored Rhai;
   test only the generic batch and FIFO seam in Rust.

For cross-run determinism, keep profile selection and state comparison in the
authored Rhai test. It reads the runner's typed parameters, selects a matching
profile from `scripts/tests/fixtures/deterministic-physics-reference.json`, and
requests only the expected row for each selected state. Rust decodes that JSON
at the `luncosim test --determinism-reference PATH` process boundary and serves
typed rows through the existing `query(...)` bridge. Rhai compares the six
selected physics checkpoints, the sparse Modelica checkpoints, articulated
checkpoints, and the explicit final stage with exact equality
(`numeric_tolerance=0`). It also deliberately alters one selected physics row
and verifies the comparison rejects it. It retains only the selected tick
numbers and the first mismatch message; it does not accumulate a state trace or
emit a result bundle.
`report_verdict` and the scene-test process exit code are the completion
contract. The Bash and PowerShell matrices in
`scripts/test-deterministic-physics-profiles.sh` and
`scripts/test-deterministic-physics-profiles.ps1` invoke the production test
for each profile and stop on a nonzero exit; they do not parse logs or compare
JSON. Rhai print/log lines remain diagnostics.

A fresh-scene startup test asserts that `on_start` observes tick 0 and the first
`on_tick` observes tick 1; startup readiness must hold the shared clock until
those callbacks can begin in order.
For USD Modelica networks, the startup gate also covers member-source resolution,
network synthesis, and generated port-surface publication; the binding epoch
must not classify authored connections while that interface is pending. Solver
compilation publishes the initialized time-zero state; do not advance Modelica
or physics clocks to prime a first exchange before opening the scenario gate.
The first live exchange uses the shared fixed tick and its normal causal barrier.
Keep the production scene-test's connection diagnostics in the pass condition.
Do not normalize startup ticks or reset elapsed time to make a trace begin at
zero. The production fixed runner completes a started causal cycle and retains
its remaining fixed time while an owner hold is active; verify startup stays at
tick 0 with no fixed elapsed time or overstep before readiness.
Never subtract a scenario start tick or translate traces to a relative tick
sequence. Match equivalent authored subjects by their USD-owned stable facts,
and omit ECS allocation and wall-clock timing from the comparison.

For observable multi-actor ordering, attach scenarios to distinct authored
hosts and assert the same-pass handoff in Rhai. Keep Rust coverage for the
generic identity key and reverse-completion commit mechanism; a Rust assertion
alone does not prove the production script path.

When a public query reports asynchronous analysis, a test may sample that
query from its test-only `on_tick` until the exact requested generation reaches
a terminal state. Assert `pending` as retryable and inspect diagnostics only
after `ready`; do not use elapsed wall time to decide which result is current.
If a running scenario needs those facts before it can initialize, declare the
owner and identity in `simulation_dependencies(...).required_inputs` and assert
from `on_start` that the committed source revision is available. The generic
scenario lifecycle test may verify Pending-to-Ready hold/release mechanics;
the authored scene gate must verify the domain key and source revision.

For transient USD curve views, use `InspectUsdCurveView` to wait until
`completed_revision == requested_revision`, fail immediately on `state ==
"failed"`, and require `applied_revision == requested_revision` for a successful
mesh. Assert `local_visibility` and result vertex count rather than reading the
unchanged USD curve seed. A route with fewer than two active points must finish
with hidden local visibility and zero generated vertices; adding the second
point must produce an applied mesh revision.

For ordered asynchronous scene checks, author the sequence with the existing
Rhai task tree (`seq`, `once`, `wait_until`, `check`, and `sel`) rather than a
numeric `phase` switch in `on_tick`. Use named `Fn("callback")` leaves when a
step needs persistent test state; the task driver binds that state's `this` to
the callback. Prefer `wait_for`/`wait_for_from` when the owner publishes the
completion event; use `wait_until` only when no suitable event exists. `wait`
uses deterministic simulation time, not wall time. For an interruptible
sequence, put a cheap state guard in `reactive_seq` and let a failed `check`
cancel its running child; event handlers can update that guard's state, which
the task kernel observes on its next deterministic pass. Keep test-only
`on_tick` for a bounded fixed-step watchdog that reports the exact condition
that timed out. Do not add a separate Rust timer callback or phase runner: task
progression already owns deterministic waits and callback cadence. Any future
task deadline must specify its clock and cancellation/failure result as part of
the task contract. The shared `auto_tests.rhai` prelude already owns assertions
and terminal verdicts; do not add a parallel test DSL unless the Rhai task
surface demonstrably cannot express a required contract.

## Production commands

Resolve the production binary once:

```bash
export LUNCOSIM_BIN="${LUNCOSIM_BIN:-luncosim}"
./scripts/run_scene_tests.sh --no-build --exact <scene-name> -j 4
```

Use `-j 1` when diagnosing ordering or nondeterminism. Keep the production
binary and API session explicit; never substitute an old sandbox executable,
`cargo run`, or a temporary Rust runner. For same-session iteration, register
or reload the Twin Rhai library, run a minimal namespaced call, then execute
the observer through the API as described by
[`test-via-api`](../test-via-api/SKILL.md).

## Review checklist

- Is this observable behavior or an authored asset? If yes, it is Rhai-owned.
- If a test loads, composes, edits, or inspects a USD/Modelica asset, author it
  as a Twin Rhai scenario. Keep Rust tests for pure, asset-free engine
  primitives and routing predicates; do not embed fixture documents or asset
  identifiers in core tests.
- Does the test load the real Twin/USD/Modelica source rather than recreate it?
- Are requirements and dimensions read from SysML/USD instead of duplicated?
- Is the component independently scoped and its evidence structured?
- If a negative case exists, is rejection or safe failure an explicit contract,
  and does it have a named, non-crashing diagnostic? Do not require a negative
  case for ordinary conformance or obsolete-implementation cleanup.
- Are units, coordinate frame, camera/time contract, and deterministic clocks
  explicit where they affect the result?
- Is canonical numeric state kept as native `f64`/USD `double`, with any
  `f32`/`float` conversion explicit and limited to a renderer/GPU boundary or a
  USD field whose schema requires it?
- Is the test short enough to reuse libraries rather than becoming a batch
  builder or a second runtime in Rhai?
- Did the run use the production binary/API and prove a real verdict?

Defer to [`sysml-requirements`](../sysml-requirements/SKILL.md) for SysML
source-set and traceability rules, [`interactive-component-authoring`](../interactive-component-authoring/SKILL.md)
for the one-component Editor loop, and [`validate-assets`](../validate-assets/SKILL.md)
for parse/lint preflight.
