---
name: hole-hunt
description: Run a multi-lens hole hunt (audit) of the dmrgpy Python layer, or add a lens, a finding or a fix cluster to an existing one. Use this whenever the user asks to audit, hole-hunt, hunt for bugs, sweep for holes, cross-check the backends against each other, or look for silently-wrong numbers, and also when they ask to record or fix a finding in docs/audit_*_hole_hunt.md. This is the repository's established audit process, run in 2026-08 and 2026-09, and it has a fixed record shape and a fixed evidence standard that a free-form bug hunt will not reproduce.
---

# Hole hunt

A hole hunt is a parallel, multi-lens search for behaviour in the dmrgpy Python
layer that is *silently* wrong: a number that is plausible and incorrect, a
dispatch that answers a question nobody asked, a kwarg with no consumer. The
previous hunts are `docs/audit_2026_08_hole_hunt.md` (five lenses, 21 findings),
`docs/audit_2026_09_hole_hunt.md` (eight lenses, 36 findings),
`docs/audit_2026_09_24_hole_hunt.md` (five lenses, 16 findings),
`docs/audit_2026_09_24b_hole_hunt.md` (five lenses, 18 findings) and
`docs/audit_2026_09_24c_hole_hunt.md` (four lenses, 18 findings), followed by
`docs/audit_2026_09_25_open_items.md`, a fix pass over ten of their open items
rather than a hunt, whose "Left open, and new leads" section belongs in the
next brief next to every record's "New leads", and by
`docs/audit_2026_09_25b_hole_hunt.md` (four lenses, 29 findings) over that fix
pass. Read the
scope section and the lens table of the most recent one before starting: a
finding already recorded there is not a new finding.

What makes this process worth following rather than improvising: every claim in
the record was *executed*, and every claim was then handed to a second agent
whose only brief was to refute it. That is what makes the record trustworthy
enough that a later fix does not have to re-derive the evidence. Predicted
output, remembered output and output from a stale `.so` all break that, so they
are the failure mode to guard against throughout.

## 1. Fix the frame

Record, and keep true for the whole hunt:

- The commit (`git rev-parse --short HEAD`) and that the tree is clean.
- Whether both compiled extensions are current. If a lens will touch C++, rebuild
  first, because a fix that lands mid-hunt and replaces `_dmrgcpp*.so` invalidates
  every other lens's measurements. The 2026-09 hunt took its one C++ fix by hand,
  separately, for exactly this reason.
- When the hunt is scoped to one commit, a compiled snapshot of its parent, so
  every candidate runs on both trees and the record can say whether the commit
  brought the defect in. `git archive` of the parent without the vendored
  `ITensor`/`TDVP` folders, those linked back to this checkout's copies (check
  they are unchanged between the two commits), then `make pybind` in each
  `mpscppN`: about a minute, and it never touches the repo's own `.so`. The
  recipe and the two-runner setup are in the 2026-09-25b record's "Shared
  helpers".
- The invocation every repro uses:

```bash
MKL_NUM_THREADS=1 OMP_NUM_THREADS=1 OPENBLAS_NUM_THREADS=1 NUMEXPR_NUM_THREADS=1 \
  PYTHONPATH=<this worktree>/src python3 <script>
```

Both halves matter. Unpinned threads make a timing claim meaningless, and a bare
import resolves to whichever checkout `site-packages` is symlinked into, so an
ad-hoc probe can silently test code that is not the code under audit.

## 2. Choose the lenses

Four to eight, each a one-line brief naming one class of problem, and as
file-disjoint as you can make them so the fix clusters afterwards can run in
parallel. Previous sets, to vary rather than repeat:

- 2026-08: silent backend and method dispatch fallbacks, dropped keyword
  arguments, feature-by-feature combinations, garbage-in-garbage-out cases that
  should raise, and cross-backend numerical disagreement on chains small enough
  for ED to be exact.
- 2026-09: `python-backend-parity`, `wavefunction-consumers`, `dispatch-matrix`,
  `pyitensor-performance`, `cpp-v3-completeness`, `ed-and-operators`,
  `recent-commits`, `docs-examples-drift`.
- 2026-09-24 (`docs/audit_2026_09_24_hole_hunt.md`, scoped to the commits since
  the previous hunt's fixes): `kpm-calibration`, `td-convention`,
  `canonical-form`, `infinite-chain`, `recent-misc`. Scoping a hunt to a commit
  window, one lens per recent change, found 16 holes in 11 commits, seven or
  eight of them introduced by the single commit that closed the previous
  hunt's open items; a hunt right after a fix pass is worth running for that
  reason alone.
- 2026-09-24b (`docs/audit_2026_09_24b_hole_hunt.md`, scoped to the one fix
  commit `30200a4`): one lens per fix cluster, `operators`, `kpm`, `realtime`,
  `kondo`, `pyitensor`; 18 findings, twelve of them older than the commit and
  reached by probing next to it.
- 2026-09-24c (`docs/audit_2026_09_24c_hole_hunt.md`, scoped to the one fix
  commit `867e2b4`): again one lens per fix cluster, `groundstate`, `kpm`,
  `realtime`, `misc`; 18 findings, six of them from the commit itself. Two things
  it taught. The brief listing what is already recorded has to carry every
  earlier record's "New leads", not only the last one's: its one refutation was
  a lead two records back that the brief left out. And a candidate turned up by
  a reviewer of a reviewer-found candidate needs its own reviewer too; a
  workflow that stops one level down leaves it as a lead.
- 2026-09-25b (`docs/audit_2026_09_25b_hole_hunt.md`, scoped to the fix-pass
  commit `e7b1196`): one lens per fix cluster, `construction`, `session`,
  `scale`, `scale_cpp`, run on `e7b1196` and on a compiled snapshot of its
  parent; 32 candidates reviewed to depth three, none refuted, 29 findings,
  three from the commit itself and four more it made reachable. What it
  taught: a fix that lowers one floor exposes the next one. Once
  `clean_threshold` stopped dropping small operators, seventeen absolute
  thresholds further down became reachable, and one of them sat in the
  previous record's "Ruled out" as unreachable for exactly the reason the fix
  removed. After a threshold or a guard changes, read the earlier "Ruled out"
  sections as candidates, not as settled.

Out of scope by construction, and stated in the record so the exclusion is on
the page rather than in someone's head: vendored ITensor (`mpscpp2/ITensor/`,
`mpscpp3/ITensor/`); the legacy bugs `CLAUDE.md` says are deliberately
reproduced (`evoloperator`'s z^3/6 term on `H2`, the `"moise"` key, the
unreachable `"tevol_fit_td"` branch); the open `docs/known_issue_*.md` items;
anything already in an earlier audit record; and gaps `ROADMAP.md` already marks as
absent. Decide explicitly whether `itensor_version="julia_live"` is in scope,
since its juliacall JIT cost dominates any lens that touches it, and say so.

## 3. Hunt, then refute

Spawn one `hole-hunter` agent per lens, all in one message so they run
concurrently. Each returns candidate findings, each carrying a repro script that
was actually run and its verbatim output.

Then spawn one `finding-reviewer` agent per candidate, briefed to refute it. A
reviewer that reproduces the repro and finds the behaviour intended, already
documented, or an artifact of the probe returns `REFUTED`, and that candidate
never enters the record. A reviewer that narrows the claim returns the narrowed
version, and the narrowed version is what gets written down. The 2026-09 record
carries several findings whose sub-claims the reviewer struck; keeping the strike
visible is part of the point.

Two practical points from the 2026-09-24 hunt. Subagents cannot write report
files (the harness refuses them), so a reviewer's report arrives only as its
hand-back text: save a condensed copy next to its scripts as it arrives, since
the conversation that holds the full text may be compacted before the record is
written. And the scratch folders every repro runs in do not outlive the session,
so the record must carry each script and output inline rather than by path;
the 2026-09-24 record was assembled by a small builder that splices the files
from disk into the markdown, which keeps the outputs byte-for-byte what was run.
A finding a reviewer turns up while reviewing another one gets its own reviewer
before it enters the record.

## 4. Write the record

`docs/audit_<YYYY_MM>_hole_hunt.md` (`<YYYY_MM_DD>` when the month already has
one), in the shape the existing records share:

````markdown
# Audit, <YYYY-MM>: <n>-lens hole hunt

<one paragraph: date, commit, tree state, that every repro was executed and
every finding handed to an independent reviewer briefed to refute it, and that
REFUTED findings are not reproduced here>

<one paragraph: this file is the evidence, not a task list; fixed entries keep
their repro and gain a **Status** line rather than being deleted>

## The <n> lenses

| Lens | Brief |
|---|---|

## Scope

<what is excluded by construction, and the pinned-threads invocation above>

## Findings

### 1. <the defect stated as a claim, with its measured size, in one sentence>

`bug` &middot; severity **HIGH** &middot; CONFIRMED &middot; lens `<lens-name>`

**Where**: `<file:line list, every site the defect reaches>`

<prose: what the code does, what the layer above believes it does, why every
existing test passes through it>

**Expected**: <what a correct implementation returns>

Repro:

```bash
<the exact pinned-threads invocation that was run>
```

Observed:

```
<verbatim output, not retyped>
```

**Reviewer (CONFIRMED)**: <the attempt to refute it, and what survived>

**Suggested fix**: <one paragraph>
````

A `**Status**` line goes directly under the classification line once the finding
has been acted on, and `**Reviewer on severity**` is the variant used where the
reviewer accepted the defect and disputed how bad it is.

Two things about the finding heading, because they are what makes the record
readable a year later: state the defect as a claim rather than a topic
("`vev(op, npow=n)` silently ignores `npow` on every ED route" beats "npow
handling"), and put the measured size in it where there is one.

## 5. Fix in file-disjoint clusters

Group the findings into clusters that do not touch the same files, and take one
cluster at a time or in parallel agents. Each cluster gets one regression file,
`tests/test_audit_<YYYY_MM>_<cluster>.py`, and each fix gets a `**Status**` line
appended to its finding: `FIXED`, `PARTIAL` (say which half landed), or the
reasoning if the behaviour turns out to be intended after all. Name the tests
that pin it, or say explicitly that none does.

Pin the property, not a golden number, wherever the property is what was wrong:
the 2026-09 TDVP fix is pinned by a test that asserts the *order* in `dt` rather
than a value. Where a fix removes a bug from an existing computation, keeping
the pre-fix reference construction verbatim inside the test is the cheapest way
to prove the new path agrees with the old one where it should.

The 2026-09-25b fix pass ran nine clusters as parallel agents, each in its own
git worktree, and four things about that shape are worth keeping:
- **Take the baseline on the machine the fix pass runs on**, with freshly built
  extensions, before any agent starts. That pass ran on a different machine
  from its hunt, and the baseline turned up a thirtieth defect: a roundoff floor
  that MKL had rounded to exact zero and OpenBLAS did not (finding 30).
- **A worktree has no compiled extension.** The `.so` files are gitignored, so
  v2/v3 silently fall back to ED and every test passes vacuously. Each agent
  copies (not symlinks) both `.so` in and asserts `cppext.available` first.
  The cluster that edits C++ builds its own by copying the untracked
  `ITensor/this_dir.mk` and `options.mk` into its worktree, which points the
  build at the main checkout's `libitensor.a`.
- **Agents amend their commits.** Merge into a separate integration worktree,
  never into the main checkout the other agents still use as their pristine
  reference. Rebuild the integration branch from the final branch heads with a
  script rather than re-merging on top, and resolve the audit record's
  Status-line conflicts by keeping both sides.
- **Decide shared formulas in the brief.** Where a fix has a Python and a C++
  half owned by two clusters (the NH Ritz window) or must happen in exactly one
  place (NH unit scaling at the Python entry, not also inside
  `Chain::nhdmrg`), write the formula and the place into both briefs.

## 6. Say when numbers change

A fix that makes a previously-returned number different is a different kind of
event from a fix that makes a crash stop, because saved results elsewhere are
now not comparable. Every such fix gets `NUMBERS CHANGE` in its `**Status**`
line, naming the old value, the new one and the exact chain it was measured on,
and the consolidated list goes into `CLAUDE.md`'s paragraph for that audit.
Other projects save results produced by this library, so this is not
bookkeeping.

## 7. Close the loop

Update `CLAUDE.md` with a paragraph for the hunt: how many lenses, how many
findings, where the regressions live, which fixes changed numbers, and which
items are open rather than fixed. An open item stays in the record with what is
known about it, the way the `kpm_energy_truncate` window problem and the
`submode="TD"`/`"TDZ"` convention items did, rather than being dropped because
it did not get fixed.
