---
name: scheduler
description: Verter scheduler — Scheduler, submit_request/submit_batch/submit_batch_atomic (atomic DAG admission via driver-drained NewRequestBatch + shared admission core + deferred DedupJoinerEvent), wait_batch (input-order), live TaskKind (Source/Analysis/Artifact), owner-affine CPU vs I/O pool routing (OwnerCommand<Cpu/Io>, bounded SchedulerCpuPool::try_submit via try_acquire_owned), host-injected SchedulerCpuPool/SchedulerIoPool + separate host-owned HostCpuPool coordinator (shared by every host batch API), account_batch_submission, Admission<T>; plus the landed-but-unwired leaf substrate (CancellationToken, opaque SchedulerCacheId newtype, caller-side DedupeHook trait) and the not-yet-implemented cache-runtime DAG design target (KeyedJob/CacheNodeDagNode/submit_dag/DAG submission)
---

# Scheduler

Concise reference for the `verter_scheduler` crate.

**Live** surface (current tree): the `submit_request` / `submit_batch` /
`submit_batch_atomic` submission API, the live `TaskKind` variant set
(`Source` / `Analysis` / `Artifact`), owner-affine CPU vs I/O pool routing
(`OwnerCommand<Cpu>` / `OwnerCommand<Io>`, bounded
`SchedulerCpuPool::try_submit` via `try_acquire_owned`), and dual
pool isolation — the host-injected scheduler stage `cpu_pool` (+ `io_pool`)
plus the separate host-owned `HostCpuPool` coordinator shared by every host
batch API, with per-batch `account_batch_submission` accounting. *Dual pool
isolation* is the authority for the live pool model.

`submit_batch` is non-atomic (N separate `Submission::NewRequest` items;
the pump may observe the batch half-admitted, `submit_count` bumped per
item). `submit_batch_atomic` lands ONE
`Submission::NewRequestBatch { requests: Vec<QueuedRequest> }` that the
driver drains as a unit and admits under a SINGLE `dag.lock()`
acquisition via `handle_new_request_batch` (generation bumps + supersede
sweeps + waiter registration for every request inside one critical
section): the pump can never observe a half-admitted batch, one batch is
ONE wake + ONE `submit_count` bump, and a source-updating batch
supersedes every file's old generation atomically. Both paths share one
admission core — `prepare_request` (pre-lock: queued-lifetime gate + live-node
lookup, dropping the `nodes` DashMap guard BEFORE locking, the AB-BA-safe
DAG-first ordering; it never creates or re-homes
a node, carrying the resolved language forward as
`PreparedRequest.requested_language` instead), `admit_prepared_under_lock`
(sole place a request bumps generation, runs the supersede sweep,
registers the waiter, admits work — including the LANGUAGE RE-HOME, which
advances a published file's generation and therefore must be atomic with
its sweep; it resolves the live `FileNode` and revalidates the queued
lifetime first, since removal/reset may have invalidated it during
preparation), and an `AdmissionPostWork`
accumulator firing deferred dedup callbacks + clearing auto-ingest
tracking AFTER the lock releases. `SchedulerDag::register_request`
returns `Option<DedupJoinerEvent>` (fired post-unlock via
`DedupJoinerEvent::fire`) instead of invoking `on_dedup_joiner` under the
DAG lock — the callback may re-enter the scheduler, so it must not run
while admission holds the mutex.

Source completion advances to Analysis only when Analysis is demanded.
Demand is the union of (1) direct Analysis or Artifact request groups,
(2) admitted DAG consumers gated on that file-generation's Analysis
identity, and (3) pre-admission Artifact blocker-registry entries carrying
that identity. A Source-only request is signalled after dependency facts are
integrated and then stops; a later Analysis/Artifact request admits the
missing stage normally. Checking only direct request waiters is incorrect:
an auto-ingested macro dependency may be required by an Artifact blocker
before it owns a direct waiter or an admitted downstream node.

Source snapshot publication and Source-stage readiness are distinct. The
worker may publish generation-coherent snapshot bytes before the driver has
integrated extracted dependencies and late blockers. `FileNode` therefore
carries an explicit ready flag plus integrated-generation fence (generation
zero is valid): request admission and the Source already-satisfied short
circuit consult `current_integrated_source`, while completion processing reads
the raw current snapshot and advances the fence only after dependency
integration under the scheduler state lock. A
request arriving in that window joins the still-live Source identity; it must
not admit Analysis early or receive an early Source result.

Late Analysis demand is handled by one lock-held
`ensure_analysis_for_demand` transition. Direct Analysis admission and both
blocker-registration paths use it. A dependency that legitimately completed a
Source-only request can therefore acquire Analysis later; the absence of a
live Source identity is not mistaken for a dead producer. Terminal producer
failure remains typed for blocker demand, while a direct request retains the
same-generation recovery path.

The blocker registry maintains a `DepKey`-keyed reverse demand refcount at
every record/replace/drain/clear/scrub/owner-removal/retirement/reset funnel.
`has_analysis_demand` probes that index directly instead of scanning all owner
blocker sets under the global scheduler state lock.

**The test for "by construction" is an enumeration, not an intuition.**
The one structural claim in this area that held — `submit` being the sole
admission primitive — held only because every writer of `by_identity` was
enumerated repository-wide and exactly one insert was found. Two sibling
claims made from intuition rather than enumeration ("no public API
widening", "the captured node is unusable by construction") were both
FALSE: the first because `pub mod dag` re-exports the module, the second
because a leading underscore suppresses a lint and is not access control.
Before writing "by construction", enumerate.

**File admission requires a live object witness.**

`WorkNodeIdentity::{FileStage, Artifact}` and the corresponding `DepKey`
variants include the process-unique `FileNode` incarnation as well as generation.
`SchedulerDag::submit_file` accepts the actual node and checks canonical,
incarnation, generation and its object-lifetime retirement marker under the
lifecycle hold. Its unchecked insertion primitive is private; raw identity
submission is test-only. Cache-node admission has a separate entry point.
Neither queued work nor a delayed worker can reconstruct authority from a
replacement node at the same generation.

Queued requests bind to a submission lifetime during submission under `dag.lock()`.
That lifetime starts with the initial object incarnation and survives language
re-home, so later queued source updates still reach the replacement. Removal/reset
ends it; a new node receives a fresh lifetime. Work identities still carry the
unique object incarnation, and retired objects cannot admit or publish work.
`prepare_request` carries only the lifetime scalar and never creates a replacement
for an obsolete inbox item. Fresh requests may load any readable SourceLoader
backing after removal; scheduler removal does not permanently suppress paths.
A missing backing file produces a terminal dependency failure.

`retire_generations_below` performs the backward sweep of file waiters, admitted
nodes, dependency waiters, blocker records and failure records. It stores no DAG
retirement floor. Admission's live-object check rejects later stale work,
including gates on stages that were never admitted.

Dependency edges have one representation, `dag::dep_edges::DepEdges`. A node's
`deps_remaining` maps each `DepKey` to the edge's sequence number; the store maps
the `DepKey` to its waiters keyed by that sequence (edge-admission order) and
indexes every gated file dependency by `(canonical, generation)` across
incarnations. Link and unlink touch both sides in one step, so `complete`/`cancel`
remove a node's edges by key (no sibling rescans) and retirement reaches waiters
of never-admitted producers through the file index alone (no global key scan).
Removal drives off the per-canonical node, blocker-owner, blocker-reference and
terminal-failure indices; no removal path scans a DAG-wide map.
`SchedulerDag::dependency_occupancy` (and `Scheduler::dependency_occupancy`)
reports the edge store, its file index and the blocker reference index as
membership counts plus backing capacity in every build — including the summed
capacity of the nested per-generation and per-canonical hash tables, which keep
their backing while any member survives; reset releases those tables' backing
storage. `TaskRegistry::wait_graph_occupancy` reports the wait-for graph the
same way, including the per-producer waiter sets.

`remove` marks the object retired, signals Shutdown, cancels work, scrubs records,
unpublishes the node and publishes its source-root Absent state under one DAG
hold. Stranded-waiter wakes run after unlock. `reset` uses the same ordering for
all members and DAG clear, excluding cooperative admission as well as the
stopped driver. Invalidation and close revalidate the sampled object after
acquiring the lock.

Scheduler tombstones and DAG retirement floors have no storage. External
artifact publication is fenced by a scheduler-minted `SourceWitness` (canonical,
node incarnation, generation), handed out only with a source snapshot by
`try_get_witnessed_source`. `commit_artifact(witness, profile, data)` publishes
only into the live, unretired node object the witness was captured from at the
witnessed generation; `remove_artifact_not_newer_than(witness, profile)` evicts
only from that object and only artifacts no newer than the witnessed
generation. `try_get_source_for_witness` gives the host's block-content
publication fence the same answer. A witness from a removed, reset or re-homed
node is rejected even when its successor serves the same content at the same
generation. The scheduler keeps no generation floors: a removed or reset file's
successor starts at generation 0. Committed `SourceSnapshot` and
`AnalysisSnapshot` carry the committing node's `incarnation`, stamped by the
driver, and `version()` returns a `SourceVersion` (incarnation, generation)
ordered incarnation-first, so a successor's versions order after its
predecessor's and never compare equal to them. `StageExecutor::execute_source`
receives the incarnation; the host's revision token (`HostSourceRevisionToken`
carries the committing host's own `SourceVersion` on every ingress, ingested
envelopes included — never the registering owner's identity), its raw-template
version rail and its upsert commit fence key on it. Internal admission,
dispatch, publication, completion and failure use full incarnation identity.

**Preparation must still name a live submission lifetime at admission.**
Preparation drops the node map guard before taking `dag.lock()`. The admission
core revalidates `PreparedRequest.prepared_lifetime` against the published node
before any mutation. Language re-home preserves this lifetime while allocating
a new object incarnation; removal/reset invalidates it. A mismatch terminalizes
the sender (`Shutdown` when absent, `Superseded` when replaced by another
lifetime). Registration precedes work admission, so a refused `submit_file`
must also terminalize the waiter rather than leave it parked.

A node published at submission can still lack admitted Source work. Its
current-generation Source admission marker distinguishes this pending producer
from a dead producer. Generation advance resets the marker so a queued reload
also remains gating before inbox admission. Both explicit blocker registration
and Source-completion dependency
integration start or track that producer before recording Analysis blockers.
The initial generation is advanced above zero before those blockers are built.
Explicit registration revalidates its captured owner incarnation/generation
under every dependency mutation hold, including iterations after inbox
backpressure has executed callbacks. No host callback or inbox send runs under
the lifecycle hold.

Generation allocation/advance uses checked arithmetic within one node object:
exhaustion refuses further advancement on that node, and leaves no history,
rather than wrapping to a generation the same object already published. Across
a removal, reset or re-home the successor object restarts its own sequence, and
safety comes from its never-reused incarnation, not from the generation:
`SourceVersion`/`SourceWitness` (and the host's registered-source file
incarnation and native-host bind check) compare the incarnation too, so a
successor's same-numbered generation never authorizes work for its predecessor.

**Generation-advance rule (both directions).** A generation advance and
its supersede sweep are ONE critical section under `dag.lock()`
(`invalidate`, `close_file`, `admit_prepared_under_lock` including the
language re-home). That covers the sweep direction only: the sweep is
purely BACKWARD-LOOKING and can never retire an identity admitted after
it ran. So admission of DERIVED work needs the matching forward gate.
`handle_stage_complete` checks on entry, then runs `extract_deps`
UNLOCKED — a real window in which an invalidate can retire the
generation. So the ONLY thing outside the lock is that executor call
(pure computation, unbounded host cost). Everything the Source completion
publishes AND everything it CONSUMES — forward edges, the destructive
deferred-blocker drain, dependency auto-ingest + admission, the Artifact
blocker registry write, Source-waiter signalling, conditional Analysis
admission, `complete(Source-G)` —
happens under ONE `dag.lock()` hold gated on `stage_completion_is_current`.
Consumption matters as much as publication: a stale completion draining a
LATER generation's deferred blockers discards them and lets that
generation's Artifact work run ungated, and stale forward edges persist
because later extraction unions rather than replaces.

The gate is **incarnation + generation + generation-coherent committed
Source snapshot**. The incarnation is `FileNode::incarnation_id()`, a
process-unique monotonic id, and it MUST be carried from dispatch on
`Submission::StageComplete { incarnation }`. Re-deriving it by map lookup
compares the live node with itself and passes vacuously — two node
objects for the same canonical can sit at the SAME generation, so the
generation check cannot catch a replacement either. When Analysis is
demanded, it is admitted BEFORE `complete(Source-G)` so the file is never
briefly without a live stage identity (a concurrent dead-producer
classification would read that as a Source-failed corpse). A Source-only
request intentionally has no downstream identity after completion. On
refusal it publishes and consumes NOTHING
and calls `refuse_stale_stage_completion`, which cancels the dequeued
identity idempotently — safe against a later generation because
`WorkNodeIdentity::FileStage` carries incarnation and generation — signals the
retired generation's waiter groups so a refusal can never strand a
request (a no-op when a sweep already drained them), requeues stranded
waiters after the lock drops, bumps the test/`semantic-observe`-only `stale_completion_refusals`, and only
THEN `debug_assert!`s.

`StageExecutor::extract_deps` is host-specific. The session host returns only
macro type dependencies as `blocker_ids`, because those are the dependencies
the scheduler's Artifact gate consumes. Ordinary imports and external `src`
edges are owned by the workspace parsed-edge graph and MUST NOT be resolved a
second time merely to populate the session scheduler's unused `forward_deps`;
session compilation is host-owned rather than `execute_artifact`-owned. The
generic scheduler still records whatever `forward_deps` another executor
returns. Priority inheritance for a completion is likewise canonical-local:
`highest_priority_for_file` reads the DAG canonical reverse-index bucket, never
the whole node map.

The same single-hold rule applies one stage down:
`admit_pending_artifacts` holds ONE lock across the profile snapshot and
every admission it drives, plus a liveness pre-check. Snapshotting,
releasing, then re-locking per profile let an invalidate bump and sweep
in the gap, admitting Artifact-G after the sweep.

Node creation must be an atomic ENSURE, never a replace: auto-ingest uses
`nodes.entry(..).or_insert_with(..)`, not `contains_key` + `insert`. The
check-then-act form let a concurrent creator's FileNode be replaced at
the same generation, orphaning the incarnation already-dispatched work
ran against.

Without these gates the stale identity is admitted and later skipped on
the dispatch-time generation-mismatch arm, which never releases the
capacity reservation parked at dispatch, so the DAG ledger — the sole
admission gate — leaks capacity on every race. `BatchHandle` carries one
`CompletionHandle` per input in submission order; `wait_batch(&self,
&BatchHandle)` returns results in INPUT order and never surfaces a
partial set. Pump discipline is preserved throughout: dispatch / wait /
parse / compile / callbacks all run outside the DAG lock, and capacity
stays reserved at dequeue time. `compile_many` IS wired onto atomic
batch admission: its source-upsert stage routes every input through
`VerterHost::upsert_many_with_priority` (the upsert engine), which lands
ONE `submit_batch_atomic` + ONE `wait_batch` for the whole batch rather
than one upsert per file. Per-call worker count is NOT a parameter of
`compile_many` — concurrency is the construction-time host-owned
`HostCpuPool` (`HostConfig::host_cpu_threads`); see *Dual pool
isolation*.

A **leaf substrate** for the cache-runtime DAG design has LANDED. G3
wired `CpuConcurrencySemaphore` as the scheduler CPU **transport bound**:
`SchedulerCpuPool::try_submit(OwnerCommand<Cpu>)` calls
`try_acquire_owned` before `rayon::spawn` and returns `Full` when
saturated (`SchedulerCpuPool::new(threads, transport_capacity)`, floored
at `threads * 4`, must dominate `dag_budget.cpu`). `Admission<T>` is the
named admission-outcome type. Still UNWIRED as a submission-path
argument: `CancellationToken` (`cancellation.rs`), the opaque
`SchedulerCacheId` newtype in `cache_id.rs`, and the caller-side
`DedupeHook` trait + `DedupeJoiner` + `NoDedupeHook` (`dedupe_hook.rs`).
The live submit entry points still return `CompletionHandle` /
`BatchHandle`, not `Admission`. Sections below describe each.

The rest of the **cache-runtime DAG design target** is still NOT on the
tree: the `submit_dag` / `CacheNodeDag` DAG surface, the `KeyedJob` /
`CacheNodeDagNode` types, the expanded `Load` / `Parse` / `CacheNode`
`TaskKind` variants, DAG semantics (dependency gating, priority
inheritance, cancellation propagation, bounded admission / backpressure),
and the wiring of `CpuConcurrencySemaphore` onto DAG *node* dispatch
(the CPU *transport* wiring is live — see above). Every section
describing those un-landed surfaces carries an explicit "Not yet
implemented" banner. The typed `SchedulerCpuPool` / `SchedulerIoPool`
pools and owner-affine `try_submit` are on the current tree.

Binding implementation spec: `.claude/skills/type-cache-architecture/SKILL.md`
(Blocks 6 and 7). When in doubt, the plan wins; this skill derives from
the plan body.

## Driver park and teardown

The submission inbox is bounded to 1024 items. Normal submitters wait for
capacity; a driver or single-threaded inline submitter consumes older inbox
items if the queue is full, then enqueues its own item. Batches still occupy
one slot and are admitted atomically. Wake notifications use `try_send`:
when the inbox is full, its existing items already cause another pump pass.
The private teardown channel remains independent of inbox capacity.
The pressure contract is `inbox_depth_max <= 1024`, with no dropped request
or stage completion while the scheduler is live. A full inbox can delay a
submitter but cannot park its own sole consumer. Stage execution returns its
terminal `StageComplete` to the caller instead of sending it: the inline pump
(and a pool worker that inline-ran a dependency) delivers it through the same
drain-capable submission path, while a pool task waits for capacity because
it is never the inbox's consumer. No parse or stage dispatch is duplicated on
that path.

The native driver loop (`driver_loop_native`) parks between wakes on three
things at once, reported as `DriverPark`: its own private teardown channel
(`Scheduler::driver_teardown`), the shared submission inbox, and an idle
re-pump deadline. The idle deadline is a missed-wake backstop for stranded
ready work, not priority aging.

**The teardown signal never travels through the inbox.** The submission inbox
is a MULTI-consumer channel — every cooperative pump drains it
(`drain_inbox_for_pump`), the driver's own dispatch loop included — so a
teardown posted there can be consumed before the driver reaches its park. The
driver would then sleep out the whole idle interval while the thread joining
it (`reset()` / `Drop`, i.e. every host `close()`) waits with it. The
`driver_teardown` channel has exactly one consumer, the parked driver, so the
signal cannot be swallowed. It has capacity one (a second signal carries no
extra meaning) and `restart_driver` drains it before spawning the next
driver, so a signal the outgoing driver never consumed cannot stop the
incoming one on arrival.

Guard: `teardown_stops_the_driver_when_its_wake_was_consumed_before_the_park`.

## Crate dependency invariant

`verter_scheduler` MUST NOT depend on any higher-level crate. Dependency
runs one-way: higher-level crates depend on `verter_scheduler`, never the
reverse. The skill never names a symbol living in a higher-level crate —
any such reference is a cycle and a violation.

Guard:
`crates/verter_scheduler/tests/cases/no_session_dep.rs::scheduler_does_not_depend_on_verter_session`
walks `crates/verter_scheduler/Cargo.toml`, every `.rs` file under
`crates/verter_scheduler/src/**` (parsed with `syn::parse_file`), AND
this skill markdown. Asserts NO mention of any higher-level crate appears
in any `use` statement, any `dependencies` / `dev-dependencies` table, OR
any skill prose substring. The guard treats the skill as a substrate
input so a relapse in this file fails the build.

## Generic dedupe-hook surface

The `DedupeHook` trait IS on the current tree, in
`crates/verter_scheduler/src/dedupe_hook.rs`. It is the **caller-side
pre-admission singleflight** hook: the calling crate implements it over
its own in-flight table and the scheduler probes it BEFORE a submission
reaches the DAG, so a caller already holding an equivalent live flight
can skip the scheduler round-trip entirely and attach as a joiner. The
scheduler owns NO in-flight cache table — the calling crate deduplicates
BEFORE submitting.

DISTINCT from the scheduler-internal post-unlock `DedupJoinerEvent`
(`crate::dag`): that is the waiter-notify fired after the DAG lock
releases, once admission has already joined a request onto an existing
waiter group. `DedupeHook` runs on the caller's side before a submission
is even constructed; `DedupJoinerEvent` runs inside admission. Two
different lifecycle points, two different types.

```rust
// crates/verter_scheduler/src/dedupe_hook.rs
pub trait DedupeHook: Send + Sync {
    /// Probe whether `identity` is already known to the caller's
    /// in-flight table. If `Some`, the caller blocks on the existing
    /// flight and the scheduler skips enqueue; if `None`, the
    /// submission proceeds to admission as usual.
    fn probe(&self, identity: &WorkNodeIdentity) -> Option<DedupeJoiner>;
}

/// Opaque handle the caller uses to attach a completion as a joiner
/// on an in-flight flight (no public fields).
#[derive(Debug)]
pub struct DedupeJoiner { /* opaque */ }

/// The genuine no-op hook used wherever a caller supplies no in-flight
/// table — `probe` always returns `None`. NOT a stub: its contract IS
/// "never deduplicate".
#[derive(Debug, Clone, Copy, Default)]
pub struct NoDedupeHook;
```

The probe key is `crate::dag::WorkNodeIdentity` — the scheduler's own
dedupe identity and the **single dedupe-identity authority**. NO parallel
`DedupKey` type: any public dedupe key is a thin wrapper/derivation of
`WorkNodeIdentity`, never a separate key, so there is one source of truth
for dedupe identity (leaf-boundary invariant H20). The trait,
`DedupeJoiner`, and `NoDedupeHook` are fully owned by `verter_scheduler`;
no method signature or struct field on any references a higher-level
crate.

The submission path probes the hook before admission. On `Some`, the
caller blocks on the existing flight and the scheduler skips enqueue
(surfaced as `Admission::DedupeJoined`, see *Admission
substrate*); on `None`, the submission proceeds to admission. The
scheduler never imports any concrete in-flight-table type from a
higher-level crate. Wiring the hook into `submit_request` / `submit_dag`
as an explicit `&dyn DedupeHook` argument on those entry points is a
future sub-block — the trait substrate is landed and unwired.

## Admission substrate

`Admission<T>` (`scheduler.rs`, LANDED) is the typed result of a
submission attempt, generic over the success-handle type `T`. Exactly
three variants — no speculative fourth case:

```rust
pub enum Admission<T> {
    /// Admitted into the DAG; carries the caller's handle.
    Admitted(T),
    /// Collapsed onto an in-flight flight by a caller-side
    /// `DedupeHook` probe; carries the opaque `DedupeJoiner`.
    DedupeJoined(crate::dedupe_hook::DedupeJoiner),
    /// Admission declined under the capacity ledger WITHOUT mutating
    /// readiness. The caller retries or blocks on capacity.
    Backpressured,
}
```

Landed substrate, UNWIRED onto the live entry points:
(`submit_request` / `submit_batch` / `submit_batch_atomic`) still return
their existing `CompletionHandle` / `BatchHandle` shapes, not
`Admission`. Routing those entry points through `Admission`
is a future sub-block.

## CancellationToken substrate

`CancellationToken` (`cancellation.rs`, LANDED) is a cheap, clonable,
thread-safe one-shot latch — a transparent `Arc<AtomicBool>`. `clone()`
is a refcount bump, `cancel()` a single `Release` store, `is_cancelled()`
a single `Acquire` load; all clones share one flag and `cancel()` is
idempotent. It is the substrate the un-landed DAG design uses for
per-node cancellation propagation
(`CacheNodeDagNode.cancellation_token`), but on the current tree it is
UNWIRED — no submission path or work node carries one yet.

## KeyedJob, DedupKey, and `CacheNodeDagNode` lifecycle

> **Not yet implemented — cache-runtime DAG design.** The `KeyedJob` /
> `DedupKey` / `CacheNodeDagNode` / `CacheNodeDag` / `submit_dag` types
> and the whole lifecycle here are the un-landed design target from
> `.claude/skills/type-cache-architecture/SKILL.md`; none are on the current
> tree. Live submission surface: `Scheduler::submit_request` /
> `submit_batch` (returning a `BatchHandle`) over the live `TaskKind` set
> `Source` / `Analysis` / `Artifact`, dispatched onto the host-injected scheduler
> `cpu_pool` via nonblocking `cpu_pool.try_submit(...)` (see *Dual pool
> isolation*). Types
> and steps below describe the intended shape.
>
> **Dedupe-identity reconciliation:** the LANDED dedupe authority is
> `crate::dag::WorkNodeIdentity` (the `DedupeHook::probe` key — see
> *Generic dedupe-hook surface*). The illustrative `DedupKey` struct
> below is an earlier draft shape; when the DAG surface lands its dedupe
> key MUST be `WorkNodeIdentity` (or a thin derivation of it), NOT a
> parallel key type. There is one dedupe-identity source of truth.

`KeyedJob` is the submission identity. `CacheNodeDagNode` is the
ready-queue envelope the driver dispatches. The inbox-level enum
`driver::Submission` is a separate type owning its own discriminator
variants.

```rust
#[derive(Clone, Debug, PartialEq, Eq, Hash)]
pub struct DedupKey {
    pub canonical: std::sync::Arc<str>,
    pub stage: TargetStage,
    pub content_hash: u64,
}

#[derive(Clone, Debug)]
pub struct KeyedJob {
    pub dedup_key: DedupKey,
    pub stage: TargetStage,
    pub priority: Priority,
    /// World generation under which the job was enqueued. Dispatch
    /// reads `node.keyed_job.generation` directly; `CacheNodeDagNode`
    /// has no `generation()` accessor.
    pub generation: u64,
}
```

`KeyedJob` carries NO `task` / `task_kind` field. The task discriminator
lives on `CacheNodeDagNode.task_kind` only — one source of truth.

> **Not yet implemented — cache-runtime DAG design (Block 7).** The
> `submit_dag` / `CacheNodeDag` / `SchedulerCpuPool` / per-task
> `cpu_concurrency_semaphore` lifecycle below is the Block 7 design
> target from `.claude/skills/type-cache-architecture/SKILL.md`; NOT on the
> current tree. On the current tree the scheduler exposes `submit_request`
> (no `submit_dag`), the live `TaskKind` set is `Source` / `Analysis` /
> `Artifact`, and CPU stage work dispatches via the host-injected scheduler
> `cpu_pool.try_submit(...)` (see *Dual pool isolation* for the authoritative
> live pool model). Steps below describe the intended DAG flow once Block
> 7 lands.

Lifecycle (Block 7 design target):

1. **Caller-side dedupe.** Cache-runtime callers consult their in-flight
   table FIRST. A matching flight short-circuits — no scheduler
   submission happens.
2. **Submit.** Caller invokes `Scheduler::submit_request(req)` or
   `Scheduler::submit_dag(dag)` (optionally passing a `DedupeHook`).
   `submit_request` lands a `driver::Submission::NewRequest` on the
   inbox. `submit_dag` constructs a `CacheNodeDag` and pushes its ready
   nodes into the bounded ready queue as upstream gates fire.
3. **Scheduler-side dedupe probe.** Driver computes `dedup_key_for(req)`
   and consults `pending_requests` (the scheduler's own per-process
   inbox-level dedupe). A duplicate `DedupKey` attaches the caller's
   `CompletionSender<RequestResult>` as a joiner on the existing flight;
   no new job enqueued.
4. **Admission.** A non-dedup submission is admitted to the priority
   ready queue
   (`Arc<crossbeam_queue::ArrayQueue<Arc<CacheNodeDagNode>>>` — the inner
   `Arc` is required because `CacheNodeDagNode` is not `Clone`: its
   `CacheNodeCompletionSender` wraps a single-use
   `tokio::sync::oneshot::Sender`, so the same node lives on both the
   ready queue and `DagState.nodes` only via `Arc`-sharing), subject to
   the bounded-admission policy below. Per-call CPU concurrency is
   enforced by the worker dispatch site (per-task
   `cpu_concurrency_semaphore.acquire()`), not by admission.
5. **Execution.** Driver pops a ready node and dispatches via `TaskKind`
   routing:
   - `Load` → `IoPool::submit`;
   - `Parse` / `CacheNode` / CPU `Analysis` / CPU `Artifact` →
     `SchedulerCpuPool::submit`.
6. **Completion.** `pending_requests` cleared; every joiner receives the
   result through their attached `CompletionSender<RequestResult>`; DAG
   dependents are re-evaluated for readiness. The worker's per-task
   `CpuConcurrencyPermit` drops via RAII immediately after the task body
   returns, releasing the semaphore counter and notifying one waiter.

## Dual pool isolation

Distinct native worker substrates cooperate so the batch-orchestration outer
wait and the scheduler's CPU stage executor cannot deadlock on the same
workers. The host constructs both roles, but the split between scheduler stage
workers and host coordinator workers is the deadlock-isolation invariant; the
separate channel-backed I/O pool preserves source-load isolation.

- **Scheduler stage pool (`cpu_pool`)** — constructed by the host from
  `SchedulerConfig::cpu_threads` and injected as an
  `Arc<SchedulerCpuPool>` into `Scheduler::with_executor` /
  `new_sync_with_executor`. The scheduler retains that handle as the ONLY pool
  for CPU stage
  execution: the driver dispatches the live `TaskKind::Source` CPU step
  (the parse folded into `Source`) plus `TaskKind::Analysis` and
  `TaskKind::Artifact` onto it via nonblocking
  `cpu_pool.try_submit(OwnerCommand::cpu(...))`. Construct with
  `SchedulerCpuPool::new(threads, transport_capacity)` (floored at
  `threads * 4`); transport capacity must dominate `dag_budget.cpu`.
  Workers register
  `CallerKind::CpuWorker` so `wait_or_drive` routes them to the
  cooperative-pump branch. The host likewise constructs and injects the
  bounded `Arc<SchedulerIoPool>` (`SchedulerConfig::io_threads`) for the
  pure-I/O step of `TaskKind::Source` (reading bytes off disk). Its transport
  capacity must dominate the same scheduler DAG's resolved I/O budget.
- **Coordinator pool (`HostCpuPool`)** —
  `crates/verter_scheduler/src/execution/host_cpu_pool.rs`. Constructed once at
  startup by the external host/runtime layer via
  `verter_scheduler::HostCpuPool::new(num_threads)` and owned THERE, as a
  sibling of the `Scheduler` — NOT passed into the scheduler and NOT a
  field on it. Shared by the outer batch coordinator of EVERY host batch
  API (batch component-meta, batch SFC compile, and any future host batch
  fan-out) for its synchronous wait points. Its workers register
  `CallerKind::External` (8 MiB stacks), so they PARK in `wait_or_drive`
  rather than inline-executing scheduler CPU tasks, and the driver's
  inline-execute branch excludes `External` — coordinator-pool workers
  therefore NEVER run scheduler CPU stage work (`TaskKind::Source` /
  `Analysis` / `Artifact`).

Native scheduler constructors receive the host-built `cpu_pool` + `io_pool` as
explicit `Arc` parameters. The coordinator pool remains entirely in the
external layer and is never passed into the scheduler:

```rust
impl Scheduler {
    pub fn new(
        config: SchedulerConfig,
        source_loader: Arc<dyn SourceLoader>,
        cpu_pool: Arc<SchedulerCpuPool>,
        io_pool: Arc<SchedulerIoPool>,
    ) -> Arc<Self> {
        Self::with_executor(
            config,
            source_loader,
            Arc::new(DefaultExecutor),
            cpu_pool,
            io_pool,
        )
    }

    pub fn with_executor(
        config: SchedulerConfig,
        source_loader: Arc<dyn SourceLoader>,
        executor: Arc<dyn StageExecutor>,
        cpu_pool: Arc<SchedulerCpuPool>,
        io_pool: Arc<SchedulerIoPool>,
    ) -> Arc<Self> {
        // Retain the injected scheduler execution pools and spawn the driver
        // thread holding `Weak<Scheduler>`. No coordinator pool here.
    }
}

pub struct Scheduler {
    #[cfg(not(target_arch = "wasm32"))]
    pub(crate) cpu_pool: Arc<SchedulerCpuPool>, // stage execution ONLY
    #[cfg(not(target_arch = "wasm32"))]
    pub(crate) io_pool: Arc<SchedulerIoPool>,
    // ... other existing state (inbox, edges, dag, overlay, source_loader,
    //     executor, deferred_blocker_ids,
    //     shutdown, driver_handle, counters, config) ...
}
```

Test corpus code may reuse these execution-pool `Arc`s across sequential fresh
scheduler shells inside one aggregate `#[test]`. It must not run multiple such
schedulers concurrently unless the shared bounded I/O transport is sized for
their aggregate admission: the ordinary host capacity rule only proves the
transport dominates one scheduler's resolved I/O ledger. Scheduler/DAG/driver
state and `HostCpuPool` coordinator ownership remain per-host.

**Single batch-coordination primitive (lives in the external host/
runtime layer, not in this crate).** Every host/runtime batch API (batch
component-meta, batch SFC compile, and any future batch fan-out) routes
its outer wait through ONE coordinator primitive owned by the external
layer, parameterised by a small per-client batch policy/context. That
primitive — not the scheduler — owns: coordinator-pool `install`; the
empty / single-item fast path; deterministic per-input ordering; a
generic per-item panic boundary (catches a panicking item and hands it to
the client's policy for domain conversion, so one item never aborts the
batch); per-batch submission accounting (when the policy carries a
scheduler handle); a per-batch tracing span; and the non-reentrant policy
below. Each client supplies only its item work and its domain
panic→result conversion. The primitive does NOT own
cancellation/shutdown — the scheduler exposes no batch-cancellation
facility today, so a batch runs to completion. The scheduler crate
exposes NO outer-fan-out API and performs NO `par_iter().install(...)`
outer wait on its `cpu_pool`; a batch's per-batch submission accounting is
a pool-free counter bump (`Scheduler::account_batch_submission`), which
the coordinator invokes once per non-empty batch.

**Non-reentrant host-batch contract.** A batch item closure may call
scalar scheduler operations, but a nested batch fan-out reached from
inside an item closure must NOT issue a fresh coordinator-pool install.
The external primitive detects re-entrancy (a per-thread marker scoped
around each item's execution) and runs the nested fan-out INLINE /
sequentially on the current coordinator worker. Stacking a second outer
wait on the same finite coordinator pool would reintroduce the starvation
class one level up.

**Deadlock-free property + new invariant.** The execution roles are distinct:
**no worker waits for a job in its OWN pool.** A coordinator-pool worker may
block on scheduler stage work
without deadlock because the scheduler's `cpu_pool` has its own
independently-proceeding worker set; a `cpu_pool` worker running
`TaskKind::Source` stage work is not a coordinator worker and does not
gate the outer coordinator's wait. The invariant in full:

> Outer API fan-out may block only on scheduler stage work;
> scheduler stage work must never require coordinator-pool workers;
> nested host-batch fan-out is rejected or collapsed inline by the
> external batch coordinator. External host/runtime layers own the
> coordinator pool(s) and the batch-coordination primitive; they must
> not run outer waits on the scheduler stage pool.

Guards live in the external host/runtime layer (not in this crate, to
preserve the one-way dependency): a watchdog-bounded regression
characterizes the starvation deadlock (cold cross-file deps + a stage
pool sized to the batch width), and the coordinator primitive's own
reentrancy test pins the inline collapse of a nested batch.

## Per-call concurrency semaphore

Current state: batch fan-out has no per-call `threads` option — the host
coordinator pool's worker count is sized once at host construction (from
the host's CPU-thread config) and reused across every batch call, and the
scheduler's stage `cpu_pool` runs at its configured concurrency.

The `CpuConcurrencySemaphore` / `OwnedCpuConcurrencyPermit` TYPES are
LANDED (`crates/verter_scheduler/src/execution/cpu_concurrency.rs`) and G3-wired as
the scheduler CPU transport: `SchedulerCpuPool::try_submit` takes an
owned permit before spawn. The per-task `acquire()` handle on a
`CacheNodeDagNode` (Block 7) is still UNWIRED.

> **Not yet implemented.** The `Scheduler::cpu_concurrency_semaphore(n)`
> constructor method and per-call concurrency capping on
> `SchedulerCpuPool` admissions (the `CpuConcurrencySemaphore` handle
> propagated through `CacheNodeDagNode.cpu_concurrency_semaphore`) are
> part of the un-landed cache-runtime DAG design target in
> `.claude/skills/type-cache-architecture/SKILL.md`. The rest of this section
> describes that intended design. Once it lands, callers attach the
> handle to every `CacheNodeDagNode.cpu_concurrency_semaphore` in the
> batch DAG:

```rust
impl Scheduler {
    /// Construct a per-batch CPU concurrency semaphore HANDLE.
    /// Returns the `Arc<CpuConcurrencySemaphore>` the calling crate
    /// attaches to every `CacheNodeDagNode.cpu_concurrency_semaphore`
    /// in the batch DAG. The worker dispatch site acquires a FRESH
    /// `CpuConcurrencyPermit` from the semaphore IMMEDIATELY BEFORE
    /// each task body runs; the permit drops on task completion.
    pub fn cpu_concurrency_semaphore(&self, n: usize)
        -> Arc<CpuConcurrencySemaphore> { /* ... */ }
}
```

`CpuConcurrencySemaphore` (LANDED) is a hand-rolled counting primitive.
Substrate: `parking_lot::Mutex<usize>` (the free-permit count) +
`parking_lot::Condvar` — the only synchronisation primitives
`parking_lot 0.12` exports (`parking_lot::Semaphore` does NOT exist in
that version; absence pinned by `tests/cases/no_parking_lot_semaphore.rs`).
`new(capacity)` PANICS on `capacity == 0` (a release-active assert: a
zero-permit semaphore would deadlock every `acquire`; the cap is
configured once at construction so the check is off the hot path).
`acquire()` BLOCKS in a predicate-rechecking `while *available == 0` loop
until a permit is free, then decrements and returns the RAII
`CpuConcurrencyPermit` (`#[must_use]`, non-`Clone` — one permit is exactly
one held slot). `Drop` increments the count and `notify_one`s a single
waiter, on BOTH the normal path AND stack-unwind on panic, so a panicking
holder still frees its slot. The `Mutex<usize>` count is the single
source of truth for available permits. Guards:
`tests/cases/cpu_concurrency_semaphore.rs` pins the capacity cap (deterministic
channel-handshake blocking proof), RAII normal-drop release, and
panic-unwind release; the `cpu_concurrency` module is
`#[cfg(not(target_arch = "wasm32"))]` (the limiter caps the native-only
scheduler CPU pool — wasm runs the scheduler inline), so the test file
compiles native-only.

Propagation model: every `CacheNodeDagNode` carries
`cpu_concurrency_semaphore: Option<Arc<CpuConcurrencySemaphore>>` — the
SEMAPHORE HANDLE, NOT a pre-acquired permit. The worker dispatch site
calls `sem.acquire()` per task immediately before the executor runs the
body; the permit drops on task completion. Cloning the
`Arc<CpuConcurrencySemaphore>` across N DAG nodes does NOT pre-acquire N
permits — only `acquire()` consumes a permit. This is the only shape that
enforces "max `capacity` concurrent CPU tasks" across the DAG. A design
propagating a shared pre-acquired `Arc<CpuConcurrencyPermit>` would
acquire ONE permit at submission and let N>capacity tasks run
concurrently.

## TaskKind routing

> **Current state.** The live `TaskKind` set is `Source` / `Analysis` /
> `Artifact`. CPU stage work (`Analysis` / `Artifact`, and the parse step
> folded into `Source`) dispatches onto the host-injected scheduler `cpu_pool`
> via `cpu_pool.try_submit(...)`; `Load`-style I/O runs on the `io_pool`. See
> *Dual pool isolation* for the authoritative live pool model.
>
> **Not yet implemented — cache-runtime DAG design (Block 7).** The
> expanded `TaskKind` shape below (`Load` / `Parse` / `CacheNode`
> variants) and the `SchedulerCpuPool::submit` dispatch form are the
> Block 7 design target from `.claude/skills/type-cache-architecture/SKILL.md`;
> NOT on the current tree. Wherever a routing bullet below says
> `SchedulerCpuPool::submit`, the current tree dispatches the equivalent
> stage work onto `cpu_pool` via `cpu_pool.try_submit(...)`. The bullets
> describe the intended Block 7 routing.

The scheduler routes (Block 7 design target):

- `TaskKind::Load { canonical }` → I/O pool (pure I/O — reads bytes off
  disk; no executor dispatch, the source loader drives the I/O directly).
- `TaskKind::Parse { canonical, source, file_language }` → stage CPU pool
  (pure CPU; payload carries the resolved `verter_language::FileLanguage`
  row so `execute_source` dispatches without re-classifying the path).
- `TaskKind::Analysis { canonical, source_snapshot }` →
  `SchedulerCpuPool::submit`. Dispatch destructures `canonical` off the
  payload and passes the snapshot reference to `execute_analysis`. The
  substrate `SourceSnapshot` has no `canonical_id()` accessor —
  `canonical` lives on the variant.
- `TaskKind::Artifact { canonical, source_snapshot, analysis_snapshot, profile_hash }`
  → `SchedulerCpuPool::submit`. Same payload-bearing shape.
- `TaskKind::CacheNode { cache_id: SchedulerCacheId, key_hash: u64 }` →
  `SchedulerCpuPool::submit`. The worker dispatches through
  `execute_cache_node(&node, &ctx) -> CacheNodeOutcome` (direct return,
  NOT `Result`-wrapped). `SchedulerCacheId` is the scheduler-local OPAQUE
  NEWTYPE `pub struct SchedulerCacheId(pub u64)` defined in
  `crates/verter_scheduler/src/cache_id.rs` (`Clone, Copy, Debug, Eq,
  Hash, Ord`). Deliberately NOT an enum — an enum would leak session
  cache-family meaning into the scheduler and create a second source of
  truth for cache identity. The scheduler stays domain-agnostic: the
  opaque `u64` is the discriminator on `WorkNodeIdentity::CacheNode`, and
  the session owns its interpretation. No `dag.rs` re-export shim for the
  type — it lives in `cache_id.rs` and is re-exported from the crate root
  only.

`TaskKind` is no longer `Copy` — payload-bearing variants carry
`Arc<str>` / `Arc<SourceSnapshot>` etc. Every existing `Copy` call site
(e.g. `supersede_old_generations` at `scheduler.rs:388`) becomes an `Arc`
clone. The discriminating test `task_kind_clone_is_cheap_arc_clone` pins
the clone cost at < 100ns p99.

Under the Block 7 design target, `TaskKind::Source` (which on the current
tree combines load + parse, with the I/O step on `io_pool` and the parse
step folded onto `cpu_pool`) is split: the source loader synthesises a
`Load → Parse` DAG edge. **On the current tree `TaskKind::Source` is the
live first stage and is NOT split or retired** — `Load` / `Parse` are not
separate variants yet. `SchedulerJobKind` (the existing non-staged
component-meta batch enum at `stage.rs:19`) is **retained** unchanged —
it discriminates `ComponentMeta { canonical_id }`. The scheduler does NOT
own the batch fan-out for it: the external host/runtime layer maps these
job items and fans them out through its own batch-coordination primitive
(see *Dual pool isolation*), calling `Scheduler::account_batch_submission`
once per non-empty batch for the O(1) submission accounting. The Block 7
`TaskKind::CacheNode` variant lives alongside it on the new ready-queue
envelope.

### StageExecutor dispatch surface

> **Not yet implemented — cache-runtime DAG design (Block 7).** The
> five-method dispatch surface, the `CacheNodeDispatchCtx` /
> `execute_cache_node` machinery, and the `Parse` / `CacheNode` / `Load`
> rows below are the Block 7 design target from
> `.claude/skills/type-cache-architecture/SKILL.md`; NOT on the current tree. On
> the current tree the `StageExecutor` dispatches the live
> `TaskKind::Source` / `Analysis` / `Artifact` stages. The surface below
> describes the intended Block 7 dispatch.

The `StageExecutor` trait exposes five dispatch methods, one per
`TaskKind` variant. Workers route through `TaskKind` at dispatch time; no
bare `executor.execute(node)` method.

| TaskKind         | StageExecutor method   | Return                                |
|------------------|------------------------|---------------------------------------|
| `Parse`          | `execute_source`       | `Result<SourceSnapshot, StageError>`  |
| `Analysis` (CPU) | `execute_analysis`     | `Result<AnalysisSnapshot, StageError>`|
| `Artifact` (CPU) | `execute_artifact`     | `Result<ArtifactSnapshot, StageError>`|
| `CacheNode`      | `execute_cache_node`   | `CacheNodeOutcome` (NOT `Result`-wrapped; errors live inside `CacheNodeOutcome::CacheNode(Err(_))`) |
| `Load` (I/O)     | (no executor; source loader directly via `IoPool::submit`)     |

The trait also requires `fn as_any(&self) -> &dyn std::any::Any`
(object-safe, no default body) for the test-support
`Scheduler::last_dispatched_task` downcast. Every concrete impl
(`DefaultExecutor`, `HostStageExecutor`, the test-support
`LastDispatchedTaskRecorder`) provides the one-line body.

The worker's `dispatch_cpu_task` constructs a `CacheNodeDispatchCtx<'_>`
(dedup key, generation, optional audit observer, cancellation token)
BEFORE the `match` and passes a non-owning borrow to `execute_cache_node`.
The default body returns `CacheNodeOutcome::stub()`; the host overrides to
drive the cache-runtime artifact / query node trait surface. The other
three CPU dispatch methods return their stage-specific result, and the
worker maps each into the unified `CacheNodeOutcome` via the
`CacheNodeOutcome::from_source` / `from_analysis` / `from_artifact`
adapters before writing it on `node.completion`.

## DAG submission semantics

> **Not yet implemented — cache-runtime DAG design (Block 7).** The
> `submit_dag` / `CacheNodeDag` / `submit_batch`-as-DAG-shim surface in
> this whole section is the Block 7 design target from
> `.claude/skills/type-cache-architecture/SKILL.md`; NOT on the current tree. On
> the current tree the scheduler exposes `submit_request` (no `submit_dag`
> and no `CacheNodeDag` envelope), and the live `TaskKind` set is
> `Source` / `Analysis` / `Artifact` dispatched onto the host-injected scheduler
> `cpu_pool` via `cpu_pool.try_submit(...)` (see *Dual pool isolation* for the
> authoritative live pool model). The DAG submission contract below
> describes the intended flow once Block 7 lands.

The DAG API (Block 7 design target) is one method, one type, one
signature, carrying every field the driver requires to dispatch each node
as an executable unit of work:

```rust
pub struct CacheNodeDag {
    pub nodes: Vec<CacheNodeDagNode>,
    pub edges: Vec<CacheNodeDagEdge>,
    pub completion_aggregator: Arc<DagCompletionAggregator>,
}

pub struct CacheNodeDagNode {
    pub id: CacheNodeId,
    pub keyed_job: KeyedJob,
    pub task_kind: TaskKind,
    pub dedup_key: DedupKey,
    pub priority: Priority,
    pub cancellation_token: CancellationToken,
    /// SEMAPHORE HANDLE (NOT a pre-acquired permit). The worker
    /// takes a fresh RAII permit on task dispatch; the permit drops
    /// on task completion.
    pub cpu_concurrency_semaphore: Option<Arc<CpuConcurrencySemaphore>>,
    /// Scheduler-local opaque wrapper
    /// (`crates/verter_execution/src/request_context.rs:103`,
    /// `pub struct OpaqueRequestContext(pub Arc<dyn RequestContextLike>)`).
    /// Calling crate wraps its concrete context inside
    /// `OpaqueRequestContext(arc as Arc<dyn RequestContextLike>)`
    /// when constructing the node.
    pub request_context: Arc<OpaqueRequestContext>,
    /// Cache-node-only completion channel. Wraps
    /// `tokio::sync::oneshot::Sender<CacheNodeOutcome>` in
    /// `Mutex<Option<...>>` so the worker dispatch site can `take()`
    /// the inner sender out of a shared `&CacheNodeDagNode` borrow.
    /// RENAMED from earlier drafts' `node::CompletionSender` to
    /// avoid collision with the substrate's `job::CompletionSender<T>`.
    pub completion: CacheNodeCompletionSender,
}

pub struct CacheNodeDagEdge {
    pub from: CacheNodeId,
    pub to: CacheNodeId,
    pub gate: EdgeGate,
}

pub enum EdgeGate {
    Sequential,
    ConditionalOnSuccess,
    ConditionalOnAdmission,
}

impl Scheduler {
    pub fn submit_dag(&self, dag: CacheNodeDag) -> DagHandle { /* ... */ }
}
```

The nine-field `CacheNodeDagNode` envelope is complete: the `task_kind`
discriminator lives on the node only (NOT on `KeyedJob`). No node enters
the ready queue without all nine fields populated; the driver does NOT
enrich nodes after submission. Guard:
`cache_node_dag_carries_required_fields_for_executable_dispatch`.

DAG contract:

- **Dependency gating.** A downstream node is not admitted to the ready
  queue until ALL of its upstream nodes have completed per their
  `EdgeGate` policy.
- **Priority inheritance.** Effective priority is
  `max(node_priority, max(root_priority for every reachable root))`.
- **Cancellation propagation.** Dropping a `DagHandle` triggers
  `CancellationToken::cancel()` on every node not yet completed;
  cancellation propagates transitively through edges.
- **Bounded admission / backpressure.** The ready queue is bounded by
  `MAX_READY_QUEUE_DEPTH = 64` (`crates/verter_scheduler/src/queue.rs`).
  When full, additional submissions either block or return
  `Admission::Backpressured` per caller preference.
  `Scheduler::ready_queue_depth()` exposes the current bounded depth for
  observability only.
- **In-flight dedupe inside a DAG.** Two nodes in the same DAG sharing a
  `dedup_key` collapse via scheduler-side `pending_requests`. Cross-DAG
  dedupe uses the consumer-side in-flight table via `DedupeHook::probe`
  BEFORE submission.

Under Block 7, `submit_batch(reqs: Vec<Request>)` becomes a thin shim
constructing a no-edge `CacheNodeDag` and calling `submit_dag`. On the
current tree it loops over `submit_request` (see the surface table below).

## MVCC source root (`source_root.rs`)

`Scheduler.nodes` is EXECUTION state only: a `FileNode` holds its
CURRENT `ArcSwap` snapshots, `bump_generation` makes the prior source
immediately unreachable, and `node_ids()` is a full map walk. Beside it
the scheduler owns `SchedulerSourceDirectory` — the epoch-indexed MVCC
authority for what `try_get_source` LOGICALLY answers.

```text
SchedulerSourceRoot { visible_epoch, root_lease }
canonical -> version history of
    { epoch, incarnation, generation, Present(whole_hash) | Absent }
```

| Surface | Contract |
|---|---|
| `Scheduler::capture_source_root() -> Arc<SchedulerSourceRoot>` | O(1) in file count: one publication-lock acquisition, one scalar read, one lease bump. Measured 37 ns @250 files, 33 ns @3,000 (release). |
| `SchedulerSourceRoot::lookup(canonical) -> SourceStateAt` | AS-OF, sealed to the root's epoch. `Unknown` / `Absent{incarnation,generation}` / `Present{…, whole_hash, semantic_hash}`. The root exposes NO path to the live directory. |
| `SchedulerSourceDirectory::publish_transition(f)` | Runs the node mutation AND the version append under ONE publication hold. |
| `SchedulerSourceDirectory::reclaim_superseded_versions()` | Root-gated GC; floor = oldest LIVE captured root capped by the current epoch. |

Rules:

- **Live generation advances have a closed type gateway.** `FileNode` stores
  its counter as the opaque child-module `LiveGenerationCounter`; the raw
  `AtomicU64` never escapes that module. Reads go through the wrapper and the
  only advance requires `&SourcePublication`, forwarded by
  `FileNode::bump_generation`. `FileNode` denies `private_interfaces`, so
  widening the private field beyond the opaque type's visibility is itself a
  compile error rather than exposing a new sibling-module mutation route.
- **Publication is atomic with the lifecycle transition.** The node
  mutation happens INSIDE the `publish_transition` closure and
  `capture_root` takes the same lock, so a capture is totally ordered
  against every transition — never a torn `(node moved, root did not)`
  pair. Publishing sites: `invalidate`, `close_file`, `remove`, `reset`
  (ONE epoch for all removed members), the `handle_new_request`
  source-update bump and language-replacement, and the Source-stage
  commit in `execute_source_stage`. Node CREATION publishes nothing —
  a fresh node has no source and an untracked canonical already reads
  `Unknown`.
- **The epoch ADDRESSES a snapshot, never validates a cache.** It is
  not a `StoreViewValidationToken` dimension and must not become one.
- **A root is a RETENTION LEASE.** GC may free a version only once it
  is invisible from the current root AND from every live captured root
  — the same reachability discipline `FileArtifactStore` applies to
  artifact versions. `HostStoreView` captures one in its pre-build read
  window and retains it by `Arc`.
- **Lock rank:** `SchedulerDag` (outer) > source-root publication >
  `nodes` / `versions` DashMap shards (inner). A publication may take a
  DashMap shard; nothing takes the publication lock while holding one,
  and nothing takes the DAG lock while holding the publication lock.
- Write-path cost: one publication is 53-59 ns (release), taking
  end-to-end `Scheduler::invalidate` from 53 ns to 103 ns.

Contract tests: `crates/verter_scheduler/src/source_root_tests.rs`
(as-of sealing, atomic publication, lease-gated reclamation, O(1)
capture) and `crates/verter_session/src/source_root_retention_tests.rs`
(the `HostStoreView` lease at the host boundary). Normative text:
`docs/contributing/path-precise-resolution-currency.md` → "An immutable root is
also a retention lease".

## Scheduler surface (current → Block 7 planned)

The right column is the Block 7 cache-runtime design target from
`.claude/skills/type-cache-architecture/SKILL.md`; NOT on the current tree. The
left column is the live surface.

| Method | Current | Block 7 (planned) |
|---|---|---|
| `submit_request(req)` | inbox + per-request `CompletionHandle` | unchanged signature; gains optional `&dyn DedupeHook` arg |
| `submit_batch(reqs)` | loop over `submit_request` | thin shim over `submit_dag` (no-edge DAG) |
| `submit_dag(dag) -> DagHandle` | absent | NEW |
| `dedup_key_for(req) -> DedupKey` | absent | NEW |
| `cpu_concurrency_semaphore(n) -> Arc<CpuConcurrencySemaphore>` | absent | NEW |
| `ready_queue_depth() -> usize` | absent | NEW |
| `register_resolved_deps` | unchanged | unchanged |

## Pool routing rules

The live `TaskKind` set is `Source` / `Analysis` / `Artifact`. (The
expanded `Load` / `Parse` / `CacheNode` shape and the
`SchedulerCpuPool::submit` dispatch form are the demarcated Block 7 design
target — see *TaskKind routing* and *Scheduler surface*; NOT current
routing rules.)

- **`io_pool`** owns the pure-I/O step of `TaskKind::Source` (reading
  bytes off disk) and any other pure-I/O work. A parse closure on the I/O
  pool is a bug
  (`pool_isolation::source_parse_runs_on_cpu_pool_not_io_pool`).
- **Scheduler stage pool (`cpu_pool`)** executes the CPU stage work — the
  parse step folded into `TaskKind::Source`, plus `TaskKind::Analysis` and
  `TaskKind::Artifact` — dispatched via `cpu_pool.try_submit(...)`. The
  host constructs it from `SchedulerConfig::cpu_threads` and injects its
  `Arc` into the constructor. It is the only pool the driver dispatches stage
  work onto.
- **Coordinator pool (`HostCpuPool`)** owns the outer batch coordinator's
  wait points for EVERY host batch API (batch component-meta, batch SFC
  compile, and any future host batch fan-out). The external host/runtime
  layer constructs it once at startup and OWNS it (a sibling of the
  `Scheduler`, never handed into the constructor). The scheduler does NOT
  reference it and NEVER dispatches tasks onto it; equally, no scheduler
  API installs an outer wait on the stage pool. The coordinator pool is
  reused across batch calls (sized once from the external layer's config).
  Guard (external layer): `two_back_to_back_compile_many_share_pool`.

## Test-support helpers (`feature = "test-support"`)

> **Not yet implemented — cache-runtime DAG design (Block 7).** The
> `feature = "test-support"` gate itself is live; its current pool-identity
> readers are `host_cpu_pool_token` and `Scheduler::test_worker_pool_ids` (see
> *Dual pool isolation*). The
> fixture catalogue below — `Scheduler::new_for_test`, `enqueue_analysis`,
> `last_dispatched_task`, `LastDispatchedTaskRecorder`, the `KeyedJob` /
> `CacheNodeDagNode` / `DedupKey` / `SchedulerCacheId` stubs, etc. — is
> the Block 7 design target from
> `.claude/skills/type-cache-architecture/SKILL.md` and is NOT on the current
> tree. The helpers below describe the intended Block 7 test surface.

The crate gates a small set of fixture helpers behind
`feature = "test-support"` so the production build never compiles them.
Integration tests under `crates/verter_scheduler/tests/` enable the
feature via
`verter_scheduler = { path = ".", features = ["test-support"] }` in their
`[dev-dependencies]`.

- `Scheduler::new_for_test() -> Arc<Self>` — single-thread pools +
  `LastDispatchedTaskRecorder` executor.
- `Scheduler::enqueue_analysis(&FileNode)` — routes through
  `submit_request` with `TargetStage::Analysis`. Drives to quiescence by
  polling
  `last_dispatched_task() -> Some((_, TaskKind::Analysis { canonical, .. })) if canonical == fixture_canonical`
  (matches specifically on `Analysis`, not on the upstream `Parse`
  dispatch the driver completes first).
- `Scheduler::last_dispatched_task() -> Option<(KeyedJob, TaskKind)>` —
  downcasts `Arc<dyn StageExecutor>` to `LastDispatchedTaskRecorder` via
  `as_any` and reads its internal cell.
- `LastDispatchedTaskRecorder` — records the last `(KeyedJob, TaskKind)`
  per dispatch.
- `NoopSourceLoader` — full `SourceLoader` impl (four methods: `load`,
  `exists`, `classify`, `realpath`).
- `FileNode::stub_with_canonical(&str)` — populates `canonical_id` from
  the argument; zero-state for other fields.
- `OpaqueRequestContext::test_stub()` — wraps a private no-op
  `RequestContextLike` impl.
- `KeyedJob::stub()`, `DedupKey::new_for_test()`,
  `CacheNodeDagNode::stub()`,
  `CacheNodeDispatchCtx::stub_with(&dedup_key, &cancellation)`,
  `SchedulerCacheId(0)` (the opaque newtype constructed directly — no
  `::Test` variant; `SchedulerCacheId` is
  `pub struct SchedulerCacheId(pub u64)`, not an enum).

See also:
- `.claude/skills/host-session/SKILL.md` — host-side ownership.
- `.claude/skills/type-cache-architecture/SKILL.md` — the substrate
  the scheduler dispatches into; defines `DedupeHook` consumers.
- `.claude/skills/type-cache-architecture/SKILL.md` — the plan that landed
  the cache-runtime + scheduler integration (Blocks 6, 7).
