Agent skill

Dynamo Kv Replay Parity

by ai-dynamo in ai-dynamo/dynamo

Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Dynamo Kv Replay Parity

skills CLI
$ npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-dynamo/dynamo dynamo-kv-replay-parity --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-dynamo/dynamo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/dynamo-kv-replay-parity .claude/skills/dynamo-kv-replay-parity && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dynamo-kv-replay-parity
GitHub stars
8.3k
Token cost
~4.9k tokens
SKILL.md length
2,535 words
Files
3 (incl. references)
Skills in repo
27
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…

  • Works in 9 steps: Pin revisions and artifacts → Establish deterministic reports → Qualify a long interaction-heavy corpus → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Required inputs, Stage 1: Pin revisions and…, Campaign concurrency and Stage 2: Establish…, plus 7 more sections
  • Calls cargo and uv

What it does

Dynamo Kv Replay Parity is an agent skill from ai-dynamo/dynamo. Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and disaggregated handoff lifecycles. It is used when validating replay refactors, routing changes, scheduler-event changes, or performance-sensitive offline simulation changes against a baseline revision.

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/internal-polynomial-golden-points.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving and Refactoring. It works with SGLang and vLLM. The repository describes itself as: A Datacenter Scale Distributed Inference Serving Framework. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve Refactoring

Example prompts

  • “/dynamo-kv-replay-parity”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Pin revisions and artifacts
  2. Establish deterministic reports
  3. Qualify a long interaction-heavy corpus
  4. Run byte parity
  5. Classify semantic differences
  6. Force rare lifecycles when needed
  7. Measure performance
  8. Decide and report
  9. Clean up

What it can do on your machine

Read from SKILL.md and the folder at commit b208989. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dynamo Kv Replay Parity loads about 4.9k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 103 tokens; SKILL.md has 2,535 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-dynamo/dynamo at commit b208989, republished under its Apache-2.0 licence (© ai-dynamo). 2,535 words, ~4,854 tokens.

Download SKILL.mdSave it as .claude/skills/dynamo-kv-replay-parity/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
dynamo-kv-replay-parity
description
Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and disaggregated handoff lifecycles. It is used when validating replay refactors, routing changes, scheduler-event changes, or performance-sensitive offline simulation changes against a baseline revision.
license
Apache-2.0
metadata.author
NVIDIA
metadata.tags
dynamo, offline-replay, kv-router, parity, performance

Dynamo KV replay parity

Compare two revisions of Dynamo offline KV-aware replay using deterministic virtual-time reports and statistically paired real-time measurements. Use the existing replay and benchmark harnesses; extend them only when a required signal is unavailable.

This campaign intentionally excludes round-robin routing. It also does not require a same-timestamp event-only progress scenario. Validate those concerns with focused tests outside this skill when a change specifically affects them.

Required inputs

Resolve and record before running anything:

  • immutable baseline and candidate commit SHAs;
  • Rust toolchain, build profile, flags, and host characteristics;
  • each artifact's exact Cargo features from the relevant manifests;
  • Mooncake trace path, SHA-256 checksum, and deterministic slice rule;
  • engine, topology, concurrency, worker counts, block sizes, and native G1 capacities;
  • the canonical-report exclusion allowlist; and
  • the performance run-order seed, CPU placement, timing scope, and invalidation rules.

Use the project root .venv/bin/python for Python analysis and uv pip for any approved installation. Do not compare against moving branches or reuse a binary after changing its checkout. Do not describe the configuration as "all features"; record explicit feature names. Feature names can differ between dynamo-mocker and dynamo-bench. The current campaign is native-G1-only; do not add removed KVBM replay features or runtime arguments to make a historical command line build.

Stage 1: Pin revisions and artifacts

  1. Confirm the baseline is an ancestor or otherwise document the comparison relationship.
  2. Create isolated checkouts for both revisions using the same host and toolchain.
  3. Apply any temporary determinism correction identically to both revisions. Keep it out of the measured semantic delta and record its patch checksum.
  4. For the current dynamo-bench harness on Linux, build one release artifact per revision with exactly replay-bench and no default features. --canonical-reports-jsonl requires replay-bench, which also selects the seeded router.
  5. Reuse that artifact across all native-G1 correctness and performance rows. Do not build separate engine, topology, or production-routing artifacts. If the named manifest feature or runtime contract no longer exists, stop and report that this protocol needs updating rather than guessing a replacement matrix.
  6. Copy the artifacts to a temporary campaign directory. Record exact features, binary SHA-256, binary size, and .text size.
  7. Return each checkout to its original branch after extracting the binaries.

Build the current benchmark artifact with:

bash
cargo build --release -p dynamo-bench --no-default-features \
  --features replay-bench --bench offline_replay_bench

Never build while collecting performance samples.

Campaign concurrency

Treat one engine/topology/frozen-configuration comparison row as the unit of node placement. When scheduler capacity permits, allocate multiple nodes and assign independent rows to them. The nodes do not need to be homogeneous because no individual comparison crosses nodes. Keep the baseline, candidate, all determinism repetitions, and lifecycle evidence for one row on the same node. Use the same prebuilt revision artifacts and inputs throughout the campaign, and record the node characteristics and CPU placement for every row.

Native replay is normally single-core: pin each process to one physical core after confirming that assumption during preflight. Keep the CPU-set size, NUMA placement, and affinity identical between baseline and candidate for a row.

For correctness, run the baseline and candidate repetitions for a row concurrently when its node has sufficient resources. Pin concurrent processes to disjoint CPU sets, give each process separate output paths, and ensure they do not share mutable state. Keep every repetition in a separate process even when several repetitions run at the same time.

For performance, keep the entire row's warmups and 60-pair measurement series on its assigned node. Execute only one performance invocation at a time on that node with a fixed CPU placement, and keep each randomized baseline/candidate pair adjacent. Different performance rows may run concurrently on different nodes, but do not pool one row's pairs across heterogeneous nodes. Do not run another replay, build, profiler, or unrelated workload concurrently on a node collecting performance samples.

Multi-node and within-node correctness parallelism are campaign throughput optimizations; they must not change workload concurrency or weaken performance isolation.

Stage 2: Establish deterministic reports

Require the harness to control every known entropy source:

  • assign stable request UUIDs during workload creation;
  • use an owned seeded RNG for temperature sampling and equal-score tie selection;
  • sort routing candidates by (worker_id, dp_rank) before deterministic selection;
  • use stable request ordinals and same-time event sequence numbers;
  • recursively sort JSON object keys;
  • sort only explicitly unordered collections such as per-request records;
  • preserve semantically ordered event and lifecycle arrays;
  • exclude exactly /summary/wall_time_ms, /summary/processed_tokens_per_s, and /summary/processed_output_tokens_per_s.

This exclusion list is closed: include every other field, and stop for review before adding another exclusion. Make semantic fields such as request UUIDs deterministic rather than dropping them. Require parity output metadata to report replay_bench: true; otherwise the seeded routing contract was not active.

Run baseline twice and candidate twice in separate processes. Each revision must produce one unique canonical digest. This is an entropy-leak check, not a statistical trial. If a revision is internally unstable, stop and diagnose it; do not increase repetitions and average the outputs.

Stage 3: Qualify a long interaction-heavy corpus

Prefer a fixed contiguous 5,000-request Mooncake window over many parity repetitions. Preserve arrival order and prefix locality. Record the starting offset, request count, and checksum. Do not randomly sample rows, duplicate a shorter trace, or silently claim a 5,000-request campaign when fewer usable requests exist.

The committed lib/bench/testdata/mooncake_trace_1000.jsonl fixture is suitable for a quick harness preflight, not the authoritative long-corpus campaign.

Qualify and freeze one configuration per comparison row, or per explicitly named configuration family when rows genuinely share every relevant parameter. For each frozen configuration, prove that it exercised every applicable path:

  • KV-overlap-sensitive routing;
  • immediate placement, plus queued placement only when queueing is explicitly enabled;
  • the row's bounded preemption or retraction band at the block-capacity edge;
  • disaggregated prefill/decode handoff;
  • terminal cleanup.

Use coverage counters, lifecycle traces, or report evidence rather than inferring these paths from successful completion. Target the bounded pressure band recorded for the qualified seed. When no seed exists, start with one to three pressure events to prove the lifecycle. For throttle-oriented disaggregated vLLM and SGLang seeds, target 10 to 20 fully readmitted pressure events so repeated scheduling is exercised. Zero means the edge was not exercised; repeated preempt/re-admit cycling, a rapidly growing pressure count, or failure to advance virtual time invalidates the fixture. Tune capacity or concurrency minimally and identically for baseline and candidate within the row or family, and back off rather than accepting a pressure flood. Never tune the revisions separately.

The offline replay CLI leaves the router queue threshold unset by default. In that mode, all route decisions are immediate and zero queued placements are expected; engine-side scheduler waiting is not router queue coverage. If queue lifecycle coverage is required, enable it through an explicit queue-capable harness or forced fixture and record the exact threshold. Account for every route decision, but never relabel scheduler waiting as queued placement.

Start from qualified internal-polynomial seeds

For the canonical 5,000-row Mooncake window with the internal polynomial mocker, read internal-polynomial golden points before searching for capacity edges. Treat those configurations and observed counters as suggested starting points, not universal constants or substitutes for qualification on the pinned baseline. Requalify one frozen configuration identically on both revisions.

Use the reference's expected preemption, retraction, reuse, worker, and handoff signals as drift detectors. Treat queue counts as expected signals only when queueing was explicitly enabled. Matching lifecycle counts do not waive an unstable canonical digest; stop correctness and performance work for that row until both internal determinism and cross-revision parity are established.

Stage 4: Run byte parity

Run the 5,000-request corpus for this authoritative matrix. Treat both disaggregated rows as the primary scheduler and handoff requalification. Keep the aggregated rows as secondary parity coverage for the corresponding engine semantics.

Engine semanticsTopologyMemory pathRouting
vLLM pass-endAggregatedNative G1KV-aware
vLLM pass-endDisaggregatedNative G1KV-aware
SGLang pass-endAggregatedNative G1KV-aware
SGLang pass-endDisaggregatedNative G1KV-aware

KVBM G2-G4 replay was removed from the current mocker. It is not an unsupported row in this matrix and must not be emulated with ignored flags. A campaign that needs historical KVBM parity must pin a historical revision together with the matching historical skill protocol. For each current row:

  1. Produce canonical baseline and candidate outputs with the frozen configuration.
  2. Compare their bytes or SHA-256 digests exactly.
  3. Verify the coverage evidence independently of the digest.
  4. Delete matching full reports and retain their digests.
  5. Preserve full reports and a focused diff only when outputs disagree.

TensorRT-LLM can be additional smoke coverage, but it is not a substitute for either authoritative engine semantic.

Stage 5: Classify semantic differences

Byte mismatch is a review gate, not an instruction to preserve incorrect behavior. Allow an intentional mismatch only when the candidate is demonstrably more faithful to the specified engine, scheduler, or routing semantics.

For every proposed exception, record:

  • the exact fields, requests, or lifecycle events that differ;
  • the baseline behavior and why it is incorrect or less faithful;
  • the candidate behavior and the semantic source of truth supporting it;
  • why the difference is caused by the intended change rather than leaked entropy;
  • a focused regression test that fails on the old behavior and passes on the correction;
  • any downstream report or API compatibility impact; and
  • the reviewer-visible disposition.

Use PASS_WITH_SEMANTIC_EXCEPTIONS only when every byte difference is covered by such a record. Unexplained, incidental, or merely convenient differences fail. Do not patch the candidate back to known-wrong behavior just to obtain identical bytes.

Show full SKILL.md (980 more words)Show less

Stage 6: Force rare lifecycles when needed

Use small deterministic fixtures only for required paths the long corpus cannot reliably trigger:

  • single-worker KV-router queueing with an explicit queue threshold;
  • a preemption-edge fixture that targets one to three preemptions and then completes;
  • scale-to-zero followed by scale-up and pending-work release;
  • backend-specific prefill/decode handoff ordering.

Assert bounded preemption, continued virtual-time progress, and the lifecycle itself. A final-completion smoke test does not prove preemption or queueing occurred.

Stage 7: Measure performance

Measure every supported row from Stage 4 with its frozen configuration. The authoritative gate reuses the replay-bench artifacts from byte parity, so routing selection is seeded and matched between revisions. Require timing metadata to report replay_bench: true. Describe the result as seeded matched-routing replay-loop parity, not production-routing performance. Production-routing timing is outside the default campaign because it requires a separate build without replay-bench; add that campaign only when the user explicitly requests it.

The primary metric is replay execution time. Start its timer after trace normalization, workload construction, and engine/runtime preparation, immediately before prepared.run(...); stop it immediately after that call returns and before collector finalization or report aggregation. Emit this value as replay_execution_ms. Record setup and end-to-end time separately as diagnostics. If the harness does not expose this exact timing boundary, extend it before running the campaign; do not substitute the existing broader wall_time_ms.

Stage binaries and trace inputs on node-local storage and verify their checksums before warmups. File transfer is campaign setup, not a sample. For each measured invocation, pass --iterations 1 and a unique node-local --timings-jsonl path. Do not pass --canonical-reports-jsonl, --report-json, or another full-report output option in a gated performance invocation. Run one fresh process per sample so lazy per-request capture, report serialization, and in-process iteration state cannot contaminate the metric.

For each row:

  1. Run five warmups per arm, alternating arms for ten warmups total.
  2. Generate and persist a fixed-seed, balanced 60-pair schedule with 30 baseline-first and 30 candidate-first pairs in randomized order.
  3. Collect all 60 measured pairs. Keep the two invocations in each pair adjacent and compute r_i = candidate_replay_execution_ms / baseline_replay_execution_ms regardless of run order.
  4. Treat the ratios as independent and identically distributed, or otherwise exchangeable, only when the campaign can justify that sampling assumption; pair adjacency does not establish it. Predeclare thresholds for serial-dependence diagnostics, including lag autocorrelation and pair-order trends against elapsed time and available temperature or frequency telemetry, and record their results.
  5. If the diagnostics breach their thresholds or exchangeability cannot be justified, report INCONCLUSIVE or use a predeclared dependence-aware method. Do not apply the order-statistic gate.
  6. Otherwise sort the 60 ratios. Conditional on the sampling assumption, use the 24th order statistic as the exact distribution-free one-sided 95% lower confidence bound for the population median ratio and the 37th order statistic as the corresponding upper bound. Do not use a Wald interval.
  7. Pass when the upper bound is at most 1.05.
  8. Fail when the lower bound is greater than 1.05.
  9. Otherwise report INCONCLUSIVE; do not add samples adaptively and claim the original confidence level.

Never remove an observation because its value looks like an outlier. Retain every attempted sample and its process status. A replay error, assertion failure, malformed timing record, or other product failure fails or blocks the row; it is not a discardable sample. Only a predeclared environmental condition, such as scheduler eviction, node-health failure, affinity violation, or independently detected competing load, may invalidate a sample. Invalidate both members of that pair, record the evidence and reason, and rerun the complete pair with the same arm order. Never retry or replace a sample silently. A paired bootstrap of log ratios may be reported as a secondary effect-size diagnostic, but it does not decide the gate.

Also fail release binary or .text growth above 5% until the added footprint is explained and narrowed.

Investigate unacceptable overhead

When a statistically meaningful regression exceeds the accepted window:

  1. Confirm baseline and candidate performed equivalent semantic work. Separate an accepted semantic correction from framework overhead when the correction intentionally adds work.
  2. Inspect the diff and hot-path structure for obvious causes: event capture, request cloning, heap allocation, admission vectors, dynamic dispatch, lock traffic, or widened generic monomorphization.
  3. If static analysis does not identify a convincing cause, rerun the representative failing configuration under Samply on a supported host. Profile baseline and candidate with equivalent release/debug-symbol settings and inputs.
  4. Compare self-time, call stacks, allocation-heavy paths, and new monomorphized functions. Attribute the regression to specific code before optimizing or requesting a waiver.
  5. Repeat the paired performance gate after any fix.

Prefer an available Samply workflow skill when one is installed. Do not profile concurrently with builds or unrelated load.

Stage 8: Decide and report

Use exactly one semantic result:

  • PASS: all authoritative canonical outputs match;
  • PASS_WITH_SEMANTIC_EXCEPTIONS: every mismatch is an evidenced improvement covered by a regression test; or
  • FAIL: any unexplained mismatch, missing lifecycle evidence, or unstable revision.

Report performance independently as pass, fail, or inconclusive. A semantic exception does not waive an unexplained performance regression.

The final report must include:

  • revision SHAs and determinism-patch checksum, if any;
  • trace checksum and 5,000-request slice specification;
  • every artifact's exact feature manifest, binary checksum, and routing-mode assertion;
  • each row's frozen configuration, node characteristics, CPU set, and NUMA placement;
  • one row per engine/topology/memory path with digests and lifecycle evidence;
  • the exact canonical exclusion allowlist;
  • every semantic exception record;
  • performance timing scopes, persisted arm-order schedule, all attempted samples and invalidations, paired ratios, sampling assumption, dependence diagnostics, and conditional order-statistic confidence bounds;
  • binary and .text sizes;
  • profiler findings for any investigated regression; and
  • skipped or unsupported coverage without overstating the result.

Stage 9: Clean up

Delete temporary full reports, traces, binaries, profiler captures, patches, and worktrees after recording the required evidence. Retain full outputs only for unresolved mismatches or performance investigations. Remove task-created Cargo targets when disk pressure matters, but never delete unrelated caches or worktrees.

© ai-dynamo, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/dynamo-kv-replay-parity of ai-dynamo/dynamo.

  • SKILL.md
  • agents/openai.yaml
  • references/internal-polynomial-golden-points.md

Open the folder on GitHubat commit b208989

Compare with similar skills

Dynamo Kv Replay Parity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dynamo Kv Replay Parity compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dynamo Kv Replay Parity this skillai-dynamo/dynamo8.3k—~4.9kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Debug InferenceNVIDIA/OpenShell16k—~1.9kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
One EvalOpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Debug Inference

    NVIDIA/OpenShell

    Official

    Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

    16k GitHub stars~1.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • One Eval

    OpenDCAI/One-Eval

    驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

    165 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed

More from ai-dynamo/dynamo

All 27 skills in this repo
  • Visual Review

    ai-dynamo/dynamo

    Create self-contained interactive HTML code-review dashboards from GitHub or GitLab pull requests, checked-out branch diffs, or supplied unified diffs, with correctness and safe-to-merge scores…

    8.3k GitHub stars~4.5k tokensUpdated today
    Auto-check passed
  • Fern Components

    ai-dynamo/dynamo

    Knowledge of Fern's built-in MDX component library (accordions, callouts, cards, steps, tabs, code blocks, API-reference snippets, and more) for authoring docs pages.

    8.3k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Fern Navigation

    ai-dynamo/dynamo

    Knowledge of Fern's site-level navigation and structure configuration — how a docs site is organized in docs.yml (and product/version .yml files) using sections, pages, folders, tabs, tab variants…

    8.3k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Dynamo Agent Harness

    ai-dynamo/dynamo

    Drives persistent Claude Code, Codex, or OpenCode agent sessions through a Dynamo OpenAI/Anthropic-compatible endpoint over Agent Client Protocol (ACP).

    8.3k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Benchmark and profile the Dynamo frontend (dynamo.frontend HTTP + tokenizer + KV router) against mock workers (dynamo.mocker).

    8.3k GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Selects and freezes a question-driven AIPerf workload, objective, load policy, and Kubernetes execution manifest for a successfully deployed Dynamo candidate.

    8.3k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Works with

Questions about Dynamo Kv Replay Parity

What does Dynamo Kv Replay Parity do?

Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…. Dynamo Kv Replay Parity is an agent skill from ai-dynamo/dynamo. Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and disaggregated handoff lifecycles.

When should I use Dynamo Kv Replay Parity?

Dynamo Kv Replay Parity fits situations like: tasks that involve LLM inference and serving; tasks that involve Refactoring.

How do I install Dynamo Kv Replay Parity in Claude Code?

Run `npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a claude-code`. Or copy the skill folder (.agents/skills/dynamo-kv-replay-parity in ai-dynamo/dynamo) into .claude/skills/dynamo-kv-replay-parity in your project. Claude Code loads it when a task matches its description.

How do I install Dynamo Kv Replay Parity in Codex?

Run `npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a codex`. Or copy the skill folder (.agents/skills/dynamo-kv-replay-parity in ai-dynamo/dynamo) into .agents/skills/dynamo-kv-replay-parity in your project. Codex loads it when a task matches its description.

Can I use Dynamo Kv Replay Parity in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dynamo-kv-replay-parity, .gemini/skills/dynamo-kv-replay-parity, .github/skills/dynamo-kv-replay-parity and .opencode/skills/dynamo-kv-replay-parity in your project.

What does Dynamo Kv Replay Parity need to run?

Going by SKILL.md and its folder, Dynamo Kv Replay Parity needs the command-line tools its instructions call (cargo and uv). Our summary lists: Python 3.

Does Dynamo Kv Replay Parity access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Dynamo Kv Replay Parity safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dynamo Kv Replay Parity use?

Dynamo Kv Replay Parity is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dynamo Kv Replay Parity use?

About 4.9k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Dynamo Kv Replay Parity?

Skills that share tags, products or a category with Dynamo Kv Replay Parity: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Debug Inference (NVIDIA/OpenShell, 16k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dynamo Kv Replay Parity?

ai-dynamo (a GitHub organization) maintains it in ai-dynamo/dynamo, which has 8,256 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 11, 2026.

Source: ai-dynamo/dynamo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.