SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…
$ npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ai-dynamo/dynamo dynamo-kv-replay-parity --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ai-dynamo/dynamo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/dynamo-kv-replay-parity .claude/skills/dynamo-kv-replay-parity && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dynamo-kv-replay-parity" agent skill from https://github.com/ai-dynamo/dynamo/tree/main/.agents/skills/dynamo-kv-replay-parity into .claude/skills/dynamo-kv-replay-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dynamo-kv-replay-parity", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ai-dynamo/dynamo/tree/main/.agents/skills/dynamo-kv-replay-parityType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ai-dynamo/dynamo dynamo-kv-replay-parity --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-dynamo/dynamo.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/dynamo-kv-replay-parity .agents/skills/dynamo-kv-replay-parity && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dynamo-kv-replay-parity" agent skill from https://github.com/ai-dynamo/dynamo/tree/main/.agents/skills/dynamo-kv-replay-parity into .agents/skills/dynamo-kv-replay-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dynamo-kv-replay-parity", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ai-dynamo/dynamo dynamo-kv-replay-parity --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-dynamo/dynamo.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/dynamo-kv-replay-parity .cursor/skills/dynamo-kv-replay-parity && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dynamo-kv-replay-parity" agent skill from https://github.com/ai-dynamo/dynamo/tree/main/.agents/skills/dynamo-kv-replay-parity into .cursor/skills/dynamo-kv-replay-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dynamo-kv-replay-parity", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ai-dynamo/dynamo.git --path .agents/skills/dynamo-kv-replay-parity--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ai-dynamo/dynamo dynamo-kv-replay-parity --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-dynamo/dynamo.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/dynamo-kv-replay-parity .gemini/skills/dynamo-kv-replay-parity && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dynamo-kv-replay-parity" agent skill from https://github.com/ai-dynamo/dynamo/tree/main/.agents/skills/dynamo-kv-replay-parity into .gemini/skills/dynamo-kv-replay-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dynamo-kv-replay-parity", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ai-dynamo/dynamo dynamo-kv-replay-parityInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ai-dynamo/dynamo.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/dynamo-kv-replay-parity .github/skills/dynamo-kv-replay-parity && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dynamo-kv-replay-parity" agent skill from https://github.com/ai-dynamo/dynamo/tree/main/.agents/skills/dynamo-kv-replay-parity into .github/skills/dynamo-kv-replay-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dynamo-kv-replay-parity", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ai-dynamo/dynamo dynamo-kv-replay-parity --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai-dynamo/dynamo.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/dynamo-kv-replay-parity .opencode/skills/dynamo-kv-replay-parity && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dynamo-kv-replay-parity" agent skill from https://github.com/ai-dynamo/dynamo/tree/main/.agents/skills/dynamo-kv-replay-parity into .opencode/skills/dynamo-kv-replay-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dynamo-kv-replay-parity", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dynamo-kv-replay-parityRuns deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…
Dynamo Kv Replay Parity is an agent skill from ai-dynamo/dynamo. Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and disaggregated handoff lifecycles. It is used when validating replay refactors, routing changes, scheduler-event changes, or performance-sensitive offline simulation changes against a baseline revision.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/internal-polynomial-golden-points.md`).
It sits in AI & LLM Engineering, covering LLM inference and serving and Refactoring. It works with SGLang and vLLM. The repository describes itself as: A Datacenter Scale Distributed Inference Serving Framework. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b208989. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
cargouvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dynamo Kv Replay Parity loads about 4.9k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 103 tokens; SKILL.md has 2,535 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ai-dynamo/dynamo at commit b208989, republished under its Apache-2.0 licence (© ai-dynamo). 2,535 words, ~4,854 tokens.
.claude/skills/dynamo-kv-replay-parity/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Compare two revisions of Dynamo offline KV-aware replay using deterministic virtual-time reports and statistically paired real-time measurements. Use the existing replay and benchmark harnesses; extend them only when a required signal is unavailable.
This campaign intentionally excludes round-robin routing. It also does not require a same-timestamp event-only progress scenario. Validate those concerns with focused tests outside this skill when a change specifically affects them.
Resolve and record before running anything:
Use the project root .venv/bin/python for Python analysis and uv pip for any approved
installation. Do not compare against moving branches or reuse a binary after changing its
checkout. Do not describe the configuration as "all features"; record explicit feature
names. Feature names can differ between dynamo-mocker and dynamo-bench. The current
campaign is native-G1-only; do not add removed KVBM replay features or runtime arguments
to make a historical command line build.
dynamo-bench harness on Linux, build one release artifact per
revision with exactly replay-bench and no default features.
--canonical-reports-jsonl requires replay-bench, which also selects the seeded
router..text size.Build the current benchmark artifact with:
cargo build --release -p dynamo-bench --no-default-features \
--features replay-bench --bench offline_replay_benchNever build while collecting performance samples.
Treat one engine/topology/frozen-configuration comparison row as the unit of node placement. When scheduler capacity permits, allocate multiple nodes and assign independent rows to them. The nodes do not need to be homogeneous because no individual comparison crosses nodes. Keep the baseline, candidate, all determinism repetitions, and lifecycle evidence for one row on the same node. Use the same prebuilt revision artifacts and inputs throughout the campaign, and record the node characteristics and CPU placement for every row.
Native replay is normally single-core: pin each process to one physical core after confirming that assumption during preflight. Keep the CPU-set size, NUMA placement, and affinity identical between baseline and candidate for a row.
For correctness, run the baseline and candidate repetitions for a row concurrently when its node has sufficient resources. Pin concurrent processes to disjoint CPU sets, give each process separate output paths, and ensure they do not share mutable state. Keep every repetition in a separate process even when several repetitions run at the same time.
For performance, keep the entire row's warmups and 60-pair measurement series on its assigned node. Execute only one performance invocation at a time on that node with a fixed CPU placement, and keep each randomized baseline/candidate pair adjacent. Different performance rows may run concurrently on different nodes, but do not pool one row's pairs across heterogeneous nodes. Do not run another replay, build, profiler, or unrelated workload concurrently on a node collecting performance samples.
Multi-node and within-node correctness parallelism are campaign throughput optimizations; they must not change workload concurrency or weaken performance isolation.
Require the harness to control every known entropy source:
(worker_id, dp_rank) before deterministic selection;/summary/wall_time_ms, /summary/processed_tokens_per_s, and
/summary/processed_output_tokens_per_s.This exclusion list is closed: include every other field, and stop for review before adding
another exclusion. Make semantic fields such as request UUIDs deterministic rather than
dropping them. Require parity output metadata to report replay_bench: true; otherwise the
seeded routing contract was not active.
Run baseline twice and candidate twice in separate processes. Each revision must produce one unique canonical digest. This is an entropy-leak check, not a statistical trial. If a revision is internally unstable, stop and diagnose it; do not increase repetitions and average the outputs.
Prefer a fixed contiguous 5,000-request Mooncake window over many parity repetitions. Preserve arrival order and prefix locality. Record the starting offset, request count, and checksum. Do not randomly sample rows, duplicate a shorter trace, or silently claim a 5,000-request campaign when fewer usable requests exist.
The committed lib/bench/testdata/mooncake_trace_1000.jsonl fixture is suitable for a
quick harness preflight, not the authoritative long-corpus campaign.
Qualify and freeze one configuration per comparison row, or per explicitly named configuration family when rows genuinely share every relevant parameter. For each frozen configuration, prove that it exercised every applicable path:
Use coverage counters, lifecycle traces, or report evidence rather than inferring these paths from successful completion. Target the bounded pressure band recorded for the qualified seed. When no seed exists, start with one to three pressure events to prove the lifecycle. For throttle-oriented disaggregated vLLM and SGLang seeds, target 10 to 20 fully readmitted pressure events so repeated scheduling is exercised. Zero means the edge was not exercised; repeated preempt/re-admit cycling, a rapidly growing pressure count, or failure to advance virtual time invalidates the fixture. Tune capacity or concurrency minimally and identically for baseline and candidate within the row or family, and back off rather than accepting a pressure flood. Never tune the revisions separately.
The offline replay CLI leaves the router queue threshold unset by default. In that mode, all route decisions are immediate and zero queued placements are expected; engine-side scheduler waiting is not router queue coverage. If queue lifecycle coverage is required, enable it through an explicit queue-capable harness or forced fixture and record the exact threshold. Account for every route decision, but never relabel scheduler waiting as queued placement.
For the canonical 5,000-row Mooncake window with the internal polynomial mocker, read internal-polynomial golden points before searching for capacity edges. Treat those configurations and observed counters as suggested starting points, not universal constants or substitutes for qualification on the pinned baseline. Requalify one frozen configuration identically on both revisions.
Use the reference's expected preemption, retraction, reuse, worker, and handoff signals as drift detectors. Treat queue counts as expected signals only when queueing was explicitly enabled. Matching lifecycle counts do not waive an unstable canonical digest; stop correctness and performance work for that row until both internal determinism and cross-revision parity are established.
Run the 5,000-request corpus for this authoritative matrix. Treat both disaggregated rows as the primary scheduler and handoff requalification. Keep the aggregated rows as secondary parity coverage for the corresponding engine semantics.
| Engine semantics | Topology | Memory path | Routing |
|---|---|---|---|
| vLLM pass-end | Aggregated | Native G1 | KV-aware |
| vLLM pass-end | Disaggregated | Native G1 | KV-aware |
| SGLang pass-end | Aggregated | Native G1 | KV-aware |
| SGLang pass-end | Disaggregated | Native G1 | KV-aware |
KVBM G2-G4 replay was removed from the current mocker. It is not an unsupported row in this matrix and must not be emulated with ignored flags. A campaign that needs historical KVBM parity must pin a historical revision together with the matching historical skill protocol. For each current row:
TensorRT-LLM can be additional smoke coverage, but it is not a substitute for either authoritative engine semantic.
Byte mismatch is a review gate, not an instruction to preserve incorrect behavior. Allow an intentional mismatch only when the candidate is demonstrably more faithful to the specified engine, scheduler, or routing semantics.
For every proposed exception, record:
Use PASS_WITH_SEMANTIC_EXCEPTIONS only when every byte difference is covered by such a
record. Unexplained, incidental, or merely convenient differences fail. Do not patch the
candidate back to known-wrong behavior just to obtain identical bytes.
Use small deterministic fixtures only for required paths the long corpus cannot reliably trigger:
Assert bounded preemption, continued virtual-time progress, and the lifecycle itself. A final-completion smoke test does not prove preemption or queueing occurred.
Measure every supported row from Stage 4 with its frozen configuration. The authoritative
gate reuses the replay-bench artifacts from byte parity, so routing selection is seeded
and matched between revisions. Require timing metadata to report replay_bench: true.
Describe the result as seeded matched-routing replay-loop parity, not production-routing
performance. Production-routing timing is outside the default campaign because it requires
a separate build without replay-bench; add that campaign only when the user explicitly
requests it.
The primary metric is replay execution time. Start its timer after trace normalization,
workload construction, and engine/runtime preparation, immediately before
prepared.run(...); stop it immediately after that call returns and before collector
finalization or report aggregation. Emit this value as replay_execution_ms. Record setup
and end-to-end time separately as diagnostics. If the harness does not expose this exact
timing boundary, extend it before running the campaign; do not substitute the existing
broader wall_time_ms.
Stage binaries and trace inputs on node-local storage and verify their checksums before
warmups. File transfer is campaign setup, not a sample. For each measured invocation, pass
--iterations 1 and a unique node-local --timings-jsonl path. Do not pass
--canonical-reports-jsonl, --report-json, or another full-report output option in a
gated performance invocation. Run one fresh process per sample so lazy per-request capture,
report serialization, and in-process iteration state cannot contaminate the metric.
For each row:
r_i = candidate_replay_execution_ms / baseline_replay_execution_ms regardless
of run order.INCONCLUSIVE or use a predeclared dependence-aware method. Do not apply the
order-statistic gate.1.05.1.05.INCONCLUSIVE; do not add samples adaptively and claim the original
confidence level.Never remove an observation because its value looks like an outlier. Retain every attempted sample and its process status. A replay error, assertion failure, malformed timing record, or other product failure fails or blocks the row; it is not a discardable sample. Only a predeclared environmental condition, such as scheduler eviction, node-health failure, affinity violation, or independently detected competing load, may invalidate a sample. Invalidate both members of that pair, record the evidence and reason, and rerun the complete pair with the same arm order. Never retry or replace a sample silently. A paired bootstrap of log ratios may be reported as a secondary effect-size diagnostic, but it does not decide the gate.
Also fail release binary or .text growth above 5% until the added footprint is explained
and narrowed.
When a statistically meaningful regression exceeds the accepted window:
Prefer an available Samply workflow skill when one is installed. Do not profile concurrently with builds or unrelated load.
Use exactly one semantic result:
PASS: all authoritative canonical outputs match;PASS_WITH_SEMANTIC_EXCEPTIONS: every mismatch is an evidenced improvement covered by
a regression test; orFAIL: any unexplained mismatch, missing lifecycle evidence, or unstable revision.Report performance independently as pass, fail, or inconclusive. A semantic exception does not waive an unexplained performance regression.
The final report must include:
.text sizes;Delete temporary full reports, traces, binaries, profiler captures, patches, and worktrees after recording the required evidence. Retain full outputs only for unresolved mismatches or performance investigations. Remove task-created Cargo targets when disk pressure matters, but never delete unrelated caches or worktrees.
© ai-dynamo, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in .agents/skills/dynamo-kv-replay-parity of ai-dynamo/dynamo.
Open the folder on GitHubat commit b208989
Dynamo Kv Replay Parity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dynamo Kv Replay Parity this skillai-dynamo/dynamo | 8.3k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Debug InferenceNVIDIA/OpenShell | 16k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| One EvalOpenDCAI/One-Eval | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
NVIDIA/OpenShell
Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
OpenDCAI/One-Eval
驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
ai-dynamo/dynamo
Create self-contained interactive HTML code-review dashboards from GitHub or GitLab pull requests, checked-out branch diffs, or supplied unified diffs, with correctness and safe-to-merge scores…
ai-dynamo/dynamo
Knowledge of Fern's built-in MDX component library (accordions, callouts, cards, steps, tabs, code blocks, API-reference snippets, and more) for authoring docs pages.
ai-dynamo/dynamo
Knowledge of Fern's site-level navigation and structure configuration — how a docs site is organized in docs.yml (and product/version .yml files) using sections, pages, folders, tabs, tab variants…
ai-dynamo/dynamo
Drives persistent Claude Code, Codex, or OpenCode agent sessions through a Dynamo OpenAI/Anthropic-compatible endpoint over Agent Client Protocol (ACP).
ai-dynamo/dynamo
Benchmark and profile the Dynamo frontend (dynamo.frontend HTTP + tokenizer + KV router) against mock workers (dynamo.mocker).
ai-dynamo/dynamo
Selects and freezes a question-driven AIPerf workload, objective, load policy, and Kubernetes execution manifest for a successfully deployed Dynamo candidate.
Categories
Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…. Dynamo Kv Replay Parity is an agent skill from ai-dynamo/dynamo. Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and disaggregated handoff lifecycles.
Dynamo Kv Replay Parity fits situations like: tasks that involve LLM inference and serving; tasks that involve Refactoring.
Run `npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a claude-code`. Or copy the skill folder (.agents/skills/dynamo-kv-replay-parity in ai-dynamo/dynamo) into .claude/skills/dynamo-kv-replay-parity in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a codex`. Or copy the skill folder (.agents/skills/dynamo-kv-replay-parity in ai-dynamo/dynamo) into .agents/skills/dynamo-kv-replay-parity in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-dynamo/dynamo --skill dynamo-kv-replay-parity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dynamo-kv-replay-parity, .gemini/skills/dynamo-kv-replay-parity, .github/skills/dynamo-kv-replay-parity and .opencode/skills/dynamo-kv-replay-parity in your project.
Going by SKILL.md and its folder, Dynamo Kv Replay Parity needs the command-line tools its instructions call (cargo and uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Dynamo Kv Replay Parity is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Dynamo Kv Replay Parity: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Debug Inference (NVIDIA/OpenShell, 16k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ai-dynamo (a GitHub organization) maintains it in ai-dynamo/dynamo, which has 8,256 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 11, 2026.
Source: ai-dynamo/dynamo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.