Vllm Ascend
ascend-ai-coding/awesome-ascend-skills
vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.
Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation.
$ npx skills add mirage-project/mirage --skill dpskv3-logistic-review -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mirage-project/mirage dpskv3-logistic-review --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/dpskv3-logistic-review .claude/skills/dpskv3-logistic-review && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dpskv3-logistic-review" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/dpskv3-logistic-review into .claude/skills/dpskv3-logistic-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dpskv3-logistic-review", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/dpskv3-logistic-reviewType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mirage-project/mirage --skill dpskv3-logistic-review -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mirage-project/mirage dpskv3-logistic-review --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/dpskv3-logistic-review .agents/skills/dpskv3-logistic-review && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dpskv3-logistic-review" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/dpskv3-logistic-review into .agents/skills/dpskv3-logistic-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dpskv3-logistic-review", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mirage-project/mirage --skill dpskv3-logistic-review -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mirage-project/mirage dpskv3-logistic-review --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/dpskv3-logistic-review .cursor/skills/dpskv3-logistic-review && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dpskv3-logistic-review" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/dpskv3-logistic-review into .cursor/skills/dpskv3-logistic-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dpskv3-logistic-review", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mirage-project/mirage.git --path .claude/skills/dpskv3-logistic-review--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mirage-project/mirage --skill dpskv3-logistic-review -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mirage-project/mirage dpskv3-logistic-review --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/dpskv3-logistic-review .gemini/skills/dpskv3-logistic-review && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dpskv3-logistic-review" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/dpskv3-logistic-review into .gemini/skills/dpskv3-logistic-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dpskv3-logistic-review", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mirage-project/mirage dpskv3-logistic-reviewInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mirage-project/mirage --skill dpskv3-logistic-review -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/dpskv3-logistic-review .github/skills/dpskv3-logistic-review && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dpskv3-logistic-review" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/dpskv3-logistic-review into .github/skills/dpskv3-logistic-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dpskv3-logistic-review", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mirage-project/mirage --skill dpskv3-logistic-review -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mirage-project/mirage dpskv3-logistic-review --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mirage-project/mirage.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/dpskv3-logistic-review .opencode/skills/dpskv3-logistic-review && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dpskv3-logistic-review" agent skill from https://github.com/mirage-project/mirage/tree/mpk/.claude/skills/dpskv3-logistic-review into .opencode/skills/dpskv3-logistic-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dpskv3-logistic-review", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dpskv3-logistic-reviewAudit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation.
Dpskv3 Logistic Review is an agent skill from mirage-project/mirage. Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation. Use after structural changes to python/mirage/mpk/models/deepseekv3/builder.py, demo/deepseekv3/demo.py, or any MLA / MoE / MTP task in python/mirage/mpk/persistentkernel.py, src/kernel/taskregister.cc, and include/mirage/persistentkernel/tasks/blackwell/mla.cuh / moe.cuh. Produces a structured drift report so the change reviewer can confirm the math/topology is still equivalent.
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with DeepSeek, Python and vLLM. The repository describes itself as: Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f9eb70c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dpskv3 Logistic Review loads about 4.6k tokens when it runs. Until then it costs about 139 tokens; SKILL.md has 1,966 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mirage-project/mirage at commit f9eb70c, republished under its Apache-2.0 licence (© mirage-project). 1,966 words, ~4,616 tokens.
.claude/skills/dpskv3-logistic-review/SKILL.md (or your agent's skills folder).This skill produces a drift report that compares the current MPK DeepSeek V3 demo chain against vLLM's reference implementation, flagging any places where the two diverge in math, KV-cache topology, weight layout, or scheduling. Use it whenever someone (Claude, Codex, human) edits the demo / builder / MLA / MoE / MTP path and you need to confirm the change preserves correctness.
The skill is read-only — it audits and reports, it does not edit. Apply fixes separately, then re-run the skill.
python/mirage/mpk/models/deepseek_v3/builder.py,
demo/deepseek_v3/demo.py, or python/mirage/mpk/persistent_kernel.py's
MLA / MoE / MTP / KV-gather wrappers.include/mirage/persistent_kernel/tasks/blackwell/mla_*.cuh,
moe_*.cuh, or linear_fp8_*.cuh.src/kernel/graph.cc or
src/kernel/task_register.cc for MLA / MoE / MTP.mpk.This skill is not a substitute for scripts/regression_test.sh —
they are complementary. Regression catches "does it run and produce
output?"; this skill catches "is what it produces actually equivalent
to the reference implementation?".
Every audit pass should cover four pillars in order. Skip a pillar only when the change being reviewed obviously cannot affect it.
demo.py weight load, FP8 absorption,
per-rank shard, MTP-layer-included absorption, weight-cache load.builder.py graph construction, tensor
lifetime, KV cache layout, dual-dispatch gates.runtime.cc,
per-iteration prepare_next_batch, AllReduce/sync-counter ordering.Each pillar has a checklist below. Cite file:line for every claim
in the report.
Reference: vLLM's DeepSeek V2 model loader
vllm/model_executor/models/deepseek_v2.py — DeepseekV2ForCausalLM
load + weight absorption.vllm/v1/spec_decode/eagle/speculator.py — MTP-specific weight
loading (eh_proj, enorm, hnorm, shared lm_head).MPK chain to audit:
demo/deepseek_v3/demo.py lines ~600-760 (selective load + absorption):absorb_layers includes num_layers (the MTP layer at idx 61) when
args.mtp > 0. Double-check on every audit.attn key in absorb_layers:q_b_proj.weight becomes the absorbed fused 576-d Q
(per absorb_kv_into_q); written back to state_dict[q_key].q_b_nope.weight
([H*128, q_lora], unabsorbed nope) and q_b_pe.weight
([H*64, q_lora]).kv_b_proj.weight is split into kv_b_k.weight ([H*128, kv_lora])
and kv_b_v.weight ([H*128, kv_lora]).o_proj is preserved as o_proj_original for the
unabsorbed prefill output; the absorbed o_proj (post-W_UV
fold-in) is what stays as o_proj.weight.pad_token_id = eos_token_id fallback in qwen3 demo (orthogonal).Audit checklist:
q_b_nope, q_b_pe, kv_b_k, kv_b_v shapes match what
the chunked-prefill kernel expects ([H * d_per_head, q_lora])?
See mla_prefill_tp8_chunked_sm100.cuh for the consumer.q_b_proj shape match what the absorbed
decode + absorbed prefill kernels expect ([H*576, q_lora])?args.mtp > 0, does layer 61 also get q_b_nope /
q_b_pe / kv_b_k / kv_b_v? Per demo.py:644
absorb_layers.append(num_layers).demo.py's SHARD_RULES regex list match
the corresponding rules in builder.py for any new parallel
linear?model{rank}-mp{world_size}.safetensors (the
DeepSeek-V3-Demo-mp8 path) is consistent with the absorb path —
meaning the same weight set is produced regardless of which
load path is taken?Common drift:
demo.py shard
rules and builder.py's SHARD_RULES.absorb_layers when adding a new layer that
needs the absorbed split (e.g., a second MTP layer).Reference: vLLM's DeepseekV2DecoderLayer (per-layer pattern)
and EagleSpeculator (MTP forward path), specifically:
vllm/model_executor/models/deepseek_v2.py:1100-1200 —
decoder layer forward.vllm/v1/worker/gpu/spec_decode/eagle/speculator.py:374-420 —
MTP first-pass + draft loop.MPK chain to audit (builder.py):
__init__ and _new_intermediate_tensors (line 811): tensor
allocation. Critical buffers:q_a_out, q_nope_pe (fused 576-d), q_nope, q_pe (split),
c_latent_out, k_pe_out, ckv_sep, kpe_sep,
prefill_k_nope, prefill_v, attn_out, attn_unabsorbed,
mla_partial_o, mla_partial_lse, contiguous_kv._use_prefill = mbt > 8 — chunked-prefill gate._direct_paged_decode_kv — eligibility for skipping the dense
KV gather copy (TP=1 single-request, TP=2, TP=4; TP=8 disabled)._build_mla_attention_layer (line 1190) — main model MLA._build_mla_attention_layer_with_prefix (line 2114) — MTP MLA.
Note: use_mtp_prefill_attention flag at line 2144 should be
True (post-fix f526c7ab). If you find False, that's a
regression._build_moe_mlp (line 1660) and _build_moe_mlp_with_prefix
(line 2463) — routed-experts group GEMM + topk + scatter +
mul_sum_add. Watch w13 (dim=1) + w2 (dim=2) shard directions._build_dense_mlp (line 1604) — gate/up + down with residual
fusion._build_mtp_layer (line 2654) — MTP draft loop. Step-0 input is
mtp_step0_input_tokens (shifted prompt + main argmax tail);
steps 1+ are autoregressive draft tokens.build_from_dict (line 3154) — top-level orchestration.Audit checklist:
_use_prefill = mbt > 8. Demo readme says >=32; trust the
code.use_mtp_prefill_attention = True at builder.py:2144. If
regressed to False, MTP's prefill attention is silently
skipped → garbage hidden states._direct_paged_decode_kv eligibility includes the TP variant
under test. TP=8 stays on dense gather until the hang is fixed.mla_kv_gather* writes to the cache the subsequent
attention task reads. Mismatched buffers = empty KV history.mla_prefill_tp8_chunked with unabsorbed
Q/K/V; MTP uses mla_prefill_absorbed with fused 576-d Q.mla_mtp_decode_tp{1,2,4,8}_layer) are
registered alongside prefill so the dual-dispatch runtime gate
can route by Q_LEN. Decode kernel returns early on Q_LEN > 8.mla_mtp_decode_tp*_reduce_layer or
mla_reduce_layer) is registered iff num_splits > 1. Single
split uses attn_out directly.linear_fp8_with_residual or
explicit allreduce_layer(residual=...)) — does TP > 1
external AllReduce only when fuse_residual is False?build_from_dict:3206-3231).main_hidden_states
(target's final hidden state), step-1+ is self.mtp_x from
previous draft.mtp_build_embed_input_layer runs once per MPK iteration
before the draft loop, populating shifted prompt tokens.Common drift:
_use_prefill threshold in builder without updating
the runtime Q_LEN gate in the kernel wrappers.MPK paths:
python/mirage/mpk/persistent_kernel.py. Each
layer-method registers a tb_graph and calls
kn_graph.register_task(tb_graph, "<task_name>", params). The
wrapper signature decides what gets bound to input_ptrs[].src/kernel/graph.cc::Graph::register_task matches
the task-name string to a TASK_* enum + register_*_task
function.src/kernel/task_register.cc::register_<task>_task emits
the C++ snippet that runs in the per-worker dispatch switch. The
snippet:task_desc->task_metadata.{request_id, kv_idx, merge_task_offset}
set by the runtime scheduler (per runtime.cc::register_mugraph).runtime_config.{qo_indptr_buffer, paged_kv_indptr_buffer, paged_kv_last_page_len_buffer, request_ids, step, prompt_length}
to compute current Q_LEN, S, KV pointer offsets.include/mirage/persistent_kernel/tasks/blackwell/<task>.cuh.*.cuh headers under tasks/blackwell/. SM100a
only — Hopper variants live under tasks/hopper/ and are
separate.Audit checklist:
tb_graph.new_input(...) count matches the
tuple in graph.cc ((num_inputs, num_outputs, TASK_*, variant_id)).
Mismatch → "Invalid global read" at runtime.store_in_dmem=True for an output (the "MPK
convention"), the tuple says (N+1, 0, ...) not (N, 1, ...).task_register.cc codegen reads input pointers from the right
indices and computes per-request offsets correctly:
qo_fp_ * (H * D) for Q, bi_ * MPK_MAX_SEQ_LENGTH * D for
KV (per-request stride), etc.Q_LEN <= 8; decode returns when
Q_LEN > 8. Off-by-one is a silent miscompare.runtime.cc::register_mugraph's metadata setup matches the
kernel's expected mapping. e.g., TASK_MLA_KV_GATHER_SM100
uses request_id = bid.x; TASK_MLA_PREFILL_TP8_CHUNKED_SM100
uses request_id = bid.x (head), kv_idx = bid.y (q_block),
merge_task_offset = bid.z (batch).runtime_header.h (enum) +
task_register.h (declaration) + task_register.cc (impl) +
graph.cc (dispatch) + runtime.cc (task_type_to_name +
metadata handler if non-default) + task_header.cuh (include) +
Python wrapper. Check ALL of these.grid_dim in the Python wrapper matches what the kernel and
the metadata setup expect. Wrong grid_dim silently produces
wrong pointer offsets per task.Invalid __shared__ write at
runtime, not compile time.Common drift:
q_nope_pe (fused) when the wrapper
passes q_nope (split), or vice versa.runtime.cc::register_mugraph's
metadata switch — request_id defaults to 0 silently..cuh file's TMA descriptor signature without updating
the corresponding task_register.cc snippet.tcgen05.alloc + bar.sync 1, 128 pattern from the existing
TP{2,4,8} kernels.MPK paths:
include/mirage/persistent_kernel/persistent_kernel.cuh — worker
loop, scheduler, NVSHMEM team setup, prepare_next_batch.src/kernel/runtime.cc::print_task_graph — emits the megakernel
test.cu, schedules tasks via dependency graph.python/mirage/mpk/persistent_kernel.py::compile — invokes nvcc,
loads the built .so via ctypes / Python ext.Audit checklist:
MPK_MAX_NUM_BATCHED_TOKENS, MPK_MAX_NUM_BATCHED_REQUESTS,
MPK_MAX_SEQ_LENGTH, MPK_PAGE_SIZE, MPK_MAX_NUM_PAGES
compile-time constants reflect the demo CLI args.-rdc=true is the default for NVSHMEM builds. MPK_RDC_FALSE=1
is an escape hatch only./usr/local/cuda-13.2/bin/nvcc. CUDA 12.8
segfaults on the post-PR674 megakernel
(project_cuda128_nvcc_segfault.md).nvshmem_team_split per layer-shape.prepare_next_batch advances step[], qo_indptr_buffer[],
paged_kv_indptr_buffer[], paged_kv_last_page_len_buffer[]
consistently across MPK iterations.Common drift:
prepare_next_batch's
dependency calculation, breaking the scheduler's task ordering.MAX_WORKER_PER_SCHEDULER or MAX_NUM_WORKERS without
re-evaluating per-worker SMEM budget.Establish baseline. Run scripts/regression_test.sh against
the unmodified code. Capture the perfetto traces and the summary
PASS/FAIL + latencies.
Apply the change under review. Do not skip this — the audit must be against the actual diff, not against an imagined diff.
Diff the perfetto trace. Compare task_type_name counts and
per-task durations between baseline and modified runs. New tasks
appearing or expected tasks missing both warrant investigation.
Walk the four pillars in order. For each item in each
checklist, either confirm file:line shows the expected pattern
or open a finding with severity:
definitely-broken — math/topology is wrong, regression test
PASS does not save you.suspicious — divergence from vLLM that may or may not be
intentional; needs author confirmation.probably-fine — divergence with a documented justification.unclear — needs more investigation.Cross-check against vLLM. Where divergent, name the vLLM file and line and explain whether the difference is performance-driven (acceptable for now), correctness-driven (must align), or just a different API surface (orthogonal).
Report. Produce a structured drift report in this format:
# DeepSeek V3 Logistic Review — <commit-or-branch>
## Change under review
<one-line summary; commit hash if any>
## Summary
<X PASS, Y SUSPICIOUS, Z BROKEN>
## Findings
### Pillar 1 — Weight pipeline
- [PASS] <one-line description, file:line>
- [SUSPICIOUS] <description, file:line, vLLM ref>
...
### Pillar 2 — Builder topology
...
### Pillar 3 — Task / kernel wiring
...
### Pillar 4 — Runtime + scheduler
...
## Cross-cuts
<findings that span multiple pillars, e.g., a layout change that
affects both weight-load and kernel>
## Recommended next actions
<ordered list, smallest-fix first>Do NOT edit while in audit mode. The skill is read-only by design; mixing audit and patching makes it impossible to know what fixed what. Hand the report to the user (or to a separate patch session) and let them apply fixes.
These are the deltas from vLLM that exist in the current codebase and have been verified. Note them in every audit report under "Probably-fine, deferred":
mtp_ckv_kpe_cache_tensor) vs vLLM's
shared KV cache. Functionally equivalent if the cache is correctly
populated; sharing would simplify the code but is a refactor.mla_unified_layer framework exists but unused. The unified
layer would fold prefill + decode into one task with internal
Q_LEN-based dispatch (vLLM-style). Not yet wired into the DeepSeek
builder; current dual-dispatch with runtime Q_LEN gates is the live
path.world_size == 8.MPK_MLA_TP4_V_SPLITS
default switches between 2 and 8 based on max_seq_length.If your audit report says "diverges from vLLM" for any of these, mark
them probably-fine, deferred, not suspicious.
Write the audit report to a fresh file (don't overwrite an old one unless the user asks). Suggested path:
outputs/dpskv3_review_<YYYYMMDD>_<HHMMSS>.mdoutputs/ is gitignored, so the report is local-only. Quote
file:line references and the relevant snippet so the user can
verify without re-deriving the audit. Length budget: 1000-2000 words
for a focused-change audit; longer reports rarely get read.
© mirage-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/dpskv3-logistic-review of mirage-project/mirage.
Open the folder on GitHubat commit f9eb70c
Dpskv3 Logistic Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dpskv3 Logistic Review this skillmirage-project/mirage | 2.5k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Vllm Ascendascend-ai-coding/awesome-ascend-skills | 174 | — | ~2.7k | Automated safety check: Pass | None | |
| External Gitcode Ascend Vllm Ascend Deployascend-ai-coding/awesome-ascend-skills | 174 | — | ~1.2k | Automated safety check: Pass | None | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| One EvalOpenDCAI/One-Eval | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Add Modelguoqingbao/xinfer | 334 | — | ~4.2k | Automated safety check: Notes | MIT |
ascend-ai-coding/awesome-ascend-skills
vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.
ascend-ai-coding/awesome-ascend-skills
昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持…
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
OpenDCAI/One-Eval
驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。
guoqingbao/xinfer
Adapt and port new LLM model architectures to this xinfer project.
Orchestra-Research/AI-Research-SKILLs
Uses the Outlines library to constrain model output to a JSON schema, Pydantic model, regex or fixed set of choices when running local models.
mirage-project/mirage
Runtime-V2 performance-iteration workflow. An agent skill from mirage-project/mirage.
mirage-project/mirage
Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).
mirage-project/mirage
A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…
mirage-project/mirage
Build or run a FAITHFUL in-MPK per-task latency gate (slowCTA at the production grid + cos) for a DeepSeek-V3 MPK decode kernel or shape.
mirage-project/mirage
A skill your agent uses when a batch of env-gated (ifdef MPKDSV3 / os.environ-controlled, default-OFF) MPK optimization levers needs to be consolidated into a single clean code path for a PR…
mirage-project/mirage
Guide for using MPK test mode to unit-test individual layers or multi-layer pipelines through the full compilation pipeline.
Categories
Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation. Dpskv3 Logistic Review is an agent skill from mirage-project/mirage. Audit the DeepSeek V3 MPK demo + builder chain end-to-end and confirm logical equivalence with vLLM's reference implementation.
Dpskv3 Logistic Review fits situations like: tasks that involve LLM inference and serving.
Run `npx skills add mirage-project/mirage --skill dpskv3-logistic-review -a claude-code`. Or copy the skill folder (.claude/skills/dpskv3-logistic-review in mirage-project/mirage) into .claude/skills/dpskv3-logistic-review in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mirage-project/mirage --skill dpskv3-logistic-review -a codex`. Or copy the skill folder (.claude/skills/dpskv3-logistic-review in mirage-project/mirage) into .agents/skills/dpskv3-logistic-review in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mirage-project/mirage --skill dpskv3-logistic-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dpskv3-logistic-review, .gemini/skills/dpskv3-logistic-review, .github/skills/dpskv3-logistic-review and .opencode/skills/dpskv3-logistic-review in your project.
SKILL.md names no scripts, command-line tools or credentials: Dpskv3 Logistic Review is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Dpskv3 Logistic Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Dpskv3 Logistic Review: Vllm Ascend (ascend-ai-coding/awesome-ascend-skills, 174 stars), External Gitcode Ascend Vllm Ascend Deploy (ascend-ai-coding/awesome-ascend-skills, 174 stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and One Eval (OpenDCAI/One-Eval, 165 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mirage-project (a GitHub organization) maintains it in mirage-project/mirage, which has 2,545 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 7, 2026.
Source: mirage-project/mirage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.