Agent skill

Qwen Mtp Gguf

by R6410418 in R6410418/Jackrong-llm-finetuning-guide

Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

MITAuto-check passedAI & LLM Engineering

Install Qwen Mtp Gguf

skills CLI
$ npx skills add R6410418/Jackrong-llm-finetuning-guide --skill qwen-mtp-gguf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install R6410418/Jackrong-llm-finetuning-guide qwen-mtp-gguf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/R6410418/Jackrong-llm-finetuning-guide.git skills-src && mkdir -p .claude/skills && cp -r skills-src/qwen-mtp-gguf .claude/skills/qwen-mtp-gguf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qwen-mtp-gguf
GitHub stars
1.7k
Token cost
~1.7k tokens
SKILL.md length
659 words
Files
18 (incl. scripts, references)
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

  • Works in 6 steps: Identify the target model and matching… → Preflight the environment, disk, RAM,… → Ask the user to choose upload strategy… → …
  • Another coding agent needs to inspect a users machine
  • SKILL.md covers Operating Principle, Required User Inputs, Upload Strategy Choice and Environment Bootstrap, plus 5 more sections
  • Runs Python and Shell scripts from its folder; calls python3 and bash; needs HF_TOKEN

What it does

Qwen Mtp Gguf is an agent skill from R6410418/Jackrong-llm-finetuning-guide. Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release. Use when Codex or another coding agent needs to inspect a user's machine, estimate disk/RAM requirements from Hugging Face model sizes and requested quant formats, bootstrap llama.cpp and Python dependencies, extract MTP heads from a matching official/base Qwen model, merge them into a fine-tuned or target safetensors model, run local HF/GGUF smoke tests with Qwen chat formatting, quantize to GGUF, and optionally upload or…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files, including scripts and reference files (for example `README.md`, `agents/openai.yaml` and `docs/Qwen-MTP-GGUF-Agent-Usage.md`).

It sits in AI & LLM Engineering, covering QA and bug reports, LLM inference and serving and Model hubs and datasets. It works with llama.cpp, Qwen, Hugging Face and Python. The licence is MIT.

When your agent uses it

  • Another coding agent needs to inspect a users machine
  • Estimate disk/RAM requirements from Hugging Face model sizes and requested quant formats
  • Bootstrap llama.cpp and Python dependencies
  • Extract MTP heads from a matching official/base Qwen model

Example prompts

  • “/qwen-mtp-gguf”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Identify the target model and matching MTP source model.
  2. Preflight the environment, disk, RAM, tools, token access, and config compatibility.
  3. Ask the user to choose upload strategy when it affects disk use.
  4. Prepare the target HF directory and inject only missing MTP/nextn tensors.
  5. Convert to a temporary F16 GGUF, smoke-test locally, then quantize.
  6. Upload with resume support, or keep local artifacts if the user chooses local-only.

What it can do on your machine

Read from SKILL.md and the folder at commit ef2b17f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python and Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Qwen Mtp Gguf loads about 1.7k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 140 tokens; SKILL.md has 659 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~140
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from R6410418/Jackrong-llm-finetuning-guide at commit ef2b17f, republished under its MIT licence (© R6410418). 659 words, ~1,676 tokens.

Download SKILL.mdSave it as .claude/skills/qwen-mtp-gguf/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
qwen-mtp-gguf
description
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release. Use when Codex or another coding agent needs to inspect a user's machine, estimate disk/RAM requirements from Hugging Face model sizes and requested quant formats, bootstrap llama.cpp and Python dependencies, extract MTP heads from a matching official/base Qwen model, merge them into a fine-tuned or target safetensors model, run local HF/GGUF smoke tests with Qwen chat formatting, quantize to GGUF, and optionally upload or resume Hugging Face releases.

Qwen MTP GGUF

Operating Principle

Run this as a staged release pipeline, not a blind conversion:

  1. Identify the target model and matching MTP source model.
  2. Preflight the environment, disk, RAM, tools, token access, and config compatibility.
  3. Ask the user to choose upload strategy when it affects disk use.
  4. Prepare the target HF directory and inject only missing MTP/nextn tensors.
  5. Convert to a temporary F16 GGUF, smoke-test locally, then quantize.
  6. Upload with resume support, or keep local artifacts if the user chooses local-only.

MTP/GGUF conversion and llama.cpp quantization do not require a GPU. GPU acceleration can make later inference tests faster, but the default smoke test should use CPU mode (-ngl 0) so it works on common machines.

Required User Inputs

Ask for missing items only when they cannot be inferred safely:

  • Target model: HF repo ID or local HF model directory.
  • Matching MTP source model: the official/base Qwen-family repo with the same architecture, size, tokenizer, and MTP layout.
  • Output mode: local-only, stream upload, or batch upload.
  • Output repo ID when uploading.
  • Quantization formats. Default full matrix: q2_k,q3_k_s,q3_k_m,q3_k_l,iq4_xs,q4_k_s,q4_k_m,q5_k_s,q5_k_m,q6_k,q8_0,bf16.

If the target and MTP source configs disagree on core architecture fields, stop and ask before continuing.

Upload Strategy Choice

When the user has not specified a strategy, explain the tradeoff briefly and ask:

  • stream: quantize one GGUF, upload it, then delete it. Lowest peak disk use and best default for large models.
  • batch: quantize everything first, then upload. Useful when network is unstable but requires much more disk.
  • local-only: prepare all GGUF files locally without uploading.

Use stream when the user asks for a full large-model release and does not care about keeping local copies.

Environment Bootstrap

Use scripts/bootstrap_qwen_mtp_env.sh when llama.cpp or Python dependencies are missing.

bash
bash scripts/bootstrap_qwen_mtp_env.sh --prefix ./qwen-mtp-env --backend cpu
source ./qwen-mtp-env/.venv/bin/activate

Backend options are cpu, cuda, metal, and vulkan. Prefer cpu unless the user explicitly wants accelerated smoke tests or already has a configured GPU toolchain.

Preflight

Always run preflight before downloading large model weights:

bash
python3 scripts/qwen_mtp_gguf_pipeline.py \
  --source-repo owner/target-qwen-finetune \
  --mtp-source-repo owner/matching-base-qwen-with-mtp \
  --output-repo owner/target-qwen-mtp-gguf \
  --work-root ./mtp-gguf-work \
  --llama-cpp ./qwen-mtp-env/llama.cpp \
  --token-env HF_TOKEN \
  --upload-strategy stream \
  --preflight-only

Review preflight_report.md before running. It reports:

  • OS, CPU cores, total RAM, detected GPU hint, Python version.
  • Required commands and packages.
  • llama.cpp converter, quantizer, and llama-cli status.
  • Target model size from HF file metadata or local files.
  • Minimal MTP shard download size.
  • Estimated F16/BF16 and quantized GGUF sizes.
  • Required free disk with a safety factor.
  • Recommended RAM and config mismatch warnings.
Show full SKILL.md (275 more words)Show less

Full Pipeline

bash
python3 scripts/qwen_mtp_gguf_pipeline.py \
  --source-repo owner/target-qwen-finetune \
  --mtp-source-repo owner/matching-base-qwen-with-mtp \
  --output-repo owner/target-qwen-mtp-gguf \
  --work-root ./mtp-gguf-work \
  --llama-cpp ./qwen-mtp-env/llama.cpp \
  --filename-prefix target-qwen-MTP \
  --token-env HF_TOKEN \
  --upload-strategy stream \
  --private \
  --smoke-test-before-upload \
  --cleanup-after-upload

The pipeline:

  • Downloads or copies the target HF model.
  • Detects existing MTP/nextn tensors and validates index mappings.
  • Downloads only the MTP source shards listed in the source index.
  • Saves mtp_heads.safetensors, updates model.safetensors.index.json, and validates all new keys.
  • Converts F16 for quantization and BF16 directly from the prepared HF directory.
  • Runs GGUF smoke tests before upload when enabled. Prefer the model's embedded chat template when llama.cpp supports it; use ChatML only as a minimal fallback for load/generation sanity checks.
  • Lists remote HF files before upload and skips completed GGUFs on resume.
  • Deletes local GGUF files only after confirmed upload when cleanup is enabled.

Smoke Tests

Use the GGUF smoke test for release validation:

bash
python3 scripts/qwen_gguf_smoke_test.py \
  --model ./mtp-gguf-work/target-qwen-MTP-GGUF/target-qwen-MTP-Q8_0.gguf \
  --llama-cli ./qwen-mtp-env/llama.cpp/build/bin/llama-cli \
  --prompt "State the capital of France in one short sentence." \
  --gpu-layers 0

Use the HF smoke test only when the machine can load the HF model:

bash
python3 scripts/qwen_hf_smoke_test.py \
  --model ./mtp-gguf-work/target-qwen-MTP-HF \
  --prompt "Write one concise sentence about MTP inference."

For Qwen-family chat formatting, use apply_chat_template(..., add_generation_prompt=True, tokenize=False) on the HF side. For GGUF, prefer the chat template embedded by llama.cpp conversion or copied from tokenizer_config.json; use raw ChatML only as a fallback smoke test, not as the quality/reasoning benchmark template.

Agent Compatibility

This skill is Codex-native, but the workflow is intentionally agent-agnostic:

  • Claude Code, Codex, OpenCode, Qwen Code, Hermes, and similar agents can run the same scripts from a shell.
  • The agent should call --preflight-only, summarize blockers, ask for the upload strategy if needed, then run the full command.
  • Keep user-specific tokens, paths, private repo names, and logs out of public PRs and model cards.

References

  • Read references/environment-and-sizing.md before changing preflight logic or resource thresholds.
  • Read references/technical-flow.md before changing extraction, injection, conversion, or upload behavior.
  • Read references/agent-integration.md when packaging this for another agent framework.
  • Read references/troubleshooting.md when conversion, tensor lookup, disk, RAM, or upload steps fail.

© R6410418, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts, references) in qwen-mtp-gguf of R6410418/Jackrong-llm-finetuning-guide.

  • SKILL.md
  • .gitignore
  • LICENSE
  • README.md
  • agents/openai.yaml
  • docs/Qwen-MTP-GGUF-Agent-Usage.md
  • docs/Qwen-MTP-GGUF-Pipeline-Guide.md
  • docs/Qwen-MTP-GGUF-Skill-README.md
  • env.example
  • jobs.example.json
  • references/agent-integration.md
  • references/environment-and-sizing.md
  • references/technical-flow.md
  • references/troubleshooting.md
  • scripts/bootstrap_qwen_mtp_env.sh
  • scripts/qwen_gguf_smoke_test.py
  • scripts/qwen_hf_smoke_test.py
  • … and 1 more

Open the folder on GitHubat commit ef2b17f

Compare with similar skills

Qwen Mtp Gguf next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Qwen Mtp Gguf compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Qwen Mtp Gguf this skillR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Test Modelguoqingbao/xinfer334—~2.6kAutomated safety check: PassMIT
Aqua Model Lifecycleoracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.0
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0

Similar skills

  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Test Model

    guoqingbao/xinfer

    Test LLM models served by xinfer for correctness, output quality, and performance.

    334 GitHub stars~2.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Aqua Model Lifecycle

    oracle/accelerated-data-science

    Official

    Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

    125 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Edge Bringup

    exeex/edge-cores

    Prepare a macOS or Ubuntu machine for edge-e3 development, diagnose missing Verilator/LLVM/Python dependencies, initialize the public repository, and answer or act on the example prompts in the root…

    110 GitHub stars~1.7k tokensUpdated 15 days ago
    AI & LLM EngineeringAuto-check: notes

More from R6410418/Jackrong-llm-finetuning-guide

  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 3 mo ago
    Auto-check passed
  • GitHub Sync Gate

    R6410418/Jackrong-llm-finetuning-guide

    Enforce this repository's local-review-first GitHub sync policy.

    1.7k GitHub stars~316 tokensUpdated 3 mo ago
    Auto-check passed
  • Qwen Mtp Gguf Release

    R6410418/Jackrong-llm-finetuning-guide

    Repository-level wrapper for the canonical Qwen MTP or nextn GGUF release workflow.

    1.7k GitHub stars~474 tokensUpdated 3 mo ago
    Auto-check passed
  • Repository Resource Portal

    R6410418/Jackrong-llm-finetuning-guide

    Maintain this repository as a growing educational LLM knowledge base.

    1.7k GitHub stars~387 tokensUpdated 3 mo ago
    Auto-check passed

Questions about Qwen Mtp Gguf

What does Qwen Mtp Gguf do?

Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release. Qwen Mtp Gguf is an agent skill from R6410418/Jackrong-llm-finetuning-guide. Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

When should I use Qwen Mtp Gguf?

Qwen Mtp Gguf fits situations like: another coding agent needs to inspect a users machine; estimate disk/RAM requirements from Hugging Face model sizes and requested quant formats; bootstrap llama.cpp and Python dependencies; extract MTP heads from a matching official/base Qwen model.

How do I install Qwen Mtp Gguf in Claude Code?

Run `npx skills add R6410418/Jackrong-llm-finetuning-guide --skill qwen-mtp-gguf -a claude-code`. Or copy the skill folder (qwen-mtp-gguf in R6410418/Jackrong-llm-finetuning-guide) into .claude/skills/qwen-mtp-gguf in your project. Claude Code loads it when a task matches its description.

How do I install Qwen Mtp Gguf in Codex?

Run `npx skills add R6410418/Jackrong-llm-finetuning-guide --skill qwen-mtp-gguf -a codex`. Or copy the skill folder (qwen-mtp-gguf in R6410418/Jackrong-llm-finetuning-guide) into .agents/skills/qwen-mtp-gguf in your project. Codex loads it when a task matches its description.

Can I use Qwen Mtp Gguf in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add R6410418/Jackrong-llm-finetuning-guide --skill qwen-mtp-gguf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qwen-mtp-gguf, .gemini/skills/qwen-mtp-gguf, .github/skills/qwen-mtp-gguf and .opencode/skills/qwen-mtp-gguf in your project.

What does Qwen Mtp Gguf need to run?

Going by SKILL.md and its folder, Qwen Mtp Gguf needs Python and a shell for the scripts in its folder, the command-line tools its instructions call (python3 and bash) and credentials named HF_TOKEN. Our summary lists: Python 3; A Bash shell.

Does Qwen Mtp Gguf access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Qwen Mtp Gguf safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Qwen Mtp Gguf use?

Qwen Mtp Gguf is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Qwen Mtp Gguf use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.

What are the alternatives to Qwen Mtp Gguf?

Skills that share tags, products or a category with Qwen Mtp Gguf: Add Model (guoqingbao/xinfer, 334 stars), Resolve (alexziskind1/model-shelf, 130 stars), Test Model (guoqingbao/xinfer, 334 stars) and Aqua Model Lifecycle (oracle/accelerated-data-science, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Qwen Mtp Gguf?

R6410418 (a GitHub user) maintains it in R6410418/Jackrong-llm-finetuning-guide, which has 1,709 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on July 11, 2026.

Source: R6410418/Jackrong-llm-finetuning-guide on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.