Qwen Mtp Gguf
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.
$ npx skills add wshobson/agents --skill quantized-export -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents quantized-export --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/quantized-export .claude/skills/quantized-export && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "quantized-export" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/quantized-export into .claude/skills/quantized-export/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quantized-export", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/quantized-exportType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill quantized-export -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents quantized-export --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/llm-finetuning/skills/quantized-export .agents/skills/quantized-export && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "quantized-export" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/quantized-export into .agents/skills/quantized-export/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quantized-export", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill quantized-export -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents quantized-export --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/llm-finetuning/skills/quantized-export .cursor/skills/quantized-export && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "quantized-export" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/quantized-export into .cursor/skills/quantized-export/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quantized-export", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/llm-finetuning/skills/quantized-export--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill quantized-export -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents quantized-export --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/llm-finetuning/skills/quantized-export .gemini/skills/quantized-export && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "quantized-export" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/quantized-export into .gemini/skills/quantized-export/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quantized-export", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents quantized-exportInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill quantized-export -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/llm-finetuning/skills/quantized-export .github/skills/quantized-export && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "quantized-export" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/quantized-export into .github/skills/quantized-export/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quantized-export", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill quantized-export -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents quantized-export --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/llm-finetuning/skills/quantized-export .opencode/skills/quantized-export && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "quantized-export" agent skill from https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/quantized-export into .opencode/skills/quantized-export/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "quantized-export", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
quantized-exportExport a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.
Quantized Export is an agent skill from wshobson/agents. Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. Use after a checkpoint passes promotion, when choosing a quantization format for a target device, or when an exported model fails its smoke test.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/export-commands.md`).
It sits in AI & LLM Engineering, covering QA and bug reports, LLM inference and serving and Fine-tuning. It works with llama.cpp. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Quantized Export loads about 2k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 1,030 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 1,030 words, ~1,979 tokens.
.claude/skills/quantized-export/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.The last stop after checkpoint-promotion
hands off a PROMOTE verdict: a checkpoint
that cleared the four-stage gate still isn't
deployed until it's exported in the right
format for its target runtime and proven to
still work post-export. A REJECT verdict
never reaches this skill — export starts only
from a promoted checkpoint.
Input: a promoted checkpoint (or LoRA adapter) plus the target deployment surface — GPU class, serving stack, and whether long-context/code/math workloads are in scope. Output format: an exported artifact in the chosen format plus a smoke-test diff report comparing 3–5 golden outputs pre-export and post-export.
Pick format by hardware and deployment shape, not by habit — the wrong pick either wastes throughput headroom or breaks silently on specific workloads (see Workload Overrides).
cvt.e2m1x2 path unless the
kernel is compiled sm_121a. Choosing
NVFP4 on a GB10 target is a regression, not
an upgrade — pick FP8 there instead.The core format-selection tradeoff, read as a lookup table for common scenarios:
| Target | Workload | Format |
|---|---|---|
| Datacenter GPU | generic chat | FP8 |
| Datacenter GPU | long-context/code/math | FP8 or W8A8 — never INT4 |
| Older GPU generation | generic | AWQ INT4 |
| Edge device / laptop | llama.cpp serving | GGUF Q4_K_M + imatrix |
| GB10 | any workload | FP8 via vLLM nightly, or GGUF via llama.cpp locally — skip NVFP4 |
# quick decision snippet — see the table above for the full map
hopper_or_newer: fp8
older_gpu: awq-int4
edge_llama_cpp: gguf-q4_k_m+imatrix
gb10_any_workload: fp8-vllm-nightly # never nvfp4 on GB10The Format Map above is a default, not a rule that survives every workload. Long-context, code, and math workloads break at INT4 — quantization error compounds across long sequences and precise token-level reasoning in ways that don't show up on short, generic prompts. For any of these three workload classes, stay on FP8 or W8A8 even if the target hardware would otherwise justify INT4 on cost grounds.
eval-harness-first, run
through the exported artifact — because
INT4 degradation on long-context, code, or
math shows up as task-specific failures
(dropped context, broken syntax, arithmetic
errors) well before it moves a knowledge
benchmark.Export bugs are silent at the file level — a malformed export still produces a loadable artifact, so file-existence checks prove nothing. The smoke test is mandatory for every export, with no exception for a format that "should just work":
eval/goldens.jsonl
eval-harness-first maintains, not a fresh
ad hoc set.references/export-commands.md's
Smoke-Test Script Skeleton.Run this as a gate, not a manual check:
python smoke_test.py "$EXPORT_PATH" \
eval/goldens.jsonl pre-export-outputs.jsonl
# non-zero exit on any pre/post mismatchWhat export bugs actually look like, not a clean pass/fail flag:
lm_head
presents as off-template or semantically
nonsensical output that still looks
fluent — the output head lost precision it
needed even though the rest of the network
quantized cleanly.Never ship an export that skipped this step —
a checkpoint's PROMOTE verdict says the
un-exported checkpoint is good; it says
nothing about the export pipeline. Re-run on
any quant-method or runtime version bump, not
only after the first export. Runnable command
sequences for every format plus the
smoke-test script skeleton:
references/export-commands.md.
checkpoint-promotion — the only valid
upstream source for this skill. A checkpoint
without a PROMOTE verdict doesn't reach
export.eval-harness-first — owns the
eval/goldens.jsonl this skill's smoke test
draws its 3–5 prompts from, and the task
evals the Workload Overrides section
requires for long-context/code/math
validation.finetuning-method-selection — its
references/model-catalog.md is the place
to check hardware-class assumptions (which
GPU generations a base model targets) before
picking a format off the Format Map above.Spark users: on GB10, GGUF via llama.cpp
works well for local serving, and FP8 serving
via vLLM nightly builds is the other proven
path — NVFP4 is the one format to avoid there
(see the Format Map exception above). Once the
dgx-spark-ops plugin is installed, defer
Spark-specific serving and thermal questions to
its skills rather than re-deriving them here.
© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in plugins/llm-finetuning/skills/quantized-export of wshobson/agents.
Open the folder on GitHubat commit 46891e7
Quantized Export next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Quantized Export this skillwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | |
| ML Research LabAnastasiyaW/codex-claude-code-config | 154 | — | ~794 | Automated safety check: Pass | MIT | |
| Monitor With HaolemeHaolemeApp/Haoleme | 157 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | |
| Aqua Deploymentoracle/accelerated-data-science | 125 | — | ~2.4k | Automated safety check: Pass | UPL-1.0 | |
| AWS AI MLaws/agent-toolkit-for-aws | 2.8k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 |
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
AnastasiyaW/codex-claude-code-config
Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.
HaolemeApp/Haoleme
Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.
oracle/accelerated-data-science
Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.
aws/agent-toolkit-for-aws
Selects, deploys, and customizes AI models on Amazon SageMaker.
ericrisco/rsc-harness
A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
Works with
Categories
Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. Quantized Export is an agent skill from wshobson/agents. Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.
Quantized Export fits situations like: tasks that involve QA and bug reports; tasks that involve LLM inference and serving; tasks that involve Fine-tuning.
Run `npx skills add wshobson/agents --skill quantized-export -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/quantized-export in wshobson/agents) into .claude/skills/quantized-export in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill quantized-export -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/quantized-export in wshobson/agents) into .agents/skills/quantized-export in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill quantized-export -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quantized-export, .gemini/skills/quantized-export, .github/skills/quantized-export and .opencode/skills/quantized-export in your project.
Going by SKILL.md and its folder, Quantized Export needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Quantized Export is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Quantized Export: Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), ML Research Lab (AnastasiyaW/codex-claude-code-config, 154 stars), Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars) and Aqua Deployment (oracle/accelerated-data-science, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,287 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.