Agent skill

Quantized Export

by wshobson in wshobson/agents

Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

MITAuto-check passedAI & LLM Engineering

Install Quantized Export

skills CLI
$ npx skills add wshobson/agents --skill quantized-export -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents quantized-export --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/quantized-export .claude/skills/quantized-export && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quantized-export
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
1,030 words
Files
2 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

  • Works in 3 steps: **Load the exported artifact in its actual → Run 3–5 golden prompts through it — → **Compare each output against the
  • Tasks that involve QA and bug reports
  • SKILL.md covers Format Map, Workload Overrides, The Smoke Test and Related Skills
  • Calls python

What it does

Quantized Export is an agent skill from wshobson/agents. Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. Use after a checkpoint passes promotion, when choosing a quantization format for a target device, or when an exported model fails its smoke test.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/export-commands.md`).

It sits in AI & LLM Engineering, covering QA and bug reports, LLM inference and serving and Fine-tuning. It works with llama.cpp. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Tasks that involve QA and bug reports
  • Tasks that involve LLM inference and serving
  • Tasks that involve Fine-tuning

Example prompts

  • “/quantized-export”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. **Load the exported artifact in its actual
  2. Run 3–5 golden prompts through it —
  3. **Compare each output against the

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quantized Export loads about 2k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 1,030 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 1,030 words, ~1,979 tokens.

Download SKILL.mdSave it as .claude/skills/quantized-export/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
quantized-export
description
Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. Use after a checkpoint passes promotion, when choosing a quantization format for a target device, or when an exported model fails its smoke test.

Quantized Export

The last stop after checkpoint-promotion hands off a PROMOTE verdict: a checkpoint that cleared the four-stage gate still isn't deployed until it's exported in the right format for its target runtime and proven to still work post-export. A REJECT verdict never reaches this skill — export starts only from a promoted checkpoint.

Input: a promoted checkpoint (or LoRA adapter) plus the target deployment surface — GPU class, serving stack, and whether long-context/code/math workloads are in scope. Output format: an exported artifact in the chosen format plus a smoke-test diff report comparing 3–5 golden outputs pre-export and post-export.

Format Map

Pick format by hardware and deployment shape, not by habit — the wrong pick either wastes throughput headroom or breaks silently on specific workloads (see Workload Overrides).

  • FP8 is the default on Hopper-class GPUs and newer. It preserves near-bf16 quality at roughly half the memory, and it's the safe first choice whenever the target GPU supports it and no edge-device constraint applies.
  • AWQ INT4 targets older GPUs that predate FP8 hardware support. GPTQ is superseded for new deployments — don't reach for it on a fresh export; AWQ has better accuracy retention at the same bit width and wider current tooling support.
  • GGUF with Q4_K_M quantization, built from an imatrix, is the edge/llama.cpp format. Use it for local or CPU-adjacent deployment, not for GPU-serving throughput — it optimizes for footprint, not tokens/sec on a datacenter GPU.
  • NVFP4 is for Blackwell-at-scale deployments only — and explicitly NOT on GB10. NVFP4 on SM121 (GB10) runs ~32% slower than FP8 because the hardware lacks a native cvt.e2m1x2 path unless the kernel is compiled sm_121a. Choosing NVFP4 on a GB10 target is a regression, not an upgrade — pick FP8 there instead.
  • Merged vs. LoRA-only is a separate axis from quant format. A merged export folds the adapter into the base weights: larger artifact, no base-model dependency at serve time. LoRA-only keeps the adapter separate: much smaller artifact, but the serving stack must load the exact same base model alongside it — a mismatched or wrong-revision base silently changes outputs. Pick merged when artifact portability matters more than storage; pick LoRA-only when disk footprint or multi-adapter serving matters more.
Worked Picks

The core format-selection tradeoff, read as a lookup table for common scenarios:

TargetWorkloadFormat
Datacenter GPUgeneric chatFP8
Datacenter GPUlong-context/code/mathFP8 or W8A8 — never INT4
Older GPU generationgenericAWQ INT4
Edge device / laptopllama.cpp servingGGUF Q4_K_M + imatrix
GB10any workloadFP8 via vLLM nightly, or GGUF via llama.cpp locally — skip NVFP4
yaml
# quick decision snippet — see the table above for the full map
hopper_or_newer: fp8
older_gpu: awq-int4
edge_llama_cpp: gguf-q4_k_m+imatrix
gb10_any_workload: fp8-vllm-nightly   # never nvfp4 on GB10

Workload Overrides

The Format Map above is a default, not a rule that survives every workload. Long-context, code, and math workloads break at INT4 — quantization error compounds across long sequences and precise token-level reasoning in ways that don't show up on short, generic prompts. For any of these three workload classes, stay on FP8 or W8A8 even if the target hardware would otherwise justify INT4 on cost grounds.

  • Don't validate this override with MMLU or similar broad-knowledge benchmarks — they don't stress the failure mode. Measure with the actual task evals — the goldens and graders from eval-harness-first, run through the exported artifact — because INT4 degradation on long-context, code, or math shows up as task-specific failures (dropped context, broken syntax, arithmetic errors) well before it moves a knowledge benchmark.
  • If a task eval regresses after an INT4 export on one of these three workload classes, the fix is switching format, not re-tuning the quantization recipe — AWQ and GPTQ variants at the same bit width share the same compounding-error failure mode on these workloads.
Show full SKILL.md (443 more words)Show less

The Smoke Test

Export bugs are silent at the file level — a malformed export still produces a loadable artifact, so file-existence checks prove nothing. The smoke test is mandatory for every export, with no exception for a format that "should just work":

  1. Load the exported artifact in its actual target runtime — vLLM for FP8/AWQ, llama.cpp for GGUF, not a quick sanity load in a different framework than the one that will serve it in production.
  2. Run 3–5 golden prompts through it — pull these from the same eval/goldens.jsonl eval-harness-first maintains, not a fresh ad hoc set.
  3. Compare each output against the pre-export generation for the same prompt, same deterministic sampling settings — greedy decoding (temperature 0) and a fixed seed, persisted and reused between the pre- and post-export runs, not just nominally identical config. For a lossless export, byte match is the gate — any diff is a bug. For a lossy (quantized) export, byte match is expected to fail; the gate is task-grader verdict agreement instead — see references/export-commands.md's Smoke-Test Script Skeleton.

Run this as a gate, not a manual check:

bash
python smoke_test.py "$EXPORT_PATH" \
    eval/goldens.jsonl pre-export-outputs.jsonl
# non-zero exit on any pre/post mismatch
Failure Signatures

What export bugs actually look like, not a clean pass/fail flag:

  • Template mismatch presents as garbled or run-on output — the chat template baked into the export doesn't match the one the checkpoint was trained and evaluated against, so turn boundaries or special tokens land in the wrong place.
  • Wrong quantization applied to lm_head presents as off-template or semantically nonsensical output that still looks fluent — the output head lost precision it needed even though the rest of the network quantized cleanly.

Never ship an export that skipped this step — a checkpoint's PROMOTE verdict says the un-exported checkpoint is good; it says nothing about the export pipeline. Re-run on any quant-method or runtime version bump, not only after the first export. Runnable command sequences for every format plus the smoke-test script skeleton: references/export-commands.md.

  • checkpoint-promotion — the only valid upstream source for this skill. A checkpoint without a PROMOTE verdict doesn't reach export.
  • eval-harness-first — owns the eval/goldens.jsonl this skill's smoke test draws its 3–5 prompts from, and the task evals the Workload Overrides section requires for long-context/code/math validation.
  • finetuning-method-selection — its references/model-catalog.md is the place to check hardware-class assumptions (which GPU generations a base model targets) before picking a format off the Format Map above.

Spark users: on GB10, GGUF via llama.cpp works well for local serving, and FP8 serving via vLLM nightly builds is the other proven path — NVFP4 is the one format to avoid there (see the Format Map exception above). Once the dgx-spark-ops plugin is installed, defer Spark-specific serving and thermal questions to its skills rather than re-deriving them here.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/llm-finetuning/skills/quantized-export of wshobson/agents.

  • SKILL.md
  • references/export-commands.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Quantized Export next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quantized Export compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quantized Export this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
ML Research LabAnastasiyaW/codex-claude-code-config154—~794Automated safety check: PassMIT
Monitor With HaolemeHaolemeApp/Haoleme157—~1.3kAutomated safety check: PassAGPL-3.0
Aqua Deploymentoracle/accelerated-data-science125—~2.4kAutomated safety check: PassUPL-1.0
AWS AI MLaws/agent-toolkit-for-aws2.8k—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • ML Research Lab

    AnastasiyaW/codex-claude-code-config

    Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

    154 GitHub stars~794 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Monitor With Haoleme

    HaolemeApp/Haoleme

    Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.

    157 GitHub stars~1.3k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Aqua Deployment

    oracle/accelerated-data-science

    Official

    Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.

    125 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • AWS AI ML

    aws/agent-toolkit-for-aws

    Official

    Selects, deploys, and customizes AI models on Amazon SageMaker.

    2.8k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Unsloth

    ericrisco/rsc-harness

    A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…

    167 GitHub stars~3.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 14 repos~473 tokens
    Auto-check passed
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Portfolio Risk Metrics

    wshobson/agents

    Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.

    40k GitHub starsUsed in 13 repos~502 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed

Works with

Questions about Quantized Export

What does Quantized Export do?

Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. Quantized Export is an agent skill from wshobson/agents. Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

When should I use Quantized Export?

Quantized Export fits situations like: tasks that involve QA and bug reports; tasks that involve LLM inference and serving; tasks that involve Fine-tuning.

How do I install Quantized Export in Claude Code?

Run `npx skills add wshobson/agents --skill quantized-export -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/quantized-export in wshobson/agents) into .claude/skills/quantized-export in your project. Claude Code loads it when a task matches its description.

How do I install Quantized Export in Codex?

Run `npx skills add wshobson/agents --skill quantized-export -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/quantized-export in wshobson/agents) into .agents/skills/quantized-export in your project. Codex loads it when a task matches its description.

Can I use Quantized Export in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill quantized-export -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quantized-export, .gemini/skills/quantized-export, .github/skills/quantized-export and .opencode/skills/quantized-export in your project.

What does Quantized Export need to run?

Going by SKILL.md and its folder, Quantized Export needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Quantized Export access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Quantized Export safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Quantized Export use?

Quantized Export is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quantized Export use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.2k tokens, read only when the agent opens those files.

What are the alternatives to Quantized Export?

Skills that share tags, products or a category with Quantized Export: Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), ML Research Lab (AnastasiyaW/codex-claude-code-config, 154 stars), Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars) and Aqua Deployment (oracle/accelerated-data-science, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quantized Export?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,287 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.