Resolve
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
Test LLM models served by xinfer for correctness, output quality, and performance.
$ npx skills add guoqingbao/xinfer --skill test-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install guoqingbao/xinfer test-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/test-model .claude/skills/test-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "test-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/test-model into .claude/skills/test-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/test-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add guoqingbao/xinfer --skill test-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install guoqingbao/xinfer test-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.cursor/skills/test-model .agents/skills/test-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "test-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/test-model into .agents/skills/test-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add guoqingbao/xinfer --skill test-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install guoqingbao/xinfer test-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.cursor/skills/test-model .cursor/skills/test-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "test-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/test-model into .cursor/skills/test-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/guoqingbao/xinfer.git --path .cursor/skills/test-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add guoqingbao/xinfer --skill test-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install guoqingbao/xinfer test-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.cursor/skills/test-model .gemini/skills/test-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "test-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/test-model into .gemini/skills/test-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install guoqingbao/xinfer test-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add guoqingbao/xinfer --skill test-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .github/skills && cp -r skills-src/.cursor/skills/test-model .github/skills/test-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "test-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/test-model into .github/skills/test-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add guoqingbao/xinfer --skill test-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install guoqingbao/xinfer test-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.cursor/skills/test-model .opencode/skills/test-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "test-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/test-model into .opencode/skills/test-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "test-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
test-modelTest LLM models served by xinfer for correctness, output quality, and performance.
Test Model is an agent skill from guoqingbao/xinfer. Test LLM models served by xinfer for correctness, output quality, and performance. Use when the user asks to test, benchmark, validate, or verify models — either from a local folder path or HuggingFace model IDs. Supports all xinfer-compatible formats: BF16, FP8, MXFP4, NVFP4, GGUF, GPTQ, AWQ, ISQ, Dense, MoE, and Multimodal architectures.
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Model hubs and datasets and LLM inference and serving. It works with llama.cpp, Hugging Face, Qwen and vLLM. The repository describes itself as: Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b88c153. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Test Model loads about 2.6k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 937 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from guoqingbao/xinfer at commit b88c153, republished under its MIT licence (© guoqingbao). 937 words, ~2,624 tokens.
.claude/skills/test-model/SKILL.md (or your agent's skills folder).Collect the models to test. The user provides one or both of:
| Input | Format | Example |
|---|---|---|
| Local folder | Absolute path to a directory containing model weights | /data/models or /data/Qwen3.5-27B-FP8 |
| HuggingFace IDs | Comma-separated model IDs | AxionML/Qwen3.5-2B-NVFP4, Qwen/Qwen3-4B |
If the user provides a parent directory (not a single model), scan it to find testable models:
# List subdirectories that look like model folders
for d in /data/*/; do
if [ -f "$d/config.json" ] || ls "$d"/*.gguf 2>/dev/null | head -1 >/dev/null; then
echo "$d"
fi
doneFor each candidate directory, determine the model type by reading config.json:
import json, os, sys, glob
def detect_model(path):
"""Detect model type and quantization from a local directory."""
config_path = os.path.join(path, "config.json")
gguf_files = glob.glob(os.path.join(path, "*.gguf"))
info = {"path": path, "name": os.path.basename(path.rstrip("/"))}
if gguf_files:
info["format"] = "gguf"
info["gguf_file"] = os.path.basename(gguf_files[0])
return info
if not os.path.exists(config_path):
return None
cfg = json.load(open(config_path))
arch = (cfg.get("architectures") or ["Unknown"])[0]
supported = [
"LlamaForCausalLM", "MistralForCausalLM", "Ministral3ForConditionalGeneration",
"Qwen2ForCausalLM", "Qwen3ForCausalLM", "Qwen3MoeForCausalLM",
"Qwen3_5ForCausalLM", "Qwen3_5MoeForCausalLM",
"Qwen3_5ForConditionalGeneration", "Qwen3_5MoeForConditionalGeneration",
"Qwen3NextForCausalLM",
"Qwen3VLForConditionalGeneration",
"Gemma3ForConditionalGeneration", "Gemma3ForCausalLM",
"Gemma4ForCausalLM", "Gemma4ForConditionalGeneration",
"Phi3ForCausalLM", "Phi4ForCausalLM",
"Glm4ForCausalLM", "Glm4MoeForCausalLM",
]
if arch not in supported:
info["skip"] = f"Unsupported architecture: {arch}"
return info
info["arch"] = arch
info["format"] = "safetensors"
qcfg = cfg.get("quantization_config", {})
qm = qcfg.get("quant_method", "")
if qm in ("fp8", "modelopt", "compressed-tensors"):
algo = qcfg.get("quant_algo", "")
fmt = qcfg.get("format", "")
if algo and ("nvfp4" in algo.lower() or "fp4" in algo.lower()):
info["quant"] = "nvfp4"
elif "nvfp4" in fmt.lower():
info["quant"] = "nvfp4"
elif "mxfp4" in fmt.lower():
info["quant"] = "mxfp4"
elif qm == "fp8":
info["quant"] = "fp8"
else:
info["quant"] = qm
elif qm in ("gptq", "awq"):
info["quant"] = qm
elif qm == "mxfp4":
info["quant"] = "mxfp4"
else:
info["quant"] = "bf16"
return infoPresent the detected models to the user as a table and confirm before proceeding.
nvidia-smi --query-gpu=index,name,memory.total,memory.free --format=csv,noheader,nounitsParse the output to get gpu_id, name, total_mb, free_mb for each GPU.
Use these rough heuristics for memory estimation (single-GPU, including KV cache overhead):
| Format | Estimate (GB) |
|---|---|
| BF16 / FP16 | params_B * 2.2 |
| FP8 | params_B * 1.2 |
| MXFP4 / NVFP4 | params_B * 0.8 |
| GGUF Q4_K_M | params_B * 0.7 |
| GGUF Q3_K_M | params_B * 0.55 |
| GGUF Q2_K | params_B * 0.45 |
| MoE (A3B active) | Use active params for compute, total params for weight memory |
Extract parameter count from the model name when possible (e.g. Qwen3.5-27B → 27B).
For MoE models with A3B in the name, the weight memory uses total params but fits better than dense.
--d <gpu_id> with the GPU that has the most free memory.--d <id1>,<id2> with the two GPUs with the most free memory.Build using build.sh:
cd <project_root>
./build.sh --install --features cuda,nccl,flashinfer,cutlassVerify the build succeeds (exit code 0). The Error: Must provide model_id or weight_path message after build is expected — it means the binary compiled correctly.
If the build fails, check and fix compilation errors before proceeding.
Create test_model.py in the project root with the following capabilities:
--port to specify the API server port--wait for server readiness timeoutthinking=false and thinking=true modesThe prompt should be a substantive multi-topic question (algorithms, data structures, etc.) padded with context tokens to reach the 1k+ input requirement. Use max_tokens: 2048 and temperature: 0.7. Set request timeout to 300s.
For thinking mode, add "extra_body": {"thinking": true} to the payload.
Quality checks:
max(10, 5% of total trigrams) timesFor each model, execute this sequence:
pkill -9 -f 'xinfer' 2>/dev/null
sleep 3Always wait 3 seconds after killing to ensure GPU memory is released.
Build the server command based on model type:
| Model source | Command pattern |
|---|---|
| Local safetensors | ./target/release/xinfer --w <path> --ui-server --d <gpus> --port 7000 |
| Local GGUF | ./target/release/xinfer --w <dir> --f <file.gguf> --ui-server --d <gpus> --port 7000 |
| HuggingFace ID | ./target/release/xinfer --m <hf_id> --ui-server --d <gpus> --port 7000 |
Run the server in the background with RUST_BACKTRACE=1 for debugging.
Poll GET /v1/models every 2-3 seconds until it returns HTTP 200, with a timeout of:
python3 test_model.py --port 7000If the server fails to start or the test script returns errors:
| Error | Likely cause | Fix |
|---|---|---|
MLX-quantized models panic | Incompatible NVFP4 packing | Skip model; use modelopt/compressed-tensors variant |
Unable to load ... projection weights | DeltaNet weights not detected as quantized | Check is_weight_quantized in deltanet.rs |
CUDA out of memory | Model too large for GPU | Try with more GPUs or skip |
| Server starts but API times out | Model too slow on prefill | Increase test timeout to 600s |
failed to fill whole buffer | Runner process crashed | Check runner logs, enable RUST_BACKTRACE=full |
Debug with unwrap: If the model crashes during inference, temporarily change guard.step() to guard.step().unwrap() in src/core/engine.rs to get a full stack trace. Revert after debugging.
If a model cannot be fixed, record the failure reason and continue to the next model.
After all models are tested, produce a summary table:
## Test Results
| # | Model | Format | GPUs | thinking=false | thinking=true | Quality |
|---|-------|--------|------|----------------|---------------|---------|
| 1 | Qwen3.5-27B-FP8 | FP8 | 1 | 1342 in / 2048 out, 42.2 tok/s | 1342 in / 2048 out, 42.2 tok/s | OK |
| 2 | ... | ... | ... | ... | ... | ... |
### Notes
- Model X: SKIPPED — reason
- Model Y: FAILED — error descriptionInclude for each model:
| File | Purpose |
|---|---|
test_model.py | OpenAI API test script (created by this skill) |
src/core/engine.rs | Engine loop; guard.step() for debug |
src/models/layers/deltanet.rs | DeltaNet layer; quantization detection |
src/models/layers/linear.rs | Linear layer loaders (FP8, MXFP4, NVFP4) |
build.sh | Build script (compiles xinfer) |
| Feature set | When to use |
|---|---|
cuda,nccl,flashinfer,cutlass | SM80+ (Ampere/Ada/Hopper), recommended |
cuda,nccl,flashattn,cutlass | Alternative to flashinfer |
cuda,nccl | V100 (SM70), no flash attention |
metal | macOS Apple Silicon |
| Flag | Purpose |
|---|---|
--w <path> | Local model weight directory |
--f <file> | GGUF filename within the weight directory |
--m <hf_id> | HuggingFace model ID (auto-downloads) |
--d <ids> | GPU device IDs (e.g. 0 or 0,1) |
--port <n> | API server port |
--disable-prefix-cache | Disable prefix caching (on by default) |
--ui-server | Enable built-in ChatGPT-like web UI |
--isq <fmt> | In-situ quantization (q2k, q3k, q4k, q5k, q6k, q8_0) |
--kvcache-dtype <mode> | KV cache quantization: fp8, turbo8, turbo4, turbo3 |
© guoqingbao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .cursor/skills/test-model of guoqingbao/xinfer.
Open the folder on GitHubat commit b88c153
Test Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Test Model this skillguoqingbao/xinfer | 333 | — | ~2.6k | Automated safety check: Pass | MIT | |
| Resolvealexziskind1/model-shelf | 130 | — | ~792 | Automated safety check: Pass | MIT | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Aqua Model Lifecycleoracle/accelerated-data-science | 125 | — | ~1.4k | Automated safety check: Pass | UPL-1.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 |
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
oracle/accelerated-data-science
Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
huggingface/skills
Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.
guoqingbao/xinfer
Adapt and port new LLM model architectures to this xinfer project.
guoqingbao/xinfer
Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.
Works with
Categories
Test LLM models served by xinfer for correctness, output quality, and performance. Test Model is an agent skill from guoqingbao/xinfer. Test LLM models served by xinfer for correctness, output quality, and performance.
Test Model fits situations like: the user asks to test; verify models — either from a local folder path; huggingFace model IDs.
Run `npx skills add guoqingbao/xinfer --skill test-model -a claude-code`. Or copy the skill folder (.cursor/skills/test-model in guoqingbao/xinfer) into .claude/skills/test-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add guoqingbao/xinfer --skill test-model -a codex`. Or copy the skill folder (.cursor/skills/test-model in guoqingbao/xinfer) into .agents/skills/test-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guoqingbao/xinfer --skill test-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-model, .gemini/skills/test-model, .github/skills/test-model and .opencode/skills/test-model in your project.
Going by SKILL.md and its folder, Test Model needs the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Test Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Test Model: Resolve (alexziskind1/model-shelf, 130 stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Aqua Model Lifecycle (oracle/accelerated-data-science, 125 stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
guoqingbao (a GitHub user) maintains it in guoqingbao/xinfer, which has 333 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 9, 2026.
Source: guoqingbao/xinfer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.