Agent skill

Test Model

by guoqingbao in guoqingbao/xinfer

Test LLM models served by xinfer for correctness, output quality, and performance.

MITAuto-check passedAI & LLM Engineering

Install Test Model

skills CLI
$ npx skills add guoqingbao/xinfer --skill test-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guoqingbao/xinfer test-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/test-model .claude/skills/test-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-model
GitHub stars
333
Token cost
~2.6k tokens
SKILL.md length
937 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Test LLM models served by xinfer for correctness, output quality, and performance.

  • Works in 6 steps: Gather Model List → Estimate GPU Requirements and Detect… → Build the Project → …
  • The user asks to test
  • SKILL.md covers Phase 0: Gather Model List, Phase 1: Estimate GPU…, Phase 2: Build the Project and Phase 3: Create the Test Script, plus 3 more sections
  • Calls python3

What it does

Test Model is an agent skill from guoqingbao/xinfer. Test LLM models served by xinfer for correctness, output quality, and performance. Use when the user asks to test, benchmark, validate, or verify models — either from a local folder path or HuggingFace model IDs. Supports all xinfer-compatible formats: BF16, FP8, MXFP4, NVFP4, GGUF, GPTQ, AWQ, ISQ, Dense, MoE, and Multimodal architectures.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets and LLM inference and serving. It works with llama.cpp, Hugging Face, Qwen and vLLM. The repository describes itself as: Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime. The licence is MIT.

When your agent uses it

  • The user asks to test
  • Verify models — either from a local folder path
  • HuggingFace model IDs

Example prompts

  • “/test-model”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Gather Model List
  2. Estimate GPU Requirements and Detect Hardware
  3. Build the Project
  4. Create the Test Script
  5. Test Each Model
  6. Summarize Results

What it can do on your machine

Read from SKILL.md and the folder at commit b88c153. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Model loads about 2.6k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 937 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from guoqingbao/xinfer at commit b88c153, republished under its MIT licence (© guoqingbao). 937 words, ~2,624 tokens.

Download SKILL.mdSave it as .claude/skills/test-model/SKILL.md (or your agent's skills folder).
name
test-model
description
Test LLM models served by xinfer for correctness, output quality, and performance. Use when the user asks to test, benchmark, validate, or verify models — either from a local folder path or HuggingFace model IDs. Supports all xinfer-compatible formats: BF16, FP8, MXFP4, NVFP4, GGUF, GPTQ, AWQ, ISQ, Dense, MoE, and Multimodal architectures.

Test Model — Validate and Benchmark LLM Models on xinfer

Phase 0: Gather Model List

Collect the models to test. The user provides one or both of:

InputFormatExample
Local folderAbsolute path to a directory containing model weights/data/models or /data/Qwen3.5-27B-FP8
HuggingFace IDsComma-separated model IDsAxionML/Qwen3.5-2B-NVFP4, Qwen/Qwen3-4B
Detecting models in a local folder

If the user provides a parent directory (not a single model), scan it to find testable models:

bash
# List subdirectories that look like model folders
for d in /data/*/; do
  if [ -f "$d/config.json" ] || ls "$d"/*.gguf 2>/dev/null | head -1 >/dev/null; then
    echo "$d"
  fi
done

For each candidate directory, determine the model type by reading config.json:

python
import json, os, sys, glob

def detect_model(path):
    """Detect model type and quantization from a local directory."""
    config_path = os.path.join(path, "config.json")
    gguf_files = glob.glob(os.path.join(path, "*.gguf"))

    info = {"path": path, "name": os.path.basename(path.rstrip("/"))}

    if gguf_files:
        info["format"] = "gguf"
        info["gguf_file"] = os.path.basename(gguf_files[0])
        return info

    if not os.path.exists(config_path):
        return None

    cfg = json.load(open(config_path))
    arch = (cfg.get("architectures") or ["Unknown"])[0]

    supported = [
        "LlamaForCausalLM", "MistralForCausalLM", "Ministral3ForConditionalGeneration",
        "Qwen2ForCausalLM", "Qwen3ForCausalLM", "Qwen3MoeForCausalLM",
        "Qwen3_5ForCausalLM", "Qwen3_5MoeForCausalLM",
        "Qwen3_5ForConditionalGeneration", "Qwen3_5MoeForConditionalGeneration",
        "Qwen3NextForCausalLM",
        "Qwen3VLForConditionalGeneration",
        "Gemma3ForConditionalGeneration", "Gemma3ForCausalLM",
        "Gemma4ForCausalLM", "Gemma4ForConditionalGeneration",
        "Phi3ForCausalLM", "Phi4ForCausalLM",
        "Glm4ForCausalLM", "Glm4MoeForCausalLM",
    ]
    if arch not in supported:
        info["skip"] = f"Unsupported architecture: {arch}"
        return info

    info["arch"] = arch
    info["format"] = "safetensors"

    qcfg = cfg.get("quantization_config", {})
    qm = qcfg.get("quant_method", "")
    if qm in ("fp8", "modelopt", "compressed-tensors"):
        algo = qcfg.get("quant_algo", "")
        fmt = qcfg.get("format", "")
        if algo and ("nvfp4" in algo.lower() or "fp4" in algo.lower()):
            info["quant"] = "nvfp4"
        elif "nvfp4" in fmt.lower():
            info["quant"] = "nvfp4"
        elif "mxfp4" in fmt.lower():
            info["quant"] = "mxfp4"
        elif qm == "fp8":
            info["quant"] = "fp8"
        else:
            info["quant"] = qm
    elif qm in ("gptq", "awq"):
        info["quant"] = qm
    elif qm == "mxfp4":
        info["quant"] = "mxfp4"
    else:
        info["quant"] = "bf16"

    return info

Present the detected models to the user as a table and confirm before proceeding.


Phase 1: Estimate GPU Requirements and Detect Hardware

Detect available GPUs
bash
nvidia-smi --query-gpu=index,name,memory.total,memory.free --format=csv,noheader,nounits

Parse the output to get gpu_id, name, total_mb, free_mb for each GPU.

Estimate model memory

Use these rough heuristics for memory estimation (single-GPU, including KV cache overhead):

FormatEstimate (GB)
BF16 / FP16params_B * 2.2
FP8params_B * 1.2
MXFP4 / NVFP4params_B * 0.8
GGUF Q4_K_Mparams_B * 0.7
GGUF Q3_K_Mparams_B * 0.55
GGUF Q2_Kparams_B * 0.45
MoE (A3B active)Use active params for compute, total params for weight memory

Extract parameter count from the model name when possible (e.g. Qwen3.5-27B → 27B). For MoE models with A3B in the name, the weight memory uses total params but fits better than dense.

GPU assignment rules
  1. If a model fits in one GPU's free memory, use --d <gpu_id> with the GPU that has the most free memory.
  2. If a model needs 2 GPUs, use --d <id1>,<id2> with the two GPUs with the most free memory.
  3. If a model exceeds all available GPU memory, report it as skipped and move to the next.
  4. For models explicitly specified as multi-GPU by the user, respect that.

Phase 2: Build the Project

Build using build.sh:

bash
cd <project_root>
./build.sh --install --features cuda,nccl,flashinfer,cutlass

Verify the build succeeds (exit code 0). The Error: Must provide model_id or weight_path message after build is expected — it means the binary compiled correctly.

If the build fails, check and fix compilation errors before proceeding.

Important, if you build on CUDA with cargo build, make sure always build xinfer binaries.

Phase 3: Create the Test Script

Create test_model.py in the project root with the following capabilities:

  • Accept --port to specify the API server port
  • Accept --wait for server readiness timeout
  • Test both thinking=false and thinking=true modes
  • Send a prompt with at least 1024 input tokens and request at least 2048 output tokens
  • Measure end-to-end throughput (completion_tokens / total_time)
  • Check output quality: detect excessive 3-gram repetition, too-short responses
  • Report prompt tokens, completion tokens, time, throughput, and quality verdict
  • Print a summary table at the end

The prompt should be a substantive multi-topic question (algorithms, data structures, etc.) padded with context tokens to reach the 1k+ input requirement. Use max_tokens: 2048 and temperature: 0.7. Set request timeout to 300s.

For thinking mode, add "extra_body": {"thinking": true} to the payload.

Quality checks:

  • Response must be at least 100 characters
  • 3-gram repetition: flag if any trigram appears more than max(10, 5% of total trigrams) times

Phase 4: Test Each Model

For each model, execute this sequence:

Step 1: Kill previous instances
bash
pkill -9 -f 'xinfer' 2>/dev/null
sleep 3

Always wait 3 seconds after killing to ensure GPU memory is released.

Step 2: Start the server

Build the server command based on model type:

Model sourceCommand pattern
Local safetensors./target/release/xinfer --w <path> --ui-server --d <gpus> --port 7000
Local GGUF./target/release/xinfer --w <dir> --f <file.gguf> --ui-server --d <gpus> --port 7000
HuggingFace ID./target/release/xinfer --m <hf_id> --ui-server --d <gpus> --port 7000

Run the server in the background with RUST_BACKTRACE=1 for debugging.

Show full SKILL.md (362 more words)Show less
Step 3: Wait for server readiness

Poll GET /v1/models every 2-3 seconds until it returns HTTP 200, with a timeout of:

  • Small models (< 10B): 120s
  • Medium models (10-40B): 300s
  • Large models (> 40B) or HF downloads: 600s
Step 4: Run the test script
bash
python3 test_model.py --port 7000
Step 5: Handle failures

If the server fails to start or the test script returns errors:

  1. Check server logs for panics or errors
  2. Common issues and fixes:
ErrorLikely causeFix
MLX-quantized models panicIncompatible NVFP4 packingSkip model; use modelopt/compressed-tensors variant
Unable to load ... projection weightsDeltaNet weights not detected as quantizedCheck is_weight_quantized in deltanet.rs
CUDA out of memoryModel too large for GPUTry with more GPUs or skip
Server starts but API times outModel too slow on prefillIncrease test timeout to 600s
failed to fill whole bufferRunner process crashedCheck runner logs, enable RUST_BACKTRACE=full
  1. Debug with unwrap: If the model crashes during inference, temporarily change guard.step() to guard.step().unwrap() in src/core/engine.rs to get a full stack trace. Revert after debugging.

  2. If a model cannot be fixed, record the failure reason and continue to the next model.


Phase 5: Summarize Results

After all models are tested, produce a summary table:

## Test Results

| # | Model | Format | GPUs | thinking=false | thinking=true | Quality |
|---|-------|--------|------|----------------|---------------|---------|
| 1 | Qwen3.5-27B-FP8 | FP8 | 1 | 1342 in / 2048 out, 42.2 tok/s | 1342 in / 2048 out, 42.2 tok/s | OK |
| 2 | ... | ... | ... | ... | ... | ... |

### Notes
- Model X: SKIPPED — reason
- Model Y: FAILED — error description

Include for each model:

  • Model name and quantization format
  • Number of GPUs used
  • Input/output token counts and throughput for both thinking modes
  • Quality verdict (OK / ISSUES / FAILED / SKIPPED)

Quick Reference

Key files
FilePurpose
test_model.pyOpenAI API test script (created by this skill)
src/core/engine.rsEngine loop; guard.step() for debug
src/models/layers/deltanet.rsDeltaNet layer; quantization detection
src/models/layers/linear.rsLinear layer loaders (FP8, MXFP4, NVFP4)
build.shBuild script (compiles xinfer)
Build features
Feature setWhen to use
cuda,nccl,flashinfer,cutlassSM80+ (Ampere/Ada/Hopper), recommended
cuda,nccl,flashattn,cutlassAlternative to flashinfer
cuda,ncclV100 (SM70), no flash attention
metalmacOS Apple Silicon
Server flags
FlagPurpose
--w <path>Local model weight directory
--f <file>GGUF filename within the weight directory
--m <hf_id>HuggingFace model ID (auto-downloads)
--d <ids>GPU device IDs (e.g. 0 or 0,1)
--port <n>API server port
--disable-prefix-cacheDisable prefix caching (on by default)
--ui-serverEnable built-in ChatGPT-like web UI
--isq <fmt>In-situ quantization (q2k, q3k, q4k, q5k, q6k, q8_0)
--kvcache-dtype <mode>KV cache quantization: fp8, turbo8, turbo4, turbo3

© guoqingbao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/test-model of guoqingbao/xinfer.

Open the folder on GitHubat commit b88c153

Compare with similar skills

Test Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Model this skillguoqingbao/xinfer333—~2.6kAutomated safety check: PassMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Aqua Model Lifecycleoracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Aqua Model Lifecycle

    oracle/accelerated-data-science

    Official

    Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

    125 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed

More from guoqingbao/xinfer

  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 29 days ago
    Auto-check: notes
  • Check Model

    guoqingbao/xinfer

    Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

    333 GitHub stars~3.8k tokensUpdated 29 days ago
    Auto-check passed

Questions about Test Model

What does Test Model do?

Test LLM models served by xinfer for correctness, output quality, and performance. Test Model is an agent skill from guoqingbao/xinfer. Test LLM models served by xinfer for correctness, output quality, and performance.

When should I use Test Model?

Test Model fits situations like: the user asks to test; verify models — either from a local folder path; huggingFace model IDs.

How do I install Test Model in Claude Code?

Run `npx skills add guoqingbao/xinfer --skill test-model -a claude-code`. Or copy the skill folder (.cursor/skills/test-model in guoqingbao/xinfer) into .claude/skills/test-model in your project. Claude Code loads it when a task matches its description.

How do I install Test Model in Codex?

Run `npx skills add guoqingbao/xinfer --skill test-model -a codex`. Or copy the skill folder (.cursor/skills/test-model in guoqingbao/xinfer) into .agents/skills/test-model in your project. Codex loads it when a task matches its description.

Can I use Test Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guoqingbao/xinfer --skill test-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-model, .gemini/skills/test-model, .github/skills/test-model and .opencode/skills/test-model in your project.

What does Test Model need to run?

Going by SKILL.md and its folder, Test Model needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Test Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Test Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Model use?

Test Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Model use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Model?

Skills that share tags, products or a category with Test Model: Resolve (alexziskind1/model-shelf, 130 stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Aqua Model Lifecycle (oracle/accelerated-data-science, 125 stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Model?

guoqingbao (a GitHub user) maintains it in guoqingbao/xinfer, which has 333 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 9, 2026.

Source: guoqingbao/xinfer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.