Agent skill

Check Model

by guoqingbao in guoqingbao/xinfer

Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

MITAuto-check passedAI & LLM Engineering

Install Check Model

skills CLI
$ npx skills add guoqingbao/xinfer --skill check-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guoqingbao/xinfer check-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/check-model .claude/skills/check-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
check-model
GitHub stars
334
Token cost
~3.8k tokens
SKILL.md length
1,342 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

  • Works in 6 steps: Gather Model Information → Parse Config and Identify Model Type → Validate Tensor Format Against… → …
  • The user asks to check
  • SKILL.md covers Phase 0: Gather Model…, Phase 1: Parse Config and…, Phase 2: Validate Tensor… and Phase 3: Multi-Rank…, plus 3 more sections
  • Reaches huggingface.co

What it does

Check Model is an agent skill from guoqingbao/xinfer. Check model compatibility with xinfer before loading. Validates config.json, weight tensor shapes and naming, quantization format correctness, and multi-rank (tensor-parallel) divisibility. Use when the user asks to check, validate, audit, or verify a model will load correctly — from a HuggingFace URL/config, local path, or pasted tensor info.

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving and Model hubs and datasets. It works with Hugging Face, Qwen, vLLM and Rust. The repository describes itself as: Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime. The licence is MIT.

When your agent uses it

  • The user asks to check
  • Verify a model will load correctly — from a HuggingFace URL/config
  • Pasted tensor info

Example prompts

  • “/check-model”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Gather Model Information
  2. Parse Config and Identify Model Type
  3. Validate Tensor Format Against Quantization Config
  4. Multi-Rank Divisibility Analysis
  5. Report Findings
  6. Common Issues Reference

What it can do on your machine

Read from SKILL.md and the folder at commit b88c153. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Check Model loads about 3.8k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 1,342 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from guoqingbao/xinfer at commit b88c153, republished under its MIT licence (© guoqingbao). 1,342 words, ~3,776 tokens.

Download SKILL.mdSave it as .claude/skills/check-model/SKILL.md (or your agent's skills folder).
name
check-model
description
Check model compatibility with xinfer before loading. Validates config.json, weight tensor shapes and naming, quantization format correctness, and multi-rank (tensor-parallel) divisibility. Use when the user asks to check, validate, audit, or verify a model will load correctly — from a HuggingFace URL/config, local path, or pasted tensor info.

Check Model — Pre-Load Compatibility Audit for xinfer

Phase 0: Gather Model Information

Collect model config and tensor info. Accept any of:

InputHow to use
HuggingFace config URLFetch config.json from the URL (e.g. https://huggingface.co/<id>/blob/main/config.json)
HuggingFace model IDFetch config from https://huggingface.co/<id>/raw/main/config.json
Local model pathRead <path>/config.json directly
Pasted config JSONParse inline
Tensor infoUser pastes tensor names/shapes/dtypes from HuggingFace safetensor viewer or provides local weights

If tensor info is missing, ask the user to provide it. They can get it by clicking any .safetensors file in the HuggingFace model page and copying the tensor tree.

For local models, extract tensor info with:

python
import json, struct, sys, glob, os
path = sys.argv[1]
for sf in sorted(glob.glob(os.path.join(path, "*.safetensors"))):
    with open(sf, "rb") as f:
        n = struct.unpack("<Q", f.read(8))[0]
        header = json.loads(f.read(n))
    for k, v in sorted(header.items()):
        if k != "__metadata__":
            print(f"{k}\t{v.get('shape')}\t{v.get('dtype')}")

Phase 1: Parse Config and Identify Model Type

Extract from config.json:

Core parameters
FieldRequiredNotes
architecturesYesDetermines model type and loader path
hidden_sizeYesOr nested under text_config for multimodal
num_attention_headsYesQ heads for full attention
num_key_value_headsYesKV heads for GQA
head_dimIf availableDefaults to hidden_size / num_attention_heads
num_hidden_layersYesTotal layer count
vocab_sizeYesEmbedding table size
Hybrid (Qwen3.5/Qwen3Next) parameters
FieldWhen presentNotes
layer_typesQwen3.5/Qwen3NextArray of "linear_attention" / "full_attention"
linear_num_key_headsHybrid modelsGDN K heads (may differ from V heads)
linear_num_value_headsHybrid modelsGDN V heads
linear_key_head_dimHybrid modelsPer-head K dimension
linear_value_head_dimHybrid modelsPer-head V dimension
linear_conv_kernel_dimHybrid modelsConv1d kernel size (typically 4)
full_attention_intervalHybrid modelsHow often full attention appears
MoE parameters
FieldWhen presentNotes
num_expertsMoE modelsExpert count per layer
num_experts_per_tokMoE modelsTop-K routing
moe_intermediate_sizeMoE modelsPer-expert FFN hidden dim
shared_expert_intermediate_sizeSome MoEShared expert dim
Quantization config
FieldNotes
quantization_config.quant_method"modelopt", "compressed-tensors", "fp8", "gptq", "awq"
quantization_config.quant_algoFor modelopt: "NVFP4", "FP4"
quantization_config.formatFor compressed-tensors: "nvfp4-pack-quantized", "mxfp4-pack-quantized"
quantization_config.config_groupsWeight/activation quant specs
quantization_config.ignoreLayers excluded from quantization (stored as BF16/FP16)
quantization_config.weight_block_sizeFP8 block dimensions (e.g. [128, 128])
Normalize quant_method

Apply the same normalization as QuantConfig::normalize_compressed_tensors():

Raw quant_methodquant_algo / formatNormalized
modeloptNVFP4 or FP4nvfp4
modelopt(detect from config_groups)nvfp4
compressed-tensorsformat contains nvfp4nvfp4
compressed-tensorsformat contains mxfp4mxfp4
fp8-fp8
gptq-gptq
awq-awq

Phase 2: Validate Tensor Format Against Quantization Config

For each layer type, check that tensor names and dtypes match the expected format.

2a. Determine which layers are quantized vs skipped

Parse the ignore list from quantization_config. Layers in the ignore list should have BF16/FP16 weights (weight tensor only). Layers NOT in the ignore list should have quantized tensors.

The ignore list supports:

  • Literal paths: "model.language_model.layers.0.linear_attn.in_proj_qkv"
  • Regex patterns: "re:.*linear_attn.*"
  • Glob-style wildcards: "model.visual*", "mtp.layers.0*"
2b. Format-specific tensor checks
Unquantized (BF16/FP16)

Expected tensors per linear layer:

  • weight — dtype BF16 or F16, shape [out_dim, in_dim]
  • bias (optional) — dtype BF16 or F16

Check: No extra scale/packed tensors should be present.

FP8 (quant_method == "fp8")

Expected tensors per linear layer:

  • weight — dtype U8 (F8_E4M3), shape [out_dim, in_dim]
  • weight_scale or weight_scale_inv — dtype F32, shape [out_dim/by, in_dim/bx] where [by, bx] = weight_block_size (default [128, 128])
  • bias (optional)

Check: weight_block_size must have exactly 2 elements. Scale dimensions must match ceil(out_dim/by) x ceil(in_dim/bx).

NVFP4 — ModelOpt format (quant_method == "modelopt" + quant_algo == "NVFP4")

Expected tensors per quantized linear layer:

  • weight — dtype U8, shape [out_dim, in_dim/2] (packed FP4, 2 values per byte)
  • weight_scale — dtype F8_E4M3 (U8), shape [out_dim, in_dim/16] (group_size=16)
  • weight_scale_2 — dtype F32, scalar (global weight scale, direct multiplier)
  • input_scale — dtype F32, scalar (activation scale)

Check: weight shape dim1 must be exactly in_dim/2. Scale dim1 must be in_dim/16.

NVFP4 — Compressed-tensors format (quant_method == "compressed-tensors" + nvfp4 format)

Expected tensors per quantized linear layer:

  • weight_packed — dtype U8, shape [out_dim, in_dim/2]
  • weight_scale — dtype F8_E4M3 (U8), shape [out_dim, in_dim/16]
  • weight_global_scale — dtype F32, scalar or [1] (divisor, inverted at load time)
  • input_global_scale — dtype F32, scalar or [1] (divisor, inverted at load time)

Check: Same shape rules as ModelOpt, but different tensor names.

MXFP4 (quant_method == "mxfp4" or compressed-tensors with mxfp4)

Expected tensors per quantized linear layer:

  • weight_packed or blocks — dtype U8, shape [out_dim, in_dim/2]
  • weight_scale or scales — dtype U8 (F8_E8M0), shape [out_dim, in_dim/32] (group_size=32)

Check: Scale dim1 must be in_dim/32.

GGUF

GGUF models are self-contained (no config.json). Weight tensor names use blk.{i} prefix mapped to model.layers.{i}. Quantization is per-tensor via GGML dtypes (Q4_K, Q6_K, Q8_0, etc.).

Check: Not applicable for safetensors checks. GGUF has its own loader path via QLinear / QMatMul.

2c. Loader path tensor name resolution

The xinfer loaders try tensor names in priority order. Verify the model's tensors match at least one:

ComponentTensor name priority (first match wins)
NVFP4/MXFP4 packed weightsweight_packed > weight > blocks
NVFP4/MXFP4 scalesweight_scale > scales
NVFP4 global scaleweight_global_scale (inverted) > weight_scale_2 (direct)
NVFP4 input scaleinput_scale (direct) > input_global_scale (inverted)
FP8 scaleweight_scale > weight_scale_inv

Flag any mismatch where the model uses a tensor name not in the priority list.

2d. Hybrid model (GDN) quantization detection

For Qwen3.5/Qwen3Next models with quantization config, the GatedDeltaNet layer has its own quantization detection (is_weight_quantized) that checks each linear_attn sublayer independently:

quant_methodDetection logic
fp8Has weight_scale or weight_scale_inv
mxfp4Has weight_packed or blocks
nvfp4(weight_packed or blocks) AND (weight_scale or scales) OR (weight_scale_2 or input_scale) AND (weight_scale or scales)

If a linear_attn sublayer is in the ignore list and has only BF16 weight, the detection returns false, and the layer loads as unquantized. Verify this matches the tensor info.


Show full SKILL.md (502 more words)Show less

Phase 3: Multi-Rank Divisibility Analysis

For each candidate world_size in [1, 2, 4, 8], check all TP-sharded dimensions.

3a. Full Attention
ComponentGlobal dimShard dimDivisibility requirement
Q projectionnum_attention_heads * head_dimdim 0num_attention_heads % world_size == 0
K/V projectionnum_kv_heads * head_dimdim 0num_kv_heads >= world_size: num_kv_heads % world_size == 0; num_kv_heads < world_size: world_size % num_kv_heads == 0 (replicated KV mode)
O projectionnum_attention_heads * head_dimdim 1Same as Q

For quantized (FP8/NVFP4/MXFP4) Q/K/V:

  • Column linear shard dim 0: per-rank out_dim / world_size must be cleanly divisible
  • For FP8: per-rank start must be aligned to weight_block_size[0] (default 128)
  • For NVFP4: per-rank output must be divisible (no block alignment needed for dim 0 shard)
3b. GatedDeltaNet (Linear Attention)
ComponentGlobal dimRequirement
num_v_headslinear_num_value_heads% world_size == 0
num_k_headslinear_num_key_heads% world_size == 0
in_proj_qkv (merged)Q=key_dim_global, K=key_dim_global, V=value_dim_globalEach chunk % world_size == 0
in_proj_zvalue_dim_global% world_size == 0
in_proj_b/anum_v_heads_global% world_size == 0
A_log / dt_biasnum_v_heads_global% world_size == 0
conv1d (Q block)key_dim_globalkey_dim / world_size channels per rank
conv1d (V block)value_dim_global% world_size == 0
out_projvalue_dim_globalRow linear dim 1 % world_size == 0

Where:

  • key_dim_global = linear_num_key_heads * linear_key_head_dim
  • value_dim_global = linear_num_value_heads * linear_value_head_dim
3c. MoE Experts
ComponentGlobal dimShard dimRequirement
gate/up_projmoe_intermediate_sizedim 0% world_size == 0
down_projmoe_intermediate_sizedim 1% world_size == 0

For NVFP4/MXFP4 MoE:

  • gate/up packed dim0: moe_intermediate_size / world_size per rank
  • down packed dim1: (moe_intermediate_size / pack_factor) / world_size per rank
3d. Shared Expert MLP

Same rules as standard MLP with shared_expert_intermediate_size:

  • Column linear (gate/up): shared_expert_intermediate_size % world_size == 0
  • Row linear (down): shared_expert_intermediate_size % world_size == 0
3e. NVFP4/MXFP4 Scale Alignment

For NVFP4 (group_size=16): after sharding, verify per_rank_in_dim % 16 == 0 for dim-1 shards. For MXFP4 (group_size=32): verify per_rank_in_dim % 32 == 0 for dim-1 shards. For FP8: verify per-rank boundaries align to weight_block_size.

3f. Embedding / LM Head
  • embed_tokens: replicated (not sharded), no divisibility constraint.
  • lm_head: replicated, no constraint. But if tie_word_embeddings is true, verify lm_head doesn't exist as a separate tensor (should reuse embed_tokens.weight).

Phase 4: Report Findings

Present results in a structured format:

Model Summary
Architecture: Qwen3_5MoeForConditionalGeneration
Model Type: qwen3_5_moe (Hybrid MoE with linear attention)
Quantization: nvfp4 (compressed-tensors format)
Layers: 48 (36 linear_attention + 12 full_attention)
Hidden size: 3072
Full attention: 32 Q heads, 2 KV heads, head_dim=256
Linear attention: 16 K heads, 64 V heads, head_dim=128
MoE: 256 experts, top-8, intermediate=1024
Shared expert: intermediate=1024
Tensor Format Validation
[OK] Linear attention layers (BF16, in ignore list)
[OK] Full attention layers (NVFP4 compressed-tensors: weight_packed + weight_scale + weight_global_scale)
[OK] MoE experts (NVFP4 compressed-tensors: per-expert weight_packed)
[OK] Shared expert MLP (NVFP4 compressed-tensors: weight_packed)
[WARN] lm_head: in ignore list, stored as BF16
Multi-Rank Compatibility
| Component | 1 GPU | 2 GPUs | 4 GPUs | 8 GPUs |
|-----------|-------|--------|--------|--------|
| Full attn Q heads (32) | OK | 16 | 8 | 4 |
| Full attn KV heads (2) | OK | 1 | repl(2) | repl(4) |
| GDN K heads (16) | OK | 8 | 4 | 2 |
| GDN V heads (64) | OK | 32 | 16 | 8 |
| MoE inter (1024) | OK | 512 | 256 | 128 |
| Overall | OK | OK | OK | OK |
Issues Found

Flag any problems:

  • [ERROR] — Will fail to load (missing tensors, wrong names, indivisible dims)
  • [WARN] — May cause issues (unusual format, edge case)
  • [INFO] — Informational (features detected, fallback paths)

Phase 5: Common Issues Reference

Known tensor name mismatches
Model sourcePacked weight namexinfer loader support
ModelOpt NVFP4weight (U8)Single-GPU: OK. Multi-GPU merged chunks: requires weight fallback in load_merged_chunks
Compressed-tensors NVFP4weight_packedOK everywhere
Legacy MXFP4/NVFP4blocksOK (final fallback)
GatedDeltaNet TP-safe loading

The in_proj_qkv tensor requires special merged-chunk loading for multi-GPU:

  • MergedParallelColumnLinear::load_merged_chunks splits Q, K, V independently
  • For quantized models (FP8/NVFP4/MXFP4), each chunk must be sharded within the quantized domain
  • For BF16 (ignore-listed layers), falls through to the unquantized path
Replicated KV heads

When num_kv_heads < world_size:

  • kv_head_shard uses replicated mode: ranks_per_kv_head = world_size / num_kv_heads
  • Each KV head is shared by ranks_per_kv_head consecutive ranks
  • Requires world_size % num_kv_heads == 0

Key Source Files

FileRelevance
src/models/layers/distributed.rsTP column/row linear, load_merged_chunks, kv_head_shard
src/models/layers/linear.rsLnFp8, LnNvfp4, LnMxfp4 loaders, tensor name resolution
src/models/layers/deltanet.rsGatedDeltaNet loading, is_weight_quantized, projection sharding
src/models/layers/attention.rsFull attention QKV loading, packed QKV for FP8
src/models/layers/moe.rsFusedMoeNvfp4, FusedMoeMxfp4, FusedMoeFp8 expert loading
src/utils/config.rsQuantConfig, normalize_compressed_tensors, should_skip_module

© guoqingbao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/check-model of guoqingbao/xinfer.

Open the folder on GitHubat commit b88c153

Compare with similar skills

Check Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Check Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Check Model this skillguoqingbao/xinfer334—~3.8kAutomated safety check: PassMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Aqua Model Lifecycleoracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.0

Similar skills

  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Aqua Model Lifecycle

    oracle/accelerated-data-science

    Official

    Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

    125 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Rtvi Byom Porting

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

    1.9k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from guoqingbao/xinfer

  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    Auto-check: notes
  • Test Model

    guoqingbao/xinfer

    Test LLM models served by xinfer for correctness, output quality, and performance.

    334 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Check Model

What does Check Model do?

Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer. Check Model is an agent skill from guoqingbao/xinfer. Check model compatibility with xinfer before loading.

When should I use Check Model?

Check Model fits situations like: the user asks to check; verify a model will load correctly — from a HuggingFace URL/config; pasted tensor info.

How do I install Check Model in Claude Code?

Run `npx skills add guoqingbao/xinfer --skill check-model -a claude-code`. Or copy the skill folder (.cursor/skills/check-model in guoqingbao/xinfer) into .claude/skills/check-model in your project. Claude Code loads it when a task matches its description.

How do I install Check Model in Codex?

Run `npx skills add guoqingbao/xinfer --skill check-model -a codex`. Or copy the skill folder (.cursor/skills/check-model in guoqingbao/xinfer) into .agents/skills/check-model in your project. Codex loads it when a task matches its description.

Can I use Check Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guoqingbao/xinfer --skill check-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/check-model, .gemini/skills/check-model, .github/skills/check-model and .opencode/skills/check-model in your project.

What does Check Model need to run?

SKILL.md names no scripts, command-line tools or credentials: Check Model is instructions for the agent only. Our summary lists: Python 3.

Does Check Model access the network?

SKILL.md names 1 domain. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Check Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Check Model use?

Check Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Check Model use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Check Model?

Skills that share tags, products or a category with Check Model: Resolve (alexziskind1/model-shelf, 130 stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Check Model?

guoqingbao (a GitHub user) maintains it in guoqingbao/xinfer, which has 334 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 9, 2026.

Source: guoqingbao/xinfer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.