Agent skill

Add Model

by guoqingbao in guoqingbao/xinfer

Adapt and port new LLM model architectures to this xinfer project.

MITAuto-check: notesAI & LLM Engineering

Install Add Model

skills CLI
$ npx skills add guoqingbao/xinfer --skill add-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guoqingbao/xinfer add-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/add-model .claude/skills/add-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-model
GitHub stars
334
Token cost
~4.2k tokens
SKILL.md length
1,477 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Adapt and port new LLM model architectures to this xinfer project.

  • Works in 8 steps: Gather Required Information → Analyze the Architecture → Check if New Kernels Are Needed → …
  • The user asks to add
  • SKILL.md covers Phase 0: Gather Required…, Phase 1: Analyze the…, Phase 2: Check if New Kernels… and Phase 3: Implement the Model, plus 5 more sections
  • Calls cargo, curl and git; reaches huggingface.co and github.com

What it does

Add Model is an agent skill from guoqingbao/xinfer. Adapt and port new LLM model architectures to this xinfer project. Use when the user asks to add, port, support, or adapt a new model (e.g. Llama, Gemma, Qwen, GPT-OSS, DeepSeek, or any HuggingFace architecture) including safetensors and GGUF formats, Dense and MoE architectures, and quantization formats (MXFP4, NVFP4, FP8, ISQ).

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving and Model hubs and datasets. It works with Hugging Face, llama.cpp, Qwen and DeepSeek. The repository describes itself as: Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime. The licence is MIT.

When your agent uses it

  • The user asks to add
  • Adapt a new model (e.g

Example prompts

  • “/add-model”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Gather Required Information
  2. Analyze the Architecture
  3. Check if New Kernels Are Needed
  4. Implement the Model
  5. Register the Model
  6. Build and Verify
  7. Check Model Compatibility
  8. Test the Model

What it can do on your machine

Read from SKILL.md and the folder at commit b88c153. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo
    • curl
    • git
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Model loads about 4.2k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 1,477 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:319
    sudo powermetrics --samplers gpu_power -i 1000 -n 1 | grep 'GPU'

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from guoqingbao/xinfer at commit b88c153, republished under its MIT licence (© guoqingbao). 1,477 words, ~4,185 tokens.

Download SKILL.mdSave it as .claude/skills/add-model/SKILL.md (or your agent's skills folder).
name
add-model
description
Adapt and port new LLM model architectures to this xinfer project. Use when the user asks to add, port, support, or adapt a new model (e.g. Llama, Gemma, Qwen, GPT-OSS, DeepSeek, or any HuggingFace architecture) including safetensors and GGUF formats, Dense and MoE architectures, and quantization formats (MXFP4, NVFP4, FP8, ISQ).

Add Model — Adapt New LLM Architectures to xinfer

Phase 0: Gather Required Information

Before starting, the agent must collect the inputs below. If any are missing, ask the user explicitly.

Resolution order (try top to bottom, stop at the first that succeeds):

  1. Local model path (preferred) — If the user provides a local directory containing safetensors, read config.json and inspect weight tensors directly from disk. If the user provides a .gguf file, extract metadata and tensor info from the GGUF header (GGUF is self-contained; there is no separate config.json). No further input is needed in either case.
  2. HuggingFace model ID — Look for the model in the local HuggingFace Hub cache (~/.cache/huggingface/hub/). If not cached, fetch config.json from the HuggingFace model repo. If that also fails, fall back to step 3.
  3. Manual input — Ask the user to provide:
    • config.json contents (for safetensors models) or GGUF metadata (for GGUF models).
    • Weight tensor info (names, shapes, dtypes). The user can obtain this by clicking a weight file in the HuggingFace model repo.

Additionally, ask the user if they can provide the Python reference implementation (modeling_<arch>.py from HuggingFace Transformers). This is not strictly required, but significantly improves accuracy — it clarifies the exact forward pass, attention variants, MoE routing, activation functions, and normalization order that config fields alone cannot fully describe.

InputHow to obtain
HuggingFace model ID (e.g. google/gemma-4-26B-A4B-it)User provides, or infer from context
Model config (config.json)Fetch from HF: https://huggingface.co/<id>/blob/main/config.json. Not needed if local model path is provided.
HF tensor info (weight names + shapes)User provides, or read from local safetensors with scripts/inspect_weights.py (create the script if it doesn't exist)
GGUF metadata + tensor info (if GGUF support needed)User provides, or extract from local .gguf with scripts/inspect_gguf.py (create the script if it doesn't exist)
Python reference implementation (optional but recommended)Fetch modeling_<arch>.py from the HuggingFace Transformers GitHub repo
Local model path (optional)User provides path containing config.json + *.safetensors or *.gguf
Extracting info from local weights

For safetensors (Python required):

python
# scripts/inspect_weights.py
import json, sys, struct
path = sys.argv[1]
with open(path, "rb") as f:
    n = struct.unpack("<Q", f.read(8))[0]
    header = json.loads(f.read(n))
for k, v in sorted(header.items()):
    if k != "__metadata__":
        print(f"{k}\t{v.get('shape')}\t{v.get('dtype')}")

For GGUF, use the gguf_helper CLI or Python gguf library to list tensors and metadata keys.


Phase 1: Analyze the Architecture

Read and compare:

  1. config.json — identify: architectures, hidden_size, num_hidden_layers, num_attention_heads, num_key_value_heads, head_dim, intermediate_size, hidden_act, vocab_size, rope_theta, partial_rotary_factor, sliding_window, tie_word_embeddings.
  2. MoE fields (if any): num_experts, num_experts_per_tok / top_k_experts, moe_intermediate_size, norm_topk_prob, routed_scaling_factor.
  3. Special features: attention_k_eq_v, layer_types array, layer_scalar, final_logit_softcapping, attn_logit_softcapping, dual head_dim, dual RoPE, per-expert scaling, shared experts, etc.
  4. Python reference — understand the forward pass, especially:
    • Attention: GQA/MQA, head dim, RoPE variant, sliding window, QK norm, V norm
    • MLP: standard SiLU-gate, GeluTanh, custom SwiGLU, etc.
    • MoE: router structure, expert computation, parallel dense+MoE, renormalization
    • Layer structure: pre-norm vs post-norm, residual connections, layer scalar

Identify which existing model in src/models/ is the closest match. Use it as a starting template.

Architecture classification
TypeCharacteristicsExample models
DenseStandard transformer, no MoELlama, Gemma3, Phi4
MoERouter + experts per layerQwen3-MoE, Gemma4-MoE
Hybrid MoEDense MLP + MoE in parallel per layerGemma4 (dense+MoE parallel)
MultimodalVision encoder + language modelGemma3-VL, Qwen3-VL

Phase 2: Check if New Kernels Are Needed

If the model requires operators not in attention.rs or src/models/layers/:

  1. Clone attention.rs to a local sibling directory (if not already present):
    bash
    cd .. && git clone https://github.com/guoqingbao/attention.rs.git
    Then switch to the commit that this project relying on.
  2. Point xinfer/Cargo.toml to local attention.rs:
    toml
    # In [dependencies], change the git URL to:
    attention-rs = { path = "../attention.rs", ... }
  3. Add both CUDA and Metal kernel implementations:
    • CUDA kernel: attention.rs/src/kernels/src/<name>.cu
    • Metal kernel: attention.rs/src/metal-kernels/src/<name>.metal
    • FFI bindings: attention.rs/src/kernels/src/ffi.rs
    • Rust wrapper: attention.rs/src/<name>.rs
    • Update attention.rs/src/lib.rs: add pub mod <name>;
    • Update attention.rs/src/kernels/build.rs: add rerun-if-changed for .cu
    • Update attention.rs/src/metal-kernels/build.rs: add to METAL_SOURCES
  4. For CUDA BF16 kernels (Ampere+ only), guard with #ifndef NO_BF16_KERNEL and provide F16 fallback or dummy stubs for older GPUs.
  5. Verify compilation: cargo check --features metal (macOS) or cargo check --features cuda (Linux).

Phase 3: Implement the Model

3a. Create model file

Create src/models/<arch>.rs. Follow the established pattern from the closest existing model:

Key struct pattern:

pub struct <Arch>DecoderLayer {
    self_attn: Attention,
    mlp: MLP,
    moe: Option<<Arch>MoE>,      // if MoE
    input_layernorm: NormX,
    post_attention_layernorm: NormX,
    // ... additional norms for hybrid MoE
}

pub struct <Arch>ForCausalLM {
    embed_tokens: candle_nn::Embedding,
    layers: Vec<<Arch>DecoderLayer>,
    norm: NormX,
    lm_head: ReplicatedLinear,
    // ...
}

Required public methods on <Arch>ForCausalLM:

rust
pub fn new(vb: &VarBuilderX, comm: Rc<Comm>, config: &Config, dtype: DType,
           is_rope_i: bool, device: &Device,
           progress_reporter: Arc<RwLock<Box<dyn ProgressLike>>>) -> Result<Self>

pub fn forward(&self, input_ids: &Tensor, positions: &Tensor,
               kv_caches: Option<&Vec<(Tensor, Tensor)>>,
               input_metadata: &InputMetadata,
               _embeded_inputs: bool) -> Result<Tensor>

pub fn forward_embedding(&self, ..same args..) -> Result<Tensor>

pub fn get_vocab_size(&self) -> usize

The 5th _embeded_inputs: bool argument is required by the model_call! macro.

3b. Weight loading paths

The model must support two VarBuilder types via vb.is_qvar_builder():

PathVarBuilderWeight prefix pattern
HF safetensorsEither::Leftlanguage_model.model.layers.{i} (multimodal) or model.layers.{i}
GGUFEither::Rightmodel.layers.{i} (maps from blk.{i})

For norm and lm_head:

ComponentHF prefixGGUF prefix
Final normlanguage_model.model.normmodel.norm
LM head (untied)lm_headmodel.output
LM head (tied)language_model.model.embed_tokensmodel.embed_tokens
3c. MoE construction

For models with MoE, the Gemma4MoE / FusedMoe selection pattern:

rust
let moe = if is_qvar_builder {
    FusedMoeGGUF::new(config, vb.clone(), comm.clone(), dtype)?
} else if quant_config == "mxfp4" {
    FusedMoeMxfp4::new(config, vb.pp("mlp"), comm.clone(), dtype)?
} else if config.quant.is_some() {
    FusedMoeISQ::new(config, vb.pp("mlp"), comm.clone(), dtype)?
} else {
    FusedMoe::new(config, vb.pp("mlp"), comm.clone(), dtype)?
};

Important: If the model's router gate is NOT at the standard mlp.gate path (e.g. Gemma4 uses router.proj), use FusedMoe::new_with_gate(config, gate_vb, experts_vb, ...) and FusedMoeISQ::new_with_gate(...).

For GGUF models with packed ffn_gate_up_exps (instead of separate ffn_gate_exps + ffn_up_exps), FusedMoeGGUF::new() auto-detects and handles this.

3d. Packed expert layout

If the model uses packed gate_up_proj with a non-standard layout, add the architecture name to the layout resolvers in src/models/layers/moe.rs:

  • resolve_packed_gate_up_layout() — InterPacked if shape is [experts, 2*intermediate, hidden]
  • resolve_packed_down_layout() — HiddenInter if shape is [experts, hidden, intermediate]

Phase 4: Register the Model

4a. src/models/mod.rs
rust
pub mod <arch>;
4b. src/utils/config.rs

Add variant to ModelType enum:

rust
pub enum ModelType {
    // ...
    <Arch>,
}
4c. src/utils/mod.rs — get_arch_rope function

Add mappings in order:

  1. GGUF canonical_arch: "<gguf_arch>" => "<HFArchitectureName>".to_string()
  2. Architecture to ModelType + chat template: Add match arm for HF architecture names
  3. Rope type selection: ("<arch_lower>", false) in the rope map
  4. HF config parsing (multimodal wrapper): If the model wraps text_config, add handler to extract nested config, MoE config, and rope_parameters
  5. GGUF extra_config_json: If the model has special GGUF metadata (sliding_window_pattern, layer_types, etc.), add extraction block
  6. GGUF MoE config: Add mod_cfg construction from GGUF expert metadata
  7. hidden_act override: If the model uses a non-Silu activation, override after config construction
  8. require_model_penalty(): Add architecture names
  9. Multimodal GGUF warning: Add if applicable
Show full SKILL.md (553 more words)Show less
4d. src/core/runner.rs
  1. Add use crate::models::<arch>::<Arch>ForCausalLM;
  2. Add <Arch>(Arc<<Arch>ForCausalLM>) to Model enum
  3. Add <Arch> => <Arch>ForCausalLM to build_model! macro
  4. Add <Arch> => EmbedInputs to graph_wrapper! macro (or ImageData for multimodal)
  5. Add <Arch> => false to both model_call! invocations (or true / image handling for multimodal)
  6. Add ModelType::<Arch> to disable_flash_attn if needed
  7. Add Model::<Arch>(model) => model.get_vocab_size() to get_vocab_size
4e. src/server/parser.rs
  1. Add ModelType::<Arch> to ToolConfig::for_model_type()
  2. Add ModelType::<Arch> to parser_name_for_model()
  3. Add ModelType::<Arch> to structured output format handling

Phase 5: Build and Verify

Build commands
PlatformCommand
macOS (Metal)cargo build --release --features metal
CUDA (basic)cargo build --release --features "cuda,flashinfer"
CUDA (full)cargo build --release --features "cuda,flashinfer,nccl"
CUDA (sm90+)Add cutlass feature: --features "cuda,flashinfer,nccl,cutlass"

If permission errors occur on target/, use: CARGO_TARGET_DIR=/tmp/xinfer-check cargo check --features metal

Verify compilation
bash
# Check xinfer compiles
cargo check --features metal   # or cuda

# Check attention.rs compiles (if modified)
cd ../attention.rs && cargo check --features metal   # or cuda

Fix all errors and warnings before proceeding.


Phase 6: Check Model Compatibility

Before testing, invoke the check-model skill (.cursor/skills/check-model/SKILL.md) to validate the new model's tensor format and multi-rank compatibility.

Provide:

  1. The model's config.json (local path or HuggingFace URL)
  2. Weight tensor info (names, shapes, dtypes) — from a local safetensors file or HuggingFace model page

The check-model skill will verify:

  • Tensor names match the loader's expected format (e.g. weight_packed vs weight vs blocks for FP4)
  • Quantized vs unquantized layer detection matches the ignore list in quantization_config
  • All TP-sharded dimensions are divisible by common GPU counts (1, 2, 4, 8)
  • FP8 block alignment and FP4 scale group alignment are correct for multi-rank

Fix any [ERROR] findings before proceeding to the test phase. [WARN] items should be reviewed but may not block loading.


Phase 7: Test the Model

Start the server

Single-GPU:

bash
# Metal (macOS)
cargo run --release --features metal -- --m <model_id_or_path> --port 8080 # or use --w to specify local model path

# CUDA
cargo run --release --features "cuda,flashinfer,cutlass" -- --m <model_id_or_path> --port 8080

Multi-GPU (CUDA):

bash
# --d used to specify device ids
xinfer --m <model_id_or_path> --port 8080 --d 0,1
Before retrying model loading

Always kill all previous instances and verify GPU memory is freed:

bash
# Kill all xinfer and runner processes
pkill -f xinfer;
sleep 2

# Check GPU memory (CUDA)
nvidia-smi --query-gpu=memory.used,memory.free --format=csv,noheader

# Check GPU memory (Metal)
sudo powermetrics --samplers gpu_power -i 1000 -n 1 | grep 'GPU'

Ensure the target GPU(s) have sufficient free memory for the model before loading.

Send test requests
bash
# Basic completion, depend on the api server port started, default 8000
curl -s http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model_id>",
    "messages": [{"role": "user", "content": "Hello, who are you?"}],
    "max_tokens": 64,
    "temperature": 0.7
  }' | python3 -m json.tool

# Streaming
curl -s http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model_id>",
    "messages": [{"role": "user", "content": "Write a haiku about Rust."}],
    "max_tokens": 64,
    "stream": true
  }'
Debugging common issues
SymptomLikely causeFix
Panic at weight loadingWrong VarBuilder prefix or tensor shape mismatchCheck HF vs GGUF prefix mapping; verify tensor shapes match config
NaN/Inf in outputMissing activation, wrong dtype, or norm misconfigurationVerify hidden_act, check if model needs GeluPytorchTanh vs Silu; check norm dtype (F32 for GGUF/ISQ)
Gibberish outputWrong RoPE params, wrong head_dim, or tied embeddings misconfiguredVerify rope_theta, partial_rotary_factor, head_dim, tie_word_embeddings
Server crash on decodeKV cache shape mismatch or sliding window misconfiguredVerify head_dim per attention layer matches KV cache allocation
CUDA OOMModel too large for GPUUse ISQ quantization (--quant q4k), or use multi-GPU with --tensor-parallel
Slow performanceFlash attention disabled or wrong featuresEnsure flashinfer feature is enabled; check disable_flash_attn isn't matching your model
Precision validation

Compare outputs against the Python reference:

  1. Run the same prompt through both the Python HF model and xinfer
  2. Compare logits for the first few tokens (top-5 should match)
  3. For MoE models, verify expert routing produces the same top-k experts

Quick Reference: Key Files

FilePurpose
src/models/<arch>.rsModel implementation
src/models/mod.rsModule registration
src/models/layers/moe.rsMoE layer (FusedMoe, FusedMoeGGUF, FusedMoeISQ, FusedMoeMxfp4)
src/models/layers/attention.rsAttention layer (handles GQA, RoPE, sliding window)
src/models/layers/mlp.rsDense MLP layer
src/models/layers/rotary_emb.rsRoPE implementations (standard, scaling, YaRN)
src/models/layers/mod.rsVarBuilderX definition
src/utils/config.rsConfig struct, ModelType enum, MoEConfig
src/utils/mod.rsConfig parsing (HF + GGUF), architecture registration, chat templates
src/core/runner.rsModel enum, build_model!, model_call!, graph_wrapper! macros
src/server/parser.rsTool parsing, structured output per model type
Cargo.tomlFeature flags: metal, cuda, flashinfer, nccl, cutlass
build.shBuild script

© guoqingbao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .cursor/skills/add-model of guoqingbao/xinfer.

Open the folder on GitHubat commit b88c153

Compare with similar skills

Add Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Model this skillguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Aqua Model Lifecycleoracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.0
Open Weightsericrisco/rsc-harness174—~4.1kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Aqua Model Lifecycle

    oracle/accelerated-data-science

    Official

    Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

    125 GitHub stars~1.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Open Weights

    ericrisco/rsc-harness

    A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.

    174 GitHub stars~4.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from guoqingbao/xinfer

  • Check Model

    guoqingbao/xinfer

    Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

    334 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Test Model

    guoqingbao/xinfer

    Test LLM models served by xinfer for correctness, output quality, and performance.

    334 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Add Model

What does Add Model do?

Adapt and port new LLM model architectures to this xinfer project. Add Model is an agent skill from guoqingbao/xinfer. Adapt and port new LLM model architectures to this xinfer project.

When should I use Add Model?

Add Model fits situations like: the user asks to add; adapt a new model (e.g.

How do I install Add Model in Claude Code?

Run `npx skills add guoqingbao/xinfer --skill add-model -a claude-code`. Or copy the skill folder (.cursor/skills/add-model in guoqingbao/xinfer) into .claude/skills/add-model in your project. Claude Code loads it when a task matches its description.

How do I install Add Model in Codex?

Run `npx skills add guoqingbao/xinfer --skill add-model -a codex`. Or copy the skill folder (.cursor/skills/add-model in guoqingbao/xinfer) into .agents/skills/add-model in your project. Codex loads it when a task matches its description.

Can I use Add Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guoqingbao/xinfer --skill add-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-model, .gemini/skills/add-model, .github/skills/add-model and .opencode/skills/add-model in your project.

What does Add Model need to run?

Going by SKILL.md and its folder, Add Model needs the command-line tools its instructions call (cargo, curl, git and python3). Our summary lists: Python 3.

Does Add Model access the network?

SKILL.md names 2 domains. In commands or code: huggingface.co and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Add Model safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Add Model use?

Add Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Model use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Model?

Skills that share tags, products or a category with Add Model: Resolve (alexziskind1/model-shelf, 130 stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Aqua Model Lifecycle (oracle/accelerated-data-science, 125 stars) and Open Weights (ericrisco/rsc-harness, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Model?

guoqingbao (a GitHub user) maintains it in guoqingbao/xinfer, which has 334 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 9, 2026.

Source: guoqingbao/xinfer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.