Resolve
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
Adapt and port new LLM model architectures to this xinfer project.
$ npx skills add guoqingbao/xinfer --skill add-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install guoqingbao/xinfer add-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/add-model .claude/skills/add-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "add-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/add-model into .claude/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/add-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add guoqingbao/xinfer --skill add-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install guoqingbao/xinfer add-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.cursor/skills/add-model .agents/skills/add-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "add-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/add-model into .agents/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add guoqingbao/xinfer --skill add-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install guoqingbao/xinfer add-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.cursor/skills/add-model .cursor/skills/add-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "add-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/add-model into .cursor/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/guoqingbao/xinfer.git --path .cursor/skills/add-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add guoqingbao/xinfer --skill add-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install guoqingbao/xinfer add-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.cursor/skills/add-model .gemini/skills/add-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "add-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/add-model into .gemini/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install guoqingbao/xinfer add-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add guoqingbao/xinfer --skill add-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .github/skills && cp -r skills-src/.cursor/skills/add-model .github/skills/add-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "add-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/add-model into .github/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add guoqingbao/xinfer --skill add-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install guoqingbao/xinfer add-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/guoqingbao/xinfer.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.cursor/skills/add-model .opencode/skills/add-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "add-model" agent skill from https://github.com/guoqingbao/xinfer/tree/main/.cursor/skills/add-model into .opencode/skills/add-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
add-modelAdapt and port new LLM model architectures to this xinfer project.
Add Model is an agent skill from guoqingbao/xinfer. Adapt and port new LLM model architectures to this xinfer project. Use when the user asks to add, port, support, or adapt a new model (e.g. Llama, Gemma, Qwen, GPT-OSS, DeepSeek, or any HuggingFace architecture) including safetensors and GGUF formats, Dense and MoE architectures, and quantization formats (MXFP4, NVFP4, FP8, ISQ).
Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving and Model hubs and datasets. It works with Hugging Face, llama.cpp, Qwen and DeepSeek. The repository describes itself as: Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b88c153. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
cargocurlgitpython3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.cogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Add Model loads about 4.2k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 1,477 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
sudo powermetrics --samplers gpu_power -i 1000 -n 1 | grep 'GPU'Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from guoqingbao/xinfer at commit b88c153, republished under its MIT licence (© guoqingbao). 1,477 words, ~4,185 tokens.
.claude/skills/add-model/SKILL.md (or your agent's skills folder).Before starting, the agent must collect the inputs below. If any are missing, ask the user explicitly.
Resolution order (try top to bottom, stop at the first that succeeds):
config.json and inspect weight tensors directly from disk. If the user provides a .gguf file, extract metadata and tensor info from the GGUF header (GGUF is self-contained; there is no separate config.json). No further input is needed in either case.~/.cache/huggingface/hub/). If not cached, fetch config.json from the HuggingFace model repo. If that also fails, fall back to step 3.config.json contents (for safetensors models) or GGUF metadata (for GGUF models).Additionally, ask the user if they can provide the Python reference implementation (modeling_<arch>.py from HuggingFace Transformers). This is not strictly required, but significantly improves accuracy — it clarifies the exact forward pass, attention variants, MoE routing, activation functions, and normalization order that config fields alone cannot fully describe.
| Input | How to obtain |
|---|---|
HuggingFace model ID (e.g. google/gemma-4-26B-A4B-it) | User provides, or infer from context |
Model config (config.json) | Fetch from HF: https://huggingface.co/<id>/blob/main/config.json. Not needed if local model path is provided. |
| HF tensor info (weight names + shapes) | User provides, or read from local safetensors with scripts/inspect_weights.py (create the script if it doesn't exist) |
| GGUF metadata + tensor info (if GGUF support needed) | User provides, or extract from local .gguf with scripts/inspect_gguf.py (create the script if it doesn't exist) |
| Python reference implementation (optional but recommended) | Fetch modeling_<arch>.py from the HuggingFace Transformers GitHub repo |
| Local model path (optional) | User provides path containing config.json + *.safetensors or *.gguf |
For safetensors (Python required):
# scripts/inspect_weights.py
import json, sys, struct
path = sys.argv[1]
with open(path, "rb") as f:
n = struct.unpack("<Q", f.read(8))[0]
header = json.loads(f.read(n))
for k, v in sorted(header.items()):
if k != "__metadata__":
print(f"{k}\t{v.get('shape')}\t{v.get('dtype')}")For GGUF, use the gguf_helper CLI or Python gguf library to list tensors and metadata keys.
Read and compare:
config.json — identify: architectures, hidden_size, num_hidden_layers, num_attention_heads, num_key_value_heads, head_dim, intermediate_size, hidden_act, vocab_size, rope_theta, partial_rotary_factor, sliding_window, tie_word_embeddings.num_experts, num_experts_per_tok / top_k_experts, moe_intermediate_size, norm_topk_prob, routed_scaling_factor.attention_k_eq_v, layer_types array, layer_scalar, final_logit_softcapping, attn_logit_softcapping, dual head_dim, dual RoPE, per-expert scaling, shared experts, etc.Identify which existing model in src/models/ is the closest match. Use it as a starting template.
| Type | Characteristics | Example models |
|---|---|---|
| Dense | Standard transformer, no MoE | Llama, Gemma3, Phi4 |
| MoE | Router + experts per layer | Qwen3-MoE, Gemma4-MoE |
| Hybrid MoE | Dense MLP + MoE in parallel per layer | Gemma4 (dense+MoE parallel) |
| Multimodal | Vision encoder + language model | Gemma3-VL, Qwen3-VL |
If the model requires operators not in attention.rs or src/models/layers/:
attention.rs to a local sibling directory (if not already present):cd .. && git clone https://github.com/guoqingbao/attention.rs.gitxinfer/Cargo.toml to local attention.rs:# In [dependencies], change the git URL to:
attention-rs = { path = "../attention.rs", ... }attention.rs/src/kernels/src/<name>.cuattention.rs/src/metal-kernels/src/<name>.metalattention.rs/src/kernels/src/ffi.rsattention.rs/src/<name>.rsattention.rs/src/lib.rs: add pub mod <name>;attention.rs/src/kernels/build.rs: add rerun-if-changed for .cuattention.rs/src/metal-kernels/build.rs: add to METAL_SOURCES#ifndef NO_BF16_KERNEL and provide F16 fallback or dummy stubs for older GPUs.cargo check --features metal (macOS) or cargo check --features cuda (Linux).Create src/models/<arch>.rs. Follow the established pattern from the closest existing model:
Key struct pattern:
pub struct <Arch>DecoderLayer {
self_attn: Attention,
mlp: MLP,
moe: Option<<Arch>MoE>, // if MoE
input_layernorm: NormX,
post_attention_layernorm: NormX,
// ... additional norms for hybrid MoE
}
pub struct <Arch>ForCausalLM {
embed_tokens: candle_nn::Embedding,
layers: Vec<<Arch>DecoderLayer>,
norm: NormX,
lm_head: ReplicatedLinear,
// ...
}Required public methods on <Arch>ForCausalLM:
pub fn new(vb: &VarBuilderX, comm: Rc<Comm>, config: &Config, dtype: DType,
is_rope_i: bool, device: &Device,
progress_reporter: Arc<RwLock<Box<dyn ProgressLike>>>) -> Result<Self>
pub fn forward(&self, input_ids: &Tensor, positions: &Tensor,
kv_caches: Option<&Vec<(Tensor, Tensor)>>,
input_metadata: &InputMetadata,
_embeded_inputs: bool) -> Result<Tensor>
pub fn forward_embedding(&self, ..same args..) -> Result<Tensor>
pub fn get_vocab_size(&self) -> usizeThe 5th _embeded_inputs: bool argument is required by the model_call! macro.
The model must support two VarBuilder types via vb.is_qvar_builder():
| Path | VarBuilder | Weight prefix pattern |
|---|---|---|
| HF safetensors | Either::Left | language_model.model.layers.{i} (multimodal) or model.layers.{i} |
| GGUF | Either::Right | model.layers.{i} (maps from blk.{i}) |
For norm and lm_head:
| Component | HF prefix | GGUF prefix |
|---|---|---|
| Final norm | language_model.model.norm | model.norm |
| LM head (untied) | lm_head | model.output |
| LM head (tied) | language_model.model.embed_tokens | model.embed_tokens |
For models with MoE, the Gemma4MoE / FusedMoe selection pattern:
let moe = if is_qvar_builder {
FusedMoeGGUF::new(config, vb.clone(), comm.clone(), dtype)?
} else if quant_config == "mxfp4" {
FusedMoeMxfp4::new(config, vb.pp("mlp"), comm.clone(), dtype)?
} else if config.quant.is_some() {
FusedMoeISQ::new(config, vb.pp("mlp"), comm.clone(), dtype)?
} else {
FusedMoe::new(config, vb.pp("mlp"), comm.clone(), dtype)?
};Important: If the model's router gate is NOT at the standard mlp.gate path (e.g. Gemma4 uses router.proj), use FusedMoe::new_with_gate(config, gate_vb, experts_vb, ...) and FusedMoeISQ::new_with_gate(...).
For GGUF models with packed ffn_gate_up_exps (instead of separate ffn_gate_exps + ffn_up_exps), FusedMoeGGUF::new() auto-detects and handles this.
If the model uses packed gate_up_proj with a non-standard layout, add the architecture name to the layout resolvers in src/models/layers/moe.rs:
resolve_packed_gate_up_layout() — InterPacked if shape is [experts, 2*intermediate, hidden]resolve_packed_down_layout() — HiddenInter if shape is [experts, hidden, intermediate]src/models/mod.rspub mod <arch>;src/utils/config.rsAdd variant to ModelType enum:
pub enum ModelType {
// ...
<Arch>,
}src/utils/mod.rs — get_arch_rope functionAdd mappings in order:
"<gguf_arch>" => "<HFArchitectureName>".to_string()("<arch_lower>", false) in the rope maptext_config, add handler to extract nested config, MoE config, and rope_parametersmod_cfg construction from GGUF expert metadatahidden_act override: If the model uses a non-Silu activation, override after config constructionrequire_model_penalty(): Add architecture namessrc/core/runner.rsuse crate::models::<arch>::<Arch>ForCausalLM;<Arch>(Arc<<Arch>ForCausalLM>) to Model enum<Arch> => <Arch>ForCausalLM to build_model! macro<Arch> => EmbedInputs to graph_wrapper! macro (or ImageData for multimodal)<Arch> => false to both model_call! invocations (or true / image handling for multimodal)ModelType::<Arch> to disable_flash_attn if neededModel::<Arch>(model) => model.get_vocab_size() to get_vocab_sizesrc/server/parser.rsModelType::<Arch> to ToolConfig::for_model_type()ModelType::<Arch> to parser_name_for_model()ModelType::<Arch> to structured output format handling| Platform | Command |
|---|---|
| macOS (Metal) | cargo build --release --features metal |
| CUDA (basic) | cargo build --release --features "cuda,flashinfer" |
| CUDA (full) | cargo build --release --features "cuda,flashinfer,nccl" |
| CUDA (sm90+) | Add cutlass feature: --features "cuda,flashinfer,nccl,cutlass" |
If permission errors occur on target/, use: CARGO_TARGET_DIR=/tmp/xinfer-check cargo check --features metal
# Check xinfer compiles
cargo check --features metal # or cuda
# Check attention.rs compiles (if modified)
cd ../attention.rs && cargo check --features metal # or cudaFix all errors and warnings before proceeding.
Before testing, invoke the check-model skill (.cursor/skills/check-model/SKILL.md) to validate the new model's tensor format and multi-rank compatibility.
Provide:
config.json (local path or HuggingFace URL)The check-model skill will verify:
weight_packed vs weight vs blocks for FP4)ignore list in quantization_configFix any [ERROR] findings before proceeding to the test phase. [WARN] items should be reviewed but may not block loading.
Single-GPU:
# Metal (macOS)
cargo run --release --features metal -- --m <model_id_or_path> --port 8080 # or use --w to specify local model path
# CUDA
cargo run --release --features "cuda,flashinfer,cutlass" -- --m <model_id_or_path> --port 8080Multi-GPU (CUDA):
# --d used to specify device ids
xinfer --m <model_id_or_path> --port 8080 --d 0,1Always kill all previous instances and verify GPU memory is freed:
# Kill all xinfer and runner processes
pkill -f xinfer;
sleep 2
# Check GPU memory (CUDA)
nvidia-smi --query-gpu=memory.used,memory.free --format=csv,noheader
# Check GPU memory (Metal)
sudo powermetrics --samplers gpu_power -i 1000 -n 1 | grep 'GPU'Ensure the target GPU(s) have sufficient free memory for the model before loading.
# Basic completion, depend on the api server port started, default 8000
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<model_id>",
"messages": [{"role": "user", "content": "Hello, who are you?"}],
"max_tokens": 64,
"temperature": 0.7
}' | python3 -m json.tool
# Streaming
curl -s http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "<model_id>",
"messages": [{"role": "user", "content": "Write a haiku about Rust."}],
"max_tokens": 64,
"stream": true
}'| Symptom | Likely cause | Fix |
|---|---|---|
| Panic at weight loading | Wrong VarBuilder prefix or tensor shape mismatch | Check HF vs GGUF prefix mapping; verify tensor shapes match config |
| NaN/Inf in output | Missing activation, wrong dtype, or norm misconfiguration | Verify hidden_act, check if model needs GeluPytorchTanh vs Silu; check norm dtype (F32 for GGUF/ISQ) |
| Gibberish output | Wrong RoPE params, wrong head_dim, or tied embeddings misconfigured | Verify rope_theta, partial_rotary_factor, head_dim, tie_word_embeddings |
| Server crash on decode | KV cache shape mismatch or sliding window misconfigured | Verify head_dim per attention layer matches KV cache allocation |
| CUDA OOM | Model too large for GPU | Use ISQ quantization (--quant q4k), or use multi-GPU with --tensor-parallel |
| Slow performance | Flash attention disabled or wrong features | Ensure flashinfer feature is enabled; check disable_flash_attn isn't matching your model |
Compare outputs against the Python reference:
| File | Purpose |
|---|---|
src/models/<arch>.rs | Model implementation |
src/models/mod.rs | Module registration |
src/models/layers/moe.rs | MoE layer (FusedMoe, FusedMoeGGUF, FusedMoeISQ, FusedMoeMxfp4) |
src/models/layers/attention.rs | Attention layer (handles GQA, RoPE, sliding window) |
src/models/layers/mlp.rs | Dense MLP layer |
src/models/layers/rotary_emb.rs | RoPE implementations (standard, scaling, YaRN) |
src/models/layers/mod.rs | VarBuilderX definition |
src/utils/config.rs | Config struct, ModelType enum, MoEConfig |
src/utils/mod.rs | Config parsing (HF + GGUF), architecture registration, chat templates |
src/core/runner.rs | Model enum, build_model!, model_call!, graph_wrapper! macros |
src/server/parser.rs | Tool parsing, structured output per model type |
Cargo.toml | Feature flags: metal, cuda, flashinfer, nccl, cutlass |
build.sh | Build script |
© guoqingbao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .cursor/skills/add-model of guoqingbao/xinfer.
Open the folder on GitHubat commit b88c153
Add Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Add Model this skillguoqingbao/xinfer | 334 | — | ~4.2k | Automated safety check: Notes | MIT | |
| Resolvealexziskind1/model-shelf | 130 | — | ~792 | Automated safety check: Pass | MIT | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Aqua Model Lifecycleoracle/accelerated-data-science | 125 | — | ~1.4k | Automated safety check: Pass | UPL-1.0 | |
| Open Weightsericrisco/rsc-harness | 174 | — | ~4.1k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 |
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
oracle/accelerated-data-science
Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.
ericrisco/rsc-harness
A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
guoqingbao/xinfer
Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.
guoqingbao/xinfer
Test LLM models served by xinfer for correctness, output quality, and performance.
Categories
Adapt and port new LLM model architectures to this xinfer project. Add Model is an agent skill from guoqingbao/xinfer. Adapt and port new LLM model architectures to this xinfer project.
Add Model fits situations like: the user asks to add; adapt a new model (e.g.
Run `npx skills add guoqingbao/xinfer --skill add-model -a claude-code`. Or copy the skill folder (.cursor/skills/add-model in guoqingbao/xinfer) into .claude/skills/add-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add guoqingbao/xinfer --skill add-model -a codex`. Or copy the skill folder (.cursor/skills/add-model in guoqingbao/xinfer) into .agents/skills/add-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guoqingbao/xinfer --skill add-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-model, .gemini/skills/add-model, .github/skills/add-model and .opencode/skills/add-model in your project.
Going by SKILL.md and its folder, Add Model needs the command-line tools its instructions call (cargo, curl, git and python3). Our summary lists: Python 3.
SKILL.md names 2 domains. In commands or code: huggingface.co and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Add Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Add Model: Resolve (alexziskind1/model-shelf, 130 stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Aqua Model Lifecycle (oracle/accelerated-data-science, 125 stars) and Open Weights (ericrisco/rsc-harness, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
guoqingbao (a GitHub user) maintains it in guoqingbao/xinfer, which has 334 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 9, 2026.
Source: guoqingbao/xinfer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.