Agent skill

Add New Model

by JakeATX in JakeATX/llamAmpere

Guided workflow for adding a new model architecture to llama.cpp.

MITAuto-check passedAI & LLM Engineering

Install Add New Model

skills CLI
$ npx skills add JakeATX/llamAmpere --skill add-new-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JakeATX/llamAmpere add-new-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/add-new-model .claude/skills/add-new-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-new-model
GitHub stars
149
Token cost
~4.1k tokens
SKILL.md length
2,371 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Guided workflow for adding a new model architecture to llama.cpp.

  • Works in 6 steps: Scope and dedup check → Convert the model to GGUF → Define the architecture in llama.cpp → …
  • The user wants to add/port a new model architecture
  • SKILL.md covers Step 0 - Scope and dedup check, Step 1 - Convert the model to…, Step 2 - Define the… and Step 3 - Build the GGML graph, plus 6 more sections
  • Calls gh and git

What it does

Add New Model is an agent skill from JakeATX/llamAmpere. Guided workflow for adding a new model architecture to llama.cpp. Use when the user wants to add/port a new model architecture.

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with llama.cpp, Qwen and CUDA. The repository describes itself as: llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): example: 95+ tok/s over a 100K-token generation at temperature 1 for qwen3.8… The licence is MIT.

When your agent uses it

  • The user wants to add/port a new model architecture
  • Tasks that involve LLM inference and serving

Example prompts

  • “/add-new-model”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Scope and dedup check
  2. Convert the model to GGUF
  3. Define the architecture in llama.cpp
  4. Build the GGML graph
  5. Optional: multimodal encoder
  6. Optional: chat template / parsing support

What it can do on your machine

Read from SKILL.md and the folder at commit 93ac427. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add New Model loads about 4.1k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 2,371 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from JakeATX/llamAmpere at commit 93ac427, republished under its MIT licence (© JakeATX). 2,371 words, ~4,100 tokens.

Download SKILL.mdSave it as .claude/skills/add-new-model/SKILL.md (or your agent's skills folder).
name
add-new-model
description
Guided workflow for adding a new model architecture to llama.cpp. Use when the user wants to add/port a new model architecture.

Add a new model architecture to llama.cpp

Fork context: this repo is the TurboQuant fork. The upstream-facing rules below (maintainer discussions, CPU-first follow-ups, PR conventions) apply to work destined for ggml-org, but most model work here is fork-internal. Read the AGENTS.md overview first if not in context - in particular, a new architecture must work with the fork's turbo KV cache types and TQ weight types, and the shared files a model touches (llama-arch.h, llama-graph.cpp, llama-context.cpp) carry turbo wiring that must not be disturbed. The fork-specific additions at the end take precedence where they conflict.

This skill walks a contributor through adding a new model architecture. AI-generated code is permitted in this project, so you may write full implementations for the steps below rather than only pointing at patterns - but follow AGENTS.md's AI usage policy throughout:

  • The contributor is 100% responsible for every line, however it was produced. They must be able to explain and defend any part of it to a reviewer. Check in with them as you go (don't silently generate everything and hand over a finished diff) so they actually absorb what was written.
  • Before writing code, make sure the contributor owns the design choices for this architecture (which reference model to follow, how non-standard bits like RoPE variants or MoE routing should be handled) - AI accelerates a design the contributor has already made, it doesn't make the design for them.
  • Disclosure is mandatory: any AI-meaningful contribution must be disclosed per the PR template. Remind the contributor of this before they open the PR.
  • Never write the PR description, commit message, GitHub issue/discussion post, or reviewer replies - those must come from the contributor. If asked to commit on their behalf, use Assisted-by: (never Co-authored-by:) and only after explicit confirmation.
  • If the requested change looks large or introduces a new pattern not covered here, pause and tell the user this kind of change is likely to need prior discussion with maintainers before a PR.
  • Keep the PR self-contained. If the work would require a lot of unconventional changes outside the new model file(s) (e.g. touching shared graph-building code, the sampler, or core APIs in ways other models don't), STOP and tell the contributor to open a discussion/issue first - invasive or excessive changes get closed without full review.
  • Do not bundle unrelated work into this PR - see Step 4 and Step 5 below for the specifics on multimodal and chat-template/parsing work.
  • Never hack around RoPE with a custom sin/cos implementation. Several past PRs tried this and were closed. If the existing ggml_rope_ext (see Step 2's RoPE tips) genuinely cannot express what this model needs, the contributor should open an issue to discuss it with maintainers first - not send a PR with a custom RoPE implementation.

Before starting, read CONTRIBUTING.md, AGENTS.md and docs/development/HOWTO-add-model.md if they are not already in context. Also run git log --oneline -- src/models and look at at least 3 recent PRs that added a model (their merge commits/diffs) - this shows current convention more reliably than the docs, which can lag behind.

Step 0 - Scope and dedup check

Ask the contributor:

  1. Which model (HF repo id or name)? Is it text-only or does it have a multimodal (vision/audio) encoder?
  2. Do they already have the HF config.json/weights available locally?
  3. Have they checked for an existing PR/issue on this model? Suggest gh search issues "<model name>" and gh search prs "<model name>" in the ggml-org/llama.cpp repo. In this fork, also check for in-flight work: git branch -r | grep <model> and gh search prs --repo TheTom/llama-cpp-turboquant "<model name>" - model work often lives in fork experiment branches (e.g. the existing origin/feat/gemma4-mtp, origin/feat/gemma4uv, origin/oscar branches). If an existing PR covers it, the contributor should comment there and collaborate rather than open a duplicate (per CONTRIBUTING.md's AI Usage Policy).
  4. What existing supported architecture is this model closest to (e.g. "Llama-like with sliding window", "MoE like DBRX", "BERT-style encoder")?

If the contributor doesn't know the closest reference architecture, you may grep conversion/*.py and src/models/*.cpp for architectures with a similar config shape (layer count, head count, MoE expert count, norm placement) and suggest 1-2 candidates - but let the contributor confirm the choice rather than picking one yourself; this choice is a design decision they need to own.

Do not proceed to Step 1 until the contributor has answered these and named a reference architecture.

Step 1 - Convert the model to GGUF

Follow HOWTO-add-model.md section 1 for the actual touch points (conversion class registration, constants.py, tensor_mapping.py, etc.) - don't re-derive them here, read them from the doc.

Skill-specific addition: for each touch point, show the contributor the equivalent code in the reference architecture they named in Step 0 before writing the new version, and check that they understand what's different about their model (e.g. non-standard tensor shapes, extra hparams) rather than just copying the pattern silently.

Step 2 - Define the architecture in llama.cpp

Follow HOWTO-add-model.md section 2 for the actual touch points (llm_arch enum, LLM_ARCH_NAMES, hparam loading, RoPE type case, etc.), including its "Tips and tricks" section for ggml_rope_ext gotchas.

Skill-specific addition: never hack around RoPE with a custom sin/cos implementation - see the RoPE rule above.

Step 3 - Build the GGML graph

Follow HOWTO-add-model.md section 3 for the actual touch points (src/models/<name>.cpp struct, llama_model_mapping registration, etc.).

Skill-specific addition: before writing src/models/<name>.cpp, read at least 10 other files under src/models/ (pick a mix, not just the one reference architecture) to confirm the struct layout, naming, and style you're about to write actually matches current convention - the pattern drifts over time and the HOWTO doc can lag behind it.

Step 4 - Optional: multimodal encoder

Only do this if the contributor flagged a vision/audio encoder in Step 0. Follow HOWTO-add-model.md section 4 and docs/multimodal.md for the actual touch points (MmprojModel subclass, clip.cpp, mtmd.cpp, encoder graph in tools/mtmd/models, etc.).

Skill-specific addition, and read this carefully: whether the multimodal encoder can be bundled into the same PR as the base text-model support depends on how conventional the change is. It's OK to bundle it if the encoder support is conventional - i.e. no new infra or logic is needed, it's just a new cgraph reusing existing preprocessing/projector machinery (e.g. siglip/pixtral/qwen with just a new projector). If it requires anything beyond that - a new preprocessor, non-standard projector logic, or changes to shared libmtmd infra/logic - STOP, tell the contributor this is non-conventional, and have them land the text model first with the encoder as a dedicated follow-up PR. Do not let this decision pass silently - call it out explicitly to the contributor before writing any clip.cpp/mtmd.cpp code.

Step 5 - Optional: chat template / parsing support

Only do this if the model needs a new built-in chat template (src/llama-chat.cpp) or a new output parser (see docs/development/parsing.md and docs/autoparser.md). If either is needed beyond what a user-supplied Jinja template already covers, treat it as its own dedicated follow-up PR, not part of the base model-support PR - call this out explicitly to the contributor rather than silently bundling it in.

Show full SKILL.md (1,218 more words)Show less

Common pitfalls (from past PR reviews)

These recur often enough in review comments on past add-model PRs that they're worth checking proactively, not just waiting for a reviewer to catch them:

  • Don't validate the same hparam/config assumption in both the Python conversion script and the C++ load path - pick one layer to own the check, duplicating it just adds maintenance surface.
  • Optional hparams that are genuinely absent from some configs (e.g. a shared-expert count) should be read with an explicit optional/fallback accessor, not assumed present.
  • Hparams that are actually load-bearing (the model produces wrong output or crashes without them, e.g. sliding_window_pattern, norm-eps) must hard-error if missing, not silently fall back to a default.
  • Don't bake a default chat template into the C++ binary - inject it into the GGUF at conversion time instead, since one llm_arch can be reused by multiple fine-tunes with different templates, and a baked-in C++ default fails silently for those.
  • Before writing a dedicated tool-call/output parser, check whether the existing autoparser already handles the template (test-chat-auto-parser <jinja> shows what it detects).
  • Marking a custom EOS/closing-tag token as eot at conversion time isn't always sufficient - in long/agentic generations a model can emit the closing sequence as literal text instead of the token, so generation never stops on EOG and raw text leaks past the parser. Verify this case, not just the token path.
  • If reusing or aliasing an existing pre-tokenizer for convenience, justify and test that choice explicitly - silent reuse is an easy source of subtle tokenizer bugs.
  • Watch for excessive graph splits caused by building per-layer view/index tensors inside the layer loop - hoist tensors that don't vary per layer out of the loop (relevant if you hit GGML_SCHED_MAX_SPLIT_INPUTS).
  • A custom KQ mask fed into flash attention must match FA's expected dtype - cast it to F16 before passing it to build_attn_mha when FA is enabled.
  • When padding a custom KV-cache size to an alignment (e.g. GGML_PAD(..., 256)), apply the padding after all other size adjustments, not before - otherwise later logic can un-align it again.
  • For non-standard cache/SWA (sliding-window-attention) semantics, override the dedicated hook (e.g. llama_model_n_swa()) rather than mutating hparams to fake the behavior - hparams may be read elsewhere for unrelated purposes.
  • Don't ship unfinished or unverified speculative-decoding (e.g. MTP) scaffolding in the base model PR - if it hasn't actually been confirmed to work, pull it out and land it as its own follow-up.
  • Conversion code should call into the base class's existing hparam logic (e.g. super().set_gguf_parameters()) rather than re-deriving it - large blocks of code that duplicate what TextModel/MmprojModel already provide will get flagged as redundant.
  • Do constant tensor modifications (e.g. norm(1 + weight)) and permutations/chunking at conversion time, not in the graph - see HOWTO-add-model.md's "Prefer conversion-time tensor modifications" tip (Gemma 3 folds its 1 + into the weights, Qwen3-Next permutes in modify_tensors). Doing these at runtime in the graph is very likely to be rejected as over-complicated; if you genuinely can't do it at conversion time, open a discussion first explaining why rather than implementing it in the graph.
    • Exception: a plain weight * scale with a constant scale is usually better applied at inference time instead of being folded into the weight at conversion. The scale conceptually applies to the activation, not the weight, so folding it in can hurt numerical stability, and it shifts the weight's value range in a way that can make quantization worse.

Validation checklist

Reference: examples/model-conversion/README.md.

  1. Convert to GGUF, then inspect/run both the original and converted tensors.
  2. Run logits verification (original vs converted). If this model is a new version of an already-supported family, verify the previous version still passes logits verification first - numerical differences may be pre-existing, not caused by the new work. The tools to perform full logits validation are available in examples/model-conversion.
  3. Quantize (including QAT variants if relevant) and re-verify.
  4. Run perplexity evaluation (simple and full).
  5. Sanity-check across tools/cli, tools/completion, tools/imatrix, tools/quantize, and tools/server.
  6. CPU backend first; other backends (CUDA, Metal, ...) can be separate follow-up PRs per CONTRIBUTING.md (relaxed in this fork - fork-internal model work may bundle backends, but CPU-first still catches the most bugs cheapest).
  7. Fork-specific: run the model with turbo KV cache types, not just f16/q8_0 - -ctk q8_0 -ctv turbo3 (and turbo2/turbo4), with flash attention. This exercises the rotation/padding path: head dims not a multiple of 128 must zero-pad correctly, and MLA/DeepSeek4 archs must use identical K/V types. Also quantize a copy to TQ4_1S and confirm it loads, decodes coherently, and runs on the CUDA kernels (TQ weights are arch-agnostic, but a new arch's graph must route them through the fused-TQ path, not the mmvq abort).
  8. Re-review every changed file against the coding/naming guidelines in AGENTS.md (and CONTRIBUTING.md's "Coding guidelines"/"Naming guidelines" sections) - this is a separate pass from functional testing and is just as important: no forced line-wrapping, no unicode punctuation, minimal/non-redundant comments, snake_case naming (kebab-case for file names), matching indentation/brace style, etc.

Fork-specific additions (TurboQuant)

Take precedence over the upstream-facing rules where they conflict:

  • Shared files carry turbo wiring: a new arch registers tensors in src/llama-arch.h/.cpp and builds its graph in src/models/<name>.cpp, but src/llama-graph.cpp also carries the fork's inverse-WHT post-processing (FA and non-FA paths) and src/llama-context.cpp carries the turbo FA auto-enable, head-dim padding, and MLA K/V equality checks. Do not refactor or reformat those blocks while adding a model; a structural change there that breaks turbo semantics is worse than a cosmetic diff.
  • Turbo cache types are the point of this fork: a new arch must be verified with turbo KV (-ctv turbo3), not just default f16. If the model's head dims are not multiples of 128, exercise the zero-padding path explicitly. If it is MLA-family, K and V cache types must match and V rotation/padding is skipped - the existing DeepSeek4 handling is the reference.
  • TQ weight types: new archs should be quantizable with llama-quantize ... TQ4_1S and the result must run (CUDA fused-TQ path; MoE models auto-disable CUDA graphs for TQ MUL_MAT_ID - do not try to re-enable).
  • Rebase hygiene: keep src/models/<name>.cpp and conversion code structurally close to upstream style so the next upstream rebase stays clean; add a fork: tag comment on any deliberately fork-divergent block (see the code-review skill's TurboQuant section).
  • Backend bundling: upstream's "CPU first, backends as follow-ups" is relaxed here for fork-internal work, but do not silently ship an unvalidated backend path - each claimed backend must have been run, not just compiled.
  • Model files are the cleanest part of the tree: the fork's model support (src/models/, conversion/, gguf-py/) mostly matches upstream; if a merge/rebase produced stacked duplicates in gguf-py/gguf/constants.py (it has before - import gguf crashes), that is a separate cleanup, not part of the model PR.

Before opening a PR

  • Run the code-review skill on the diff first - it catches the convention and scope issues reviewers flag most often, and it's recommended to do this locally before pushing the PR.
  • Confirm the contributor can explain every changed line to a reviewer and is prepared to be asked about any of it - this is required regardless of how much of the code was AI-generated.
  • Confirm they did a comprehensive manual review of the full diff, not just a skim.
  • Fill in the AI-disclosure section of .github/pull_request_template.md describing how AI was used (do not omit or understate this).
  • Do not write the PR description, commit message, GitHub issue/discussion text, or any reviewer replies yourself - the contributor writes these.

© JakeATX, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/add-new-model of JakeATX/llamAmpere.

Open the folder on GitHubat commit 93ac427

Compare with similar skills

Add New Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add New Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add New Model this skillJakeATX/llamAmpere149—~4.1kAutomated safety check: PassMIT
Market Datazhongkaifu/TensorSharp559—~1kAutomated safety check: PassBSD-3-Clause
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT

Similar skills

  • Market Data

    zhongkaifu/TensorSharp

    Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).

    559 GitHub stars~1k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 29 days ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from JakeATX/llamAmpere

  • Code Review

    JakeATX/llamAmpere

    Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.

    149 GitHub stars~5.2k tokensUpdated 2 days ago
    Auto-check passed
  • App

    JakeATX/llamAmpere

    Opinionated app components building on top of ./ui primitives

    149 GitHub stars~146 tokensUpdated 2 days ago
    Auto-check passed

Questions about Add New Model

What does Add New Model do?

Guided workflow for adding a new model architecture to llama.cpp. Add New Model is an agent skill from JakeATX/llamAmpere.cpp.

When should I use Add New Model?

Add New Model fits situations like: the user wants to add/port a new model architecture; tasks that involve LLM inference and serving.

How do I install Add New Model in Claude Code?

Run `npx skills add JakeATX/llamAmpere --skill add-new-model -a claude-code`. Or copy the skill folder (skills/add-new-model in JakeATX/llamAmpere) into .claude/skills/add-new-model in your project. Claude Code loads it when a task matches its description.

How do I install Add New Model in Codex?

Run `npx skills add JakeATX/llamAmpere --skill add-new-model -a codex`. Or copy the skill folder (skills/add-new-model in JakeATX/llamAmpere) into .agents/skills/add-new-model in your project. Codex loads it when a task matches its description.

Can I use Add New Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JakeATX/llamAmpere --skill add-new-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-new-model, .gemini/skills/add-new-model, .github/skills/add-new-model and .opencode/skills/add-new-model in your project.

What does Add New Model need to run?

Going by SKILL.md and its folder, Add New Model needs the command-line tools its instructions call (gh and git). Our summary lists: Python 3.

Does Add New Model access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Add New Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add New Model use?

Add New Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add New Model use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add New Model?

Skills that share tags, products or a category with Add New Model: Market Data (zhongkaifu/TensorSharp, 559 stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars) and Hugging Face Local Models (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add New Model?

JakeATX (a GitHub user) maintains it in JakeATX/llamAmpere, which has 149 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.

Source: JakeATX/llamAmpere on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.