Market Data
zhongkaifu/TensorSharp
Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).
Guided workflow for adding a new model architecture to llama.cpp.
$ npx skills add JakeATX/llamAmpere --skill add-new-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install JakeATX/llamAmpere add-new-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/add-new-model .claude/skills/add-new-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "add-new-model" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/add-new-model into .claude/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/JakeATX/llamAmpere/tree/main/skills/add-new-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add JakeATX/llamAmpere --skill add-new-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install JakeATX/llamAmpere add-new-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/add-new-model .agents/skills/add-new-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "add-new-model" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/add-new-model into .agents/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add JakeATX/llamAmpere --skill add-new-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install JakeATX/llamAmpere add-new-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/add-new-model .cursor/skills/add-new-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "add-new-model" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/add-new-model into .cursor/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/JakeATX/llamAmpere.git --path skills/add-new-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add JakeATX/llamAmpere --skill add-new-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install JakeATX/llamAmpere add-new-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/add-new-model .gemini/skills/add-new-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "add-new-model" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/add-new-model into .gemini/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install JakeATX/llamAmpere add-new-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add JakeATX/llamAmpere --skill add-new-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/add-new-model .github/skills/add-new-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "add-new-model" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/add-new-model into .github/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add JakeATX/llamAmpere --skill add-new-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install JakeATX/llamAmpere add-new-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/add-new-model .opencode/skills/add-new-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "add-new-model" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/add-new-model into .opencode/skills/add-new-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-new-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
add-new-modelGuided workflow for adding a new model architecture to llama.cpp.
Add New Model is an agent skill from JakeATX/llamAmpere. Guided workflow for adding a new model architecture to llama.cpp. Use when the user wants to add/port a new model architecture.
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with llama.cpp, Qwen and CUDA. The repository describes itself as: llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): example: 95+ tok/s over a 100K-token generation at temperature 1 for qwen3.8… The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 93ac427. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghgitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Add New Model loads about 4.1k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 2,371 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from JakeATX/llamAmpere at commit 93ac427, republished under its MIT licence (© JakeATX). 2,371 words, ~4,100 tokens.
.claude/skills/add-new-model/SKILL.md (or your agent's skills folder).Fork context: this repo is the TurboQuant fork. The upstream-facing rules below (maintainer discussions, CPU-first follow-ups, PR conventions) apply to work destined for ggml-org, but most model work here is fork-internal. Read the AGENTS.md overview first if not in context - in particular, a new architecture must work with the fork's turbo KV cache types and TQ weight types, and the shared files a model touches (llama-arch.h, llama-graph.cpp, llama-context.cpp) carry turbo wiring that must not be disturbed. The fork-specific additions at the end take precedence where they conflict.
This skill walks a contributor through adding a new model architecture. AI-generated code is permitted in this project, so you may write full implementations for the steps below rather than only pointing at patterns - but follow AGENTS.md's AI usage policy throughout:
Assisted-by: (never Co-authored-by:) and only after explicit confirmation.ggml_rope_ext (see Step 2's RoPE tips) genuinely cannot express what this model needs, the contributor should open an issue to discuss it with maintainers first - not send a PR with a custom RoPE implementation.Before starting, read CONTRIBUTING.md, AGENTS.md and docs/development/HOWTO-add-model.md if they are not already in context. Also run git log --oneline -- src/models and look at at least 3 recent PRs that added a model (their merge commits/diffs) - this shows current convention more reliably than the docs, which can lag behind.
Ask the contributor:
config.json/weights available locally?gh search issues "<model name>" and gh search prs "<model name>" in the ggml-org/llama.cpp repo. In this fork, also check for in-flight work: git branch -r | grep <model> and gh search prs --repo TheTom/llama-cpp-turboquant "<model name>" - model work often lives in fork experiment branches (e.g. the existing origin/feat/gemma4-mtp, origin/feat/gemma4uv, origin/oscar branches). If an existing PR covers it, the contributor should comment there and collaborate rather than open a duplicate (per CONTRIBUTING.md's AI Usage Policy).If the contributor doesn't know the closest reference architecture, you may grep conversion/*.py and src/models/*.cpp for architectures with a similar config shape (layer count, head count, MoE expert count, norm placement) and suggest 1-2 candidates - but let the contributor confirm the choice rather than picking one yourself; this choice is a design decision they need to own.
Do not proceed to Step 1 until the contributor has answered these and named a reference architecture.
Follow HOWTO-add-model.md section 1 for the actual touch points (conversion class registration, constants.py, tensor_mapping.py, etc.) - don't re-derive them here, read them from the doc.
Skill-specific addition: for each touch point, show the contributor the equivalent code in the reference architecture they named in Step 0 before writing the new version, and check that they understand what's different about their model (e.g. non-standard tensor shapes, extra hparams) rather than just copying the pattern silently.
Follow HOWTO-add-model.md section 2 for the actual touch points (llm_arch enum, LLM_ARCH_NAMES, hparam loading, RoPE type case, etc.), including its "Tips and tricks" section for ggml_rope_ext gotchas.
Skill-specific addition: never hack around RoPE with a custom sin/cos implementation - see the RoPE rule above.
Follow HOWTO-add-model.md section 3 for the actual touch points (src/models/<name>.cpp struct, llama_model_mapping registration, etc.).
Skill-specific addition: before writing src/models/<name>.cpp, read at least 10 other files under src/models/ (pick a mix, not just the one reference architecture) to confirm the struct layout, naming, and style you're about to write actually matches current convention - the pattern drifts over time and the HOWTO doc can lag behind it.
Only do this if the contributor flagged a vision/audio encoder in Step 0. Follow HOWTO-add-model.md section 4 and docs/multimodal.md for the actual touch points (MmprojModel subclass, clip.cpp, mtmd.cpp, encoder graph in tools/mtmd/models, etc.).
Skill-specific addition, and read this carefully: whether the multimodal encoder can be bundled into the same PR as the base text-model support depends on how conventional the change is. It's OK to bundle it if the encoder support is conventional - i.e. no new infra or logic is needed, it's just a new cgraph reusing existing preprocessing/projector machinery (e.g. siglip/pixtral/qwen with just a new projector). If it requires anything beyond that - a new preprocessor, non-standard projector logic, or changes to shared libmtmd infra/logic - STOP, tell the contributor this is non-conventional, and have them land the text model first with the encoder as a dedicated follow-up PR. Do not let this decision pass silently - call it out explicitly to the contributor before writing any clip.cpp/mtmd.cpp code.
Only do this if the model needs a new built-in chat template (src/llama-chat.cpp) or a new output parser (see docs/development/parsing.md and docs/autoparser.md). If either is needed beyond what a user-supplied Jinja template already covers, treat it as its own dedicated follow-up PR, not part of the base model-support PR - call this out explicitly to the contributor rather than silently bundling it in.
These recur often enough in review comments on past add-model PRs that they're worth checking proactively, not just waiting for a reviewer to catch them:
sliding_window_pattern, norm-eps) must hard-error if missing, not silently fall back to a default.llm_arch can be reused by multiple fine-tunes with different templates, and a baked-in C++ default fails silently for those.test-chat-auto-parser <jinja> shows what it detects).eot at conversion time isn't always sufficient - in long/agentic generations a model can emit the closing sequence as literal text instead of the token, so generation never stops on EOG and raw text leaks past the parser. Verify this case, not just the token path.GGML_SCHED_MAX_SPLIT_INPUTS).build_attn_mha when FA is enabled.GGML_PAD(..., 256)), apply the padding after all other size adjustments, not before - otherwise later logic can un-align it again.llama_model_n_swa()) rather than mutating hparams to fake the behavior - hparams may be read elsewhere for unrelated purposes.super().set_gguf_parameters()) rather than re-deriving it - large blocks of code that duplicate what TextModel/MmprojModel already provide will get flagged as redundant.norm(1 + weight)) and permutations/chunking at conversion time, not in the graph - see HOWTO-add-model.md's "Prefer conversion-time tensor modifications" tip (Gemma 3 folds its 1 + into the weights, Qwen3-Next permutes in modify_tensors). Doing these at runtime in the graph is very likely to be rejected as over-complicated; if you genuinely can't do it at conversion time, open a discussion first explaining why rather than implementing it in the graph.weight * scale with a constant scale is usually better applied at inference time instead of being folded into the weight at conversion. The scale conceptually applies to the activation, not the weight, so folding it in can hurt numerical stability, and it shifts the weight's value range in a way that can make quantization worse.Reference: examples/model-conversion/README.md.
examples/model-conversion.tools/cli, tools/completion, tools/imatrix, tools/quantize, and tools/server.CONTRIBUTING.md (relaxed in this fork - fork-internal model work may bundle backends, but CPU-first still catches the most bugs cheapest).-ctk q8_0 -ctv turbo3 (and turbo2/turbo4), with flash attention. This exercises the rotation/padding path: head dims not a multiple of 128 must zero-pad correctly, and MLA/DeepSeek4 archs must use identical K/V types. Also quantize a copy to TQ4_1S and confirm it loads, decodes coherently, and runs on the CUDA kernels (TQ weights are arch-agnostic, but a new arch's graph must route them through the fused-TQ path, not the mmvq abort).AGENTS.md (and CONTRIBUTING.md's "Coding guidelines"/"Naming guidelines" sections) - this is a separate pass from functional testing and is just as important: no forced line-wrapping, no unicode punctuation, minimal/non-redundant comments, snake_case naming (kebab-case for file names), matching indentation/brace style, etc.Take precedence over the upstream-facing rules where they conflict:
src/llama-arch.h/.cpp and builds its graph in src/models/<name>.cpp, but src/llama-graph.cpp also carries the fork's inverse-WHT post-processing (FA and non-FA paths) and src/llama-context.cpp carries the turbo FA auto-enable, head-dim padding, and MLA K/V equality checks. Do not refactor or reformat those blocks while adding a model; a structural change there that breaks turbo semantics is worse than a cosmetic diff.-ctv turbo3), not just default f16. If the model's head dims are not multiples of 128, exercise the zero-padding path explicitly. If it is MLA-family, K and V cache types must match and V rotation/padding is skipped - the existing DeepSeek4 handling is the reference.llama-quantize ... TQ4_1S and the result must run (CUDA fused-TQ path; MoE models auto-disable CUDA graphs for TQ MUL_MAT_ID - do not try to re-enable).src/models/<name>.cpp and conversion code structurally close to upstream style so the next upstream rebase stays clean; add a fork: tag comment on any deliberately fork-divergent block (see the code-review skill's TurboQuant section).src/models/, conversion/, gguf-py/) mostly matches upstream; if a merge/rebase produced stacked duplicates in gguf-py/gguf/constants.py (it has before - import gguf crashes), that is a separate cleanup, not part of the model PR.code-review skill on the diff first - it catches the convention and scope issues reviewers flag most often, and it's recommended to do this locally before pushing the PR..github/pull_request_template.md describing how AI was used (do not omit or understate this).© JakeATX, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/add-new-model of JakeATX/llamAmpere.
Open the folder on GitHubat commit 93ac427
Add New Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Add New Model this skillJakeATX/llamAmpere | 149 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Market Datazhongkaifu/TensorSharp | 559 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Model Serving MinefieldBlackwellboy/model-serving-minefield | 135 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Hugging Face Local Modelshuggingface/skills | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 | |
| Add Modelguoqingbao/xinfer | 334 | — | ~4.2k | Automated safety check: Notes | MIT |
zhongkaifu/TensorSharp
Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
huggingface/skills
Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.
guoqingbao/xinfer
Adapt and port new LLM model architectures to this xinfer project.
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
JakeATX/llamAmpere
Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.
JakeATX/llamAmpere
Opinionated app components building on top of ./ui primitives
Categories
Guided workflow for adding a new model architecture to llama.cpp. Add New Model is an agent skill from JakeATX/llamAmpere.cpp.
Add New Model fits situations like: the user wants to add/port a new model architecture; tasks that involve LLM inference and serving.
Run `npx skills add JakeATX/llamAmpere --skill add-new-model -a claude-code`. Or copy the skill folder (skills/add-new-model in JakeATX/llamAmpere) into .claude/skills/add-new-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add JakeATX/llamAmpere --skill add-new-model -a codex`. Or copy the skill folder (skills/add-new-model in JakeATX/llamAmpere) into .agents/skills/add-new-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JakeATX/llamAmpere --skill add-new-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-new-model, .gemini/skills/add-new-model, .github/skills/add-new-model and .opencode/skills/add-new-model in your project.
Going by SKILL.md and its folder, Add New Model needs the command-line tools its instructions call (gh and git). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Add New Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Add New Model: Market Data (zhongkaifu/TensorSharp, 559 stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars) and Hugging Face Local Models (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
JakeATX (a GitHub user) maintains it in JakeATX/llamAmpere, which has 149 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.
Source: JakeATX/llamAmpere on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.