Market Data
zhongkaifu/TensorSharp
Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).
Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.
$ npx skills add JakeATX/llamAmpere --skill code-review -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install JakeATX/llamAmpere code-review --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/code-review .claude/skills/code-review && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "code-review" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/code-review into .claude/skills/code-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-review", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/JakeATX/llamAmpere/tree/main/skills/code-reviewType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add JakeATX/llamAmpere --skill code-review -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install JakeATX/llamAmpere code-review --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/code-review .agents/skills/code-review && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "code-review" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/code-review into .agents/skills/code-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-review", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add JakeATX/llamAmpere --skill code-review -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install JakeATX/llamAmpere code-review --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/code-review .cursor/skills/code-review && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "code-review" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/code-review into .cursor/skills/code-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-review", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/JakeATX/llamAmpere.git --path skills/code-review--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add JakeATX/llamAmpere --skill code-review -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install JakeATX/llamAmpere code-review --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/code-review .gemini/skills/code-review && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "code-review" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/code-review into .gemini/skills/code-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-review", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install JakeATX/llamAmpere code-reviewInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add JakeATX/llamAmpere --skill code-review -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/code-review .github/skills/code-review && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "code-review" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/code-review into .github/skills/code-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-review", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add JakeATX/llamAmpere --skill code-review -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install JakeATX/llamAmpere code-review --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/code-review .opencode/skills/code-review && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "code-review" agent skill from https://github.com/JakeATX/llamAmpere/tree/main/skills/code-review into .opencode/skills/code-review/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "code-review", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
code-reviewReview llama.cpp changes against project conventions and common reviewer pitfalls before a PR.
Code Review is an agent skill from JakeATX/llamAmpere. Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. Use when the user wants to review a diff, branch, or PR.
Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving and Code review. It works with llama.cpp, Qwen and CUDA. The repository describes itself as: llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): example: 95+ tok/s over a 100K-token generation at temperature 1 for qwen3.8… The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 93ac427. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ghgitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Code Review loads about 5.2k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 2,933 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from JakeATX/llamAmpere at commit 93ac427, republished under its MIT licence (© JakeATX). 2,933 words, ~5,248 tokens.
.claude/skills/code-review/SKILL.md (or your agent's skills folder).This skill reviews changes against llama.cpp's conventions and the pitfalls that reviewers flag most often, so the contributor can fix them before a maintainer has to. It has two modes:
git diff feature/turboquant-kv-cache...HEAD plus any uncommitted changes (use the upstream merge-base of the fork as the base when the diff against the feature branch is noisy).Fork context: this repo is the TurboQuant fork, not upstream llama.cpp. Most changes here are fork-internal and never go to ggml-org; the upstream-facing rules below (issue-first, quick-reject gates, maintainer approval expectations) still apply to the shared upstream code, but the fork-specific checklist at the end takes precedence for TurboQuant paths. The AGENTS.md overview is required context - read it first if not already in context; its "Known pitfalls" list is the first place to check on any turbo regression.
In both modes the output is private review notes for the user to read and act on - it is never something to post. This is a hard rule from AGENTS.md: an agent must NEVER write, or help write, a PR comment, a review comment, or a reply to a reviewer, by any means including gh. Do not offer to. If the user asks you to post the notes, refuse and point them at that rule. Present findings in the conversation only.
Before starting, read AGENTS.md and CONTRIBUTING.md if not already in context - the "Coding guidelines", "Naming guidelines", and AI usage sections are the baseline this review enforces. For a diff that adds a new model architecture, also read docs/development/HOWTO-add-model.md and consider the dedicated add-new-model skill.
Identify what actually changed and which area checklists below apply. Run git diff --stat (or gh pr view <n> --json files for PR mode) and bucket the touched paths:
conversion/, gguf-py/, src/models/, src/llama-arch.* -> New model / architectureggml/ (any backend, op, or ggml.h) -> ggml / backendggml-turbo-quant.c, turbo/TQ weight or cache types, GGML_OP_TURBO_WHT, turbo kernels in any backend -> TurboQuant / fork-specific (in addition to ggml / backend)include/llama.h and other public headers -> Public APItools/server/ -> ServerAlways run the Scope and quick-reject gate, the Security review, and the General checklist. Run each area checklist whose paths were touched. Additionally, if the diff introduces a new component, subsystem, or piece of infrastructure (a new file/class/module, a new abstraction, or hand-rolled machinery), run the Approach and design review. Tell the user which checklists you're running and why.
These are the patterns that get PRs closed without a full review. Check them first - a finding here is more important than any code nit, because it can mean the change shouldn't be a PR in its current form at all.
CONTRIBUTING.md). If this is a nontrivial feature with no linked issue, flag it and suggest opening one first.gh search prs / gh search issues for the feature. Many closed PRs were duplicates of something already queued.CONTRIBUTING.md). Flag CUDA/Metal/Vulkan/etc. changes bundled into a feature's first PR.ggml_type / quantization type? That carries a disproportionate maintenance burden and needs the full justification package (GGUF sample upload, perplexity vs FP16/BF16 and similar sizes, KL-divergence data, CPU perf numbers). Absent that, it will be rejected regardless of code quality.Mandatory on every review; any finding here is blocking. Rule of thumb: GGUF metadata, tensor shapes, tokenizer/grammar input, and all server/RPC fields are attacker-controlled - bound them before use.
ne[i]*nb[i]/nbytes can overflow on crafted dims into an undersized alloc then heap overflow. Overflow checks must run BEFORE the arithmetic they guard - padding/alignment macros wrap to 0 near SIZE_MAX, so a guard after the pad passes.[i+1], [0..2]).gguf_get_arr_data() or tensor->data to float */int32_t * needs an element-type check first (gguf_get_kv_type() == GGUF_TYPE_ARRAY then gguf_get_arr_type(); type == GGML_TYPE_F32 for tensors). A UINT8 array or I8 tensor passes every length check, then gets read 4 bytes per element - a nearby length check is not a type check.GGML_ASSERT on a file-derived value aborts the process; throw instead where the caller already catches (vocab, model loader, clip).LLAMA_MAX_* array) before indexing; watch checks that only fire when an optional key is present.size_t->int32_t) and signed/unsigned mixing that can bypass a length check and copy past a buffer.stoi/atoi results and catch parse throws; never use a default or derived token id (EOS/BOS/...) as an index without a bounds check.reserve() then index-by-assumed-size, and header fields read before their length is checked.Run this whenever the diff adds a new component, subsystem, or piece of infrastructure. Reviews too often stop at "does it work" - a diff can be correct and still be the wrong approach, and a messy design costs more long-term than a bug. Evaluate the approach, not just the behavior; raising a cleaner one is a high-value finding, not a nit. If you see a better design, describe it concretely rather than just calling the current one bad.
See the add-new-model skill and docs/development/HOWTO-add-model.md for the full workflow; this is the review-time subset that reviewers most often catch:
model.arch when the real dependency is a config/capability value - gate on the hparam/capability, not the architecture enum.src/models/<name>.cpp will be asked to merge with its sibling.tensor_mapping.py, not ad-hoc name matching.ggml_view, not the weight tensor; rely on ggml broadcasting instead of manually duplicating tensors.build_lora_mm and the existing helpers, matching convention - don't leave raw matmuls copied from another arch.ggml_rope_ext genuinely can't express it, that's an issue for discussion, not a PR.-ctk/-ctv q8_0), not just default f16 - new speculative/attention features silently break there.supports_op (and any dispatch/gating condition) must be scoped exactly to the cases being changed - a condition meant for a few quant types must not silently disable or enable everything else.ggml_cuda_get_physical_warp_size() (32 on CUDA, 64 on HIP/ROCm) and the portable helpers.docs/ops.md and the relevant docs/ops/*.csv for the touched backend.test-backend-ops cases, and (per CONTRIBUTING.md) consistency across at least two backends.ggml/ changes, not a sign something is wrong.include/llama.h)Public API changes carry a higher bar than internal ones (CONTRIBUTING.md). Review for:
cb_eval, existing batch/sampler knobs) suffice? If it does, the change likely shouldn't add public surface. This is the single most common reason these PRs are rejected.llama-ext.h), not in llama.h.llama-cpp.h stays a thin convenience layer.int32_t, size_t for sizes/offsets); snake_case; <class>_<method> = <class>_<action>_<noun>; enum values upper-case and prefixed with the enum name; _t suffix for opaque types. Avoid gratuitous signature/ABI changes to existing exported functions.server, embedding, perplexity, etc.tools/server/)tools/server/README-dev.md - out-of-scope features get declined.X-Forwarded-For) or add footguns; things like IP allowlisting belong at a reverse proxy unless there's a trusted-proxy design.tools/mtmd/)v., a., mm. or a.mm. (legacy naming doesn't follow this convention - this is expected, but new code should follow it).ggml_rope_ext instead, see HOWTO-add-model.md. If it can't express the needed behavior, that's a design discussion, not a PR.build_vit should be enough to build the transformer graph for vision models. Do not add a loop to build the transformer graph manually, unless you have a very good reason to do so. If you do, please explain why in the PR description.mtmd.h, open a discussion first.tools/mtmd/README-dev.mdEnforce the AGENTS.md / CONTRIBUTING.md coding and naming guidelines on every changed line - this is a distinct pass from checking that the code works, and matters just as much for review speed:
x, ... used as unicode; use -, ->, x, ... ASCII equivalents.snake_case names; kebab-case (lowercase-with-dashes) file names for C/C++, .h headers; Python files lowercase-with-underscores. Naming optimizes for longest common prefix (number_small, not small_number).void * ptr, int & a, no trailing whitespace; match the surrounding style.for loops are fine here.Co-authored-by: must be reserved for human co-authors; AI contributions (claude, cursor, codex, etc.) must use Assisted-by:; if this point is violated, it's a blocking finding.AGENTS.md for why.Run this in addition to the General checklist whenever the diff touches turbo code, and instead of the upstream-facing gates where they conflict. The AGENTS.md "Known pitfalls" list is the regression checklist - each entry there cost a real bug; verify the diff doesn't disturb those invariants.
GGML_TYPE_TURBO2_0=43, TURBO3_0=44, TQ3_1S=45, TQ4_1S=46, TURBO4_0=47 and GGML_OP_TURBO_WHT are baked into GGUF files and cross-backend dispatch - never renumber, reorder, or repurpose. New fork types go after 47.ggml/src/ggml-turbo-quant.c must stay byte-identical to fork tip conventions; any change needs test-turbo-quant passing (turbo3 MSE=0/Cosine=1.0, turbo4 Cosine=0.9956) plus test-quantize-fns (TQ3_1S/TQ4_1S cases).[[host_name]] kernel instantiations dropped (NULL-pipeline deref), Vulkan SET_ROWS pipeline registration missing TURBO types or the require_full_subgroups=true, subgroup_size=32 flags (abort), CUDA TQ exclusion from the mmvq path dropped (abort with GGML_TQ_NATIVE=1).__byte_perm in centroid LUTs (mmvq-tq.cu) - plain shifts only; the permute chain silently produces garbage on some toolchains. MoE TQ MUL_MAT_ID must keep disabling CUDA graphs (stream sync requirement).-ctk/-ctv turbo3), which require flash attention (auto-enabled). Respect the 128-element rotation block: zero-padding of head dims, and no V rotation/padding for MLA models. K/V types must stay identical for MLA/DeepSeek4.TURBO_LAYER_ADAPTIVE, TURBO_AUTO_ASYMMETRIC, TURBO_SPARSE_V, GGML_TQ_NATIVE, LLAMA_ATTN_ROT_* semantics in docs/KV-cache-quantization.md are the contract - changing behavior without updating that doc is a finding.src/, ggml/, common/) should stay structurally close to upstream so the next upstream rebase does not produce stacked duplicates - the gguf-py/gguf/constants.py duplicate-model-tensor crash is the canonical example. Fork-only logic in shared files needs a fork: tag in the comment so rebase conflict resolution can find it.test-backend-ops sweeps that cover TQ3_1S/TQ4_1S in all_types and the turbo3/4 FA cases (on CUDA), not just the default matrix. llama-bench with -ctk/-ctv turboN is the perf gate for cache-type work.Group findings by severity so the user knows what actually blocks a merge:
For each finding, point to the file and line and say concretely what to change and why. Do not rewrite the whole diff unprompted; let the contributor make the fixes so they own and understand them. And do not draft any PR text, commit message, or reviewer reply - that is the contributor's to write.
© JakeATX, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/code-review of JakeATX/llamAmpere.
Open the folder on GitHubat commit 93ac427
Code Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Code Review this skillJakeATX/llamAmpere | 148 | — | ~5.2k | Automated safety check: Pass | MIT | |
| Market Datazhongkaifu/TensorSharp | 557 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | |
| Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Model Serving MinefieldBlackwellboy/model-serving-minefield | 135 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Add Modelguoqingbao/xinfer | 333 | — | ~4.2k | Automated safety check: Notes | MIT | |
| Hugging Face Local Modelshuggingface/skills | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 |
zhongkaifu/TensorSharp
Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).
R6410418/Jackrong-llm-finetuning-guide
Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
guoqingbao/xinfer
Adapt and port new LLM model architectures to this xinfer project.
huggingface/skills
Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
JakeATX/llamAmpere
Guided workflow for adding a new model architecture to llama.cpp.
JakeATX/llamAmpere
Opinionated app components building on top of ./ui primitives
Categories
Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. Code Review is an agent skill from JakeATX/llamAmpere.cpp changes against project conventions and common reviewer pitfalls before a PR.
Code Review fits situations like: the user wants to review a diff; tasks that involve LLM inference and serving; tasks that involve Code review.
Run `npx skills add JakeATX/llamAmpere --skill code-review -a claude-code`. Or copy the skill folder (skills/code-review in JakeATX/llamAmpere) into .claude/skills/code-review in your project. Claude Code loads it when a task matches its description.
Run `npx skills add JakeATX/llamAmpere --skill code-review -a codex`. Or copy the skill folder (skills/code-review in JakeATX/llamAmpere) into .agents/skills/code-review in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JakeATX/llamAmpere --skill code-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-review, .gemini/skills/code-review, .github/skills/code-review and .opencode/skills/code-review in your project.
Going by SKILL.md and its folder, Code Review needs the command-line tools its instructions call (gh and git).
SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Code Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Code Review: Market Data (zhongkaifu/TensorSharp, 557 stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars) and Add Model (guoqingbao/xinfer, 333 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
JakeATX (a GitHub user) maintains it in JakeATX/llamAmpere, which has 148 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.
Source: JakeATX/llamAmpere on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.