Agent skill

Code Review

by JakeATX in JakeATX/llamAmpere

Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.

MITAuto-check passedAI & LLM Engineering

Install Code Review

skills CLI
$ npx skills add JakeATX/llamAmpere --skill code-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JakeATX/llamAmpere code-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JakeATX/llamAmpere.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/code-review .claude/skills/code-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-review
GitHub stars
148
Token cost
~5.2k tokens
SKILL.md length
2,933 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.

  • Works in 3 steps: Blocking - quick-reject/scope issues and… → Will slow the review -… → Nits - minor style, optional cleanups.
  • The user wants to review a diff
  • SKILL.md covers Step 0 - Scope the diff and…, Scope and quick-reject gate…, Security review (mandatory) and Approach and design (when a…, plus 8 more sections
  • Calls gh and git

What it does

Code Review is an agent skill from JakeATX/llamAmpere. Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. Use when the user wants to review a diff, branch, or PR.

Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving and Code review. It works with llama.cpp, Qwen and CUDA. The repository describes itself as: llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): example: 95+ tok/s over a 100K-token generation at temperature 1 for qwen3.8… The licence is MIT.

When your agent uses it

  • The user wants to review a diff
  • Tasks that involve LLM inference and serving
  • Tasks that involve Code review

Example prompts

  • “/code-review”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Blocking - quick-reject/scope issues and correctness bugs; these can sink the PR regardless of everything else.
  2. Will slow the review - convention/naming/comment violations, missing tests/docs/perf data, missing API justification or example.
  3. Nits - minor style, optional cleanups.

What it can do on your machine

Read from SKILL.md and the folder at commit 93ac427. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Review loads about 5.2k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 2,933 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from JakeATX/llamAmpere at commit 93ac427, republished under its MIT licence (© JakeATX). 2,933 words, ~5,248 tokens.

Download SKILL.mdSave it as .claude/skills/code-review/SKILL.md (or your agent's skills folder).
name
code-review
description
Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. Use when the user wants to review a diff, branch, or PR.

Review llama.cpp changes

This skill reviews changes against llama.cpp's conventions and the pitfalls that reviewers flag most often, so the contributor can fix them before a maintainer has to. It has two modes:

  • Self-review (default): review the contributor's own local changes (uncommitted work, or a branch vs its base) as a pre-PR pass. Ask which if it's ambiguous; default to git diff feature/turboquant-kv-cache...HEAD plus any uncommitted changes (use the upstream merge-base of the fork as the base when the diff against the feature branch is noisy).
  • Read-only review of a PR/file: if the user points at a PR number or specific files (including code they didn't write), review those and report findings.

Fork context: this repo is the TurboQuant fork, not upstream llama.cpp. Most changes here are fork-internal and never go to ggml-org; the upstream-facing rules below (issue-first, quick-reject gates, maintainer approval expectations) still apply to the shared upstream code, but the fork-specific checklist at the end takes precedence for TurboQuant paths. The AGENTS.md overview is required context - read it first if not already in context; its "Known pitfalls" list is the first place to check on any turbo regression.

In both modes the output is private review notes for the user to read and act on - it is never something to post. This is a hard rule from AGENTS.md: an agent must NEVER write, or help write, a PR comment, a review comment, or a reply to a reviewer, by any means including gh. Do not offer to. If the user asks you to post the notes, refuse and point them at that rule. Present findings in the conversation only.

Before starting, read AGENTS.md and CONTRIBUTING.md if not already in context - the "Coding guidelines", "Naming guidelines", and AI usage sections are the baseline this review enforces. For a diff that adds a new model architecture, also read docs/development/HOWTO-add-model.md and consider the dedicated add-new-model skill.

Step 0 - Scope the diff and pick the checklists

Identify what actually changed and which area checklists below apply. Run git diff --stat (or gh pr view <n> --json files for PR mode) and bucket the touched paths:

  • conversion/, gguf-py/, src/models/, src/llama-arch.* -> New model / architecture
  • ggml/ (any backend, op, or ggml.h) -> ggml / backend
  • ggml-turbo-quant.c, turbo/TQ weight or cache types, GGML_OP_TURBO_WHT, turbo kernels in any backend -> TurboQuant / fork-specific (in addition to ggml / backend)
  • include/llama.h and other public headers -> Public API
  • tools/server/ -> Server
  • anything else, plus all of the above -> General (always runs)

Always run the Scope and quick-reject gate, the Security review, and the General checklist. Run each area checklist whose paths were touched. Additionally, if the diff introduces a new component, subsystem, or piece of infrastructure (a new file/class/module, a new abstraction, or hand-rolled machinery), run the Approach and design review. Tell the user which checklists you're running and why.

Scope and quick-reject gate (always)

These are the patterns that get PRs closed without a full review. Check them first - a finding here is more important than any code nit, because it can mean the change shouldn't be a PR in its current form at all.

  • Is there a prior issue/discussion for this? Features are supposed to start as an issue, not a PR (CONTRIBUTING.md). If this is a nontrivial feature with no linked issue, flag it and suggest opening one first.
  • Is it a duplicate of existing/in-flight work? Suggest gh search prs / gh search issues for the feature. Many closed PRs were duplicates of something already queued.
  • Is it self-contained and single-purpose? Multiple unrelated changes/optimizations bundled together get sent back to be split. Flag unrelated changes and suggest separate PRs.
  • Does it touch multiple ggml backends at once? Initial support should be CPU-only, other backends as follow-ups (CONTRIBUTING.md). Flag CUDA/Metal/Vulkan/etc. changes bundled into a feature's first PR.
  • Does it add a new ggml_type / quantization type? That carries a disproportionate maintenance burden and needs the full justification package (GGUF sample upload, perplexity vs FP16/BF16 and similar sizes, KL-divergence data, CPU perf numbers). Absent that, it will be rejected regardless of code quality.
  • Is it invasive - new subsystem, core-API reshaping, changes to shared graph/sampler code that other models don't need? Flag it and suggest a discussion with maintainers before investing further.
  • Is it niche/vendor-specific in a way that adds a maintenance burden nobody will own long-term? Flag the maintenance-ownership question.
  • Is the change semantically correct, or a plausible-looking "fix" that misunderstands the code? Sanity-check the actual behavior, not just that it compiles.
  • AI-disclosure: if AI meaningfully contributed, is the PR template's disclosure section filled in? Remind the user. Never suggest writing the PR description or commit message for them.

Security review (mandatory)

Mandatory on every review; any finding here is blocking. Rule of thumb: GGUF metadata, tensor shapes, tokenizer/grammar input, and all server/RPC fields are attacker-controlled - bound them before use.

  • Sizes/counts from tensor dims: validate before allocating. Products like ne[i]*nb[i]/nbytes can overflow on crafted dims into an undersized alloc then heap overflow. Overflow checks must run BEFORE the arithmetic they guard - padding/alignment macros wrap to 0 near SIZE_MAX, so a guard after the pad passes.
  • GGUF strings/arrays: cap declared lengths and element counts before using them to size a loop or buffer; validate element type and length before casting an array to a pointer or reading fixed indices ([i+1], [0..2]).
  • Element-type confusion: casting gguf_get_arr_data() or tensor->data to float */int32_t * needs an element-type check first (gguf_get_kv_type() == GGUF_TYPE_ARRAY then gguf_get_arr_type(); type == GGML_TYPE_F32 for tensors). A UINT8 array or I8 tensor passes every length check, then gets read 4 bytes per element - a nearby length check is not a type check.
  • Loaders: GGML_ASSERT on a file-derived value aborts the process; throw instead where the caller already catches (vocab, model loader, clip).
  • File-supplied counts indexing fixed arrays: bound any count (e.g. layer/block count into a LLAMA_MAX_* array) before indexing; watch checks that only fire when an optional key is present.
  • Declared vs actual array length: check the declared length of a GGUF array against the count actually read, not just against a buffer size.
  • Bounds comparisons: flag narrowing casts (size_t->int32_t) and signed/unsigned mixing that can bypass a length check and copy past a buffer.
  • Parsed/derived indices: range-check stoi/atoi results and catch parse throws; never use a default or derived token id (EOS/BOS/...) as an index without a bounds check.
  • Reused/reserved buffers: recheck bounds after a buffer is shrunk or reused; watch reserve() then index-by-assumed-size, and header fields read before their length is checked.
  • Server JSON ints: clamp client-supplied integers (token/discard counts, offsets) to non-negative and an upper bound before they reach index/pointer arithmetic.
  • RPC-deserialized fields: treat every field (type/buffer/data/ne/nb/op_params) as hostile - validate before use. Null/zero buffers skipping validation, attacker data pointers, out-of-range type indices, and negative strides sign-extending past a corner-only assert all give arbitrary read/write.
  • Lifetime/UAF: flag stored raw pointers to caller/temporary storage, cached pointers to buffers a later free releases, async ops whose source may drop before completion, and structures not invalidated on free/realloc. Null-check conditionally-built or "not required" tensors before dereferencing.

Approach and design (when a new component/infra is introduced)

Run this whenever the diff adds a new component, subsystem, or piece of infrastructure. Reviews too often stop at "does it work" - a diff can be correct and still be the wrong approach, and a messy design costs more long-term than a bug. Evaluate the approach, not just the behavior; raising a cleaner one is a high-value finding, not a nit. If you see a better design, describe it concretely rather than just calling the current one bad.

  • Simpler approach upstream: the biggest win is often a different data model or design that removes whole subsystems, not tweaks to the code as written. Complexity must be justified by the problem, not by the first thing that worked.
  • Reuse over reinvention: grep for an existing helper, library, object, or mechanism before adding a new one. Reimplementing what the codebase already has reintroduces solved bugs and adds maintenance surface.
  • Clear ownership/lifetime: prefer RAII and obvious ownership over manual liveness flags, hand-tracked pointers, and "is it still alive?" checks - manual lifetime tracking is a recurring source of subtle bugs.
  • Right-sized machinery: flag redundant, overkill, or heavier-than-needed primitives and abstractions; use the minimum the design actually needs.
  • Right structure and fit: a new type should earn its place (split it if it serves two roles); follow existing patterns, idioms, and naming, and avoid constructs the project shuns.
  • Root cause vs symptom: fixes layered on fixes signal a design to correct, not guard around.

New model / architecture

See the add-new-model skill and docs/development/HOWTO-add-model.md for the full workflow; this is the review-time subset that reviewers most often catch:

  • Don't branch on model.arch when the real dependency is a config/capability value - gate on the hparam/capability, not the architecture enum.
  • If the model is a close variant of an existing arch, is the delta justified? Prefer reusing or subclassing the existing arch/model class over duplicating it. A near-duplicate class or src/models/<name>.cpp will be asked to merge with its sibling.
  • New tensor names go through tensor_mapping.py, not ad-hoc name matching.
  • For QKV, split the activation with ggml_view, not the weight tensor; rely on ggml broadcasting instead of manually duplicating tensors.
  • New graph inputs are declared at the top of the graph-build function, not inline where first used.
  • Hparams that the model can't run correctly without must be mandatory (hard-error if missing), not read with a silent default fallback. Only genuinely-optional-across-configs values get a fallback accessor.
  • New/optional weight tensors (scales, etc.) must route through build_lora_mm and the existing helpers, matching convention - don't leave raw matmuls copied from another arch.
  • Don't hack RoPE with a custom sin/cos implementation. If ggml_rope_ext genuinely can't express it, that's an issue for discussion, not a PR.
  • Test the quantized-KV path (-ctk/-ctv q8_0), not just default f16 - new speculative/attention features silently break there.
  • Preserve existing explanatory comments about model-specific quirks when copying code; note the provenance ("copied from X, with Y added").
  • Remove dead code/branches left over from adapting a reference implementation.

ggml / backend

  • supports_op (and any dispatch/gating condition) must be scoped exactly to the cases being changed - a condition meant for a few quant types must not silently disable or enable everything else.
  • No hardcoded warp/lane size - use ggml_cuda_get_physical_warp_size() (32 on CUDA, 64 on HIP/ROCm) and the portable helpers.
  • Strip leftover debug/profiling/logging code before review.
  • New or changed op? Update docs/ops.md and the relevant docs/ops/*.csv for the touched backend.
  • New op or operator change needs corresponding test-backend-ops cases, and (per CONTRIBUTING.md) consistency across at least two backends.
  • New kernels are expected to come with concrete perf data (throughput across realistic tensor shapes), not just correctness.
  • Don't have a backend mutate the cgraph as a shortcut - that's an unresolved architectural question, not something to slip in.
  • Expect this to need two maintainer approvals; that's normal for ggml/ changes, not a sign something is wrong.
  • For CUDA: Avoid excessively templating kernels, only add this where it shows visible performance gain.
Show full SKILL.md (1,127 more words)Show less

Public API (include/llama.h)

Public API changes carry a higher bar than internal ones (CONTRIBUTING.md). Review for:

  • Justification: why doesn't an existing mechanism (e.g. cb_eval, existing batch/sampler knobs) suffice? If it does, the change likely shouldn't add public surface. This is the single most common reason these PRs are rejected.
  • Experimental or stop-gap surface belongs in a side header (llama-ext.h), not in llama.h.
  • Keep it minimal and general: prefer one general call over several narrow convenience wrappers; make new calls forward-compatible (e.g. mixed-modality batches) rather than assuming today's shape.
  • The C API is the first-class, stable, ABI-defining surface - don't propose a parallel C++ API as a replacement. llama-cpp.h stays a thin convenience layer.
  • Types and naming: sized integer types (int32_t, size_t for sizes/offsets); snake_case; <class>_<method> = <class>_<action>_<noun>; enum values upper-case and prefixed with the enum name; _t suffix for opaque types. Avoid gratuitous signature/ABI changes to existing exported functions.
  • Every new API needs a working example/tool exercising it in the same PR - reviewers find real bugs by requiring it to be wired into server, embedding, perplexity, etc.

Server (tools/server/)

  • Is the feature within server's defined scope? Check tools/server/README-dev.md - out-of-scope features get declined.
  • Security: don't trust client-supplied headers (e.g. X-Forwarded-For) or add footguns; things like IP allowlisting belong at a reverse proxy unless there's a trusted-proxy design.
  • Wire new behavior into the existing request/response and checkpoint paths correctly; watch for resource leaks across requests.

Multimodal (tools/mtmd/)

  • Tensor names must be prefixed by v., a., mm. or a.mm. (legacy naming doesn't follow this convention - this is expected, but new code should follow it).
  • Do not use explicit sin/cos for RoPE; use ggml_rope_ext instead, see HOWTO-add-model.md. If it can't express the needed behavior, that's a design discussion, not a PR.
  • New GGML ops must not be introduced in the same PR, you must push it as a separate PR.
  • In most cases, build_vit should be enough to build the transformer graph for vision models. Do not add a loop to build the transformer graph manually, unless you have a very good reason to do so. If you do, please explain why in the PR description.
  • If you need a dedicated preprocessor, there is a high chance that it can be a derived class from one of the existing preprocessors. Check carefully before adding a new preprocessor class.
  • If the model need a new public API in mtmd.h, open a discussion first.
  • For audio generation models, see tools/mtmd/README-dev.md

General (always)

Enforce the AGENTS.md / CONTRIBUTING.md coding and naming guidelines on every changed line - this is a distinct pass from checking that the code works, and matters just as much for review speed:

  • ASCII only in code and comments - no emdash, unicode arrows, x, ... used as unicode; use -, ->, x, ... ASCII equivalents.
  • Comments are concise and explain non-obvious why, not what. Flag verbose comments, comments that restate the code, comments that reference the current task/PR, and comments hard-wrapped to a fixed column width.
  • Do not force-wrap prose/comments to a fixed character count or split a sentence across lines.
  • snake_case names; kebab-case (lowercase-with-dashes) file names for C/C++, .h headers; Python files lowercase-with-underscores. Naming optimizes for longest common prefix (number_small, not small_number).
  • 4-space indentation, brackets on the same line, void * ptr, int & a, no trailing whitespace; match the surrounding style.
  • Reuse existing infrastructure over introducing new components; no new third-party dependencies, extra headers, or files unless clearly justified.
  • Keep it simple: a simpler change doing 90% is often preferable to a complex one doing 100%. Flag unnecessary templates/fancy STL; basic for loops are fine here.
  • Every added line should be something the contributor can explain and defend to a reviewer without AI help - flag anything that looks copied-in without understanding.
  • Co-authored-by: must be reserved for human co-authors; AI contributions (claude, cursor, codex, etc.) must use Assisted-by:; if this point is violated, it's a blocking finding.
  • Any mentions of Minja must be treated as blocking; see AGENTS.md for why.

TurboQuant / fork-specific (takes precedence on fork paths)

Run this in addition to the General checklist whenever the diff touches turbo code, and instead of the upstream-facing gates where they conflict. The AGENTS.md "Known pitfalls" list is the regression checklist - each entry there cost a real bug; verify the diff doesn't disturb those invariants.

  • Type/enum stability: GGML_TYPE_TURBO2_0=43, TURBO3_0=44, TQ3_1S=45, TQ4_1S=46, TURBO4_0=47 and GGML_OP_TURBO_WHT are baked into GGUF files and cross-backend dispatch - never renumber, reorder, or repurpose. New fork types go after 47.
  • Codec fidelity: ggml/src/ggml-turbo-quant.c must stay byte-identical to fork tip conventions; any change needs test-turbo-quant passing (turbo3 MSE=0/Cosine=1.0, turbo4 Cosine=0.9956) plus test-quantize-fns (TQ3_1S/TQ4_1S cases).
  • Backend coverage matrix: a change to turbo KV or TQ weight behavior must keep every backend in sync - CPU, CUDA/HIP, Metal, Vulkan, SYCL. The classic rebase regressions, in order: Metal [[host_name]] kernel instantiations dropped (NULL-pipeline deref), Vulkan SET_ROWS pipeline registration missing TURBO types or the require_full_subgroups=true, subgroup_size=32 flags (abort), CUDA TQ exclusion from the mmvq path dropped (abort with GGML_TQ_NATIVE=1).
  • CUDA TQ kernels: no __byte_perm in centroid LUTs (mmvq-tq.cu) - plain shifts only; the permute chain silently produces garbage on some toolchains. MoE TQ MUL_MAT_ID must keep disabling CUDA graphs (stream sync requirement).
  • KV cache types: changes to cache-type handling must be exercised with turbo types (-ctk/-ctv turbo3), which require flash attention (auto-enabled). Respect the 128-element rotation block: zero-padding of head dims, and no V rotation/padding for MLA models. K/V types must stay identical for MLA/DeepSeek4.
  • Env knobs: TURBO_LAYER_ADAPTIVE, TURBO_AUTO_ASYMMETRIC, TURBO_SPARSE_V, GGML_TQ_NATIVE, LLAMA_ATTN_ROT_* semantics in docs/KV-cache-quantization.md are the contract - changing behavior without updating that doc is a finding.
  • Rebase hygiene: shared upstream files (src/, ggml/, common/) should stay structurally close to upstream so the next upstream rebase does not produce stacked duplicates - the gguf-py/gguf/constants.py duplicate-model-tensor crash is the canonical example. Fork-only logic in shared files needs a fork: tag in the comment so rebase conflict resolution can find it.
  • Test gates: run test-backend-ops sweeps that cover TQ3_1S/TQ4_1S in all_types and the turbo3/4 FA cases (on CUDA), not just the default matrix. llama-bench with -ctk/-ctv turboN is the perf gate for cache-type work.
  • Upstream-facing sections below still apply to changes in shared code, but do not block fork-internal work on upstream acceptance criteria (issue-first, two-maintainer approvals, upstream docs) that do not apply here.

Reporting

Group findings by severity so the user knows what actually blocks a merge:

  1. Blocking - quick-reject/scope issues and correctness bugs; these can sink the PR regardless of everything else.
  2. Will slow the review - convention/naming/comment violations, missing tests/docs/perf data, missing API justification or example.
  3. Nits - minor style, optional cleanups.

For each finding, point to the file and line and say concretely what to change and why. Do not rewrite the whole diff unprompted; let the contributor make the fixes so they own and understand them. And do not draft any PR text, commit message, or reviewer reply - that is the contributor's to write.

© JakeATX, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/code-review of JakeATX/llamAmpere.

Open the folder on GitHubat commit 93ac427

Compare with similar skills

Code Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Review this skillJakeATX/llamAmpere148—~5.2kAutomated safety check: PassMIT
Market Datazhongkaifu/TensorSharp557—~1kAutomated safety check: PassBSD-3-Clause
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0

Similar skills

  • Market Data

    zhongkaifu/TensorSharp

    Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).

    557 GitHub stars~1k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check: notes
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from JakeATX/llamAmpere

  • Add New Model

    JakeATX/llamAmpere

    Guided workflow for adding a new model architecture to llama.cpp.

    148 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • App

    JakeATX/llamAmpere

    Opinionated app components building on top of ./ui primitives

    148 GitHub stars~146 tokensUpdated yesterday
    Auto-check passed

Questions about Code Review

What does Code Review do?

Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. Code Review is an agent skill from JakeATX/llamAmpere.cpp changes against project conventions and common reviewer pitfalls before a PR.

When should I use Code Review?

Code Review fits situations like: the user wants to review a diff; tasks that involve LLM inference and serving; tasks that involve Code review.

How do I install Code Review in Claude Code?

Run `npx skills add JakeATX/llamAmpere --skill code-review -a claude-code`. Or copy the skill folder (skills/code-review in JakeATX/llamAmpere) into .claude/skills/code-review in your project. Claude Code loads it when a task matches its description.

How do I install Code Review in Codex?

Run `npx skills add JakeATX/llamAmpere --skill code-review -a codex`. Or copy the skill folder (skills/code-review in JakeATX/llamAmpere) into .agents/skills/code-review in your project. Codex loads it when a task matches its description.

Can I use Code Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JakeATX/llamAmpere --skill code-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-review, .gemini/skills/code-review, .github/skills/code-review and .opencode/skills/code-review in your project.

What does Code Review need to run?

Going by SKILL.md and its folder, Code Review needs the command-line tools its instructions call (gh and git).

Does Code Review access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Code Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Code Review use?

Code Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Review use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Code Review?

Skills that share tags, products or a category with Code Review: Market Data (zhongkaifu/TensorSharp, 557 stars), Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars) and Add Model (guoqingbao/xinfer, 333 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Review?

JakeATX (a GitHub user) maintains it in JakeATX/llamAmpere, which has 148 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.

Source: JakeATX/llamAmpere on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.