Agent skill

Gemma4 Local Deploy

by majiayu000 in majiayu000/spellbook

在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4KM、64K/128K 长上下文、QAT Q40 @ 256K、左右对比演示之间选择,配置 tmux…

MITAuto-check: notesAI & LLM Engineering

Install Gemma4 Local Deploy

skills CLI
$ npx skills add majiayu000/spellbook --skill gemma4-local-deploy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/spellbook gemma4-local-deploy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gemma4-local-deploy .claude/skills/gemma4-local-deploy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemma4-local-deploy
GitHub stars
287
Token cost
~875 tokens
SKILL.md length
187 words
Files
5 (incl. references)
Skills in repo
97
Repo updated
First seen
Licence
MIT

At a glance

在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4KM、64K/128K 长上下文、QAT Q40 @ 256K、左右对比演示之间选择,配置 tmux…

  • Works in 3 steps: 搜索并确认现状 → 执行一条部署路线 → 验证并报告
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Operating Contract, 默认选择, Profile 选择 and 执行流程, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Gemma4 Local Deploy is an agent skill from majiayu000/spellbook. 在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4KM、64K/128K 长上下文、QAT Q40 @ 256K、左右对比演示之间选择,配置 tmux 后台运行,验证健康检查、问答接口、资源占用和常见故障。当用户说部署 Gemma 4、Gemma 4 12B、本地大模型、长上下文、QAT、量化、llama-server、Ollama、GGUF、Mac 本地模型服务时使用。

Its SKILL.md is about 880 tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `agents/openai.yaml`, `references/llama-cpp.md` and `references/ollama.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with llama.cpp, Ollama, tmux and OpenAI. The repository describes itself as: Cross-runtime skills for Claude Code, Codex, and multi-agent workflows. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/gemma4-local-deploy”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Read, WebSearch, WebFetch

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. 搜索并确认现状
  2. 执行一条部署路线
  3. 验证并报告

What it can do on your machine

Read from SKILL.md and the folder at commit ed52af7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • WebSearch
    • WebFetch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemma4 Local Deploy loads about 875 tokens when it runs, and up to ~3.2k if it reads all its reference files. Until then it costs about 82 tokens; SKILL.md has 187 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~875
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, WebSearch, WebFetch

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/spellbook at commit ed52af7, republished under its MIT licence (© majiayu000). 187 words, ~875 tokens.

Download SKILL.mdSave it as .claude/skills/gemma4-local-deploy/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
gemma4-local-deploy
description
在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4_K_M、64K/128K 长上下文、QAT Q4_0 @ 256K、左右对比演示之间选择,配置 tmux 后台运行,验证健康检查、问答接口、资源占用和常见故障。当用户说部署 Gemma 4、Gemma 4 12B、本地大模型、长上下文、QAT、量化、llama-server、Ollama、GGUF、Mac 本地模型服务时使用。
allowed-tools
Bash, Read, WebSearch, WebFetch
metadata.argument-hint
[模型量化/端口/是否后台运行]

Gemma 4 12B 本地部署

把 Gemma 4 12B 的 GGUF 版本部署成本机模型服务。默认使用 llama.cpp / llama-server、Apple Metal、Q4_K_M 和 tmux,只监听 loopback;用户明确要求 QAT、256K、对比演示或 Ollama 时才切换路线。

Operating Contract

  • Direct actions: 读取本机硬件、磁盘、端口、进程和模型缓存;在用户已要求本地部署时,安装或升级明确的软件包、下载选定模型、创建专用模型目录和 tmux 会话,并只绑定 127.0.0.1。
  • Escalate before: 停止不属于本 Skill 的现有进程、覆盖已有模型或配置、删除用户数据、监听公网地址、改变防火墙,或下载用户未选择的大型模型变体。
  • Evidence-backed pushback: 如果用户指定的模型标签、上下文、内存预算或本机能力与当前可验证状态冲突,先展示命令输出并提出可运行的 profile,不伪造支持状态。
  • Feedback loop: 现状检查 → 选择并复述 profile → 执行一条部署路线 → 当前会话完成健康、模型和聊天验证 → 报告端点、资源与限制。

默认选择

  • 默认模型仓库:ggml-org/gemma-4-12B-it-GGUF
  • 默认量化:Q4_K_M
  • 默认模型名:gemma-4-12b-it
  • 默认端点:http://127.0.0.1:8080
  • 默认上下文:32768
  • 12B 长上下文:用户明确要求时选择 65536 或 131072
  • QAT 仓库:google/gemma-4-12B-it-qat-q4_0-gguf
  • QAT profile:Q4_0、262144 上下文
  • 默认后台会话:gemma4-12b
  • 默认关闭 thinking:--reasoning off,避免 OpenAI API 的 message.content 为空
  • Ollama:只在用户明确要求 Ollama 或需要 Ollama 生态时使用

QAT 是训练时模拟量化,不等于无损。关键任务仍要用当前会话的真实响应验证。用户明确要更高质量时,优先建议 Q6_K 或 Q8_0;除非用户接受更高内存和更慢加载,不默认使用 bf16。

Profile 选择

Profile适用场景Model / quantContextPort / alias
daily-q4km-32k默认日常聊天、编码、低风险本地 APIggml-org/...:Q4_K_M327688080 / gemma-4-12b-it
long-q4km-128k明确需要更长上下文,但保留默认 GGUF 路线ggml-org/...:Q4_K_M65536 或 1310728080 / gemma-4-12b-it
qat-q4_0-256k明确要求 QAT、Q4_0、256K 或低内存长上下文google/...qat-q4_0-gguf:Q4_02621448080 / gemma-4-12b-it-qat-q4_0
compare-32k-vs-256k录屏、演示或 A/B 比较资源与速度左 Q4_K_M,右 QAT Q4_032768 + 2621448080 + 8081

最终回复必须说明选定 profile、端口、上下文和选择依据。不要把 256K 当作日常默认值。

执行流程

1. 搜索并确认现状

先检查已有安装、进程、端口、缓存、硬件和磁盘,避免重复部署:

bash
command -v llama-server || true
llama-server --version || true
tmux has-session -t gemma4-12b 2>/dev/null && tmux display-message -p -t gemma4-12b '#S #{pane_pid}' || true
lsof -nP -iTCP:8080 -sTCP:LISTEN || true
ls -lh "$HOME/Library/Caches/llama.cpp/"*gemma-4-12B-it*Q4_K_M*.gguf 2>/dev/null || true
find "$HOME/Library/Caches/llama.cpp" "$HOME/Models" \( -name '*gemma-4-12b-it-qat-q4_0*.gguf' -o -name '*gemma-4-12B-it-qat-q4_0*.gguf' \) 2>/dev/null || true
system_profiler SPHardwareDataType | sed -n '1,30p'
df -h "$HOME"

这些 || true 只用于允许“尚未安装/尚未运行”这一预期发现结果;必须展示实际输出,不能把查询失败描述成部署成功。

2. 执行一条部署路线
  • daily-q4km-32k、long-q4km-128k、qat-q4_0-256k 或 compare-32k-vs-256k:先读并执行 llama.cpp 部署路线。
  • 用户明确要求 Ollama:先读并执行 Ollama 部署路线。
  • 不要同时混用两条路线,也不要在没有端口检查的情况下启动第二个服务。
3. 验证并报告

部署后必须读取并执行 验证、资源与排障。成功至少需要当前会话证明:

  • /health 返回健康状态
  • /v1/models 或 Ollama 模型列表包含选定模型
  • 用户要求长上下文时,运行时报告实际 n_ctx
  • 一次聊天响应的正文非空
  • 端点仍只监听预期的本机地址和端口

Final response shape

默认用中文回答,并包含:

  • 实际 endpoint URL 和 model id
  • 选定 profile、量化与上下文
  • tmux/session 管理命令
  • 当前会话的验证结果
  • 实际资源摘要、失败项和限制

没有验证数据时写“未验证”,不能用计划值代替运行值。

Cross-check

部署计划涉及超出默认 profile 的模型、上下文或资源判断时,使用 agents/openai.yaml 做独立复核;复核不能替代当前会话的本机验证。

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/gemma4-local-deploy of majiayu000/spellbook.

  • SKILL.md
  • agents/openai.yaml
  • references/llama-cpp.md
  • references/ollama.md
  • references/verification.md

Open the folder on GitHubat commit ed52af7

Compare with similar skills

Gemma4 Local Deploy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemma4 Local Deploy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemma4 Local Deploy this skillmajiayu000/spellbook287—~875Automated safety check: NotesMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Local Modelsglebis/claude-skills391—~1.4kAutomated safety check: PassMIT
Vllmmagnus919/agent-skills115—~4.1kAutomated safety check: NotesMIT
Llama Cppmagnus919/agent-skills115—~2.3kAutomated safety check: PassMIT
Perfupraullenchai/Rapid-MLX4k—~1.6kAutomated safety check: NotesCustom licence

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Local Models

    glebis/claude-skills

    Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama.

    391 GitHub stars~1.4k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Vllm

    magnus919/agent-skills

    Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

    115 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Llama Cpp

    magnus919/agent-skills

    Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

    115 GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Perfup

    raullenchai/Rapid-MLX

    Autonomous performance optimization: research, PoC, benchmark, implement, review, PR

    4k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from majiayu000/spellbook

All 97 skills in this repo
  • Skill Ecosystem Doctor

    majiayu000/spellbook

    Audits and repairs how coding-agent Skills are owned, copied and exposed across runtimes, from canonical sources to quarantine and retirement.

    287 GitHub stars~3k tokensUpdated 2 days ago
    Auto-check passed
  • AGENTS.md Scaffold

    majiayu000/spellbook

    Scans a repository for real evidence and proposes, or on request writes, a small stack of root and scoped AGENTS.md files with validation commands and generated-file boundaries.

    287 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Product Demo Builder

    majiayu000/spellbook

    Plans, produces or diagnoses evidence-backed product demo videos: script, capture plan, pacing checks and verified final media built on real product behavior.

    287 GitHub stars~3.3k tokensUpdated 2 days ago
    Auto-check passed
  • Flowguard Task Guard

    majiayu000/spellbook

    Single entry point that routes long or ambiguous agent tasks, checks live state, bounds autonomous loops and leaves a resumable handoff.

    287 GitHub stars~2.1k tokensUpdated 2 days ago
    Auto-check passed
  • npm Supply Chain Check

    majiayu000/spellbook

    Scans a repository, its lockfiles and node_modules for known malicious npm package versions and install-time indicators, using a read-only Python scanner.

    287 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Product Manager Toolkit

    majiayu000/spellbook

    Product management helpers: a RICE scoring script, an interview transcript analyzer and PRD templates for prioritizing features, synthesizing research and writing requirements.

    287 GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed

Questions about Gemma4 Local Deploy

What does Gemma4 Local Deploy do?

在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4KM、64K/128K 长上下文、QAT Q40 @ 256K、左右对比演示之间选择,配置 tmux…. Gemma4 Local Deploy is an agent skill from majiayu000/spellbook.

When should I use Gemma4 Local Deploy?

Gemma4 Local Deploy fits situations like: tasks that involve LLM inference and serving.

How do I install Gemma4 Local Deploy in Claude Code?

Run `npx skills add majiayu000/spellbook --skill gemma4-local-deploy -a claude-code`. Or copy the skill folder (skills/gemma4-local-deploy in majiayu000/spellbook) into .claude/skills/gemma4-local-deploy in your project. Claude Code loads it when a task matches its description.

How do I install Gemma4 Local Deploy in Codex?

Run `npx skills add majiayu000/spellbook --skill gemma4-local-deploy -a codex`. Or copy the skill folder (skills/gemma4-local-deploy in majiayu000/spellbook) into .agents/skills/gemma4-local-deploy in your project. Codex loads it when a task matches its description.

Can I use Gemma4 Local Deploy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/spellbook --skill gemma4-local-deploy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemma4-local-deploy, .gemini/skills/gemma4-local-deploy, .github/skills/gemma4-local-deploy and .opencode/skills/gemma4-local-deploy in your project.

What does Gemma4 Local Deploy need to run?

SKILL.md names no scripts, command-line tools or credentials: Gemma4 Local Deploy is instructions for the agent only. Its frontmatter pre-approves these tools: Bash, Read, WebSearch, WebFetch.

Does Gemma4 Local Deploy access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gemma4 Local Deploy safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Gemma4 Local Deploy use?

Gemma4 Local Deploy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemma4 Local Deploy use?

About 875 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Gemma4 Local Deploy?

Skills that share tags, products or a category with Gemma4 Local Deploy: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Local Models (glebis/claude-skills, 391 stars), Vllm (magnus919/agent-skills, 115 stars) and Llama Cpp (magnus919/agent-skills, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemma4 Local Deploy?

majiayu000 (a GitHub user) maintains it in majiayu000/spellbook, which has 287 GitHub stars. The repository holds 97 skills in this directory. The repository was last updated on October 8, 2026.

Source: majiayu000/spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.