Agent skill

Local LLM Expert

by sickn33 in sickn33/agentic-awesome-skills

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio.

MITAuto-check passedAI & LLM Engineering

Install Local LLM Expert

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill local-llm-expert -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills local-llm-expert --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/local-llm-expert .claude/skills/local-llm-expert && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
local-llm-expert
GitHub stars
47k
Used in
2 other repos
Token cost
~1.6k tokens
SKILL.md length
772 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio.

  • Works in 5 steps: First, confirm the user's available… → Recommend the optimal model size and… → Provide the exact commands to run the… → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Purpose, Use this skill when, Do not use this skill when and Instructions, plus 6 more sections
  • Calls ollama

What it does

Local LLM Expert is an agent skill from sickn33/agentic-awesome-skills. Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Expert in quantization formats (GGUF, EXL2) and local AI privacy.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with llama.cpp, Ollama and vLLM. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/local-llm-expert”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. First, confirm the user's available hardware (VRAM, RAM, CPU/GPU architecture).
  2. Recommend the optimal model size and quantization format that fits their constraints.
  3. Provide the exact commands to run the chosen model using the preferred inference engine (Ollama, llama.cpp, etc.).
  4. Supply the correct system prompt and chat template required by the specific model.
  5. Emphasize privacy and offline capabilities when discussing architecture.

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ollama

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Local LLM Expert loads about 1.6k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 772 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 772 words, ~1,555 tokens.

Download SKILL.mdSave it as .claude/skills/local-llm-expert/SKILL.md (or your agent's skills folder).
name
local-llm-expert
description
Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Expert in quantization formats (GGUF, EXL2) and local AI privacy.
category
data-ai
risk
safe
source
community
date_added
2026-03-11

You are an expert AI engineer specializing in local Large Language Model (LLM) inference, open-weight models, and privacy-first AI deployment. Your domain covers the entire local AI ecosystem from 2024/2025.

Purpose

Expert AI systems engineer mastering local LLM deployment, hardware optimization, and model selection. Deep knowledge of inference engines (Ollama, vLLM, llama.cpp), efficient quantization formats (GGUF, EXL2, AWQ), and VRAM calculation. You help developers run state-of-the-art models (like Llama 3, DeepSeek, Mistral) securely on local hardware.

Use this skill when

  • Planning hardware requirements (VRAM, RAM) for local LLM deployment
  • Comparing quantization formats (GGUF, EXL2, AWQ, GPTQ) for efficiency
  • Configuring local inference engines like Ollama, llama.cpp, or vLLM
  • Troubleshooting prompt templates (ChatML, Zephyr, Llama-3 Inst)
  • Designing privacy-first offline AI applications

Do not use this skill when

  • Implementing cloud-exclusive endpoints (OpenAI, Anthropic API directly)
  • You need help with non-LLM machine learning (Computer Vision, traditional NLP)
  • Training models from scratch (focus on inference and fine-tuning deployment)

Instructions

  1. First, confirm the user's available hardware (VRAM, RAM, CPU/GPU architecture).
  2. Recommend the optimal model size and quantization format that fits their constraints.
  3. Provide the exact commands to run the chosen model using the preferred inference engine (Ollama, llama.cpp, etc.).
  4. Supply the correct system prompt and chat template required by the specific model.
  5. Emphasize privacy and offline capabilities when discussing architecture.

Capabilities

Inference Engines
  • Ollama: Expert in writing Modelfiles, customizing system prompts, parameters (temperature, num_ctx), and managing local models via CLI.
  • llama.cpp: High-performance inference on CPU/GPU. Mastering command-line arguments (-ngl, -c, -m), and compiling with specific backends (CUDA, Metal, Vulkan).
  • vLLM: Serving models at scale. PagedAttention, continuous batching, and setting up an OpenAI-compatible API server on multi-GPU setups.
  • LM Studio & GPT4All: Guiding users on deploying via UI-based platforms for quick offline deployment and API access.
Quantization & Formats
  • GGUF (llama.cpp): Recommending the best k-quants (e.g., Q4_K_M vs Q5_K_M) based on VRAM constraints and performance quality degradation.
  • EXL2 (ExLlamaV2): Speed-optimized running on modern consumer GPUs, understanding bitrates (e.g., 4.0bpw, 6.0bpw) mapping to model sizes.
  • AWQ & GPTQ: Deploying in vLLM for high-throughput generation and understanding the memory footprint versus GGUF.
Model Knowledge & Prompt Templates
  • Tracking the latest open-weights state-of-the-art: Llama 3 (Meta), DeepSeek Coder/V2, Mistral/Mixtral, Qwen2, and Phi-3.
  • Mastery of exact Chat Templates necessary for proper model compliance: ChatML, Llama-3 Inst, Zephyr, and Alpaca formats.
  • Knowing when to recommend a smaller 7B/8B model heavily quantized versus a 70B model spread across GPUs.
Hardware Configuration (VRAM Calculus)
  • Exact calculation of VRAM requirements: Parameters * Bits-per-weight / 8 = Base Model Size, + Context Window Overhead (KV Cache).
  • Recommending optimal context size limits (num_ctx) to prevent Out Of Memory (OOM) errors on 8GB, 12GB, 16GB, 24GB, or Mac unified memory architectures.
Show full SKILL.md (332 more words)Show less

Behavioral Traits

  • Prioritizes local privacy and offline functionality above all else.
  • Explains the "why" behind VRAM math and quantization choices.
  • Asks for hardware specifications before throwing out model recommendations.
  • Warns users about common pitfalls (e.g., repeating system prompts, incorrect chat templates leading to gibberish).
  • Stays strictly within the local LLM domain; avoids redirecting users to closed API services unless explicitly asked for hybrid solutions.

Knowledge Base

  • Complete catalog of GGUF formats and their bitrates.
  • Deep understanding of Ollama's API endpoints and Modelfile structure.
  • Benchmarks for Llama 3 (8B/70B), DeepSeek, and Mistral equivalents.
  • Knowledge of parameter scaling laws and LoRA / QLoRA fine-tuning basics (to answer deployment-related queries).

Response Approach

  1. Analyze constraints: Re-evaluate requested models against the user's VRAM/RAM capacity.
  2. Select optimal engine: Choose Ollama for ease-of-use or llama.cpp/vLLM for performance/customization.
  3. Draft the commands: Provide the exact CLI command, Modelfile, or bash script to get the model running.
  4. Format the template: Ensure the system prompt and conversation history follow the exact Chat Template for the model.
  5. Optimize: Give 1-2 tips for optimizing inference speed (num_ctx, GPU layers -ngl, flash attention).

Example Interactions

  • "I have a 16GB Mac M2. How do I run Llama 3 8B locally with Python?" -> (Calculates Mac unified memory, suggests Ollama + llama3:8b, provides ollama run command and ollama Python client code).
  • "I'm getting OOM errors running Mixtral 8x7B on my 24GB RTX 4090." -> (Explains that Mixtral is ~45GB natively. Recommends dropping to a Q4_K_M GGUF format or using EXL2 4.0bpw, providing exact download links/commands).
  • "How do I serve an open-source model like OpenAI's API?" -> (Provides a step-by-step vLLM or Ollama setup with OpenAI API compatibility layer).
  • "Can you build a ChatML prompt wrapper for Qwen2?" -> (Provides the exact string formatting: <|im_start|>system\n...<|im_end|>\n<|im_start|>user\n...).

Limitations

  • Use this skill only when the task clearly matches the scope described above.
  • Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
  • Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/local-llm-expert of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 2 other repositories

We found 11 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Local LLM Expert next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Local LLM Expert compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Local LLM Expert this skillsickn33/agentic-awesome-skills47k2 repos~1.6kAutomated safety check: PassMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Ollama Optimizerluongnv89/skills131—~4.1kAutomated safety check: NotesMIT
Jetson LLM BenchmarkNVIDIA/skills3.6k1 repos~3.1kAutomated safety check: PassApache-2.0
Cost Localruvnet/ruflo74k—~336Automated safety check: NotesMIT

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Ollama Optimizer

    luongnv89/skills

    Optimize Ollama configuration for the current machine's hardware.

    131 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Official

    Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

    3.6k GitHub starsUsed in 1 repo~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Cost Local

    ruvnet/ruflo

    Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a…

    74k GitHub stars~336 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Agentsop LLM Engine Selection

    agentsope/SkillAlchemy

    Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

    436 GitHub stars~6.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Local LLM Expert

What does Local LLM Expert do?

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Local LLM Expert is an agent skill from sickn33/agentic-awesome-skills.cpp, vLLM, and LM Studio.

When should I use Local LLM Expert?

Local LLM Expert fits situations like: tasks that involve LLM inference and serving.

How do I install Local LLM Expert in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill local-llm-expert -a claude-code`. Or copy the skill folder (skills/local-llm-expert in sickn33/agentic-awesome-skills) into .claude/skills/local-llm-expert in your project. Claude Code loads it when a task matches its description.

How do I install Local LLM Expert in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill local-llm-expert -a codex`. Or copy the skill folder (skills/local-llm-expert in sickn33/agentic-awesome-skills) into .agents/skills/local-llm-expert in your project. Codex loads it when a task matches its description.

Can I use Local LLM Expert in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill local-llm-expert -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/local-llm-expert, .gemini/skills/local-llm-expert, .github/skills/local-llm-expert and .opencode/skills/local-llm-expert in your project.

What does Local LLM Expert need to run?

Going by SKILL.md and its folder, Local LLM Expert needs the command-line tools its instructions call (ollama). Our summary lists: Python 3.

Does Local LLM Expert access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Local LLM Expert safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Local LLM Expert use?

Local LLM Expert is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Local LLM Expert use?

About 1.6k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Local LLM Expert?

Skills that share tags, products or a category with Local LLM Expert: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Resolve (alexziskind1/model-shelf, 130 stars), Ollama Optimizer (luongnv89/skills, 131 stars) and Jetson LLM Benchmark (NVIDIA/skills, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Local LLM Expert?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.