Agent skill

Ollama Optimizer

by luongnv89 in luongnv89/skills

Optimize Ollama configuration for the current machine's hardware.

MITAuto-check: notesAI & LLM Engineering

Install Ollama Optimizer

skills CLI
$ npx skills add luongnv89/skills --skill ollama-optimizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install luongnv89/skills ollama-optimizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/luongnv89/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ollama-optimizer .claude/skills/ollama-optimizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ollama-optimizer
GitHub stars
131
Token cost
~4.1k tokens
SKILL.md length
1,955 words
Files
9 (incl. scripts, references)
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Optimize Ollama configuration for the current machine's hardware.

  • Works in 4 steps: System Detection → Analyze and Recommend → Generate Optimization Plan → …
  • Asked to speed up Ollama
  • SKILL.md covers When to Use, Safety Rules, Workflow and Edge Cases, plus 6 more sections
  • Runs Python scripts from its folder; calls ollama, python3 and docker

What it does

Ollama Optimizer is an agent skill from luongnv89/skills. Optimize Ollama configuration for the current machine's hardware. Use when asked to speed up Ollama, tune local LLM performance, or pick models that fit available GPU/RAM. Don't use for LM Studio, llama.cpp, vLLM, or hosted-API LLM providers.

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `docs/README.md`, `evals/evals.json` and `references/environment_variables.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Ollama, llama.cpp and vLLM. The repository describes itself as: Supercharge your AI agents/bots with reusable skills. The licence is MIT.

When your agent uses it

  • Asked to speed up Ollama
  • Tune local LLM performance
  • Pick models that fit available GPU/RAM
  • Hosted-API LLM providers

Example prompts

  • “/ollama-optimizer”

Requirements

  • Python 3
  • Docker

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. System Detection
  2. Analyze and Recommend
  3. Generate Optimization Plan
  4. Verification

What it can do on your machine

Read from SKILL.md and the folder at commit 891c720. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • ollama
    • python3
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ollama Optimizer loads about 4.1k tokens when it runs, and up to ~8.6k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 1,955 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:127
    zz-ollama-optimizer.conf`, written with `sudo tee`; never `systemctl revert`, which deletes the user's own drop-ins too
  • NoteRuns commands with sudoSKILL.md:139
    serve`, quit and reopen Ollama.app, run `sudo systemctl daemon-reload && sudo systemctl restart ollama`, quit and reopen

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from luongnv89/skills at commit 891c720, republished under its MIT licence (© luongnv89). 1,955 words, ~4,114 tokens.

Download SKILL.mdSave it as .claude/skills/ollama-optimizer/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
ollama-optimizer
description
Optimize Ollama configuration for the current machine's hardware. Use when asked to speed up Ollama, tune local LLM performance, or pick models that fit available GPU/RAM. Don't use for LM Studio, llama.cpp, vLLM, or hosted-API LLM providers.
license
MIT
effort
medium
metadata.version
1.3.0
metadata.author
Luong NGUYEN <luongnv89@gmail.com>

Ollama Optimizer

Optimize Ollama configuration based on system hardware analysis.

When to Use

Use this skill when the user asks to optimize Ollama, configure Ollama, speed up Ollama, fix Ollama running slow, set up a local LLM, tune inference speed, reduce memory usage, or select models that fit their GPU/RAM. The skill analyzes hardware (GPU, VRAM, RAM, CPU) and produces tailored recommendations.

Do not use for LM Studio, llama.cpp, vLLM, or hosted-API LLM providers (OpenAI, Anthropic) — those use different runtimes and tuning surfaces.

Safety Rules

  • Phases 1 and 2 are read-only. They run scripts/detect_system.py, ollama --version, ollama list, and ollama ps, and they change nothing.
  • An applied change is any write to the machine: an env-var write, a service or app restart, ollama pull, a Modelfile ollama create, or any sudo command. Before each applied change, show its exact command and wait for the user's yes for that change.
  • Saving the guide to the path the user chose is not an applied change.
  • If the user asks for recommendations only, apply nothing and deliver the guide. A recommendations-only request is not a decline.
  • Back up each config location before the first write to it (Choose where the env vars go).
  • Never delete a model, and never run ollama rm. List oversized models in the guide instead.
  • Never set OLLAMA_HOST=0.0.0.0 unless the user asks for network access. If they ask, warn that it exposes the Ollama API to the network without authentication.

Workflow

Fast path (opt-in only): only skip full hardware analysis if the user explicitly asks to. Otherwise always run Phases 1-4 and follow the tier-based recommendation — do not apply shortcuts by default, and do not let them override a tier decision already made. For the per-platform shortcut commands and env vars, see Platform-Specific Setup and Environment Variables.

Phase 1: System Detection

Run the detection script to gather hardware information:

bash
python3 scripts/detect_system.py

The script prints JSON on stdout and exits 0. On an unexpected failure it prints Error: ... on stderr and exits 1. If it exits 1, show that message, follow its fix hint, and run it once more. If the second run also exits 1, stop with BLOCKED (Final Report).

Parse the JSON output to identify:

  • OS and version
  • CPU model and core count
  • Total RAM / unified memory
  • GPU type, VRAM, and driver version
  • Current Ollama installation and environment variables
  • hardware_tier — the script's computed category, max_model_size, and recommended_quant

If ollama.installed is false, continue to Phase 2. The plan then starts with the install step from Platform-Specific Setup. Installing Ollama is an applied change.

Phase 2: Analyze and Recommend

Use hardware_tier from Phase 1 as the tier decision. Do not re-derive it; the table below explains what each tier means and which optimizations it implies. Override the script only with an explicit reason (e.g. VRAM shared with a display), and state that reason in the report.

Unknown VRAM. The script tiers a GPU only from vram_gb. An AMD ROCm GPU, an NVIDIA GPU whose VRAM reads [N/A], and the Windows WMIC fallback report no vram_gb, so the script returns low_vram. An Intel Mac with a discrete GPU reports an empty gpu list, so the script returns cpu_only. In these cases, ask the user for the VRAM size. If they give it, override the tier from the table and state the reason. If they do not, keep the script's tier and list it under Uncertainty: as a conservative tier.

Hardware Tier Classification:

Tier (category)Script bandMax ModelKey Optimizations
cpu_onlyNo GPU detected3Bnum_thread tuning, Q4_K_M quant
low_vram<6GB VRAM3BFlash attention, KV cache q4_0
entry6-10GB VRAM8BFlash attention, KV cache q8_0
prosumer10-16GB VRAM14BFlash attention, full offload
workstation16-48GB VRAM32BStandard config, Q5_K_M option
high_end48GB+ VRAM70B+Multiple models, Q5/Q6 quants

Apple Silicon Special Case:

  • Unified memory = shared CPU/GPU RAM; the script tiers it directly from total unified memory
  • 8GB Mac → entry
  • 16GB Mac → prosumer
  • 32GB Mac → workstation; 64GB+ Mac → high_end

Nothing to change. current_env_vars shows only the environment the script ran in. For Ollama.app or the systemd service, also read the values Ollama uses (launchctl getenv VAR, systemctl show ollama -p Environment). If every recommended env var already has the recommended value there and every installed model fits the tier, the guide says so and the run proposes no applied change.

Phase 3: Generate Optimization Plan

Read Report Templates, then create the guide with these sections:

1. System Overview

Present detected hardware specs and highlight constraints (e.g., "8GB unified memory limits to 8B models").

2. Dependency Assessment

List what's needed based on the platform:

  • macOS: Ollama only (Metal automatic)
  • Linux NVIDIA: Ollama + NVIDIA driver 450+
  • Linux AMD: Ollama + ROCm 5.0+
  • Windows: Ollama + NVIDIA driver 452+
3. Configuration Recommendations

Essential environment variables:

bash
# Always recommended
export OLLAMA_FLASH_ATTENTION=1

# Memory-constrained systems (<12GB)
export OLLAMA_KV_CACHE_TYPE=q8_0  # or q4_0 for severe constraints

Set OLLAMA_KV_CACHE_TYPE when the GPU has under 12GB of VRAM or unified memory: q4_0 for low_vram, q8_0 otherwise. OLLAMA_KV_CACHE_TYPE requires OLLAMA_FLASH_ATTENTION=1.

Model selection guidance:

  • Recommend specific models from ollama list output
  • Suggest appropriate quantization (Q4_K_M default, Q5_K_M if headroom exists)
  • Warn if current models exceed hardware capacity

Modelfile tuning (when needed):

PARAMETER num_gpu <layers>    # Partial offload for limited VRAM
PARAMETER num_thread <cores>  # CPU threads (physical cores, not hyperthreads)
PARAMETER num_ctx <size>      # Reduce context for memory savings
4. Choose where the env vars go

A shell init file reaches only an ollama serve started from that shell. Find how Ollama runs: on macOS, check for a running Ollama.app (pgrep -x Ollama); on Linux, run systemctl is-active ollama; on Windows, use the Windows row; if docker ps lists an Ollama container, use the Docker row. If the checks are inconclusive, ask the user. Then use the matching row; the commands are in Environment Variables → Setting Variables Permanently.

How Ollama runsWhere the env vars goBackup before the writeOne-command rollback
ollama serve from a terminal (macOS, Linux)the shell init file that $SHELL usescp "$RC" "$RC.ollama-bak"cp "$RC.ollama-bak" "$RC"
Ollama.app (macOS)launchctl setenv VAR value, then quit and reopen the app; the value does not survive a reboot, so list that under Uncertainty:record launchctl getenv VARlaunchctl unsetenv VAR for each var, chained on one line
systemd service (Linux)Environment= lines in a dedicated drop-in, /etc/systemd/system/ollama.service.d/zz-ollama-optimizer.conf, written with sudo tee; never systemctl revert, which deletes the user's own drop-ins toosystemctl cat ollama > ~/ollama.service.baksudo rm /etc/systemd/system/ollama.service.d/zz-ollama-optimizer.conf && sudo systemctl daemon-reload && sudo systemctl restart ollama
Windowsuser env vars via SetEnvironmentVariable(..., "User")record the current valueset each var to $null at "User" scope, chained on one line
Docker-e flags or the compose environment: listcopy the compose filerestore the copy and recreate the container
5. Execution Checklist

Provide copy-paste commands in order. Run a command only after the user approves it (Safety Rules):

  1. Back up the config location from step 4 and write the env vars. For the shell init file ($SHELL decides: ~/.zshrc, ~/.bashrc, or ~/.bash_profile):
    bash
    RC=~/.zshrc  # or ~/.bashrc / ~/.bash_profile, matching $SHELL
    cp "$RC" "$RC.ollama-bak"
    printf '\n# ollama-optimizer start\nexport OLLAMA_FLASH_ATTENTION=1\n<KV cache + other export lines from section 3, per tier>\n# ollama-optimizer end\n' >> "$RC"
  2. Restart Ollama the way the step 4 row runs it: restart ollama serve, quit and reopen Ollama.app, run sudo systemctl daemon-reload && sudo systemctl restart ollama, quit and reopen Ollama from the Windows taskbar, or recreate the container
  3. Pull recommended models
  4. Test with ollama run <model> --verbose
  5. Rollback (one command, same location as step 1): for the shell init file, cp "$RC.ollama-bak" "$RC" — then restart Ollama. Other locations use the rollback column in step 4.

If an approved command exits non-zero, stop the checklist, show the error, and keep the backup. Offer the rollback command, and run it only after the user approves it.

Show full SKILL.md (733 more words)Show less
Phase 4: Verification

Run Phase 4 only when, at this point, ollama --version exits 0 and ollama list shows at least one model. Otherwise, skip it and record the reason.

bash
# Benchmark current performance
python3 scripts/benchmark_ollama.py --model <model>
# Expected output: tokens/s and generation latency — record as the post-tuning baseline.

# Check GPU memory usage (NVIDIA only)
nvidia-smi

# Verify config is applied
ollama run <model> "test" --verbose 2>&1 | head -20

benchmark_ollama.py exits 1 with a JSON error on stderr when Ollama is missing or not running, when no model is installed, or when the requested model is absent; it exits 2 on an invalid argument. If every run reports "success": false, the model has no averages, the script prints a Warning: line on stderr, and verification failed. If no change was applied, the numbers are the current baseline, not a post-tuning result; say so under Uncertainty:.

Edge Cases

When several rows match, the first-match status rules in Final Report decide.

CaseHandlingStatus
Ollama not installed (ollama.installed: false), and not installed during the runPlan with the install step first; skip Phase 4PARTIAL
No model in ollama list, and none pulled during the runRecommend models for the tier; skip Phase 4PARTIAL
GPU with unknown VRAMAsk for VRAM; otherwise keep the conservative tierper the status rules; an Uncertainty line when the tier stays conservative
User wants recommendations onlyApply nothing; deliver the guideCOMPLETE
User declines an applied changeMark it not applied in the guide; continue with the restPARTIAL
Approved command fails, or verification failsStop the checklist; offer rollbackPARTIAL
Nothing to changeSay so in the guide; benchmark the current setupCOMPLETE
detect_system.py exits 1 twiceStop before any recommendationBLOCKED

Final Report

End every run, early stops included, with this summary. Fill rules: Report Templates → Final summary fill rules.

Result: COMPLETE | PARTIAL | BLOCKED — <tier>; <what changed>
Evidence: <checks that ran, with observed values>
Uncertainty: <untested or assumed items, or "none">
Decision: <approval needed, or "No approval needed.">
Remaining action: <one user action per line; omit when none>
Guide: <saved path, or "printed inline">

Choose the status with the first rule that matches:

  1. BLOCKED — no guide was delivered: detect_system.py exited 1 twice, or the user stopped the run before the guide was delivered.
  2. PARTIAL — the guide was delivered, and at least one of these holds: Phase 4 was skipped because Ollama or a model was missing or the user stopped the run before it, a recommended change was not applied because the user declined it or stopped the run, an approved command failed, or Phase 4 ran and failed.
  3. COMPLETE — the guide was delivered, and every applied change the user approved succeeded and was verified in Phase 4. A recommendations-only run and a nothing-to-change run are COMPLETE.

A user declining a change never produces BLOCKED.

Example

An 8GB Apple Silicon Mac running Ollama.app, where the user approved both env vars:

Result: COMPLETE — entry; OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0 applied
Evidence: detect_system.py exit 0 (entry, 8GB unified); launchctl setenv exit 0 for both vars; benchmark_ollama.py llama3.1:8b avg 21.7 tokens/s after reopening Ollama.app
Uncertainty: launchctl values reset at reboot; not re-checked after a reboot
Decision: No approval needed.
Remaining action: after a reboot, re-run the two launchctl setenv commands from the guide
Guide: ~/.config/ollama/optimization-guide.md

Acceptance Criteria

A run passes when all of the following are true:

  • Hardware tier (CPU-only / Low-VRAM / Entry / Prosumer / Workstation / High-end) is identified explicitly in the report.
  • Recommended model size fits within detected VRAM/unified-memory budget (no recommending a 14B model on an 8GB Mac).
  • Each applied env var is written to the location the running Ollama reads (Choose where the env vars go), after a backup and the user's approval. A recommendations-only run lists these commands without running them.
  • Apple Silicon special case is applied when detected — unified memory is not double-counted as separate VRAM + RAM.
  • Verification step runs ollama run <model> with --verbose and captures the actual offload/cache numbers, or the report states why Phase 4 was skipped.
  • Rollback instructions are included so the user can revert all env changes with one command.
  • The final summary passes the four reader checks in Report Templates → Reader checks: the status is on the first line, facts and assumptions are separated, each claim is traceable to a check that ran, and the next decision is explicit. Without responsive human feedback, human understanding stays unconfirmed.

Step Completion Reports

After completing each major step, output a status report in this format:

◆ [Step Name] ([step N of M] — [context])
··································································
  [Check 1]:          √ pass
  [Check 2]:          √ pass (note if relevant)
  [Check 3]:          × fail — [reason]
  [Check 4]:          √ pass
  [Criteria]:         √ N/M met
  ____________________________
  Result:             PASS | FAIL | PARTIAL

Adapt the check names to match what the step actually validates. Use √ for pass, × for fail, and — to add brief context. The "Criteria" line summarizes how many acceptance criteria were met. The "Result" line gives the overall verdict. One example per phase (Detection, Analysis, Plan, Verification): Report Templates → Step Completion Report examples.

Reference Files

Expected Output

Generate an ollama-optimization-guide.md file from the guide template in Report Templates. Ask the user where to save it (suggest ~/.config/ollama/optimization-guide.md or current directory). If the user declines a file, print the guide in the reply. Then print the final summary (Final Report).

© luongnv89, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/ollama-optimizer of luongnv89/skills.

  • SKILL.md
  • docs/README.md
  • evals/evals.json
  • references/environment_variables.md
  • references/platform_specific.md
  • references/report-template.md
  • references/vram_requirements.md
  • scripts/benchmark_ollama.py
  • scripts/detect_system.py

Open the folder on GitHubat commit 891c720

Compare with similar skills

Ollama Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ollama Optimizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ollama Optimizer this skillluongnv89/skills131—~4.1kAutomated safety check: NotesMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Local LLM Expertsickn33/agentic-awesome-skills47k2 repos~1.6kAutomated safety check: PassMIT
Jetson LLM BenchmarkNVIDIA/skills3.5k1 repos~3.1kAutomated safety check: PassApache-2.0
Cost Localruvnet/ruflo74k—~336Automated safety check: NotesMIT

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Local LLM Expert

    sickn33/agentic-awesome-skills

    Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio.

    47k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

    3.5k GitHub starsUsed in 1 repo~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Cost Local

    ruvnet/ruflo

    Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a…

    74k GitHub stars~336 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Agentsop LLM Engine Selection

    agentsope/SkillAlchemy

    Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

    466 GitHub stars~6.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from luongnv89/skills

All 37 skills in this repo
  • Dont Make Me Think

    luongnv89/skills

    Review UI usability using Steve Krug's principles and produce a scannable report.

    131 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Herdr Agent

    luongnv89/skills

    Manage AI agent fleets in Herdr: tile root + sub-agents in one tab, start/prompt/wait/read/monitor via the herdr agent CLI, steer any pane; help lists every operation.

    131 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Security Setup

    luongnv89/skills

    Install local-first security hardening: pre-commit secret detection, offline dependency scans, static analysis, reports, and gated free CI.

    131 GitHub stars~4.5k tokensUpdated today
    Auto-check passed
  • SEO AI Optimizer

    luongnv89/skills

    Audit and optimize websites for technical SEO, content SEO, and AI bot accessibility.

    131 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Tasks Generator

    luongnv89/skills

    Generate sprint-based development tasks from a PRD. An agent skill from luongnv89/skills.

    131 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Tmux Agent Comms

    luongnv89/skills

    Manage AI agents in tmux: spawn sessions, send messages, wait, capture replies, inspect fleets, and tear down safely.

    131 GitHub stars~3.4k tokensUpdated today
    Auto-check passed

Questions about Ollama Optimizer

What does Ollama Optimizer do?

Optimize Ollama configuration for the current machine's hardware. Ollama Optimizer is an agent skill from luongnv89/skills. Optimize Ollama configuration for the current machine's hardware.

When should I use Ollama Optimizer?

Ollama Optimizer fits situations like: asked to speed up Ollama; tune local LLM performance; pick models that fit available GPU/RAM; hosted-API LLM providers.

How do I install Ollama Optimizer in Claude Code?

Run `npx skills add luongnv89/skills --skill ollama-optimizer -a claude-code`. Or copy the skill folder (skills/ollama-optimizer in luongnv89/skills) into .claude/skills/ollama-optimizer in your project. Claude Code loads it when a task matches its description.

How do I install Ollama Optimizer in Codex?

Run `npx skills add luongnv89/skills --skill ollama-optimizer -a codex`. Or copy the skill folder (skills/ollama-optimizer in luongnv89/skills) into .agents/skills/ollama-optimizer in your project. Codex loads it when a task matches its description.

Can I use Ollama Optimizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add luongnv89/skills --skill ollama-optimizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ollama-optimizer, .gemini/skills/ollama-optimizer, .github/skills/ollama-optimizer and .opencode/skills/ollama-optimizer in your project.

What does Ollama Optimizer need to run?

Going by SKILL.md and its folder, Ollama Optimizer needs Python for the scripts in its folder and the command-line tools its instructions call (ollama, python3 and docker). Our summary lists: Python 3; Docker.

Does Ollama Optimizer access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Ollama Optimizer safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ollama Optimizer use?

Ollama Optimizer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ollama Optimizer use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Ollama Optimizer?

Skills that share tags, products or a category with Ollama Optimizer: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Resolve (alexziskind1/model-shelf, 130 stars), Local LLM Expert (sickn33/agentic-awesome-skills, 47k stars) and Jetson LLM Benchmark (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ollama Optimizer?

luongnv89 (a GitHub user) maintains it in luongnv89/skills, which has 131 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on October 9, 2026.

Source: luongnv89/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.