Agent skill

LLM Cost Optimization

by sickn33 in sickn33/agentic-awesome-skills

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

MITAuto-check passedAI & LLM Engineering

Install LLM Cost Optimization

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill llm-cost-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills llm-cost-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-cost-optimization .claude/skills/llm-cost-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-cost-optimization
GitHub stars
47k
Used in
1 other repo
Token cost
~2.5k tokens
SKILL.md length
276 words
Files
1
Skills in repo
1,354
Repo updated
First seen
Licence
MIT

At a glance

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

  • Tasks that involve LLM cost and token optimization
  • SKILL.md covers When to Use This Skill, Cost Levers by Impact, Track Costs First and Model Right-Sizing, plus 9 more sections
  • Calls git and kubectl
  • Tasks that involve Caching

What it does

LLM Cost Optimization is an agent skill from sickn33/agentic-awesome-skills. Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and…

It sits in AI & LLM Engineering, covering LLM cost and token optimization, Caching and LLM inference and serving. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM cost and token optimization
  • Tasks that involve Caching
  • Tasks that involve LLM inference and serving

Example prompts

  • “/llm-cost-optimization”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit ec02547. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

LLM Cost Optimization loads about 2.5k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 276 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit ec02547, republished under its MIT licence (© sickn33). 276 words, ~2,481 tokens.

Download SKILL.mdSave it as .claude/skills/llm-cost-optimization/SKILL.md (or your agent's skills folder).
name
llm-cost-optimization
description
Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.
compatibility
Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

LLM Cost Optimization

Cut LLM costs by 50–90% with the right combination of caching, model selection, prompt optimization, and self-hosting.

When to Use This Skill

Use this skill when:

  • LLM API spend is growing faster than revenue
  • You need to attribute AI costs to teams, products, or customers
  • Implementing caching to avoid redundant LLM calls
  • Deciding when to switch from API providers to self-hosted models
  • Optimizing prompt length without sacrificing quality

Cost Levers by Impact

StrategyTypical SavingsEffort
Semantic caching20–50%Low
Model right-sizing30–70%Low
Prompt compression10–30%Medium
Provider caching (prompt cache)10–25%Low
Batching offline workloads50% (Batch API)Medium
Self-hosting 7–8B models80–95% at scaleHigh
Quantization30–50% VRAM costMedium

Track Costs First

python
# Use LiteLLM's cost tracking (automatic per-model pricing)
import litellm

response = litellm.completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
)
cost = litellm.completion_cost(response)
print(f"Cost: ${cost:.6f}")

# Add custom cost callbacks
def log_cost(kwargs, completion_response, start_time, end_time):
    cost = kwargs.get("response_cost", 0)
    model = kwargs.get("model")
    user = kwargs.get("user")
    # Send to your analytics DB
    db.record_cost(user=user, model=model, cost=cost)

litellm.success_callback = [log_cost]

Model Right-Sizing

python
# Route by task complexity — don't use GPT-4o for everything
def get_model_for_task(task_type: str) -> str:
    routing = {
        "classification":     "gpt-4o-mini",      # ~30× cheaper than gpt-4o
        "summarization":      "gpt-4o-mini",
        "extraction":         "gpt-4o-mini",
        "simple_qa":          "gpt-4o-mini",
        "complex_reasoning":  "gpt-4o",
        "code_generation":    "claude-sonnet-4-6",
        "creative_writing":   "claude-opus-4-6",
    }
    return routing.get(task_type, "gpt-4o-mini")

# Cost comparison (per 1M tokens, 2025 approx.)
# gpt-4o-mini:          input $0.15 / output $0.60
# gpt-4o:               input $2.50 / output $10.00
# claude-sonnet-4-6:    input $3.00 / output $15.00
# llama-3.1-8b (self):  ~$0.05–0.10 all-in (GPU amortized)

Prompt Caching (Provider-Side)

python
# Anthropic — cache long system prompts (saves 90% on cached tokens)
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": "You are a helpful assistant.",
        },
        {
            "type": "text",
            "text": open("large-context.txt").read(),  # large doc
            "cache_control": {"type": "ephemeral"},     # cache this!
        }
    ],
    messages=[{"role": "user", "content": "Summarize the key points."}],
)
# First call: full price. Subsequent calls: 90% discount on cached part.
print(f"Cache read tokens: {response.usage.cache_read_input_tokens}")

# OpenAI — prompt caching is automatic for repeated prefixes >1024 tokens
# No code change needed; check usage.prompt_tokens_details.cached_tokens

Batching with OpenAI Batch API (50% Discount)

python
import json
from openai import OpenAI

client = OpenAI()

# Prepare batch requests
requests = [
    {
        "custom_id": f"task-{i}",
        "method": "POST",
        "url": "/v1/chat/completions",
        "body": {
            "model": "gpt-4o-mini",
            "messages": [{"role": "user", "content": f"Classify: {text}"}],
            "max_tokens": 50,
        }
    }
    for i, text in enumerate(texts)
]

# Write JSONL file
with open("batch.jsonl", "w") as f:
    for req in requests:
        f.write(json.dumps(req) + "\n")

# Upload and create batch
batch_file = client.files.create(file=open("batch.jsonl", "rb"), purpose="batch")
batch = client.batches.create(
    input_file_id=batch_file.id,
    endpoint="/v1/chat/completions",
    completion_window="24h",
)
print(f"Batch ID: {batch.id}")  # poll status with client.batches.retrieve(batch.id)

Semantic Caching

python
import hashlib
import json
import redis
import numpy as np
from sentence_transformers import SentenceTransformer

r = redis.Redis(host="localhost", port=6379)
embed_model = SentenceTransformer("BAAI/bge-small-en-v1.5")

SIMILARITY_THRESHOLD = 0.92
CACHE_TTL = 3600 * 24  # 24 hours

def cached_llm_call(prompt: str, llm_fn) -> str:
    # 1. Exact match (free)
    exact_key = f"exact:{hashlib.sha256(prompt.encode()).hexdigest()}"
    if cached := r.get(exact_key):
        return cached.decode()

    # 2. Semantic match
    query_vec = embed_model.encode(prompt)
    cached_keys = r.keys("sem:*")
    for key in cached_keys:
        data = json.loads(r.get(key))
        similarity = np.dot(query_vec, data["embedding"]) / (
            np.linalg.norm(query_vec) * np.linalg.norm(data["embedding"])
        )
        if similarity >= SIMILARITY_THRESHOLD:
            return data["response"]

    # 3. Cache miss — call LLM
    response = llm_fn(prompt)

    # Store exact match
    r.setex(exact_key, CACHE_TTL, response)

    # Store semantic embedding
    sem_key = f"sem:{hashlib.sha256(prompt.encode()).hexdigest()}"
    r.setex(sem_key, CACHE_TTL, json.dumps({
        "embedding": query_vec.tolist(),
        "response": response,
        "prompt": prompt,
    }))
    return response

Prompt Compression

python
# LLMLingua — compress long prompts by 3–20× with minimal quality loss
from llmlingua import PromptCompressor

compressor = PromptCompressor(
    model_name="microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank",
    device_map="cpu",
)

compressed = compressor.compress_prompt(
    long_context,
    ratio=0.5,       # keep 50% of tokens
    rank_method="longllmlingua",
)
print(f"Original: {len(long_context.split())} words")
print(f"Compressed: {len(compressed['compressed_prompt'].split())} words")
print(f"Savings: {compressed['saving']}")

Self-Hosting Break-Even Calculator

python
def break_even_analysis(
    monthly_api_spend_usd: float,
    gpu_cost_per_hour_usd: float = 2.50,   # e.g., A10G on AWS
    utilization: float = 0.70,             # 70% GPU utilization
) -> dict:
    monthly_gpu_cost = gpu_cost_per_hour_usd * 24 * 30 * utilization
    break_even = monthly_gpu_cost / monthly_api_spend_usd
    recommendation = (
        "Self-host now — strong ROI" if break_even < 0.5 else
        "Self-host if traffic grows 2×" if break_even < 0.8 else
        "Stick with API — not enough scale yet"
    )
    return {
        "monthly_gpu_cost": f"${monthly_gpu_cost:.0f}",
        "monthly_api_spend": f"${monthly_api_spend_usd:.0f}",
        "gpu_as_pct_of_api": f"{break_even*100:.0f}%",
        "recommendation": recommendation,
    }

# Example: $5k/month on OpenAI, $2.50/hr A10G
print(break_even_analysis(5000))
# → gpu_cost ~$1,260/mo = 25% of API spend → self-host now

Cost Dashboard (Grafana)

python
# Emit cost metrics to Prometheus
from prometheus_client import Counter, Histogram

llm_cost_total = Counter(
    "llm_cost_usd_total",
    "Total LLM spend in USD",
    ["model", "team", "task_type"],
)
llm_tokens_total = Counter(
    "llm_tokens_total",
    "Total tokens used",
    ["model", "token_type"],  # token_type: prompt, completion, cached
)

def track_call(model, team, task_type, response):
    cost = calculate_cost(model, response.usage)
    llm_cost_total.labels(model=model, team=team, task_type=task_type).inc(cost)
    llm_tokens_total.labels(model=model, token_type="prompt").inc(
        response.usage.prompt_tokens)
    llm_tokens_total.labels(model=model, token_type="completion").inc(
        response.usage.completion_tokens)

Best Practices

  • Use gpt-4o-mini or claude-haiku for 80% of tasks — they're 10–30× cheaper.
  • Enable prompt caching for system prompts >1,024 tokens (Anthropic) or >1,024 tokens (OpenAI).
  • Audit your top 5 prompts by token count — compress or cache them.
  • Set hard budget limits with LiteLLM virtual keys before costs spiral.
  • Self-host 7B–8B models when monthly API spend exceeds $2k/month.
  • llm-gateway (llm-gateway) - Centralized cost control
  • llm-caching (llm-caching) - Semantic caching patterns
  • vllm-server (vllm-server) - Self-hosted inference
  • agent-observability (agent-observability) - Token and cost telemetry

Limitations

  • Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
  • Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
Example
bash
git status && git diff --stat
kubectl diff -f manifest.yaml

Adapted from BagelHole/DevOps-Security-Agent-Skills (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/llm-cost-optimization of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit ec02547

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LLM Cost Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Cost Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Cost Optimization this skillsickn33/agentic-awesome-skills47k1 repos~2.5kAutomated safety check: PassMIT
Anth Performance Tuningjeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: PassMIT
LLM Cost Latency Budgetmohitagw15856/pm-claude-skills1.4k—~985Automated safety check: PassMIT
Prefix Cache Replaybenchflow-ai/skillsbench1.8k—~2.3kAutomated safety check: PassApache-2.0
AIbutterbase-ai/butterbase-skills534—~1.1kAutomated safety check: PassMIT
SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT

Similar skills

  • Anth Performance Tuning

    jeremylongshore/tons-of-skills-marketplace

    Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques.

    2.8k GitHub stars~1.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Latency Budget

    mohitagw15856/pm-claude-skills

    Model the cost and latency of an LLM feature before it ships and surprises the bill.

    1.4k GitHub stars~985 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Prefix Cache Replay

    benchflow-ai/skillsbench

    Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics.

    1.8k GitHub stars~2.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • AI

    butterbase-ai/butterbase-skills

    A skill your agent uses when calling the app's AI gateway from agent tools — chat completions, embeddings, listing models, configuring defaults or BYOK, reading token/cost usage

    534 GitHub stars~1.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 3 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Prefix Cache Bench

    vllm-project/vllm-skills

    This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

    103 GitHub stars~1.4k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,354 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Content Creator

    sickn33/agentic-awesome-skills

    Drafts and reviews audience-specific content from supplied brand examples, with local scripts for brand voice and SEO diagnostics, channel templates and a content calendar.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed

Questions about LLM Cost Optimization

What does LLM Cost Optimization do?

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies. LLM Cost Optimization is an agent skill from sickn33/agentic-awesome-skills. Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

When should I use LLM Cost Optimization?

LLM Cost Optimization fits situations like: tasks that involve LLM cost and token optimization; tasks that involve Caching; tasks that involve LLM inference and serving.

How do I install LLM Cost Optimization in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill llm-cost-optimization -a claude-code`. Or copy the skill folder (skills/llm-cost-optimization in sickn33/agentic-awesome-skills) into .claude/skills/llm-cost-optimization in your project. Claude Code loads it when a task matches its description.

How do I install LLM Cost Optimization in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill llm-cost-optimization -a codex`. Or copy the skill folder (skills/llm-cost-optimization in sickn33/agentic-awesome-skills) into .agents/skills/llm-cost-optimization in your project. Codex loads it when a task matches its description.

Can I use LLM Cost Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill llm-cost-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-cost-optimization, .gemini/skills/llm-cost-optimization, .github/skills/llm-cost-optimization and .opencode/skills/llm-cost-optimization in your project.

What does LLM Cost Optimization need to run?

Going by SKILL.md and its folder, LLM Cost Optimization needs the command-line tools its instructions call (git and kubectl). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires the relevant platform CLIs (kubectl, helm, terraform, git, CI runners) and authorized access to the target environment. Docs-only; helper scripts and templates not bundled..

Does LLM Cost Optimization access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is LLM Cost Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Cost Optimization use?

LLM Cost Optimization is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Cost Optimization use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Cost Optimization?

Skills that share tags, products or a category with LLM Cost Optimization: Anth Performance Tuning (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), LLM Cost Latency Budget (mohitagw15856/pm-claude-skills, 1.4k stars), Prefix Cache Replay (benchflow-ai/skillsbench, 1.8k stars) and AI (butterbase-ai/butterbase-skills, 534 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Cost Optimization?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,343 GitHub stars. The repository holds 1,354 skills in this directory. The repository was last updated on October 7, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.