Agent skill

LLM Cost Optimization

by BagelHole in BagelHole/DevOps-Security-Agent-Skills

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

MITAuto-check passedAI & LLM Engineering

Install LLM Cost Optimization

skills CLI
$ npx skills add BagelHole/DevOps-Security-Agent-Skills --skill llm-cost-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BagelHole/DevOps-Security-Agent-Skills llm-cost-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BagelHole/DevOps-Security-Agent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/devops/ai/llm-cost-optimization .claude/skills/llm-cost-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-cost-optimization
GitHub stars
1.1k
Token cost
~2.2k tokens
SKILL.md length
217 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

  • Tasks that involve LLM cost and token optimization
  • SKILL.md covers When to Use This Skill, Cost Levers by Impact, Track Costs First and Model Right-Sizing, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Caching

What it does

LLM Cost Optimization is an agent skill from BagelHole/DevOps-Security-Agent-Skills. Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies. Track spend by team and model, set budgets, and implement cost-aware routing.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM cost and token optimization, Caching and LLM inference and serving. The repository describes itself as: Agent-ready DevOps, security, infrastructure, and compliance knowledge base with 80+ skills across Kubernetes, Terraform, AWS/Azure/GCP, AI platform operations, container… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM cost and token optimization
  • Tasks that involve Caching
  • Tasks that involve LLM inference and serving

Example prompts

  • “/llm-cost-optimization”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 0365f57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Cost Optimization loads about 2.2k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 217 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from BagelHole/DevOps-Security-Agent-Skills at commit 0365f57, republished under its MIT licence (© BagelHole). 217 words, ~2,241 tokens.

Download SKILL.mdSave it as .claude/skills/llm-cost-optimization/SKILL.md (or your agent's skills folder).
name
llm-cost-optimization
description
Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies. Track spend by team and model, set budgets, and implement cost-aware routing.
license
MIT
metadata.author
devops-skills
metadata.version
1.0

LLM Cost Optimization

Cut LLM costs by 50–90% with the right combination of caching, model selection, prompt optimization, and self-hosting.

When to Use This Skill

Use this skill when:

  • LLM API spend is growing faster than revenue
  • You need to attribute AI costs to teams, products, or customers
  • Implementing caching to avoid redundant LLM calls
  • Deciding when to switch from API providers to self-hosted models
  • Optimizing prompt length without sacrificing quality

Cost Levers by Impact

StrategyTypical SavingsEffort
Semantic caching20–50%Low
Model right-sizing30–70%Low
Prompt compression10–30%Medium
Provider caching (prompt cache)10–25%Low
Batching offline workloads50% (Batch API)Medium
Self-hosting 7–8B models80–95% at scaleHigh
Quantization30–50% VRAM costMedium

Track Costs First

python
# Use LiteLLM's cost tracking (automatic per-model pricing)
import litellm

response = litellm.completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
)
cost = litellm.completion_cost(response)
print(f"Cost: ${cost:.6f}")

# Add custom cost callbacks
def log_cost(kwargs, completion_response, start_time, end_time):
    cost = kwargs.get("response_cost", 0)
    model = kwargs.get("model")
    user = kwargs.get("user")
    # Send to your analytics DB
    db.record_cost(user=user, model=model, cost=cost)

litellm.success_callback = [log_cost]

Model Right-Sizing

python
# Route by task complexity — don't use GPT-4o for everything
def get_model_for_task(task_type: str) -> str:
    routing = {
        "classification":     "gpt-4o-mini",      # ~30× cheaper than gpt-4o
        "summarization":      "gpt-4o-mini",
        "extraction":         "gpt-4o-mini",
        "simple_qa":          "gpt-4o-mini",
        "complex_reasoning":  "gpt-4o",
        "code_generation":    "claude-sonnet-4-6",
        "creative_writing":   "claude-opus-4-6",
    }
    return routing.get(task_type, "gpt-4o-mini")

# Cost comparison (per 1M tokens, 2025 approx.)
# gpt-4o-mini:          input $0.15 / output $0.60
# gpt-4o:               input $2.50 / output $10.00
# claude-sonnet-4-6:    input $3.00 / output $15.00
# llama-3.1-8b (self):  ~$0.05–0.10 all-in (GPU amortized)

Prompt Caching (Provider-Side)

python
# Anthropic — cache long system prompts (saves 90% on cached tokens)
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": "You are a helpful assistant.",
        },
        {
            "type": "text",
            "text": open("large-context.txt").read(),  # large doc
            "cache_control": {"type": "ephemeral"},     # cache this!
        }
    ],
    messages=[{"role": "user", "content": "Summarize the key points."}],
)
# First call: full price. Subsequent calls: 90% discount on cached part.
print(f"Cache read tokens: {response.usage.cache_read_input_tokens}")

# OpenAI — prompt caching is automatic for repeated prefixes >1024 tokens
# No code change needed; check usage.prompt_tokens_details.cached_tokens

Batching with OpenAI Batch API (50% Discount)

python
import json
from openai import OpenAI

client = OpenAI()

# Prepare batch requests
requests = [
    {
        "custom_id": f"task-{i}",
        "method": "POST",
        "url": "/v1/chat/completions",
        "body": {
            "model": "gpt-4o-mini",
            "messages": [{"role": "user", "content": f"Classify: {text}"}],
            "max_tokens": 50,
        }
    }
    for i, text in enumerate(texts)
]

# Write JSONL file
with open("batch.jsonl", "w") as f:
    for req in requests:
        f.write(json.dumps(req) + "\n")

# Upload and create batch
batch_file = client.files.create(file=open("batch.jsonl", "rb"), purpose="batch")
batch = client.batches.create(
    input_file_id=batch_file.id,
    endpoint="/v1/chat/completions",
    completion_window="24h",
)
print(f"Batch ID: {batch.id}")  # poll status with client.batches.retrieve(batch.id)

Semantic Caching

python
import hashlib
import json
import redis
import numpy as np
from sentence_transformers import SentenceTransformer

r = redis.Redis(host="localhost", port=6379)
embed_model = SentenceTransformer("BAAI/bge-small-en-v1.5")

SIMILARITY_THRESHOLD = 0.92
CACHE_TTL = 3600 * 24  # 24 hours

def cached_llm_call(prompt: str, llm_fn) -> str:
    # 1. Exact match (free)
    exact_key = f"exact:{hashlib.sha256(prompt.encode()).hexdigest()}"
    if cached := r.get(exact_key):
        return cached.decode()

    # 2. Semantic match
    query_vec = embed_model.encode(prompt)
    cached_keys = r.keys("sem:*")
    for key in cached_keys:
        data = json.loads(r.get(key))
        similarity = np.dot(query_vec, data["embedding"]) / (
            np.linalg.norm(query_vec) * np.linalg.norm(data["embedding"])
        )
        if similarity >= SIMILARITY_THRESHOLD:
            return data["response"]

    # 3. Cache miss — call LLM
    response = llm_fn(prompt)

    # Store exact match
    r.setex(exact_key, CACHE_TTL, response)

    # Store semantic embedding
    sem_key = f"sem:{hashlib.sha256(prompt.encode()).hexdigest()}"
    r.setex(sem_key, CACHE_TTL, json.dumps({
        "embedding": query_vec.tolist(),
        "response": response,
        "prompt": prompt,
    }))
    return response

Prompt Compression

python
# LLMLingua — compress long prompts by 3–20× with minimal quality loss
from llmlingua import PromptCompressor

compressor = PromptCompressor(
    model_name="microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank",
    device_map="cpu",
)

compressed = compressor.compress_prompt(
    long_context,
    ratio=0.5,       # keep 50% of tokens
    rank_method="longllmlingua",
)
print(f"Original: {len(long_context.split())} words")
print(f"Compressed: {len(compressed['compressed_prompt'].split())} words")
print(f"Savings: {compressed['saving']}")

Self-Hosting Break-Even Calculator

python
def break_even_analysis(
    monthly_api_spend_usd: float,
    gpu_cost_per_hour_usd: float = 2.50,   # e.g., A10G on AWS
    utilization: float = 0.70,             # 70% GPU utilization
) -> dict:
    monthly_gpu_cost = gpu_cost_per_hour_usd * 24 * 30 * utilization
    break_even = monthly_gpu_cost / monthly_api_spend_usd
    recommendation = (
        "Self-host now — strong ROI" if break_even < 0.5 else
        "Self-host if traffic grows 2×" if break_even < 0.8 else
        "Stick with API — not enough scale yet"
    )
    return {
        "monthly_gpu_cost": f"${monthly_gpu_cost:.0f}",
        "monthly_api_spend": f"${monthly_api_spend_usd:.0f}",
        "gpu_as_pct_of_api": f"{break_even*100:.0f}%",
        "recommendation": recommendation,
    }

# Example: $5k/month on OpenAI, $2.50/hr A10G
print(break_even_analysis(5000))
# → gpu_cost ~$1,260/mo = 25% of API spend → self-host now

Cost Dashboard (Grafana)

python
# Emit cost metrics to Prometheus
from prometheus_client import Counter, Histogram

llm_cost_total = Counter(
    "llm_cost_usd_total",
    "Total LLM spend in USD",
    ["model", "team", "task_type"],
)
llm_tokens_total = Counter(
    "llm_tokens_total",
    "Total tokens used",
    ["model", "token_type"],  # token_type: prompt, completion, cached
)

def track_call(model, team, task_type, response):
    cost = calculate_cost(model, response.usage)
    llm_cost_total.labels(model=model, team=team, task_type=task_type).inc(cost)
    llm_tokens_total.labels(model=model, token_type="prompt").inc(
        response.usage.prompt_tokens)
    llm_tokens_total.labels(model=model, token_type="completion").inc(
        response.usage.completion_tokens)

Best Practices

  • Use gpt-4o-mini or claude-haiku for 80% of tasks — they're 10–30× cheaper.
  • Enable prompt caching for system prompts >1,024 tokens (Anthropic) or >1,024 tokens (OpenAI).
  • Audit your top 5 prompts by token count — compress or cache them.
  • Set hard budget limits with LiteLLM virtual keys before costs spiral.
  • Self-host 7B–8B models when monthly API spend exceeds $2k/month.

© BagelHole, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in devops/ai/llm-cost-optimization of BagelHole/DevOps-Security-Agent-Skills.

Open the folder on GitHubat commit 0365f57

Compare with similar skills

LLM Cost Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Cost Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Cost Optimization this skillBagelHole/DevOps-Security-Agent-Skills1.1k—~2.2kAutomated safety check: PassMIT
LLM Cost Optimizationsickn33/agentic-awesome-skills47k1 repos~2.5kAutomated safety check: PassMIT
Anth Performance Tuningjeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: PassMIT
LLM Cost Latency Budgetmohitagw15856/pm-claude-skills1.4k—~985Automated safety check: PassMIT
Prefix Cache Replaybenchflow-ai/skillsbench1.8k—~2.3kAutomated safety check: PassApache-2.0
AIbutterbase-ai/butterbase-skills534—~1.1kAutomated safety check: PassMIT

Similar skills

  • LLM Cost Optimization

    sickn33/agentic-awesome-skills

    Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Anth Performance Tuning

    jeremylongshore/tons-of-skills-marketplace

    Optimize Claude API performance with prompt caching, model selection, streaming, and latency reduction techniques.

    2.8k GitHub stars~1.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Latency Budget

    mohitagw15856/pm-claude-skills

    Model the cost and latency of an LLM feature before it ships and surprises the bill.

    1.4k GitHub stars~985 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Prefix Cache Replay

    benchflow-ai/skillsbench

    Replay an LLM inference request trace (Mooncake / vLLM / SGLang hashids format) against a block-level KV prefix cache and compute hit statistics.

    1.8k GitHub stars~2.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • AI

    butterbase-ai/butterbase-skills

    A skill your agent uses when calling the app's AI gateway from agent tools — chat completions, embeddings, listing models, configuring defaults or BYOK, reading token/cost usage

    534 GitHub stars~1.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from BagelHole/DevOps-Security-Agent-Skills

All 44 skills in this repo
  • Hashicorp Vault

    BagelHole/DevOps-Security-Agent-Skills

    Manage secrets and PKI with HashiCorp Vault. An agent skill from BagelHole/DevOps-Security-Agent-Skills.

    1.1k GitHub stars~2k tokensUpdated 4 mo ago
    Auto-check passed
  • Incident Response

    BagelHole/DevOps-Security-Agent-Skills

    Handle security incidents with IR playbooks and procedures. An agent skill from BagelHole/DevOps-Security-Agent-Skills.

    1.1k GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Kubernetes Ops

    BagelHole/DevOps-Security-Agent-Skills

    Deploy, scale, and manage Kubernetes workloads. An agent skill from BagelHole/DevOps-Security-Agent-Skills.

    1.1k GitHub stars~2.3k tokensUpdated 4 mo ago
    Auto-check passed
  • Linux Hardening

    BagelHole/DevOps-Security-Agent-Skills

    Apply CIS benchmarks and secure Linux servers. An agent skill from BagelHole/DevOps-Security-Agent-Skills.

    1.1k GitHub stars~662 tokensUpdated 4 mo ago
    Auto-check: notes
  • Prometheus Grafana

    BagelHole/DevOps-Security-Agent-Skills

    Set up metrics collection and visualization with Prometheus and Grafana.

    1.1k GitHub stars~2.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Vulnerability Scanning

    BagelHole/DevOps-Security-Agent-Skills

    Scan systems and dependencies for CVEs and security vulnerabilities.

    1.1k GitHub stars~2.4k tokensUpdated 4 mo ago
    Auto-check passed

Questions about LLM Cost Optimization

What does LLM Cost Optimization do?

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies. LLM Cost Optimization is an agent skill from BagelHole/DevOps-Security-Agent-Skills. Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

When should I use LLM Cost Optimization?

LLM Cost Optimization fits situations like: tasks that involve LLM cost and token optimization; tasks that involve Caching; tasks that involve LLM inference and serving.

How do I install LLM Cost Optimization in Claude Code?

Run `npx skills add BagelHole/DevOps-Security-Agent-Skills --skill llm-cost-optimization -a claude-code`. Or copy the skill folder (devops/ai/llm-cost-optimization in BagelHole/DevOps-Security-Agent-Skills) into .claude/skills/llm-cost-optimization in your project. Claude Code loads it when a task matches its description.

How do I install LLM Cost Optimization in Codex?

Run `npx skills add BagelHole/DevOps-Security-Agent-Skills --skill llm-cost-optimization -a codex`. Or copy the skill folder (devops/ai/llm-cost-optimization in BagelHole/DevOps-Security-Agent-Skills) into .agents/skills/llm-cost-optimization in your project. Codex loads it when a task matches its description.

Can I use LLM Cost Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BagelHole/DevOps-Security-Agent-Skills --skill llm-cost-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-cost-optimization, .gemini/skills/llm-cost-optimization, .github/skills/llm-cost-optimization and .opencode/skills/llm-cost-optimization in your project.

What does LLM Cost Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: LLM Cost Optimization is instructions for the agent only. Our summary lists: Python 3.

Does LLM Cost Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is LLM Cost Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Cost Optimization use?

LLM Cost Optimization is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Cost Optimization use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Cost Optimization?

Skills that share tags, products or a category with LLM Cost Optimization: LLM Cost Optimization (sickn33/agentic-awesome-skills, 47k stars), Anth Performance Tuning (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), LLM Cost Latency Budget (mohitagw15856/pm-claude-skills, 1.4k stars) and Prefix Cache Replay (benchflow-ai/skillsbench, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Cost Optimization?

BagelHole (a GitHub user) maintains it in BagelHole/DevOps-Security-Agent-Skills, which has 1,148 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on May 22, 2026.

Source: BagelHole/DevOps-Security-Agent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.