Agent skill

Cost Aware LLM Pipeline

by aAAaqwq in aAAaqwq/AGI-Super-Team

Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching.

MITAuto-check passedAI & LLM Engineering

Install Cost Aware LLM Pipeline

skills CLI
$ npx skills add aAAaqwq/AGI-Super-Team --skill cost-aware-llm-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aAAaqwq/AGI-Super-Team cost-aware-llm-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aAAaqwq/AGI-Super-Team.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cost-aware-llm-pipeline .claude/skills/cost-aware-llm-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cost-aware-llm-pipeline
GitHub stars
105
Used in
5 other repos
Token cost
~1.4k tokens
SKILL.md length
324 words
Files
1
Skills in repo
167
Repo updated
First seen
Licence
MIT

At a glance

Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching.

  • Works in 4 steps: Model Routing by Task Complexity → Immutable Cost Tracking → Narrow Retry Logic → …
  • Tasks that involve LLM cost and token optimization
  • SKILL.md covers When to Activate, Core Concepts, Composition and Pricing Reference (2025-2026), plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cost Aware LLM Pipeline is an agent skill from aAAaqwq/AGI-Super-Team. Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM cost and token optimization, Model routing and gateways and Error handling. The repository describes itself as: An installable, cross-framework AI organization: C-suite agents, expert subagents, curated skills, independent review, and one-command setup across 18 AI client/runtime adapters. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM cost and token optimization
  • Tasks that involve Model routing and gateways
  • Tasks that involve Error handling

Example prompts

  • “/cost-aware-llm-pipeline”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Model Routing by Task Complexity
  2. Immutable Cost Tracking
  3. Narrow Retry Logic
  4. Prompt Caching

What it can do on your machine

Read from SKILL.md and the folder at commit 7cefd81. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cost Aware LLM Pipeline loads about 1.4k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 324 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aAAaqwq/AGI-Super-Team at commit 7cefd81, republished under its MIT licence (© aAAaqwq). 324 words, ~1,427 tokens.

Download SKILL.mdSave it as .claude/skills/cost-aware-llm-pipeline/SKILL.md (or your agent's skills folder).
name
cost-aware-llm-pipeline
description
Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching.
origin
ECC

Cost-Aware LLM Pipeline

Patterns for controlling LLM API costs while maintaining quality. Combines model routing, budget tracking, retry logic, and prompt caching into a composable pipeline.

When to Activate

  • Building applications that call LLM APIs (Claude, GPT, etc.)
  • Processing batches of items with varying complexity
  • Need to stay within a budget for API spend
  • Optimizing cost without sacrificing quality on complex tasks

Core Concepts

1. Model Routing by Task Complexity

Automatically select cheaper models for simple tasks, reserving expensive models for complex ones.

python
MODEL_SONNET = "claude-sonnet-4-6"
MODEL_HAIKU = "claude-haiku-4-5-20251001"

_SONNET_TEXT_THRESHOLD = 10_000  # chars
_SONNET_ITEM_THRESHOLD = 30     # items

def select_model(
    text_length: int,
    item_count: int,
    force_model: str | None = None,
) -> str:
    """Select model based on task complexity."""
    if force_model is not None:
        return force_model
    if text_length >= _SONNET_TEXT_THRESHOLD or item_count >= _SONNET_ITEM_THRESHOLD:
        return MODEL_SONNET  # Complex task
    return MODEL_HAIKU  # Simple task (3-4x cheaper)
2. Immutable Cost Tracking

Track cumulative spend with frozen dataclasses. Each API call returns a new tracker — never mutates state.

python
from dataclasses import dataclass

@dataclass(frozen=True, slots=True)
class CostRecord:
    model: str
    input_tokens: int
    output_tokens: int
    cost_usd: float

@dataclass(frozen=True, slots=True)
class CostTracker:
    budget_limit: float = 1.00
    records: tuple[CostRecord, ...] = ()

    def add(self, record: CostRecord) -> "CostTracker":
        """Return new tracker with added record (never mutates self)."""
        return CostTracker(
            budget_limit=self.budget_limit,
            records=(*self.records, record),
        )

    @property
    def total_cost(self) -> float:
        return sum(r.cost_usd for r in self.records)

    @property
    def over_budget(self) -> bool:
        return self.total_cost > self.budget_limit
3. Narrow Retry Logic

Retry only on transient errors. Fail fast on authentication or bad request errors.

python
from anthropic import (
    APIConnectionError,
    InternalServerError,
    RateLimitError,
)

_RETRYABLE_ERRORS = (APIConnectionError, RateLimitError, InternalServerError)
_MAX_RETRIES = 3

def call_with_retry(func, *, max_retries: int = _MAX_RETRIES):
    """Retry only on transient errors, fail fast on others."""
    for attempt in range(max_retries):
        try:
            return func()
        except _RETRYABLE_ERRORS:
            if attempt == max_retries - 1:
                raise
            time.sleep(2 ** attempt)  # Exponential backoff
    # AuthenticationError, BadRequestError etc. → raise immediately
4. Prompt Caching

Cache long system prompts to avoid resending them on every request.

python
messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "text",
                "text": system_prompt,
                "cache_control": {"type": "ephemeral"},  # Cache this
            },
            {
                "type": "text",
                "text": user_input,  # Variable part
            },
        ],
    }
]

Composition

Combine all four techniques in a single pipeline function:

python
def process(text: str, config: Config, tracker: CostTracker) -> tuple[Result, CostTracker]:
    # 1. Route model
    model = select_model(len(text), estimated_items, config.force_model)

    # 2. Check budget
    if tracker.over_budget:
        raise BudgetExceededError(tracker.total_cost, tracker.budget_limit)

    # 3. Call with retry + caching
    response = call_with_retry(lambda: client.messages.create(
        model=model,
        messages=build_cached_messages(system_prompt, text),
    ))

    # 4. Track cost (immutable)
    record = CostRecord(model=model, input_tokens=..., output_tokens=..., cost_usd=...)
    tracker = tracker.add(record)

    return parse_result(response), tracker

Pricing Reference (2025-2026)

ModelInput ($/1M tokens)Output ($/1M tokens)Relative Cost
Haiku 4.5$0.80$4.001x
Sonnet 4.6$3.00$15.00~4x
Opus 4.5$15.00$75.00~19x

Best Practices

  • Start with the cheapest model and only route to expensive models when complexity thresholds are met
  • Set explicit budget limits before processing batches — fail early rather than overspend
  • Log model selection decisions so you can tune thresholds based on real data
  • Use prompt caching for system prompts over 1024 tokens — saves both cost and latency
  • Never retry on authentication or validation errors — only transient failures (network, rate limit, server error)

Anti-Patterns to Avoid

  • Using the most expensive model for all requests regardless of complexity
  • Retrying on all errors (wastes budget on permanent failures)
  • Mutating cost tracking state (makes debugging and auditing difficult)
  • Hardcoding model names throughout the codebase (use constants or config)
  • Ignoring prompt caching for repetitive system prompts

When to Use

  • Any application calling Claude, OpenAI, or similar LLM APIs
  • Batch processing pipelines where cost adds up quickly
  • Multi-model architectures that need intelligent routing
  • Production systems that need budget guardrails

© aAAaqwq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cost-aware-llm-pipeline of aAAaqwq/AGI-Super-Team.

Open the folder on GitHubat commit 7cefd81

Used in 5 other repositories

We found 11 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in aAAaqwq/AGI-Super-Team, which our catalogue first saw on October 9, 2026.

Compare with similar skills

Cost Aware LLM Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cost Aware LLM Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cost Aware LLM Pipeline this skillaAAaqwq/AGI-Super-Team1055 repos~1.4kAutomated safety check: PassMIT
Anth Cost Tuningjeremylongshore/tons-of-skills-marketplace2.8k—~2.1kAutomated safety check: PassMIT
ClawRouter LLM GatewayBlockRunAI/ClawRouter6.6k—~6.8kAutomated safety check: PassMIT
AIbutterbase-ai/butterbase-skills534—~1.1kAutomated safety check: PassMIT
Cost TrackingHabitat-Thinking/ai-literacy-superpowers114—~1.7kAutomated safety check: PassCustom licence
LLM Cost Optimizationsickn33/agentic-awesome-skills47k1 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Anth Cost Tuning

    jeremylongshore/tons-of-skills-marketplace

    Optimize Anthropic Claude API costs with model routing, prompt caching, batching, and spend monitoring.

    2.8k GitHub stars~2.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • ClawRouter LLM Gateway

    BlockRunAI/ClawRouter

    Describes ClawRouter, a local proxy that forwards each LLM request to the blockrun.ai gateway, which routes to a cheaper capable model, paid by USDC wallet or API key credit.

    6.6k GitHub stars~6.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • AI

    butterbase-ai/butterbase-skills

    A skill your agent uses when calling the app's AI gateway from agent tools — chat completions, embeddings, listing models, configuring defaults or BYOK, reading token/cost usage

    534 GitHub stars~1.1k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Cost Tracking

    Habitat-Thinking/ai-literacy-superpowers

    A skill your agent uses when the user wants to capture AI tool costs, review spending trends, set cost budgets, or integrate cost data into health snapshots — guides quarterly cost capture, records…

    114 GitHub stars~1.7k tokensUpdated 20 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Optimization

    sickn33/agentic-awesome-skills

    Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Token Optimizer

    LeoYeAI/openclaw-master-skills

    Reduce OpenClaw token usage and API costs through smart model routing, heartbeat optimization, budget tracking, and multi-provider fallbacks.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from aAAaqwq/AGI-Super-Team

All 167 skills in this repo
  • Content Creator

    aAAaqwq/AGI-Super-Team

    Create SEO-optimized marketing content with consistent brand voice.

    105 GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • Financial Calculator

    aAAaqwq/AGI-Super-Team

    Advanced financial calculator with future value tables, present value, discount calculations, markup pricing, and compound interest.

    105 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Bankr Signals

    aAAaqwq/AGI-Super-Team

    Transaction-verified trading signals on Base blockchain. An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub starsUsed in 2 repos~3.3k tokens
    Auto-check passed
  • Erc 8004

    aAAaqwq/AGI-Super-Team

    Register AI agents on Ethereum mainnet using ERC-8004 (Trustless Agents).

    105 GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Frontend Design Ultimate

    aAAaqwq/AGI-Super-Team

    Create distinctive, production-grade static sites with React, Tailwind CSS, and shadcn/ui — no mockups needed.

    105 GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Zsxq Smart Publish

    aAAaqwq/AGI-Super-Team

    Publish and manage content on 知识星球 (zsxq.com). An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed

Questions about Cost Aware LLM Pipeline

What does Cost Aware LLM Pipeline do?

Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching. Cost Aware LLM Pipeline is an agent skill from aAAaqwq/AGI-Super-Team. Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching.

When should I use Cost Aware LLM Pipeline?

Cost Aware LLM Pipeline fits situations like: tasks that involve LLM cost and token optimization; tasks that involve Model routing and gateways; tasks that involve Error handling.

How do I install Cost Aware LLM Pipeline in Claude Code?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill cost-aware-llm-pipeline -a claude-code`. Or copy the skill folder (skills/cost-aware-llm-pipeline in aAAaqwq/AGI-Super-Team) into .claude/skills/cost-aware-llm-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Cost Aware LLM Pipeline in Codex?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill cost-aware-llm-pipeline -a codex`. Or copy the skill folder (skills/cost-aware-llm-pipeline in aAAaqwq/AGI-Super-Team) into .agents/skills/cost-aware-llm-pipeline in your project. Codex loads it when a task matches its description.

Can I use Cost Aware LLM Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aAAaqwq/AGI-Super-Team --skill cost-aware-llm-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cost-aware-llm-pipeline, .gemini/skills/cost-aware-llm-pipeline, .github/skills/cost-aware-llm-pipeline and .opencode/skills/cost-aware-llm-pipeline in your project.

What does Cost Aware LLM Pipeline need to run?

SKILL.md names no scripts, command-line tools or credentials: Cost Aware LLM Pipeline is instructions for the agent only. Our summary lists: Python 3.

Does Cost Aware LLM Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cost Aware LLM Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cost Aware LLM Pipeline use?

Cost Aware LLM Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cost Aware LLM Pipeline use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cost Aware LLM Pipeline?

Skills that share tags, products or a category with Cost Aware LLM Pipeline: Anth Cost Tuning (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), ClawRouter LLM Gateway (BlockRunAI/ClawRouter, 6.6k stars), AI (butterbase-ai/butterbase-skills, 534 stars) and Cost Tracking (Habitat-Thinking/ai-literacy-superpowers, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cost Aware LLM Pipeline?

aAAaqwq (a GitHub user) maintains it in aAAaqwq/AGI-Super-Team, which has 105 GitHub stars. The repository holds 167 skills in this directory. The repository was last updated on October 8, 2026.

Source: aAAaqwq/AGI-Super-Team on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.