Agent skill

Cost Aware LLM Pipeline

by xu-xiang in xu-xiang/everything-claude-code-zh

LLM API 使用成本优化模式——基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示词缓存. An agent skill from xu-xiang/everything-claude-code-zh.

MITAuto-check passedAI & LLM Engineering

Install Cost Aware LLM Pipeline

skills CLI
$ npx skills add xu-xiang/everything-claude-code-zh --skill cost-aware-llm-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xu-xiang/everything-claude-code-zh cost-aware-llm-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xu-xiang/everything-claude-code-zh.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cost-aware-llm-pipeline .claude/skills/cost-aware-llm-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cost-aware-llm-pipeline
GitHub stars
2k
Token cost
~1.1k tokens
SKILL.md length
120 words
Files
1
Skills in repo
78
Repo updated
First seen
Licence
MIT

At a glance

LLM API 使用成本优化模式——基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示词缓存. An agent skill from xu-xiang/everything-claude-code-zh.

  • Works in 4 steps: 基于任务复杂度的模型路由 (Model Routing) → 不可变成本跟踪 (Immutable Cost Tracking) → 精细化重试逻辑 (Narrow Retry Logic) → …
  • Tasks that involve LLM API integration
  • SKILL.md covers 何时启用, 核心概念, 组合使用 and 价格参考 (2025-2026), plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cost Aware LLM Pipeline is an agent skill from xu-xiang/everything-claude-code-zh. LLM API 使用成本优化模式——基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示词缓存。

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM API integration and LLM cost and token optimization. The repository describes itself as: everything-claude-code 中文翻译项目:完整的 Claude Code 配置集合(agents, skills, hooks, commands, rules, MCPs)。源自 Anthropic 黑客松获胜者的实战配置,助力中文工程师高效理解与使用 Claude Code。 The licence is MIT.

When your agent uses it

  • Tasks that involve LLM API integration
  • Tasks that involve LLM cost and token optimization

Example prompts

  • “/cost-aware-llm-pipeline”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 基于任务复杂度的模型路由 (Model Routing)
  2. 不可变成本跟踪 (Immutable Cost Tracking)
  3. 精细化重试逻辑 (Narrow Retry Logic)
  4. 提示词缓存 (Prompt Caching)

What it can do on your machine

Read from SKILL.md and the folder at commit dfbf946. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cost Aware LLM Pipeline loads about 1.1k tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 120 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from xu-xiang/everything-claude-code-zh at commit dfbf946, republished under its MIT licence (© xu-xiang). 120 words, ~1,103 tokens.

Download SKILL.mdSave it as .claude/skills/cost-aware-llm-pipeline/SKILL.md (or your agent's skills folder).
name
cost-aware-llm-pipeline
description
LLM API 使用成本优化模式——基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示词缓存。
origin
ECC

成本感知型 LLM 流水线 (Cost-Aware LLM Pipeline)

在保持质量的同时控制 LLM API 成本的模式。将模型路由 (Model Routing)、预算跟踪 (Budget Tracking)、重试逻辑 (Retry Logic) 和提示词缓存 (Prompt Caching) 组合成一个可复用的流水线。

何时启用

  • 构建调用 LLM API(Claude、GPT 等)的应用
  • 处理具有不同复杂度的批量项目
  • 需要将 API 支出控制在预算范围内
  • 在不牺牲复杂任务质量的前提下优化成本

核心概念

1. 基于任务复杂度的模型路由 (Model Routing)

为简单任务自动选择更便宜的模型,将昂贵的模型留给复杂任务。

python
MODEL_SONNET = "claude-sonnet-4-6"
MODEL_HAIKU = "claude-haiku-4-5-20251001"

_SONNET_TEXT_THRESHOLD = 10_000  # 字符数阈值
_SONNET_ITEM_THRESHOLD = 30     # 项目数阈值

def select_model(
    text_length: int,
    item_count: int,
    force_model: str | None = None,
) -> str:
    """根据任务复杂度选择模型。"""
    if force_model is not None:
        return force_model
    if text_length >= _SONNET_TEXT_THRESHOLD or item_count >= _SONNET_ITEM_THRESHOLD:
        return MODEL_SONNET  # 复杂任务
    return MODEL_HAIKU  # 简单任务 (便宜 3-4 倍)
2. 不可变成本跟踪 (Immutable Cost Tracking)

使用冻结的数据类 (Frozen Dataclasses) 跟踪累计支出。每次 API 调用都会返回一个新的跟踪器——绝不修改原始状态。

python
from dataclasses import dataclass

@dataclass(frozen=True, slots=True)
class CostRecord:
    model: str
    input_tokens: int
    output_tokens: int
    cost_usd: float

@dataclass(frozen=True, slots=True)
class CostTracker:
    budget_limit: float = 1.00
    records: tuple[CostRecord, ...] = ()

    def add(self, record: CostRecord) -> "CostTracker":
        """返回添加了新记录的新跟踪器 (绝不修改自身状态)。"""
        return CostTracker(
            budget_limit=self.budget_limit,
            records=(*self.records, record),
        )

    @property
    def total_cost(self) -> float:
        return sum(r.cost_usd for r in self.records)

    @property
    def over_budget(self) -> bool:
        return self.total_cost > self.budget_limit
3. 精细化重试逻辑 (Narrow Retry Logic)

仅在瞬时错误 (Transient Errors) 时重试。对身份验证或错误请求执行快速失败 (Fail Fast)。

python
from anthropic import (
    APIConnectionError,
    InternalServerError,
    RateLimitError,
)

_RETRYABLE_ERRORS = (APIConnectionError, RateLimitError, InternalServerError)
_MAX_RETRIES = 3

def call_with_retry(func, *, max_retries: int = _MAX_RETRIES):
    """仅在瞬时错误时重试,其他错误立即报错。"""
    for attempt in range(max_retries):
        try:
            return func()
        except _RETRYABLE_ERRORS:
            if attempt == max_retries - 1:
                raise
            time.sleep(2 ** attempt)  # 指数退避 (Exponential backoff)
    # AuthenticationError, BadRequestError 等 -> 立即抛出异常
4. 提示词缓存 (Prompt Caching)

缓存较长的系统提示词 (System Prompts),避免在每次请求时重复发送。

python
messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "text",
                "text": system_prompt,
                "cache_control": {"type": "ephemeral"},  # 缓存此内容
            },
            {
                "type": "text",
                "text": user_input,  # 变量部分
            },
        ],
    }
]

组合使用

在单个流水线函数中组合所有四项技术:

python
def process(text: str, config: Config, tracker: CostTracker) -> tuple[Result, CostTracker]:
    # 1. 路由模型
    model = select_model(len(text), estimated_items, config.force_model)

    # 2. 检查预算
    if tracker.over_budget:
        raise BudgetExceededError(tracker.total_cost, tracker.budget_limit)

    # 3. 带重试与缓存的调用
    response = call_with_retry(lambda: client.messages.create(
        model=model,
        messages=build_cached_messages(system_prompt, text),
    ))

    # 4. 跟踪成本 (不可变模式)
    record = CostRecord(model=model, input_tokens=..., output_tokens=..., cost_usd=...)
    tracker = tracker.add(record)

    return parse_result(response), tracker

价格参考 (2025-2026)

模型输入 ($/1M tokens)输出 ($/1M tokens)相对成本
Haiku 4.5$0.80$4.001x
Sonnet 4.6$3.00$15.00~4x
Opus 4.5$15.00$75.00~19x

最佳实践

  • 从最便宜的模型开始,仅在达到复杂度阈值时路由到昂贵模型。
  • 在批量处理前设置明确的预算限制——宁愿提前失败也不要超支。
  • 记录模型选择决策,以便根据真实数据调整阈值。
  • 为超过 1024 tokens 的系统提示词使用提示词缓存——既能节省成本又能降低延迟。
  • 绝不在身份验证或校验错误时重试——仅重试瞬时故障(网络、频率限制、服务器错误)。

应避免的反模式 (Anti-Patterns)

  • 不分复杂度,对所有请求都使用最昂贵的模型。
  • 对所有错误进行重试(在永久性失败上浪费预算)。
  • 修改成本跟踪状态(增加调试和审计难度)。
  • 在整个代码库中硬编码模型名称(应使用常量或配置)。
  • 忽视对重复系统提示词的提示词缓存。

使用场景

  • 任何调用 Claude、OpenAI 或类似 LLM API 的应用。
  • 成本累积迅速的批量处理流水线。
  • 需要智能路由的多模型架构。
  • 需要预算安全护栏 (Guardrails) 的生产系统。

© xu-xiang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cost-aware-llm-pipeline of xu-xiang/everything-claude-code-zh.

Open the folder on GitHubat commit dfbf946

Compare with similar skills

Cost Aware LLM Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cost Aware LLM Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cost Aware LLM Pipeline this skillxu-xiang/everything-claude-code-zh2k—~1.1kAutomated safety check: PassMIT
AIbutterbase-ai/butterbase-skills534—~1.1kAutomated safety check: PassMIT
LLM Cost Optimizationsickn33/agentic-awesome-skills47k1 repos~2.5kAutomated safety check: PassMIT
Cost Aware LLM PipelineaAAaqwq/AGI-Super-Team1055 repos~1.4kAutomated safety check: PassMIT
LLM Cost Optimizerborghei/Claude-Skills891—~1.1kAutomated safety check: PassMIT
Claude APIkid-sid/claude-spellbook190—~2.7kAutomated safety check: PassMIT

Similar skills

  • AI

    butterbase-ai/butterbase-skills

    A skill your agent uses when calling the app's AI gateway from agent tools — chat completions, embeddings, listing models, configuring defaults or BYOK, reading token/cost usage

    534 GitHub stars~1.1k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Optimization

    sickn33/agentic-awesome-skills

    Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Cost Aware LLM Pipeline

    aAAaqwq/AGI-Super-Team

    Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching.

    105 GitHub starsUsed in 5 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Cost Optimizer

    borghei/Claude-Skills

    This skill should be used when the user asks to "estimate LLM costs", "count tokens in prompts", "optimize prompt token usage", "compare model pricing", or "reduce LLM API costs".

    891 GitHub stars~1.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Claude API

    kid-sid/claude-spellbook

    A skill your agent uses when building or debugging apps that call the Claude API — implementing tool use, streaming, vision, prompt caching, batch processing, extended thinking, or an agentic loop…

    190 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Anth Cost Tuning

    jeremylongshore/tons-of-skills-marketplace

    Optimize Anthropic Claude API costs with model routing, prompt caching, batching, and spend monitoring.

    2.8k GitHub stars~2.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from xu-xiang/everything-claude-code-zh

All 78 skills in this repo
  • Configure Ecc

    xu-xiang/everything-claude-code-zh

    Everything Claude Code 的交互式安装程序 — 引导用户选择并安装技能和规则到用户级或项目级目录,验证路径,并可选择优化已安装文件。

    2k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Continuous Learning V2

    xu-xiang/everything-claude-code-zh

    基于本能(Instinct)的学习系统,通过钩子(hooks)观察会话,创建带有置信度评分的原子本能,并将其演化为技能(Skills)、命令(Commands)或智能体(Agents)。v2.1 版本增加了项目作用域(project-scoped)的本能,以防止跨项目污染。

    2k GitHub stars~2.1k tokensUpdated 7 mo ago
    Auto-check passed
  • API Design

    xu-xiang/everything-claude-code-zh

    生产级 API 的 REST API 设计模式,包括资源命名、状态码、分页、过滤、错误响应、版本控制和速率限制. An agent skill from xu-xiang/everything-claude-code-zh.

    2k GitHub stars~2.7k tokensUpdated 7 mo ago
    Auto-check passed
  • Backend Patterns

    xu-xiang/everything-claude-code-zh

    后端架构模式、API 设计、数据库优化以及适用于 Node.js、Express 和 Next.js API 路由的服务端最佳实践。

    2k GitHub stars~3.2k tokensUpdated 7 mo ago
    Auto-check passed
  • Backend Patterns

    xu-xiang/everything-claude-code-zh

    后端架构模式、API 设计、数据库优化以及 Node.js、Express 和 Next.js API 路由的服务端最佳实践。

    2k GitHub stars~3.1k tokensUpdated 7 mo ago
    Auto-check passed
  • Backend Patterns

    xu-xiang/everything-claude-code-zh

    后端架构模式、API 设计、数据库优化以及针对 Node.js、Express 和 Next.js API 路由的服务端最佳实践。

    2k GitHub stars~3.2k tokensUpdated 7 mo ago
    Auto-check passed

Questions about Cost Aware LLM Pipeline

What does Cost Aware LLM Pipeline do?

LLM API 使用成本优化模式——基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示词缓存. An agent skill from xu-xiang/everything-claude-code-zh. Cost Aware LLM Pipeline is an agent skill from xu-xiang/everything-claude-code-zh.

When should I use Cost Aware LLM Pipeline?

Cost Aware LLM Pipeline fits situations like: tasks that involve LLM API integration; tasks that involve LLM cost and token optimization.

How do I install Cost Aware LLM Pipeline in Claude Code?

Run `npx skills add xu-xiang/everything-claude-code-zh --skill cost-aware-llm-pipeline -a claude-code`. Or copy the skill folder (skills/cost-aware-llm-pipeline in xu-xiang/everything-claude-code-zh) into .claude/skills/cost-aware-llm-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Cost Aware LLM Pipeline in Codex?

Run `npx skills add xu-xiang/everything-claude-code-zh --skill cost-aware-llm-pipeline -a codex`. Or copy the skill folder (skills/cost-aware-llm-pipeline in xu-xiang/everything-claude-code-zh) into .agents/skills/cost-aware-llm-pipeline in your project. Codex loads it when a task matches its description.

Can I use Cost Aware LLM Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xu-xiang/everything-claude-code-zh --skill cost-aware-llm-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cost-aware-llm-pipeline, .gemini/skills/cost-aware-llm-pipeline, .github/skills/cost-aware-llm-pipeline and .opencode/skills/cost-aware-llm-pipeline in your project.

What does Cost Aware LLM Pipeline need to run?

SKILL.md names no scripts, command-line tools or credentials: Cost Aware LLM Pipeline is instructions for the agent only. Our summary lists: Python 3.

Does Cost Aware LLM Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cost Aware LLM Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cost Aware LLM Pipeline use?

Cost Aware LLM Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cost Aware LLM Pipeline use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cost Aware LLM Pipeline?

Skills that share tags, products or a category with Cost Aware LLM Pipeline: AI (butterbase-ai/butterbase-skills, 534 stars), LLM Cost Optimization (sickn33/agentic-awesome-skills, 47k stars), Cost Aware LLM Pipeline (aAAaqwq/AGI-Super-Team, 105 stars) and LLM Cost Optimizer (borghei/Claude-Skills, 891 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cost Aware LLM Pipeline?

xu-xiang (a GitHub user) maintains it in xu-xiang/everything-claude-code-zh, which has 1,978 GitHub stars. The repository holds 78 skills in this directory. The repository was last updated on March 5, 2026.

Source: xu-xiang/everything-claude-code-zh on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.