Agent skill

Cost Aware LLM Pipeline

by affaan-m in affaan-m/ECC

LLM API 使用成本优化模式 —— 基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示缓存. An agent skill from affaan-m/ECC.

MITAuto-check passedAI & LLM Engineering

Install Cost Aware LLM Pipeline

skills CLI
$ npx skills add affaan-m/ECC --skill cost-aware-llm-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC cost-aware-llm-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/zh-CN/skills/cost-aware-llm-pipeline .claude/skills/cost-aware-llm-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cost-aware-llm-pipeline
GitHub stars
277k
Used in
3 other repos
Token cost
~1.1k tokens
SKILL.md length
107 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

LLM API 使用成本优化模式 —— 基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示缓存. An agent skill from affaan-m/ECC.

  • Works in 4 steps: 根据任务复杂度进行模型路由 → 不可变的成本跟踪 → 窄范围重试逻辑 → …
  • Tasks that involve LLM API integration
  • SKILL.md covers 何时激活, 核心概念, 组合 and 价格参考(2026), plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cost Aware LLM Pipeline is an agent skill from affaan-m/ECC. LLM API 使用成本优化模式 —— 基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示缓存。

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM API integration. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM API integration

Example prompts

  • “/cost-aware-llm-pipeline”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. 根据任务复杂度进行模型路由
  2. 不可变的成本跟踪
  3. 窄范围重试逻辑
  4. 提示词缓存

What it can do on your machine

Read from SKILL.md and the folder at commit 2d515e4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cost Aware LLM Pipeline loads about 1.1k tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 107 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 2d515e4, republished under its MIT licence (© affaan-m). 107 words, ~1,119 tokens.

Download SKILL.mdSave it as .claude/skills/cost-aware-llm-pipeline/SKILL.md (or your agent's skills folder).
name
cost-aware-llm-pipeline
description
LLM API 使用成本优化模式 —— 基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示缓存。
origin
ECC

成本感知型 LLM 流水线

在保持质量的同时控制 LLM API 成本的模式。将模型路由、预算跟踪、重试逻辑和提示词缓存组合成一个可组合的流水线。

何时激活

  • 构建调用 LLM API(Claude、GPT 等)的应用程序时
  • 处理具有不同复杂度的批量项目时
  • 需要将 API 支出控制在预算范围内时
  • 需要在复杂任务上优化成本而不牺牲质量时

核心概念

1. 根据任务复杂度进行模型路由

自动为简单任务选择更便宜的模型,为复杂任务保留昂贵的模型。

python
MODEL_SONNET = "claude-sonnet-5"
MODEL_HAIKU = "claude-haiku-4-5-20251001"

_SONNET_TEXT_THRESHOLD = 10_000  # chars
_SONNET_ITEM_THRESHOLD = 30     # items

def select_model(
    text_length: int,
    item_count: int,
    force_model: str | None = None,
) -> str:
    """Select model based on task complexity."""
    if force_model is not None:
        return force_model
    if text_length >= _SONNET_TEXT_THRESHOLD or item_count >= _SONNET_ITEM_THRESHOLD:
        return MODEL_SONNET  # Complex task
    return MODEL_HAIKU  # Simple task (3-4x cheaper)
2. 不可变的成本跟踪

使用冻结的数据类跟踪累计支出。每个 API 调用都会返回一个新的跟踪器 —— 永不改变状态。

python
from dataclasses import dataclass

@dataclass(frozen=True, slots=True)
class CostRecord:
    model: str
    input_tokens: int
    output_tokens: int
    cost_usd: float

@dataclass(frozen=True, slots=True)
class CostTracker:
    budget_limit: float = 1.00
    records: tuple[CostRecord, ...] = ()

    def add(self, record: CostRecord) -> "CostTracker":
        """Return new tracker with added record (never mutates self)."""
        return CostTracker(
            budget_limit=self.budget_limit,
            records=(*self.records, record),
        )

    @property
    def total_cost(self) -> float:
        return sum(r.cost_usd for r in self.records)

    @property
    def over_budget(self) -> bool:
        return self.total_cost > self.budget_limit
3. 窄范围重试逻辑

仅在暂时性错误时重试。对于认证或错误请求错误,快速失败。

python
from anthropic import (
    APIConnectionError,
    InternalServerError,
    RateLimitError,
)

_RETRYABLE_ERRORS = (APIConnectionError, RateLimitError, InternalServerError)
_MAX_RETRIES = 3

def call_with_retry(func, *, max_retries: int = _MAX_RETRIES):
    """Retry only on transient errors, fail fast on others."""
    for attempt in range(max_retries):
        try:
            return func()
        except _RETRYABLE_ERRORS:
            if attempt == max_retries - 1:
                raise
            time.sleep(2 ** attempt)  # Exponential backoff
    # AuthenticationError, BadRequestError etc. → raise immediately
4. 提示词缓存

缓存长的系统提示词,以避免在每个请求上重新发送它们。

python
messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "text",
                "text": system_prompt,
                "cache_control": {"type": "ephemeral"},  # Cache this
            },
            {
                "type": "text",
                "text": user_input,  # Variable part
            },
        ],
    }
]

组合

将所有四种技术组合到一个流水线函数中:

python
def process(text: str, config: Config, tracker: CostTracker) -> tuple[Result, CostTracker]:
    # 1. Route model
    model = select_model(len(text), estimated_items, config.force_model)

    # 2. Check budget
    if tracker.over_budget:
        raise BudgetExceededError(tracker.total_cost, tracker.budget_limit)

    # 3. Call with retry + caching
    response = call_with_retry(lambda: client.messages.create(
        model=model,
        messages=build_cached_messages(system_prompt, text),
    ))

    # 4. Track cost (immutable)
    record = CostRecord(model=model, input_tokens=..., output_tokens=..., cost_usd=...)
    tracker = tracker.add(record)

    return parse_result(response), tracker

价格参考(2026)

模型输入(美元/百万令牌)输出(美元/百万令牌)相对成本
Haiku 3.5 (legacy)$0.80$4.000.8x
Haiku 4.5$1.00$5.001x
Sonnet 5$2.00$10.002x
Sonnet 4.6$3.00$15.003x
Opus 4.8$5.00$25.005x
Fable 5 / Mythos 5$10.00$50.0010x
Opus 4.0 / 4.1 (legacy)$15.00$75.0015x

最佳实践

  • 从最便宜的模型开始,仅在达到复杂度阈值时才路由到昂贵的模型
  • 在处理批次之前设置明确的预算限制 —— 尽早失败而不是超支
  • 记录模型选择决策,以便您可以根据实际数据调整阈值
  • 对于超过 1024 个令牌的系统提示词,使用提示词缓存 —— 既能节省成本,又能降低延迟
  • 切勿在认证或验证错误时重试 —— 仅针对暂时性故障(网络、速率限制、服务器错误)重试

应避免的反模式

  • 无论复杂度如何,对所有请求都使用最昂贵的模型
  • 对所有错误都进行重试(在永久性故障上浪费预算)
  • 改变成本跟踪状态(使调试和审计变得困难)
  • 在整个代码库中硬编码模型名称(使用常量或配置)
  • 对重复的系统提示词忽略提示词缓存

适用场景

  • 任何调用 Claude、OpenAI 或类似 LLM API 的应用程序
  • 成本快速累积的批处理流水线
  • 需要智能路由的多模型架构
  • 需要预算护栏的生产系统

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/zh-CN/skills/cost-aware-llm-pipeline of affaan-m/ECC.

Open the folder on GitHubat commit 2d515e4

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Cost Aware LLM Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cost Aware LLM Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cost Aware LLM Pipeline this skillaffaan-m/ECC277k3 repos~1.1kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~833Automated safety check: PassApache-2.0
Agnes Free Textkangarooking/agnes-free-model-skills199—~630Automated safety check: PassMIT
Openai Docstheowenyoung/home1151 repos~1.4kAutomated safety check: PassApache-2.0
AIbutterbase-ai/butterbase-skills534—~1.1kAutomated safety check: PassMIT
Openagents APIOpenAgentsInc/openagents455—~430Automated safety check: PassApache-2.0

Similar skills

  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~833 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Agnes Free Text

    kangarooking/agnes-free-model-skills

    Call the free Agnes text model API for chat completions, streaming answers, coding help, tool-calling experiments, and OpenAI-compatible text generation.

    199 GitHub stars~630 tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Openai Docs

    theowenyoung/home

    A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or…

    115 GitHub starsUsed in 1 repo~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • AI

    butterbase-ai/butterbase-skills

    A skill your agent uses when calling the app's AI gateway from agent tools — chat completions, embeddings, listing models, configuring defaults or BYOK, reading token/cost usage

    534 GitHub stars~1.1k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Openagents API

    OpenAgentsInc/openagents

    Call the OpenAgents API (Open Responses and Chat Completions, many models, one key) and read its docs.

    455 GitHub stars~430 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Building Streamlit Chat UI

    iusztinpaul/designing-real-world-ai-agents-workshop

    Building chat interfaces in Streamlit. An agent skill from iusztinpaul/designing-real-world-ai-agents-workshop.

    512 GitHub stars~1.4k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Questions about Cost Aware LLM Pipeline

What does Cost Aware LLM Pipeline do?

LLM API 使用成本优化模式 —— 基于任务复杂度的模型路由、预算跟踪、重试逻辑和提示缓存. An agent skill from affaan-m/ECC. Cost Aware LLM Pipeline is an agent skill from affaan-m/ECC.

When should I use Cost Aware LLM Pipeline?

Cost Aware LLM Pipeline fits situations like: tasks that involve LLM API integration.

How do I install Cost Aware LLM Pipeline in Claude Code?

Run `npx skills add affaan-m/ECC --skill cost-aware-llm-pipeline -a claude-code`. Or copy the skill folder (docs/zh-CN/skills/cost-aware-llm-pipeline in affaan-m/ECC) into .claude/skills/cost-aware-llm-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Cost Aware LLM Pipeline in Codex?

Run `npx skills add affaan-m/ECC --skill cost-aware-llm-pipeline -a codex`. Or copy the skill folder (docs/zh-CN/skills/cost-aware-llm-pipeline in affaan-m/ECC) into .agents/skills/cost-aware-llm-pipeline in your project. Codex loads it when a task matches its description.

Can I use Cost Aware LLM Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill cost-aware-llm-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cost-aware-llm-pipeline, .gemini/skills/cost-aware-llm-pipeline, .github/skills/cost-aware-llm-pipeline and .opencode/skills/cost-aware-llm-pipeline in your project.

What does Cost Aware LLM Pipeline need to run?

SKILL.md names no scripts, command-line tools or credentials: Cost Aware LLM Pipeline is instructions for the agent only. Our summary lists: Python 3.

Does Cost Aware LLM Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cost Aware LLM Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cost Aware LLM Pipeline use?

Cost Aware LLM Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cost Aware LLM Pipeline use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cost Aware LLM Pipeline?

Skills that share tags, products or a category with Cost Aware LLM Pipeline: Dingo Verify (MigoXLab/dingo, 757 stars), Agnes Free Text (kangarooking/agnes-free-model-skills, 199 stars), Openai Docs (theowenyoung/home, 115 stars) and AI (butterbase-ai/butterbase-skills, 534 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cost Aware LLM Pipeline?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,673 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 11, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.