Agent skill

Cost Aware LLM Pipeline

by affaan-m in affaan-m/ECC

LLM APIの使用量のコスト最適化パターン — タスクの複雑さによるモデルルーティング、予算追跡、リトライロジック、プロンプトキャッシング。

MITAuto-check passedAI & LLM Engineering

Install Cost Aware LLM Pipeline

skills CLI
$ npx skills add affaan-m/ECC --skill cost-aware-llm-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC cost-aware-llm-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/ja-JP/skills/cost-aware-llm-pipeline .claude/skills/cost-aware-llm-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cost-aware-llm-pipeline
GitHub stars
276k
Token cost
~1.2k tokens
SKILL.md length
91 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

LLM APIの使用量のコスト最適化パターン — タスクの複雑さによるモデルルーティング、予算追跡、リトライロジック、プロンプトキャッシング。

  • Works in 4 steps: タスクの複雑さによるモデルルーティング → 不変のコスト追跡 → 狭いリトライロジック → …
  • Tasks that involve LLM API integration
  • SKILL.md covers 起動条件, コアコンセプト, 合成 and 価格リファレンス(2026年), plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cost Aware LLM Pipeline is an agent skill from affaan-m/ECC. LLM APIの使用量のコスト最適化パターン — タスクの複雑さによるモデルルーティング、予算追跡、リトライロジック、プロンプトキャッシング。

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM API integration. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM API integration

Example prompts

  • “/cost-aware-llm-pipeline”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. タスクの複雑さによるモデルルーティング
  2. 不変のコスト追跡
  3. 狭いリトライロジック
  4. プロンプトキャッシング

What it can do on your machine

Read from SKILL.md and the folder at commit 4eb71d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cost Aware LLM Pipeline loads about 1.2k tokens when it runs. Until then it costs about 24 tokens; SKILL.md has 91 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~24
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 4eb71d9, republished under its MIT licence (© affaan-m). 91 words, ~1,170 tokens.

Download SKILL.mdSave it as .claude/skills/cost-aware-llm-pipeline/SKILL.md (or your agent's skills folder).
name
cost-aware-llm-pipeline
description
LLM APIの使用量のコスト最適化パターン — タスクの複雑さによるモデルルーティング、予算追跡、リトライロジック、プロンプトキャッシング。
origin
ECC

コスト認識LLMパイプライン

品質を維持しながらLLM APIのコストをコントロールするためのパターン。モデルルーティング、予算追跡、リトライロジック、プロンプトキャッシングを組み合わせた合成可能なパイプライン。

起動条件

  • LLM APIを呼び出すアプリケーションの構築(Claude、GPTなど)
  • 複雑さが異なるアイテムのバッチ処理
  • API支出の予算内に収める必要がある場合
  • 複雑なタスクの品質を犠牲にせずにコストを最適化する場合

コアコンセプト

1. タスクの複雑さによるモデルルーティング

シンプルなタスクには自動的に安価なモデルを選択し、複雑なタスクのために高価なモデルを予約します。

python
MODEL_SONNET = "claude-sonnet-5"
MODEL_HAIKU = "claude-haiku-4-5-20251001"

_SONNET_TEXT_THRESHOLD = 10_000  # 文字数
_SONNET_ITEM_THRESHOLD = 30     # アイテム数

def select_model(
    text_length: int,
    item_count: int,
    force_model: str | None = None,
) -> str:
    """タスクの複雑さに基づいてモデルを選択。"""
    if force_model is not None:
        return force_model
    if text_length >= _SONNET_TEXT_THRESHOLD or item_count >= _SONNET_ITEM_THRESHOLD:
        return MODEL_SONNET  # 複雑なタスク
    return MODEL_HAIKU  # シンプルなタスク(3〜4倍安価)
2. 不変のコスト追跡

凍結データクラスで累積支出を追跡します。各API呼び出しは新しいトラッカーを返します — 状態を変更しません。

python
from dataclasses import dataclass

@dataclass(frozen=True, slots=True)
class CostRecord:
    model: str
    input_tokens: int
    output_tokens: int
    cost_usd: float

@dataclass(frozen=True, slots=True)
class CostTracker:
    budget_limit: float = 1.00
    records: tuple[CostRecord, ...] = ()

    def add(self, record: CostRecord) -> "CostTracker":
        """追加されたレコードで新しいトラッカーを返す(selfは変更しない)。"""
        return CostTracker(
            budget_limit=self.budget_limit,
            records=(*self.records, record),
        )

    @property
    def total_cost(self) -> float:
        return sum(r.cost_usd for r in self.records)

    @property
    def over_budget(self) -> bool:
        return self.total_cost > self.budget_limit
3. 狭いリトライロジック

一時的なエラーのみリトライします。認証やリクエストエラーでは素早く失敗します。

python
from anthropic import (
    APIConnectionError,
    InternalServerError,
    RateLimitError,
)

_RETRYABLE_ERRORS = (APIConnectionError, RateLimitError, InternalServerError)
_MAX_RETRIES = 3

def call_with_retry(func, *, max_retries: int = _MAX_RETRIES):
    """一時的なエラーのみリトライし、それ以外はすぐに失敗する。"""
    for attempt in range(max_retries):
        try:
            return func()
        except _RETRYABLE_ERRORS:
            if attempt == max_retries - 1:
                raise
            time.sleep(2 ** attempt)  # 指数バックオフ
    # AuthenticationError、BadRequestErrorなど → 即座に例外発生
4. プロンプトキャッシング

長いシステムプロンプトをキャッシュして、リクエストごとに再送信しないようにします。

python
messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "text",
                "text": system_prompt,
                "cache_control": {"type": "ephemeral"},  # これをキャッシュ
            },
            {
                "type": "text",
                "text": user_input,  # 可変部分
            },
        ],
    }
]

合成

4つのテクニックすべてを単一のパイプライン関数に組み合わせます:

python
def process(text: str, config: Config, tracker: CostTracker) -> tuple[Result, CostTracker]:
    # 1. モデルをルーティング
    model = select_model(len(text), estimated_items, config.force_model)

    # 2. 予算を確認
    if tracker.over_budget:
        raise BudgetExceededError(tracker.total_cost, tracker.budget_limit)

    # 3. リトライ + キャッシングで呼び出し
    response = call_with_retry(lambda: client.messages.create(
        model=model,
        messages=build_cached_messages(system_prompt, text),
    ))

    # 4. コストを追跡(不変)
    record = CostRecord(model=model, input_tokens=..., output_tokens=..., cost_usd=...)
    tracker = tracker.add(record)

    return parse_result(response), tracker

価格リファレンス(2026年)

モデル入力($/1Mトークン)出力($/1Mトークン)相対コスト
Haiku 3.5 (legacy)$0.80$4.000.8x
Haiku 4.5$1.00$5.001x
Sonnet 5$2.00$10.002x
Sonnet 4.6$3.00$15.003x
Opus 4.8$5.00$25.005x
Fable 5 / Mythos 5$10.00$50.0010x
Opus 4.0 / 4.1 (legacy)$15.00$75.0015x

ベストプラクティス

  • 最も安価なモデルから始める、複雑さの閾値が満たされた場合にのみ高価なモデルにルーティングする
  • バッチ処理の前に明示的な予算制限を設定する — 過剰支出より早期に失敗する
  • モデル選択の決定をログに記録する、実際のデータに基づいて閾値を調整できるように
  • 1024トークンを超えるシステムプロンプトにはプロンプトキャッシングを使用する — コストとレイテンシーの両方を節約
  • 認証またはバリデーションエラーではリトライしない — 一時的な失敗のみ(ネットワーク、レート制限、サーバーエラー)

避けるべきアンチパターン

  • 複雑さに関わらずすべてのリクエストに最も高価なモデルを使用すること
  • すべてのエラーでリトライすること(永続的な失敗で予算を無駄にする)
  • コスト追跡の状態を変更すること(デバッグと監査が困難になる)
  • コードベース全体にモデル名をハードコードすること(定数または設定を使用する)
  • 繰り返しのシステムプロンプトでプロンプトキャッシングを無視すること

使用すべき場合

  • Claude、OpenAI、または同様のLLM APIを呼び出すすべてのアプリケーション
  • コストが積み上がるバッチ処理パイプライン
  • インテリジェントルーティングが必要なマルチモデルアーキテクチャ
  • 予算ガードレールが必要な本番システム

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/ja-JP/skills/cost-aware-llm-pipeline of affaan-m/ECC.

Open the folder on GitHubat commit 4eb71d9

Compare with similar skills

Cost Aware LLM Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cost Aware LLM Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cost Aware LLM Pipeline this skillaffaan-m/ECC276k—~1.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~833Automated safety check: PassApache-2.0
Agnes Free Textkangarooking/agnes-free-model-skills199—~630Automated safety check: PassMIT
Openai Docstheowenyoung/home1151 repos~1.4kAutomated safety check: PassApache-2.0
AIbutterbase-ai/butterbase-skills534—~1.1kAutomated safety check: PassMIT
Openagents APIOpenAgentsInc/openagents455—~430Automated safety check: PassApache-2.0

Similar skills

  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~833 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Agnes Free Text

    kangarooking/agnes-free-model-skills

    Call the free Agnes text model API for chat completions, streaming answers, coding help, tool-calling experiments, and OpenAI-compatible text generation.

    199 GitHub stars~630 tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Openai Docs

    theowenyoung/home

    A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or…

    115 GitHub starsUsed in 1 repo~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • AI

    butterbase-ai/butterbase-skills

    A skill your agent uses when calling the app's AI gateway from agent tools — chat completions, embeddings, listing models, configuring defaults or BYOK, reading token/cost usage

    534 GitHub stars~1.1k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Openagents API

    OpenAgentsInc/openagents

    Call the OpenAgents API (Open Responses and Chat Completions, many models, one key) and read its docs.

    455 GitHub stars~430 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Building Streamlit Chat UI

    iusztinpaul/designing-real-world-ai-agents-workshop

    Building chat interfaces in Streamlit. An agent skill from iusztinpaul/designing-real-world-ai-agents-workshop.

    512 GitHub stars~1.4k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Questions about Cost Aware LLM Pipeline

What does Cost Aware LLM Pipeline do?

LLM APIの使用量のコスト最適化パターン — タスクの複雑さによるモデルルーティング、予算追跡、リトライロジック、プロンプトキャッシング。. Cost Aware LLM Pipeline is an agent skill from affaan-m/ECC.

When should I use Cost Aware LLM Pipeline?

Cost Aware LLM Pipeline fits situations like: tasks that involve LLM API integration.

How do I install Cost Aware LLM Pipeline in Claude Code?

Run `npx skills add affaan-m/ECC --skill cost-aware-llm-pipeline -a claude-code`. Or copy the skill folder (docs/ja-JP/skills/cost-aware-llm-pipeline in affaan-m/ECC) into .claude/skills/cost-aware-llm-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Cost Aware LLM Pipeline in Codex?

Run `npx skills add affaan-m/ECC --skill cost-aware-llm-pipeline -a codex`. Or copy the skill folder (docs/ja-JP/skills/cost-aware-llm-pipeline in affaan-m/ECC) into .agents/skills/cost-aware-llm-pipeline in your project. Codex loads it when a task matches its description.

Can I use Cost Aware LLM Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill cost-aware-llm-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cost-aware-llm-pipeline, .gemini/skills/cost-aware-llm-pipeline, .github/skills/cost-aware-llm-pipeline and .opencode/skills/cost-aware-llm-pipeline in your project.

What does Cost Aware LLM Pipeline need to run?

SKILL.md names no scripts, command-line tools or credentials: Cost Aware LLM Pipeline is instructions for the agent only. Our summary lists: Python 3.

Does Cost Aware LLM Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cost Aware LLM Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cost Aware LLM Pipeline use?

Cost Aware LLM Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cost Aware LLM Pipeline use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cost Aware LLM Pipeline?

Skills that share tags, products or a category with Cost Aware LLM Pipeline: Dingo Verify (MigoXLab/dingo, 757 stars), Agnes Free Text (kangarooking/agnes-free-model-skills, 199 stars), Openai Docs (theowenyoung/home, 115 stars) and AI (butterbase-ai/butterbase-skills, 534 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cost Aware LLM Pipeline?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,111 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 10, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.