Topic · AI & LLM Engineering
Best LLM cost and token optimization skills, page 4
LLM cost and token optimization skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | 145.Repo Map Ranked symbol map of a codebase within a token budget — a compact "what matters in this repo" before reading files. | AnastasiyaW/ | 154 | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 146 | Master the four operations of context engineering — Write, Select, Compress, Isolate. | rohitg00/ | 2.9k | — | ~1.6k | Automated safety check: Pass | No licence | 8 days ago |
| 147 | A skill your agent uses when the user asks to "optimize CLAUDE.md", "create a new skill", "write a custom agent", "configure hooks", "manage context window", "set up MCP servers", "scaffold a skill… | borghei/ | 881 | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 148 | Send images, audio, video, or documents into an AG2 beta Agent alongside text. | ag2ai/ | 252 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 149 | Investigate LLM spend in PostHog — total cost over time, cost by model, provider, user, trace, or custom dimension, token and cache-hit economics, and cost regressions. | PostHog/ | 40k | — | ~2.9k | Automated safety check: Pass | Unknown | today |
| 150 | Configure and optimize .geminiignore files for AI context window efficiency and token cost reduction (FinOps). | sickn33/ | 47k | 1 repo | ~1.4k | Automated safety check: Notes | MIT | today |
| 151 | Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies. | sickn33/ | 47k | 1 repo | ~2.5k | Automated safety check: Pass | MIT | today |
| 152 | 152.LLM Gateway Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking. | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | today |
| 153 | Optimizes AI agent performance by pruning redundant context, managing token usage, and enforcing ultra-concise, direct-to-value responses. | sickn33/ | 47k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 154 | 154.Zipai Optimizer Ultra-dense token optimizer skill for prompt caching, log pruning, AST-based inspection, and minified JSON payloads. | sickn33/ | 47k | 1 repo | ~1.3k | Automated safety check: Pass | MIT | today |
| 155 | A skill your agent uses for debugging DSPy programs, inspecthistory, tracing LLM calls, custom callbacks, observability, monitoring, and cost tracking. | OmidZamani/ | 124 | 1 repo | ~2.1k | Automated safety check: Warn | MIT | 3 mo ago |
| 156 | 156.Cost Tracking ローカルのコスト追跡データベースからClaude Codeのトークン使用量、支出、予算を追跡・レポートします。コスト、支出、使用量、トークン、予算、またはプロジェクト、ツール、セッション、日付によるコスト内訳について質問する場合に使用します。 | affaan-m/ | 275k | — | ~829 | Automated safety check: Pass | MIT | 3 days ago |
| 157 | 回答する前に、どれだけの回答深度を消費するかについてユーザーに情報に基づいた選択を提供する。ユーザーが回答の長さ、深さ、またはトークンバジェットを明示的に制御したい場合にこのスキルを使用する。トリガー条件:"token budget", "token count", "token usage", "token limit", "response length", "answer depth"… | affaan-m/ | 275k | — | ~910 | Automated safety check: Pass | MIT | 3 days ago |
| 158 | 在回答前,为用户提供关于消耗多少响应深度的知情选择。当用户明确希望控制响应长度、深度或令牌预算时使用此技能。触发条件:"token budget", "token count", "token usage", "token limit", "response length", "answer depth", "short version", "brief answer", "detailed… | affaan-m/ | 275k | — | ~927 | Automated safety check: Pass | MIT | 3 days ago |
| 159 | Use proactively whenever LLM API costs come up -- or should. | alirezarezvani/ | 28k | 1 repo | ~2.9k | Automated safety check: Pass | MIT | 1 mo ago |
| 160 | 160.Workflow Mastery Claude Code workflow mastery for .NET developers. An agent skill from codewithmukesh/dotnet-claude-kit. | codewithmukesh/ | 751 | — | ~3.5k | Automated safety check: Pass | MIT | 2 mo ago |
| 161 | Analyze the most expensive users in AI observability and explain why they cost so much. | PostHog/ | 40k | — | ~3.9k | Automated safety check: Pass | Unknown | today |
| 162 | 162.RAG Architect A skill your agent uses when the user asks to design a RAG pipeline, choose a chunking strategy or embedding model, pick a vector database, or evaluate retrieval quality (precision@k, recall@k, NDCG). | alirezarezvani/ | 28k | — | ~1.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 163 | OpenCode doesn't set promptcachekey for Mistral models, missing the documented 10% cached-token discount. | OnlyTerp/ | 114 | — | ~638 | Automated safety check: Pass | Unknown | 1 mo ago |
| 164 | Read accumulated cost-tracking spend + budget config, compute utilization, emit 50/75/90/100% alert ladder | ruvnet/ | 74k | — | ~645 | Automated safety check: Notes | MIT | today |
| 165 | 165.Cost Burn Burn-rate trend over time with optional drift-alert exit code. | ruvnet/ | 74k | — | ~824 | Automated safety check: Notes | MIT | today |
| 166 | 166.Cost Track Auto-capture per-session token usage from the Claude Code session jsonl and persist to the cost-tracking namespace | ruvnet/ | 74k | — | ~773 | Automated safety check: Notes | MIT | today |
| 167 | Cost optimization patterns for LLM API usage — model routing by task complexity, budget tracking, retry logic, and prompt caching. | majiayu000/ | 666 | 6 repos | ~1.4k | Automated safety check: Pass | MIT | today |
| 168 | Intercept the response flow to offer the user a choice about response depth before Claude answers Offers the user an informed choice about how much response depth to consume before answering. | aAAaqwq/ | 105 | 4 repos | ~1.5k | Automated safety check: Pass | MIT | 10 days ago |
| 169 | 169.Map Structural codebase index generator. An agent skill from SethGammon/Citadel. | SethGammon/ | 922 | — | ~1.8k | Automated safety check: Pass | MIT | 6 days ago |
| 170 | Analyze and reduce LLM spend: read usage breakdowns by call site, model, and inference profile, understand single-winner profile resolution, and pin call sites to managed profiles (Balanced /… | vellum-ai/ | 1.4k | — | ~3.9k | Automated safety check: Pass | MIT | today |
| 171 | Monitor and optimize LLM costs using Langfuse analytics and dashboards. | jeremylongshore/ | 2.8k | 1 repo | ~2.4k | Automated safety check: Pass | MIT | today |
| 172 | When agent sessions generate millions of tokens of conversation history, compression becomes mandatory. | aiskillstore/ | 430 | 5 repos | ~3.1k | Automated safety check: Pass | No licence | today |
| 173 | Calculate and display the cost of an Output SDK workflow execution run. | growthxai/ | 440 | — | ~1.4k | Automated safety check: Notes | Apache-2.0 | today |
| 174 | Guide to the providerOptions structure in .prompt files — decision tree for where an option goes, common mistakes, per-provider quick reference, and Anthropic prompt caching. | growthxai/ | 440 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 175 | 175.Claude API Anthropic Claude API patterns for Python and TypeScript. An agent skill from majiayu000/claude-skill-registry. | majiayu000/ | 666 | 3 repos | ~2.1k | Automated safety check: Pass | MIT | today |
| 176 | 176.Cv Usage A skill your agent uses when the user asks about usage analytics, statistics, token usage, or cost summary — e.g. | tombelieber/ | 110 | — | ~3.1k | Automated safety check: Pass | MIT | 7 days ago |
| 177 | Per-conversation cost view — list every session in cost-tracking with started-at, message count, top model, and total cost | ruvnet/ | 74k | — | ~407 | Automated safety check: Notes | MIT | today |
| 178 | A skill your agent uses when the user says 'token optimization', 'save tokens', 'context window', 'reduce tokens', 'token stack', or 'TokenStack', or asks about extending context window capacity. | cwinvestments/ | 423 | — | ~1.4k | Automated safety check: Pass | Proprietary | 11 days ago |
| 179 | [omh] Context window or token budget at risk: plan compact context, token/cost budgets, summarization checkpoints, and overflow recovery before long agent work. | rlaope/ | 3.2k | — | ~2k | Automated safety check: Pass | MIT | today |
| 180 | 180.Token Optimizer Reduce OpenClaw token usage and API costs through smart model routing, heartbeat optimization, budget tracking, and multi-provider fallbacks. | LeoYeAI/ | 2.2k | — | ~4.2k | Automated safety check: Pass | MIT | 2 mo ago |
| 181 | Apply Sun Tzu's Art of War to AI agent orchestration. An agent skill from LeoYeAI/openclaw-master-skills. | LeoYeAI/ | 2.2k | — | ~5k | Automated safety check: Pass | MIT | 2 mo ago |
| 182 | Builds generative AI applications on Amazon Bedrock. An agent skill from aws/agent-toolkit-for-aws. | aws/ | 2.8k | — | ~8.6k | Automated safety check: Pass | Apache-2.0 | today |
| 183 | This skill should be used when the user asks to "estimate LLM costs", "count tokens in prompts", "optimize prompt token usage", "compare model pricing", or "reduce LLM API costs". | borghei/ | 881 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 184 | Query and analyze a Dynatrace tenant's ACTUAL billing and usage data with DQL against dt.system.events — DPS consumption breakdown, cost-normalized spend ranking, included volume deduction… | Dynatrace/ | 161 | — | ~5.7k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 185 | [omh] Tracking cost, tokens, latency, or service health: prepare an operations command-board for wrapper-safe token, cost, latency, run history, queue, failure-mode, external metric-provider, and… | rlaope/ | 3.2k | — | ~2k | Automated safety check: Pass | MIT | today |
| 186 | 186.Ulw Perf [omh] Software slowness, memory leaks, or cost spikes: find where a system is actually slow, leaking, or expensive across runtime, memory, token cost, storage, rendering, inference, CI, and query… | rlaope/ | 3.2k | — | ~2.4k | Automated safety check: Pass | MIT | today |
| 187 | A Senior AI Engineer interviewer that simulates a technical interview focused on prompt engineering and LLM architecture at scale. | PrepLabsAI/ | 112 | — | ~5k | Automated safety check: Pass | MIT | yesterday |
| 188 | 188.Oc Doctor Runs a comprehensive 11-section health check on local OpenClaw installations. | LeoYeAI/ | 2.2k | — | ~4.1k | Automated safety check: Pass | MIT | 2 mo ago |
| 189 | 189.LLM Caching Implement multi-layer LLM caching with exact match, semantic similarity, and provider-side prompt caching. | majiayu000/ | 666 | 3 repos | ~2.6k | Automated safety check: Pass | MIT | today |
| 190 | 190.Prompt Caching Caching strategies for LLM prompts including Anthropic prompt caching, response caching, and CAG (Cache Augmented Generation) | majiayu000/ | 666 | 3 repos | ~3.4k | Automated safety check: Pass | MIT | today |
| 191 | A skill your agent uses when a complex task will span turns or sessions and needs structured working notes to survive context compression or handoff; short single-turn work does not trigger it. | Peiiii/ | 260 | — | ~509 | Automated safety check: Pass | MIT | today |
| 192 | Instrument an application with Sentry — detect the platform, install and initialize the SDK if needed, and wire up any signal — error monitoring, tracing/performance, logging, metrics, profiling… | getsentry/ | 268 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | today |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- LLM evaluation308
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23