Agent skill

Openrouter Context Optimization

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Optimize context window usage for OpenRouter models to reduce cost and improve quality.

MITAuto-check passedAI & LLM Engineering

Install Openrouter Context Optimization

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-context-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace openrouter-context-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/openrouter-context-optimization .claude/skills/openrouter-context-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
openrouter-context-optimization
GitHub stars
2.8k
Token cost
~2.4k tokens
SKILL.md length
529 words
Files
10 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Optimize context window usage for OpenRouter models to reduce cost and improve quality.

  • Works in 6 steps: Run the Query Context Limits one-liner —… → Estimate input size (~4 characters per… → Keep long conversations inside budget… → …
  • Hitting context limits
  • SKILL.md covers Overview, Prerequisites, Instructions and Query Context Limits, plus 9 more sections
  • Calls curl and jq; reaches openrouter.ai; needs OPENROUTER_API_KEY

What it does

Openrouter Context Optimization is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize context window usage for OpenRouter models to reduce cost and improve quality. Use when hitting context limits, managing long conversations, or building RAG systems. Triggers: 'openrouter context', 'context window', 'openrouter token limit', 'reduce tokens openrouter'.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `references/context-recycling.md`, `references/context-truncation.md` and `references/efficient-message-patterns.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Model routing and gateways, Context engineering and LLM cost and token optimization. It works with OpenRouter. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Hitting context limits
  • Managing long conversations
  • Building RAG systems

Example prompts

  • “openrouter context”
  • “context window”
  • “openrouter token limit”
  • “/openrouter-context-optimization”

Requirements

  • Python 3
  • Node.js
  • A credential in OPENROUTER_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Bash(python3:*), Bash(node:*), Bash(curl:*), Bash(jq:*)

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Run the Query Context Limits one-liner — it returns context_length and prompt price per 1M tokens for each candidate model, so you know…
  2. Estimate input size (~4 characters per token, or exactly with tiktoken per the references) and pick a model with…
  3. Keep long conversations inside budget with trim_conversation() per Conversation Trimming: system prompt plus the last N messages, with a…
  4. For documents that exceed any window, use chunk_and_process() per Chunking for Large Documents — 8,000-char chunks with 500-char overlap…
  5. Mark large static blocks with cache_control: {"type": "ephemeral"} per Prompt Caching for Repeated Context to cut repeated input cost by…
  6. Monitor prompt_tokens on every response (Enterprise Considerations) to catch context bloat before it becomes a 400 context_length_exceeded.

What it can do on your machine

Read from SKILL.md and the folder at commit 80f86df. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Bash(python3:*)
    • Bash(node:*)
    • Bash(curl:*)
    • Bash(jq:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Openrouter Context Optimization loads about 2.4k tokens when it runs, and up to ~6.6k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 529 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit 80f86df, republished under its MIT licence (© jeremylongshore). 529 words, ~2,407 tokens.

Download SKILL.mdSave it as .claude/skills/openrouter-context-optimization/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
openrouter-context-optimization
description
Optimize context window usage for OpenRouter models to reduce cost and improve quality. Use when hitting context limits, managing long conversations, or building RAG systems. Triggers: 'openrouter context', 'context window', 'openrouter token limit', 'reduce tokens openrouter'.
allowed-tools
Read, Write, Edit, Grep, Bash(python3:*), Bash(node:*), Bash(curl:*), Bash(jq:*)
compatibility
Designed for Claude Code
version
1.20.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, openrouter, optimization, context-window

OpenRouter Context Optimization

Overview

OpenRouter models have varying context windows (4K to 1M+ tokens). Since pricing is per-token, stuffing unnecessary context wastes money and can degrade output quality. This skill covers context window lookup, token estimation, conversation trimming, chunking strategies, and Anthropic prompt caching for large contexts.

Prerequisites

  • An OpenRouter API key (sk-or-v1-...) exported as OPENROUTER_API_KEY — see the openrouter-install-auth skill for setup
  • Python 3.8+ with the OpenAI SDK and requests for model-metadata lookup; tiktoken for exact token counting per the references
  • curl and jq to query context windows and pricing from /api/v1/models
  • Node.js 18+ if you use the TypeScript context-budget calculator in the references

Instructions

  1. Run the Query Context Limits one-liner — it returns context_length and prompt price per 1M tokens for each candidate model, so you know the real budget before writing code.
  2. Estimate input size (~4 characters per token, or exactly with tiktoken per the references) and pick a model with select_model_for_context() from Context-Aware Model Selection — it applies an 80% safety margin and falls back through gpt-4o-mini (128K) → Claude 3.5 Sonnet (200K) → Gemini 2.0 Flash (1M).
  3. Keep long conversations inside budget with trim_conversation() per Conversation Trimming: system prompt plus the last N messages, with a trim-marker note injected where history was dropped.
  4. For documents that exceed any window, use chunk_and_process() per Chunking for Large Documents — 8,000-char chunks with 500-char overlap, analyzed independently at temperature=0 and then synthesized.
  5. Mark large static blocks with cache_control: {"type": "ephemeral"} per Prompt Caching for Repeated Context to cut repeated input cost by 90% on Anthropic models.
  6. Monitor prompt_tokens on every response (Enterprise Considerations) to catch context bloat before it becomes a 400 context_length_exceeded.

Query Context Limits

bash
# Check context window for specific models
curl -s https://openrouter.ai/api/v1/models | jq '[.data[] | select(
  .id == "anthropic/claude-3.5-sonnet" or
  .id == "openai/gpt-4o" or
  .id == "google/gemini-2.0-flash-001" or
  .id == "meta-llama/llama-3.1-70b-instruct"
) | {id, context_length, prompt_per_M: ((.pricing.prompt|tonumber)*1000000)}]'

Context-Aware Model Selection

python
import os, requests
from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
    default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)

# Cache model metadata at startup
MODELS = {m["id"]: m for m in requests.get("https://openrouter.ai/api/v1/models").json()["data"]}

def estimate_tokens(text: str) -> int:
    """Rough estimate: 1 token ~ 4 characters for English text."""
    return len(text) // 4

def select_model_for_context(messages: list, preferred: str = "anthropic/claude-3.5-sonnet") -> str:
    """Pick a model that fits the context, falling back to larger windows."""
    estimated_tokens = sum(len(m.get("content", "")) for m in messages) // 4

    FALLBACK_CHAIN = [
        ("openai/gpt-4o-mini", 128_000),
        ("anthropic/claude-3.5-sonnet", 200_000),
        ("google/gemini-2.0-flash-001", 1_000_000),
    ]

    # Try preferred model first
    preferred_ctx = MODELS.get(preferred, {}).get("context_length", 0)
    if estimated_tokens < preferred_ctx * 0.8:  # 80% safety margin
        return preferred

    for model_id, ctx in FALLBACK_CHAIN:
        if estimated_tokens < ctx * 0.8:
            return model_id

    raise ValueError(f"Content too large ({estimated_tokens} est. tokens)")

Conversation Trimming

python
def trim_conversation(
    messages: list[dict],
    max_tokens: int = 100_000,
    keep_system: bool = True,
    keep_last_n: int = 4,
) -> list[dict]:
    """Trim conversation history to fit context window.

    Strategy: Keep system prompt + last N messages.
    If still too large, reduce to last 2 messages.
    """
    system = [m for m in messages if m["role"] == "system"] if keep_system else []
    non_system = [m for m in messages if m["role"] != "system"]

    kept = non_system[-keep_last_n:]
    trimmed = non_system[:-keep_last_n] if len(non_system) > keep_last_n else []

    total_est = sum(estimate_tokens(m.get("content", "")) for m in system + kept)
    if total_est > max_tokens and keep_last_n > 2:
        kept = non_system[-2:]

    result = system + kept
    if trimmed:
        summary_note = {
            "role": "system",
            "content": f"[Previous {len(trimmed)} messages trimmed for context limits]",
        }
        result = system + [summary_note] + kept

    return result

Chunking for Large Documents

python
def chunk_and_process(document: str, question: str, model: str = "openai/gpt-4o-mini",
                      chunk_size: int = 8000, overlap: int = 500) -> str:
    """Process a large document in overlapping chunks, then synthesize."""
    chunks = []
    start = 0
    while start < len(document):
        chunks.append(document[start:start + chunk_size])
        start += chunk_size - overlap

    results = []
    for i, chunk in enumerate(chunks):
        response = client.chat.completions.create(
            model=model,
            messages=[
                {"role": "system", "content": f"Analyzing chunk {i+1}/{len(chunks)}."},
                {"role": "user", "content": f"Document:\n{chunk}\n\nQuestion: {question}"},
            ],
            max_tokens=1024, temperature=0,
        )
        results.append(response.choices[0].message.content)

    # Synthesize
    synthesis = client.chat.completions.create(
        model=model,
        messages=[
            {"role": "system", "content": "Synthesize these partial analyses."},
            {"role": "user", "content": f"Question: {question}\n\nResults:\n" + "\n---\n".join(results)},
        ],
        max_tokens=2048, temperature=0,
    )
    return synthesis.choices[0].message.content

Prompt Caching for Repeated Context

python
# Anthropic models support prompt caching -- mark large static blocks
# Subsequent requests with same cached block cost 90% less for input tokens
response = client.chat.completions.create(
    model="anthropic/claude-3.5-sonnet",
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": large_reference_document,  # 50K+ tokens
                    "cache_control": {"type": "ephemeral"},
                }
            ],
        },
        {"role": "user", "content": "Summarize section 3."},
    ],
    max_tokens=1024,
)
# First request: cache_creation_input_tokens at 1.25x rate
# Subsequent: cache_read_input_tokens at 0.1x rate (90% savings)
Show full SKILL.md (237 more words)Show less

Output

  • A jq-formatted listing of model IDs with context_length and per-1M prompt pricing from /api/v1/models
  • A model ID selected to fit the estimated token count within an 80% safety margin, or a ValueError when nothing fits
  • A trimmed message list containing the system prompt, a [Previous N messages trimmed for context limits] note, and the most recent turns
  • A single synthesized answer assembled from per-chunk analyses of an oversized document

Examples

Multi-turn chat with the references' prune_conversation() holding a 2,000-token budget — oldest messages drop as the conversation grows:

text
[Pruned] 9 -> 7 messages (1876 tokens)
Q: What about class-based decorators?...
Tokens: 412

The pruner always keeps the system message and removes the oldest non-system turns first. More worked examples: references/examples.md.

Error Handling

ErrorCauseFix
400 context_length_exceededInput + max_tokens > model limitTrim messages or use larger-context model
400 max_tokens too largemax_tokens alone exceeds limitReduce max_tokens
Slow responsesVery large contextUse streaming; consider chunking
Degraded qualityToo much irrelevant contextTrim to relevant content only

Enterprise Considerations

  • Query /api/v1/models at startup to cache context limits -- don't hardcode (they change)
  • Use max_tokens on every request to prevent runaway completion costs on large contexts
  • Implement conversation trimming as middleware so all calls respect limits
  • Use Anthropic prompt caching for RAG contexts that repeat across requests (90% input savings)
  • Route large-context tasks to cost-effective models (Gemini Flash for 1M context at low cost)
  • Monitor prompt_tokens in responses to detect context bloat before it hits limits

References

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (references) in skills/.curated/openrouter-context-optimization of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/context-recycling.md
  • references/context-truncation.md
  • references/efficient-message-patterns.md
  • references/errors.md
  • references/examples.md
  • references/model-specific-optimization.md
  • references/prompt-optimization.md
  • references/response-length-control.md
  • references/token-estimation.md

Open the folder on GitHubat commit 80f86df

Compare with similar skills

Openrouter Context Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Openrouter Context Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Openrouter Context Optimization this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.4kAutomated safety check: PassMIT
Openrouter Trending ModelsMadAppGang/claude-code285—~3.6kAutomated safety check: PassMIT
FreeRide Free Model ManagerShaivpidadi/FreeRide2372 repos~1.1kAutomated safety check: PassNone
Hyper Jevdisler/ten-levels-of-jev213—~1.7kAutomated safety check: PassMIT
Claudish UsageMadAppGang/claudish1k—~9kAutomated safety check: PassNone
Jev Model Routingkerpopule/hermes-jev-skills1.1k—~2.7kAutomated safety check: PassMIT

Similar skills

  • Openrouter Trending Models

    MadAppGang/claude-code

    Fetch trending programming models from OpenRouter rankings. An agent skill from MadAppGang/claude-code.

    285 GitHub stars~3.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • FreeRide Free Model Manager

    Shaivpidadi/FreeRide

    Configures OpenClaw to use free OpenRouter models, setting the best one as primary and adding ranked fallbacks so rate limits do not interrupt work.

    237 GitHub starsUsed in 2 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hyper Jev

    disler/ten-levels-of-jev

    Integrate and use Jev, TypeSafe AI's System One decision model, in production codebases.

    213 GitHub stars~1.7k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Claudish Usage

    MadAppGang/claudish

    CRITICAL - Guide for using Claudish CLI ONLY through sub-agents to run Claude Code with any AI model (OpenRouter, Gemini, OpenAI, local models).

    1k GitHub stars~9k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Jev Model Routing

    kerpopule/hermes-jev-skills

    Routes a turn or delegated task to the cheapest model and effort lane that will still do it right, using the Jev decision model to classify difficulty and escalate only when needed.

    1.1k GitHub stars~2.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Run Deep Swe

    sickn33/agentic-awesome-skills

    Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.

    47k GitHub starsUsed in 1 repo~1.2k tokens
    AI & LLM EngineeringAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated yesterday
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Openrouter Context Optimization

What does Openrouter Context Optimization do?

Optimize context window usage for OpenRouter models to reduce cost and improve quality. Openrouter Context Optimization is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize context window usage for OpenRouter models to reduce cost and improve quality.

When should I use Openrouter Context Optimization?

Openrouter Context Optimization fits situations like: hitting context limits; managing long conversations; building RAG systems.

How do I install Openrouter Context Optimization in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-context-optimization -a claude-code`. Or copy the skill folder (skills/.curated/openrouter-context-optimization in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/openrouter-context-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Openrouter Context Optimization in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-context-optimization -a codex`. Or copy the skill folder (skills/.curated/openrouter-context-optimization in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/openrouter-context-optimization in your project. Codex loads it when a task matches its description.

Can I use Openrouter Context Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-context-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/openrouter-context-optimization, .gemini/skills/openrouter-context-optimization, .github/skills/openrouter-context-optimization and .opencode/skills/openrouter-context-optimization in your project.

What does Openrouter Context Optimization need to run?

Going by SKILL.md and its folder, Openrouter Context Optimization needs the command-line tools its instructions call (curl and jq) and credentials named OPENROUTER_API_KEY. Our summary lists: Python 3; Node.js; A credential in OPENROUTER_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Bash(python3:*), Bash(node:*), Bash(curl:*), Bash(jq:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Openrouter Context Optimization access the network?

SKILL.md names 1 domain. In commands or code: openrouter.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Openrouter Context Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Openrouter Context Optimization use?

Openrouter Context Optimization is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Openrouter Context Optimization use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.

What are the alternatives to Openrouter Context Optimization?

Skills that share tags, products or a category with Openrouter Context Optimization: Openrouter Trending Models (MadAppGang/claude-code, 285 stars), FreeRide Free Model Manager (Shaivpidadi/FreeRide, 237 stars), Hyper Jev (disler/ten-levels-of-jev, 213 stars) and Claudish Usage (MadAppGang/claudish, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Openrouter Context Optimization?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,825 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 9, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.