Agent skill

Openrouter Performance Tuning

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Optimize OpenRouter request latency and throughput. An agent skill from jeremylongshore/tons-of-skills-marketplace.

MITAuto-check passedAI & LLM Engineering

Install Openrouter Performance Tuning

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-performance-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace openrouter-performance-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/openrouter-performance-tuning .claude/skills/openrouter-performance-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
openrouter-performance-tuning
GitHub stars
2.8k
Token cost
~2.4k tokens
SKILL.md length
582 words
Files
11 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Optimize OpenRouter request latency and throughput. An agent skill from jeremylongshore/tons-of-skills-marketplace.

  • Works in 6 steps: Establish a baseline: run… → Check the results against the Model… → Switch user-facing paths to… → …
  • Building real-time applications
  • SKILL.md covers Overview, Prerequisites, Instructions and Benchmark Latency, plus 10 more sections
  • Reaches openrouter.ai; needs OPENROUTER_API_KEY

What it does

Openrouter Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize OpenRouter request latency and throughput. Use when building real-time applications, reducing TTFT, or scaling request volume. Triggers: 'openrouter performance', 'openrouter latency', 'openrouter speed', 'optimize openrouter throughput'.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `references/batch-processing.md`, `references/best-practices-summary.md` and `references/caching-strategies.md`). Compatibility notes: Designed for Claude Code

It sits in AI & LLM Engineering, covering Model routing and gateways. It works with OpenRouter and OpenAI. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Building real-time applications
  • Scaling request volume

Example prompts

  • “openrouter performance”
  • “openrouter latency”
  • “openrouter speed”
  • “/openrouter-performance-tuning”

Requirements

  • Python 3
  • A credential in OPENROUTER_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Bash(python3:*)

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Establish a baseline: run benchmark_model() from Benchmark Latency against your candidate models (e.g. openai/gpt-4o-mini vs…
  2. Check the results against the Model Speed Tiers table to confirm each candidate sits in the right tier for your latency budget (200-500ms…
  3. Switch user-facing paths to stream_completion() per Streaming for Lower TTFT and verify ttft_ms drops (typically 2-10x).
  4. Move batch workloads to parallel_completions() per Parallel Request Processing, capping concurrency with asyncio.Semaphore…
  5. Apply Connection Optimization — one shared client with timeout=30.0 and max_retries=2 instead of a new client per request.
  6. Work through the Performance Optimization Checklist (set max_tokens, shrink prompts, consider :nitro variants and provider routing), then…

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Bash(python3:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Openrouter Performance Tuning loads about 2.4k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 582 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 582 words, ~2,429 tokens.

Download SKILL.mdSave it as .claude/skills/openrouter-performance-tuning/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
openrouter-performance-tuning
description
Optimize OpenRouter request latency and throughput. Use when building real-time applications, reducing TTFT, or scaling request volume. Triggers: 'openrouter performance', 'openrouter latency', 'openrouter speed', 'optimize openrouter throughput'.
allowed-tools
Read, Write, Edit, Grep, Bash(python3:*)
compatibility
Designed for Claude Code
version
1.20.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, openrouter, performance, latency, optimization

OpenRouter Performance Tuning

Overview

OpenRouter adds minimal overhead (~50-100ms) to direct provider calls. Most latency comes from the upstream model. Key levers: model selection (smaller = faster), streaming (lower TTFT), parallel requests, prompt size reduction, and provider routing to faster infrastructure. This skill covers benchmarking, streaming optimization, concurrent processing, and connection tuning.

Prerequisites

  • An OpenRouter API key (sk-or-v1-...) exported as OPENROUTER_API_KEY — see the openrouter-install-auth skill for setup
  • Python 3.8+ with the OpenAI SDK (openai package) — the examples use both the sync OpenAI client and AsyncOpenAI for parallel processing
  • Credits on the key if you benchmark paid models like anthropic/claude-3.5-sonnet; a :free model is enough to validate the benchmark harness itself
  • HTTP-Referer / X-Title header values for your app (set in every client constructor here)

Instructions

  1. Establish a baseline: run benchmark_model() from Benchmark Latency against your candidate models (e.g. openai/gpt-4o-mini vs anthropic/claude-3.5-sonnet) and record p50/p95.
  2. Check the results against the Model Speed Tiers table to confirm each candidate sits in the right tier for your latency budget (200-500ms TTFT fastest tier; 5-30s for reasoning models).
  3. Switch user-facing paths to stream_completion() per Streaming for Lower TTFT and verify ttft_ms drops (typically 2-10x).
  4. Move batch workloads to parallel_completions() per Parallel Request Processing, capping concurrency with asyncio.Semaphore (max_concurrent=5-10).
  5. Apply Connection Optimization — one shared client with timeout=30.0 and max_retries=2 instead of a new client per request.
  6. Work through the Performance Optimization Checklist (set max_tokens, shrink prompts, consider :nitro variants and provider routing), then re-run the benchmark to quantify each change.

Benchmark Latency

python
import os, time, statistics
from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
    default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)

def benchmark_model(model: str, prompt: str = "Say hello", n: int = 5) -> dict:
    """Benchmark a model's latency over N requests."""
    latencies = []
    for _ in range(n):
        start = time.monotonic()
        response = client.chat.completions.create(
            model=model,
            messages=[{"role": "user", "content": prompt}],
            max_tokens=50,
        )
        latencies.append((time.monotonic() - start) * 1000)

    return {
        "model": model,
        "p50_ms": round(statistics.median(latencies)),
        "p95_ms": round(sorted(latencies)[int(len(latencies) * 0.95)]),
        "avg_ms": round(statistics.mean(latencies)),
        "min_ms": round(min(latencies)),
        "max_ms": round(max(latencies)),
    }

# Compare fast vs slow models
for model in ["openai/gpt-4o-mini", "anthropic/claude-3-haiku", "anthropic/claude-3.5-sonnet"]:
    result = benchmark_model(model)
    print(f"{result['model']}: p50={result['p50_ms']}ms p95={result['p95_ms']}ms")

Streaming for Lower TTFT

python
def stream_completion(messages, model="openai/gpt-4o-mini", **kwargs):
    """Stream response for lower time-to-first-token."""
    start = time.monotonic()
    first_token_time = None
    full_content = []

    stream = client.chat.completions.create(
        model=model, messages=messages, stream=True,
        stream_options={"include_usage": True},  # Get token counts at end
        **kwargs,
    )

    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            if first_token_time is None:
                first_token_time = (time.monotonic() - start) * 1000
            full_content.append(chunk.choices[0].delta.content)

    total_time = (time.monotonic() - start) * 1000
    return {
        "content": "".join(full_content),
        "ttft_ms": round(first_token_time or 0),
        "total_ms": round(total_time),
    }

Parallel Request Processing

python
import asyncio
from openai import AsyncOpenAI

async def parallel_completions(prompts: list[str], model="openai/gpt-4o-mini",
                                max_concurrent=10, **kwargs):
    """Process multiple prompts concurrently."""
    semaphore = asyncio.Semaphore(max_concurrent)
    client = AsyncOpenAI(
        base_url="https://openrouter.ai/api/v1",
        api_key=os.environ["OPENROUTER_API_KEY"],
        default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
    )

    async def process(prompt):
        async with semaphore:
            response = await client.chat.completions.create(
                model=model,
                messages=[{"role": "user", "content": prompt}],
                **kwargs,
            )
            return response.choices[0].message.content

    return await asyncio.gather(*[process(p) for p in prompts])

# 10 requests in parallel instead of sequential
results = asyncio.run(parallel_completions(
    ["Summarize: " + text for text in documents],
    max_concurrent=5,
    max_tokens=200,
))

Performance Optimization Checklist

OptimizationImpactEffort
Use streamingTTFT drops 2-10xLow
Use smaller models for simple tasks2-5x fasterLow
Reduce prompt sizeProportional to reductionMedium
Set max_tokensCaps response timeLow
Parallel requestsN requests in ~1 request timeMedium
Use :nitro variantFaster inference (where available)Low
Provider routing to fastest10-30% latency reductionLow
Connection keep-aliveSaves TCP/TLS handshakeLow

Model Speed Tiers

SpeedModelsTypical TTFT
Fastestopenai/gpt-4o-mini, anthropic/claude-3-haiku200-500ms
Fastopenai/gpt-4o, google/gemini-2.0-flash-001500ms-1s
Standardanthropic/claude-3.5-sonnet1-3s
Slowopenai/o1, reasoning models5-30s

Connection Optimization

text
# Reuse client instance (connection pooling)
# BAD: creating new client per request
for prompt in prompts:
    c = OpenAI(base_url="https://openrouter.ai/api/v1", ...)  # New TCP connection each time
    c.chat.completions.create(...)

# GOOD: reuse single client
client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
    timeout=30.0,           # Set appropriate timeout
    max_retries=2,          # Built-in retry with backoff
    default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)
for prompt in prompts:
    client.chat.completions.create(...)  # Reuses HTTP connection
Show full SKILL.md (234 more words)Show less

Output

  • A latency benchmark table per model from benchmark_model(): p50_ms, p95_ms, avg_ms, min_ms, max_ms over N sample requests
  • Streaming metrics from stream_completion(): the full content plus ttft_ms and total_ms for each request
  • A list of completions from parallel_completions() produced in roughly one request's wall-clock time instead of N sequential round-trips
  • A prioritized tuning plan drawn from the Performance Optimization Checklist (lever, expected impact, effort)

Examples

Benchmark two fastest-tier candidates before committing to one:

python
for model in ["openai/gpt-4o-mini", "anthropic/claude-3-haiku"]:
    r = benchmark_model(model, n=5)
    print(f"{r['model']}: p50={r['p50_ms']}ms p95={r['p95_ms']}ms avg={r['avg_ms']}ms")
# openai/gpt-4o-mini: p50=430ms p95=610ms avg=455ms
# anthropic/claude-3-haiku: p50=395ms p95=580ms avg=418ms

Both land in the fastest tier (200-500ms typical TTFT), so choose on cost or quality — then stream_completion() cuts perceived latency further for user-facing paths. More worked examples: references/examples.md.

Error Handling

ErrorCauseFix
High TTFT (>5s)Model cold-starting or overloadedSwitch to :nitro variant or different provider
Timeout errorsmax_tokens too high or model too slowReduce max_tokens; use streaming; increase timeout
Throughput bottleneckSequential processingUse async + semaphore for concurrent requests
Inconsistent latencyProvider load variesUse provider.order to pin to fastest provider

Enterprise Considerations

  • Benchmark models in your infrastructure, not just locally -- network path matters
  • Use streaming for all user-facing requests to minimize perceived latency
  • Set max_tokens on every request to bound response time and cost
  • Reuse client instances to benefit from HTTP connection pooling
  • Use asyncio.Semaphore to control concurrency and avoid overwhelming the API
  • Monitor P95 latency, not just average -- tail latencies indicate provider issues
  • Consider :nitro model variants for latency-critical paths

References

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/.curated/openrouter-performance-tuning of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/batch-processing.md
  • references/best-practices-summary.md
  • references/caching-strategies.md
  • references/configuration-tuning.md
  • references/errors.md
  • references/examples.md
  • references/latency-optimization.md
  • references/model-selection-for-speed.md
  • references/monitoring-&-profiling.md
  • references/request-optimization.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Openrouter Performance Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Openrouter Performance Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Openrouter Performance Tuning this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.4kAutomated safety check: PassMIT
Embeddings via 9Routerdecolua/9router31k—~604Automated safety check: PassMIT
Using Ccproxy Inspectorstarbaser/ccproxy350—~2.7kAutomated safety check: PassCustom licence
Mecatl Model Router Configstacklok/mecatl254—~2.7kAutomated safety check: PassApache-2.0
Using Ccproxy APIstarbaser/ccproxy350—~4kAutomated safety check: PassCustom licence
Configuring Visionoxbshw/watch-skill470—~509Automated safety check: NotesMIT

Similar skills

  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    31k GitHub stars~604 tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy Inspector

    starbaser/ccproxy

    Operates the ccproxy inspector MITM system for intercepting, inspecting, and transforming LLM API traffic.

    350 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Interviews you about provider, cost, openness and image needs, then designs the models section of a mecatl settings file with aliases, slots and router categories.

    254 GitHub stars~2.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Using Ccproxy API

    starbaser/ccproxy

    Guides users through ccproxy as an OpenAI-compatible and Anthropic-compatible LLM API server with SDK integration, OAuth authentication, sentinel key substitution, model routing, and troubleshooting.

    350 GitHub stars~4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Configuring Vision

    oxbshw/watch-skill

    The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models.

    470 GitHub stars~509 tokensUpdated 26 days ago
    AI & LLM EngineeringAuto-check: notes
  • Quorum

    Detrol/quorum-cli

    Run a structured debate between agent CLIs (claude, codex, agy, grok) and the user's configured API or local models (OpenAI, Anthropic, Google, xAI, OpenRouter, Ollama and more) through the Quorum…

    119 GitHub stars~807 tokensUpdated 14 days ago
    AI & LLM EngineeringAuto-check: notes

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Openrouter Performance Tuning

What does Openrouter Performance Tuning do?

Optimize OpenRouter request latency and throughput. An agent skill from jeremylongshore/tons-of-skills-marketplace. Openrouter Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize OpenRouter request latency and throughput.

When should I use Openrouter Performance Tuning?

Openrouter Performance Tuning fits situations like: building real-time applications; scaling request volume.

How do I install Openrouter Performance Tuning in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-performance-tuning -a claude-code`. Or copy the skill folder (skills/.curated/openrouter-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/openrouter-performance-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Openrouter Performance Tuning in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-performance-tuning -a codex`. Or copy the skill folder (skills/.curated/openrouter-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/openrouter-performance-tuning in your project. Codex loads it when a task matches its description.

Can I use Openrouter Performance Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill openrouter-performance-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/openrouter-performance-tuning, .gemini/skills/openrouter-performance-tuning, .github/skills/openrouter-performance-tuning and .opencode/skills/openrouter-performance-tuning in your project.

What does Openrouter Performance Tuning need to run?

Going by SKILL.md and its folder, Openrouter Performance Tuning needs credentials named OPENROUTER_API_KEY. Our summary lists: Python 3; A credential in OPENROUTER_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Bash(python3:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Openrouter Performance Tuning access the network?

SKILL.md names 1 domain. In commands or code: openrouter.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Openrouter Performance Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Openrouter Performance Tuning use?

Openrouter Performance Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Openrouter Performance Tuning use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Openrouter Performance Tuning?

Skills that share tags, products or a category with Openrouter Performance Tuning: Embeddings via 9Router (decolua/9router, 31k stars), Using Ccproxy Inspector (starbaser/ccproxy, 350 stars), Mecatl Model Router Config (stacklok/mecatl, 254 stars) and Using Ccproxy API (starbaser/ccproxy, 350 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Openrouter Performance Tuning?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.