Agent skill

LLM Gateway

by sickn33 in sickn33/agentic-awesome-skills

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

MITAuto-check passedBackend & APIs

Install LLM Gateway

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill llm-gateway -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills llm-gateway --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-gateway .claude/skills/llm-gateway && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-gateway
GitHub stars
47k
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
310 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

  • Tasks that involve Model routing and gateways
  • SKILL.md covers When to Use This Skill, Prerequisites, LiteLLM Proxy — Quick Start and LiteLLM Configuration, plus 8 more sections
  • Calls curl and docker; needs OPENAI_API_KEY and LITELLM_MASTER_KEY
  • Tasks that involve Rate limiting

What it does

LLM Gateway is an agent skill from sickn33/agentic-awesome-skills. Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

It sits in Backend & APIs, covering Model routing and gateways, Rate limiting and Cloud networking. It works with Docker, vLLM and OpenAI. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Model routing and gateways
  • Tasks that involve Rate limiting
  • Tasks that involve Cloud networking

Example prompts

  • “/llm-gateway”

Requirements

  • Docker
  • A credential in OPENAI_API_KEY
  • A credential in ANTHROPIC_API_KEY
  • Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • LITELLM_MASTER_KEY
    • ANTHROPIC_API_KEY
    • LANGFUSE_PUBLIC_KEY
    • LANGFUSE_SECRET_KEY
    • POSTGRES_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

LLM Gateway loads about 2.1k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 310 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 310 words, ~2,109 tokens.

Download SKILL.mdSave it as .claude/skills/llm-gateway/SKILL.md (or your agent's skills folder).
name
llm-gateway
description
Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.
compatibility
Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

LLM Gateway

A unified API gateway that routes LLM requests across providers and self-hosted models — with rate limiting, cost tracking, caching, and failover.

When to Use This Skill

Use this skill when:

  • Running multiple LLM backends (OpenAI, Anthropic, vLLM, Ollama) behind a single endpoint
  • Enforcing per-team or per-user rate limits and spend budgets
  • Implementing automatic fallback when a provider is down
  • Adding semantic caching to reduce API costs by 20–50%
  • Centralizing API key management instead of distributing keys to every app

Prerequisites

  • Docker and Docker Compose
  • A PostgreSQL or SQLite database (for LiteLLM state)
  • LLM API keys (OpenAI, Anthropic, etc.) or self-hosted vLLM endpoints
  • Optional: Redis for caching and rate limiting

LiteLLM Proxy — Quick Start

LiteLLM is the de facto open-source LLM gateway with OpenAI-compatible API.

bash
# Run with Docker
docker run -d \
  --name litellm-proxy \
  -p 4000:4000 \
  -e OPENAI_API_KEY=$OPENAI_API_KEY \
  -e ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \
  -v $(pwd)/litellm-config.yaml:/app/config.yaml \
  ghcr.io/berriai/litellm:main-latest \
  --config /app/config.yaml \
  --detailed_debug

LiteLLM Configuration

yaml
# litellm-config.yaml
model_list:
  # OpenAI models
  - model_name: gpt-4o
    litellm_params:
      model: openai/gpt-4o
      api_key: os.environ/OPENAI_API_KEY
      rpm: 10000
      tpm: 2000000

  - model_name: gpt-4o-mini
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: os.environ/OPENAI_API_KEY

  # Anthropic
  - model_name: claude-sonnet-4-6
    litellm_params:
      model: anthropic/claude-sonnet-4-6
      api_key: os.environ/ANTHROPIC_API_KEY

  # Self-hosted vLLM instances (load balanced)
  - model_name: llama-3.1-8b
    litellm_params:
      model: openai/meta-llama/Llama-3.1-8B-Instruct
      api_base: http://vllm-1:8000/v1
      api_key: fake                    # vLLM key
  - model_name: llama-3.1-8b
    litellm_params:
      model: openai/meta-llama/Llama-3.1-8B-Instruct
      api_base: http://vllm-2:8000/v1  # second replica — auto load balanced
      api_key: fake

  # Fallback: cheap model if primary fails
  - model_name: gpt-4o
    litellm_params:
      model: openai/gpt-4o-mini        # fallback to cheaper model
      api_key: os.environ/OPENAI_API_KEY

router_settings:
  routing_strategy: least-busy         # or: latency-based, simple-shuffle
  num_retries: 3
  retry_after: 5
  allowed_fails: 2
  cooldown_time: 60

  # Fallback configuration
  fallbacks:
    - gpt-4o: [claude-sonnet-4-6]
    - claude-sonnet-4-6: [gpt-4o]

litellm_settings:
  # Semantic caching
  cache: true
  cache_params:
    type: redis
    host: redis
    port: 6379
    similarity_threshold: 0.90        # cache if >90% semantic similarity

  # Logging
  success_callback: ["langfuse"]
  failure_callback: ["langfuse"]
  langfuse_public_key: os.environ/LANGFUSE_PUBLIC_KEY
  langfuse_secret_key: os.environ/LANGFUSE_SECRET_KEY

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: postgresql://litellm:password@postgres:5432/litellm
  store_model_in_db: true

Docker Compose: Full Gateway Stack

yaml
services:
  litellm:
    image: ghcr.io/berriai/litellm:main-latest
    command: ["--config", "/app/config.yaml", "--port", "4000"]
    volumes:
      - ./litellm-config.yaml:/app/config.yaml
    ports:
      - "4000:4000"
    environment:
      - OPENAI_API_KEY=${OPENAI_API_KEY}
      - ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY}
      - LITELLM_MASTER_KEY=${LITELLM_MASTER_KEY}
      - DATABASE_URL=postgresql://litellm:password@postgres:5432/litellm
    depends_on:
      postgres:
        condition: service_healthy
      redis:
        condition: service_started
    restart: unless-stopped

  postgres:
    image: postgres:16-alpine
    environment:
      POSTGRES_DB: litellm
      POSTGRES_USER: litellm
      POSTGRES_PASSWORD: password
    volumes:
      - postgres-data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U litellm"]
      interval: 5s
      retries: 5
    restart: unless-stopped

  redis:
    image: redis:7-alpine
    command: redis-server --maxmemory 2gb --maxmemory-policy allkeys-lru
    volumes:
      - redis-data:/data
    restart: unless-stopped

volumes:
  postgres-data:
  redis-data:

Virtual Keys & Rate Limiting

bash
# Create a virtual API key for a team (via LiteLLM API)
curl -X POST http://localhost:4000/key/generate \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "team_id": "team-backend",
    "key_alias": "backend-team-key",
    "models": ["gpt-4o-mini", "llama-3.1-8b"],
    "max_budget": 100,              # USD limit
    "budget_duration": "monthly",
    "rpm_limit": 100,               # requests per minute
    "tpm_limit": 500000             # tokens per minute
  }'

# View spend
curl http://localhost:4000/spend/keys \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY"

Nginx Load Balancer (Alternative/Complement)

nginx
# nginx.conf — round-robin across vLLM replicas
upstream vllm_backends {
    least_conn;
    server vllm-1:8000 max_fails=3 fail_timeout=30s;
    server vllm-2:8000 max_fails=3 fail_timeout=30s;
    server vllm-3:8000 max_fails=3 fail_timeout=30s;
    keepalive 32;
}

server {
    listen 80;
    server_name llm-api.internal;

    # Rate limiting
    limit_req_zone $http_authorization zone=per_key:10m rate=100r/m;
    limit_req zone=per_key burst=20 nodelay;

    location /v1/ {
        proxy_pass http://vllm_backends;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        proxy_set_header Host $host;
        proxy_read_timeout 300s;        # long timeout for streaming
        proxy_buffering off;            # required for SSE streaming
        proxy_cache_bypass 1;
    }
}

Monitoring Gateway Health

bash
# Check LiteLLM health
curl http://localhost:4000/health

# Model-level health
curl http://localhost:4000/health/liveliness

# Spend by model
curl http://localhost:4000/spend/models \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY"

# Active virtual keys
curl http://localhost:4000/key/list \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY"

Common Issues

IssueCauseFix
ConnectionRefusedError to backendBackend not reachableCheck api_base URL; verify backend is healthy
Rate limit errors (429)Budget/RPM exceededIncrease limits or rotate to fallback model
Slow streaming responsesproxy_buffering enabledSet proxy_buffering off in Nginx
Cache miss rate highThreshold too strictLower similarity_threshold to 0.85
Postgres connection errorsDB not readyAdd depends_on with condition: service_healthy

Best Practices

  • Use virtual keys per team/app — never expose raw provider API keys.
  • Enable cache: true with Redis for repeated or similar queries; can cut costs 30–50%.
  • Set num_retries: 3 with fallbacks to handle provider outages gracefully.
  • Log all requests to Langfuse or OpenTelemetry for cost attribution and debugging.
  • Use least-busy routing strategy for self-hosted models to avoid GPU saturation.
  • vllm-server (vllm-server) - Backend inference server
  • llm-inference-scaling (llm-inference-scaling) - Auto-scaling backends
  • llm-caching (llm-caching) - Semantic cache patterns
  • llm-cost-optimization (llm-cost-optimization) - Cost management

Limitations

  • Infrastructure commands can disrupt services: confirm target host/scope and have backups/snapshots before mutating state.
  • Docs-only import: upstream scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/llm-gateway of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LLM Gateway next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Gateway compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Gateway this skillsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
LLM GatewayBagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT
Rtvi Vlm Customize ModelNVIDIA/skills3.6k—~5kAutomated safety check: NotesApache-2.0
Deepstream SopNVIDIA/skills3.6k—~4.7kAutomated safety check: NotesApache-2.0
System Designninehills/skills280—~4.7kAutomated safety check: PassMIT
System Designwondelai/skills2.4k—~4kAutomated safety check: PassMIT

Similar skills

  • LLM Gateway

    BagelHole/DevOps-Security-Agent-Skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    1.2k GitHub stars~2k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Backend & APIsAuto-check: notes
  • Deepstream Sop

    NVIDIA/skills

    Official

    A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Backend & APIsAuto-check: notes
  • System Design

    ninehills/skills

    Design scalable distributed systems using structured approaches for load balancing, caching, database scaling, and message queues.

    280 GitHub stars~4.7k tokensUpdated 3 mo ago
    Backend & APIsAuto-check passed
  • System Design

    wondelai/skills

    Design scalable distributed systems using structured approaches for load balancing, caching, database scaling, and message queues.

    2.4k GitHub stars~4k tokensUpdated 29 days ago
    Backend & APIsAuto-check passed
  • Openrouter Rate Limits

    jeremylongshore/tons-of-skills-marketplace

    Understand and handle OpenRouter rate limits. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2.4k tokensUpdated today
    Backend & APIsAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about LLM Gateway

What does LLM Gateway do?

Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking. LLM Gateway is an agent skill from sickn33/agentic-awesome-skills. Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

When should I use LLM Gateway?

LLM Gateway fits situations like: tasks that involve Model routing and gateways; tasks that involve Rate limiting; tasks that involve Cloud networking.

How do I install LLM Gateway in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill llm-gateway -a claude-code`. Or copy the skill folder (skills/llm-gateway in sickn33/agentic-awesome-skills) into .claude/skills/llm-gateway in your project. Claude Code loads it when a task matches its description.

How do I install LLM Gateway in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill llm-gateway -a codex`. Or copy the skill folder (skills/llm-gateway in sickn33/agentic-awesome-skills) into .agents/skills/llm-gateway in your project. Codex loads it when a task matches its description.

Can I use LLM Gateway in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill llm-gateway -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-gateway, .gemini/skills/llm-gateway, .github/skills/llm-gateway and .opencode/skills/llm-gateway in your project.

What does LLM Gateway need to run?

Going by SKILL.md and its folder, LLM Gateway needs the command-line tools its instructions call (curl and docker) and credentials named OPENAI_API_KEY, LITELLM_MASTER_KEY, ANTHROPIC_API_KEY and LANGFUSE_PUBLIC_KEY. Our summary lists: Docker; A credential in OPENAI_API_KEY; A credential in ANTHROPIC_API_KEY. Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled..

Does LLM Gateway access the network?

SKILL.md contains no URLs. Its commands use curl and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is LLM Gateway safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Gateway use?

LLM Gateway is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Gateway use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Gateway?

Skills that share tags, products or a category with LLM Gateway: LLM Gateway (BagelHole/DevOps-Security-Agent-Skills, 1.2k stars), Rtvi Vlm Customize Model (NVIDIA/skills, 3.6k stars), Deepstream Sop (NVIDIA/skills, 3.6k stars) and System Design (ninehills/skills, 280 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Gateway?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.