Agent skill

Routing Architecture

by majiayu000 in majiayu000/litellm-rs

LiteLLM-RS Routing Architecture. An agent skill from majiayu000/litellm-rs.

MITAuto-check passedDevOps & Cloud

Install Routing Architecture

skills CLI
$ npx skills add majiayu000/litellm-rs --skill routing-architecture -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/litellm-rs routing-architecture --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/litellm-rs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/routing-architecture .claude/skills/routing-architecture && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
routing-architecture
GitHub stars
118
Token cost
~1.9k tokens
SKILL.md length
638 words
Files
4
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

LiteLLM-RS Routing Architecture. An agent skill from majiayu000/litellm-rs.

  • Works in 7 steps: SimpleShuffle (runtime default) → RoundRobin (gateway YAML default) → LeastBusy → …
  • Tuning a routing strategy
  • SKILL.md covers Overview, Selection Flow, Routing Strategies and References
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Routing Architecture is an agent skill from majiayu000/litellm-rs. LiteLLM-RS Routing Architecture. Covers 7 routing strategies over immutable routing snapshots (ArcSwap) with atomic/DashMap state, health-aware deployment selection with cooldown circuit breaker, model-keyed fallback chains, and load balancing. Use when selecting or tuning a routing strategy, or configuring failover, health checks, and load balancing.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `reference/health-and-fallbacks.md`, `reference/performance-and-practices.md` and `reference/router-configuration.md`).

It sits in DevOps & Cloud, covering Cloud networking, Model routing and gateways and Backup and disaster recovery. It works with OpenAI and Rust. The repository describes itself as: Self-hosted Rust LLM gateway with OpenAI-compatible APIs, load balancing, failover, and a reusable Rust kernel. The licence is MIT.

When your agent uses it

  • Tuning a routing strategy
  • Configuring failover

Example prompts

  • “/routing-architecture”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. SimpleShuffle (runtime default)
  2. RoundRobin (gateway YAML default)
  3. LeastBusy
  4. LatencyBased
  5. PriorityBased
  6. UsageBased
  7. RateLimitAware

What it can do on your machine

Read from SKILL.md and the folder at commit ed3f4d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are rust).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Routing Architecture loads about 1.9k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 638 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/litellm-rs at commit ed3f4d9, republished under its MIT licence (© majiayu000). 638 words, ~1,853 tokens.

Download SKILL.mdSave it as .claude/skills/routing-architecture/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
routing-architecture
description
LiteLLM-RS Routing Architecture. Covers 7 routing strategies over immutable routing snapshots (ArcSwap) with atomic/DashMap state, health-aware deployment selection with cooldown circuit breaker, model-keyed fallback chains, and load balancing. Use when selecting or tuning a routing strategy, or configuring failover, health checks, and load balancing.

Routing Architecture Guide

Overview

The router (src/core/router/) selects among deployments — concrete provider+model pairs registered via Router::add_deployment / Router::set_model_list — using one of 7 strategies. There is no per-strategy router struct, no Router trait, and no create_router factory: strategy dispatch is a match on the RoutingStrategy enum calling free functions in src/core/router/strategy_impl.rs.

Key Design Principles
  • Snapshot isolation: deployments, the model index, and aliases live in an immutable RoutingSnapshot, published through ArcSwap (src/core/router/unified.rs:61,308). Readers load one generation lock-free; writers clone-modify-store under a parking_lot::Mutex that only serializes writers (unified.rs:312,379-398).
  • Atomic deployment state: per-deployment runtime state (DeploymentState) is plain atomics with Relaxed ordering (deployment.rs:210-252); RoundRobin uses a DashMap<String, AtomicUsize> of per-model counters (unified.rs:321).
  • Health-aware: candidates are filtered by cooldown, health status, parallel limits, and RPM/TPM limits before a strategy picks among them.
  • Fallback chains: the built-in execution path tries the model-keyed General fallback list after retries are exhausted. Typed context-window, content-policy, and rate-limit lists require explicit caller lookup/wiring today.

Selection Flow

Entry point Router::select_deployment_lease (selection.rs:95) delegates to select_deployment_matching (selection.rs:226). The real flow:

  1. Load the current snapshot; resolve model aliases via resolve_model_name (max MAX_ALIAS_HOPS = 16 hops, unified.rs:26).
  2. Look up the resolved model in snapshot.model_index to get candidate DeploymentIds.
  3. One filter pass builds Vec<RoutingContext> (strategy_impl.rs:17), skipping deployments that are in cooldown (is_in_cooldown), unhealthy (is_healthy), at their max_parallel_requests limit, or at their rpm_limit / tpm_limit (selection.rs:287-344). Each context copies weight, priority, active_requests, tpm_current/tpm_limit, rpm_current/rpm_limit, avg_latency_us.
  4. Router::select_from_routing_contexts dispatches on the configured strategy (selection.rs:195-224):
rust
match strategy {
    RoutingStrategy::SimpleShuffle => strategy_impl::weighted_random_from_context(routing_contexts),
    RoutingStrategy::LeastBusy     => strategy_impl::least_busy_from_context(routing_contexts),
    RoutingStrategy::UsageBased    => strategy_impl::lowest_usage_from_context(routing_contexts),
    RoutingStrategy::LatencyBased  => strategy_impl::lowest_latency_from_context(routing_contexts),
    RoutingStrategy::PriorityBased => strategy_impl::lowest_priority_from_context(routing_contexts),
    RoutingStrategy::RateLimitAware => strategy_impl::rate_limit_aware_from_context(routing_contexts),
    RoutingStrategy::RoundRobin    => strategy_impl::round_robin_from_context(
        model_name, routing_contexts, round_robin_counters),
}
  1. The winner is reserved by try_reserve_deployment — an atomic increment of active_requests (CAS loop when max_parallel_requests is set). If another caller wins the last slot, that candidate is removed from the contexts and selection retries (selection.rs:378-415).
  2. Returns a DeploymentLease; dropping it decrements active_requests (RAII release, selection.rs:64-70). The deprecated ID-returning select_deployment still exists but converts the lease to an ID without release-on-drop.

Routing Strategies

Enum RoutingStrategy (src/core/router/config.rs:22-39) serializes as snake_case (simple_shuffle, round_robin, least_busy, latency_based, priority_based, usage_based, rate_limit_aware); PriorityBased also accepts the serde alias "cost_based".

1. SimpleShuffle (runtime default)

Weighted random selection: draws a point in 0..total_weight and walks candidates until the cumulative weight covers it; uniform random when total weight is 0 (weighted_random_from_context, strategy_impl.rs:56-85). Weights come from DeploymentConfig.weight (default 1).

Use when: general traffic where deployments have different capacity.

Show full SKILL.md (273 more words)Show less
2. RoundRobin (gateway YAML default)

Per-model counter in round_robin_counters: DashMap<String, AtomicUsize> cycles through candidate order (round_robin_from_context, strategy_impl.rs:247-274). Note the defaults diverge: GatewayRouterConfig defaults to round_robin (src/config/models/router.rs:48-50) while runtime RouterConfig defaults to SimpleShuffle.

Use when: predictable distribution needed, debugging provider issues.

3. LeastBusy

Single pass for the fewest active_requests; ties are broken randomly with reservoir sampling so equal-load deployments share traffic (least_busy_from_context, strategy_impl.rs:88-116).

Use when: high concurrency, need to prevent deployment overload.

4. LatencyBased

Lowest avg_latency_us wins. Deployments reporting 0 latency (no success yet) inherit the average of non-zero latencies in the pool, so new deployments neither always win nor starve (lowest_latency_from_context, strategy_impl.rs:145-182). Latency is recorded per request via record_success into DeploymentState.avg_latency_us.

Use when: response time is critical, deployments have varying latencies.

5. PriorityBased

Lowest priority value wins (lower = higher priority; u32, default 0). This is tier ordering, not cost — despite the legacy "cost_based" serde alias (lowest_priority_from_context, strategy_impl.rs:185-203).

Use when: primary/backup tiering, e.g. production vs backup deployments (see gateway.yaml.example provider priority).

6. UsageBased

Lowest TPM usage percentage wins: (tpm_current * 100) / tpm_limit; deployments with no limit count as 0% usage (lowest_usage_from_context, strategy_impl.rs:119-142).

Use when: spreading load relative to token budgets, avoiding TPM exhaustion. Requires tpm to be configured on deployments to be meaningful.

7. RateLimitAware

Picks the deployment furthest from its rate limits: score is the minimum of remaining TPM fraction and remaining RPM fraction; unlimited axes score 1.0 (rate_limit_aware_from_context, strategy_impl.rs:206-241).

Use when: high request volume against deployments with strict TPM/RPM limits.


References

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in .claude/skills/routing-architecture of majiayu000/litellm-rs.

  • SKILL.md
  • reference/health-and-fallbacks.md
  • reference/performance-and-practices.md
  • reference/router-configuration.md

Open the folder on GitHubat commit ed3f4d9

Compare with similar skills

Routing Architecture next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Routing Architecture compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Routing Architecture this skillmajiayu000/litellm-rs118—~1.9kAutomated safety check: PassMIT
Azure Resource Manager Mysql Dotnetmicrosoft/skills3.1k5 repos~3.5kAutomated safety check: PassMIT
Azure Resource Manager Postgresql Dotnetmicrosoft/skills3.1k5 repos~4kAutomated safety check: PassMIT
LLM Gatewaysickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
Openrouter Fallback Configjeremylongshore/tons-of-skills-marketplace2.8k—~2.4kAutomated safety check: PassMIT
LLM GatewayBagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT

Similar skills

  • Official

    Azure MySQL Flexible Server SDK for .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~3.5k tokens
    DevOps & CloudAuto-check passed
  • Azure PostgreSQL Flexible Server SDK for .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~4k tokens
    DevOps & CloudAuto-check passed
  • LLM Gateway

    sickn33/agentic-awesome-skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Backend & APIsAuto-check passed
  • Openrouter Fallback Config

    jeremylongshore/tons-of-skills-marketplace

    Configure automatic model fallbacks for high availability on OpenRouter.

    2.8k GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Gateway

    BagelHole/DevOps-Security-Agent-Skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    1.2k GitHub stars~2k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • GreptimeDB Dev Docker Image

    GreptimeTeam/greptimedb

    Packages a locally built GreptimeDB debug binary into a development-only Docker image for local-cluster testing, with an optional push to a dev registry.

    6.7k GitHub stars~4k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes

More from majiayu000/litellm-rs

All 9 skills in this repo
  • Auth Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Authentication Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Caching Architecture

    majiayu000/litellm-rs

    LiteLLM-RS response caching architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Config Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Configuration Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Error Handling

    majiayu000/litellm-rs

    LiteLLM-RS Error Handling Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Observability Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Provider Architecture

    majiayu000/litellm-rs

    LiteLLM-RS provider system in two tiers - data-driven OpenAI-compatible catalog entries auto-routed through OpenAILikeProvider, plus code-based provider modules implementing the LLMProvider trait…

    118 GitHub stars~4.9k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Routing Architecture

What does Routing Architecture do?

LiteLLM-RS Routing Architecture. An agent skill from majiayu000/litellm-rs. Routing Architecture is an agent skill from majiayu000/litellm-rs. LiteLLM-RS Routing Architecture.

When should I use Routing Architecture?

Routing Architecture fits situations like: tuning a routing strategy; configuring failover.

How do I install Routing Architecture in Claude Code?

Run `npx skills add majiayu000/litellm-rs --skill routing-architecture -a claude-code`. Or copy the skill folder (.claude/skills/routing-architecture in majiayu000/litellm-rs) into .claude/skills/routing-architecture in your project. Claude Code loads it when a task matches its description.

How do I install Routing Architecture in Codex?

Run `npx skills add majiayu000/litellm-rs --skill routing-architecture -a codex`. Or copy the skill folder (.claude/skills/routing-architecture in majiayu000/litellm-rs) into .agents/skills/routing-architecture in your project. Codex loads it when a task matches its description.

Can I use Routing Architecture in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/litellm-rs --skill routing-architecture -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/routing-architecture, .gemini/skills/routing-architecture, .github/skills/routing-architecture and .opencode/skills/routing-architecture in your project.

What does Routing Architecture need to run?

SKILL.md names no scripts, command-line tools or credentials: Routing Architecture is instructions for the agent only.

Does Routing Architecture access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Routing Architecture safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Routing Architecture use?

Routing Architecture is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Routing Architecture use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Routing Architecture?

Skills that share tags, products or a category with Routing Architecture: Azure Resource Manager Mysql Dotnet (microsoft/skills, 3.1k stars), Azure Resource Manager Postgresql Dotnet (microsoft/skills, 3.1k stars), LLM Gateway (sickn33/agentic-awesome-skills, 47k stars) and Openrouter Fallback Config (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Routing Architecture?

majiayu000 (a GitHub user) maintains it in majiayu000/litellm-rs, which has 118 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 11, 2026.

Source: majiayu000/litellm-rs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.