Agent skill

Fastllm Routing

by azrtydxb in azrtydxb/Fastllm-proxy

Route requests across models in FastLLM — create, inspect or change frontend models, their default targets, weighted splits, and routing rules (by caller, prompt size, streaming, headers, budget…

Apache-2.0Auto-check passedDevOps & Cloud

Install Fastllm Routing

skills CLI
$ npx skills add azrtydxb/Fastllm-proxy --skill fastllm-routing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install azrtydxb/Fastllm-proxy fastllm-routing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/azrtydxb/Fastllm-proxy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fastllm-routing .claude/skills/fastllm-routing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fastllm-routing
GitHub stars
108
Token cost
~2.7k tokens
SKILL.md length
1,407 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Route requests across models in FastLLM — create, inspect or change frontend models, their default targets, weighted splits, and routing rules (by caller, prompt size, streaming, headers, budget…

  • Asked to expose a model under a client-facing name
  • SKILL.md covers Auth, Test before you apply, Actions and Three levels, and they answer…, plus 3 more sections
  • Calls curl
  • Decide which backend a request should hit

What it does

Fastllm Routing is an agent skill from azrtydxb/Fastllm-proxy. Route requests across models in FastLLM — create, inspect or change frontend models, their default targets, weighted splits, and routing rules (by caller, prompt size, streaming, headers, budget, time of day, semantic class, or backend load for local/cloud spillover). Use when asked to expose a model under a client-facing name, decide which backend a request should hit, set up failover or canary traffic, or explain why a request went where it did. Not for registering models or backends (fastllm-models) or for…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Deployment and Backup and disaster recovery. The repository describes itself as: The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O… The licence is Apache-2.0.

When your agent uses it

  • Asked to expose a model under a client-facing name
  • Decide which backend a request should hit
  • Set up failover
  • Explain why a request went where it did

Example prompts

  • “/fastllm-routing”

What it can do on your machine

Read from SKILL.md and the folder at commit 5d53db8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fastllm Routing loads about 2.7k tokens when it runs. Until then it costs about 144 tokens; SKILL.md has 1,407 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from azrtydxb/Fastllm-proxy at commit 5d53db8, republished under its Apache-2.0 licence (© azrtydxb). 1,407 words, ~2,727 tokens.

Download SKILL.mdSave it as .claude/skills/fastllm-routing/SKILL.md (or your agent's skills folder).
name
fastllm-routing
description
Route requests across models in FastLLM — create, inspect or change frontend models, their default targets, weighted splits, and routing rules (by caller, prompt size, streaming, headers, budget, time of day, semantic class, or backend load for local/cloud spillover). Use when asked to expose a model under a client-facing name, decide which backend a request should hit, set up failover or canary traffic, or explain why a request went where it did. Not for registering models or backends (fastllm-models) or for sending inference requests (fastllm-gateway).

FastLLM routing

A frontend model is a client-facing name backed by an ordered list of rules. The first rule whose conditions all hold wins and commits to its own targets; if none match, the defaults are used. Everything is pre-resolved into the snapshot, so the request path does no I/O to route.

Auth

Admin endpoints need a session cookie, not a bearer token. The gateway master key is not an admin credential.

bash
curl -sk -c /tmp/ck -X POST https://192.168.10.129:4001/login \
  -H 'content-type: application/json' -d '{"name":"<user>","password":"<pw>"}'
curl -sk -b /tmp/ck https://192.168.10.129:4001/admin/frontend-models

Test before you apply

POST /admin/routing/dry-run answers "which model would this request hit, and which rule decided" without changing anything. Use it before and after any change here — it is how you check a rule does what you meant.

It cannot evaluate max_inflight_per_backend. The dry-run runs on the control plane, whose registry has no in-flight counters and no engine scrape, so every backend looks idle and a spill rule always reports as still matching. Only real traffic exercises that condition.

<!-- BEGIN GENERATED: endpoints -->
MethodPathSummaryBody fields
DELETE/admin/frontend-model-defaults/{id}Delete frontend-model-defaults id—
GET/admin/frontend-modelsRead frontend-models—
POST/admin/frontend-modelsCreate frontend-modelsname, description*
PATCH/admin/frontend-models/{id}Change how a frontend model chooses between its targetsname, description
DELETE/admin/frontend-models/{id}Delete frontend-models id—
POST/admin/frontend-models/{id}/defaultsCreate frontend-models id defaultsprovider_model_id, model_pool_id, weight*, position
POST/admin/frontend-models/{id}/rulesAdd a routing rule. First match wins, and the matching rule decides everything — every action is terminalposition, action, deny_status, deny_message, jump_to, tag*, match_condition
DELETE/admin/model-pool-members/{id}Remove one model from its pool—
GET/admin/model-poolsNamed groups of provider models, each with one policy for choosing between them—
POST/admin/model-poolsCreate a pool. Its name must not collide with a provider model or a frontend model, because routing resolves targets by namename, description, policy
PATCH/admin/model-pools/{id}Rename a pool, change its description, or change how it chooses between its members. A rename carries onto the targets pointing at itname, description, policy*
DELETE/admin/model-pools/{id}Delete a pool. Refused while any rule target or default still routes to it, rather than cascading into rules that point at a name which no longer resolves—
POST/admin/model-pools/{id}/membersAdd a member to a pool: a whole provider model, or one of its attachments when model_backend_id is given. The same attachment twice is refused, and so is the same model twice at model grainprovider_model_id, model_backend_id, weight, position*
POST/admin/routing/dry-runWhich rule would decide, and what the chain resolves to, without dispatchingmodel, principal_id, streaming, prompt_tokens, max_tokens, headers, class, class_refines*
DELETE/admin/rule-targets/{id}Delete rule-targets id—
PATCH/admin/rules/{id}Change how a rule chooses among its targets, or where it sits in the order. Its conditions are not editable: delete and recreate instead of letting a rule change meaning while keeping the position that makes it firstposition, match_condition, action, deny_status, deny_message, jump_to, tag*
DELETE/admin/rules/{id}Delete rules id—
POST/admin/rules/{id}/targetsCreate rules id targetsprovider_model_id, model_pool_id, weight*, position

* optional field

<!-- END GENERATED: endpoints -->

Actions

Every rule has one, and every one is terminal — no rule contributes and passes on, so dry-run names a single deciding rule.

  • route (default) — the rule's targets, ordered by its policy.
  • deny — refuse. Needs deny_status, 4xx only (a 5xx would have every client library retrying something that is never going to be allowed).
  • jump — continue in another frontend model's chain (jump_to), so shared policy is written once. A loop is refused at write time; a destination that was deleted is treated as no match and falls through.

tag is a field, not an action: it labels the usage rows the rule produced, so spend is answerable per decision. A rule that tags still routes.

A deny shows up in dry-run as a denied object with the status and message, and an empty candidates — which is a different thing from "nothing is serving", and the distinction is the point.

Cost as a condition (min/max_request_cost_micros) is priced at the cheapest model the frontend model can reach — one number per request, the same for every rule, so it does not depend on which rule is asking. Pairs with deny ("refuse anything over $0.50 however I route it"). Sending the expensive ones somewhere cheaper is not this; order the chain with the cheaper target first. All targets unpriced means no cost, and the rule declines.

Three levels, and they answer different questions

A rule's targets are an ordered failover chain. No policy, no weighted split — tried in the order written. POST /admin/rules/{id}/targets takes either provider_model_id or model_pool_id, never both.

A pool (/admin/model-pools) is a named group of provider models with one policy: cache-affinity, least-loaded, lowest-latency, round-robin, or unset for the weighted split by member weight. Cost is not one of them — a price is fixed, so it never balances; use a rule's min/max_request_cost_micros. Two pools may hold the same members and differ only in policy — that is what naming it buys. A pool expands in place: its chosen member, then its others, then the chain's next target.

A provider model has no policy. It is a model running on endpoints. Where several endpoints serve it they all go in its pool and the deployment's --policy spreads across them; there is nothing to set per model.

routing_rules.policy and frontend_models.policy no longer exist; a rule that balanced across its targets is a rule pointing at a pool.

Show full SKILL.md (549 more words)Show less

Traps

A provider model and a frontend model may share a name, and normally do. This used to be a 409. It is not ambiguous: resolve_target_models looks in frontend models first and falls through to a provider model only when there is none of that name, so the frontend model wins deterministically. Migration 0034 depends on it — every provider model gets a frontend model of the same name, so it stays callable once frontend models are the only addressable surface. Renaming the provider model out of the way instead would revoke every grant naming it. Pinned by a_provider_model_and_a_frontend_model_may_share_a_name.

frontend_model_defaults.position is NOT NULL with no default. An INSERT that omits it fails. The API sets it; hand-written SQL must too.

weight is a relative share, not a percentage. Two targets at 1 and 1 split evenly; 1 and 3 split 25/75. They need not sum to 100, so adding a third target never forces you to rebalance the other two.

First match wins, and a matching rule commits. If its targets resolve to nothing routable the request fails — it does not fall through to the next rule or to the defaults. Falling through would make "first match wins" a lie that depends on backend health the rule author cannot see.

The weighted split is deterministic, not random. It hashes the same request prefix the backend router hashes, so a multi-turn conversation stays on one side of a canary instead of flipping per request.

max_inflight_per_backend is the only condition that is not a pure function of the request. It reads live in-flight counters, so two identical requests a second apart can route differently. That is the price of local/cloud spillover; the field is named for the mechanism rather than the intent for that reason.

It counts the engine's requests, not this proxy's. Each proxy scrapes every backend's Prometheus /metrics every two seconds (--engine-scrape-interval, 0 disables) and compares the ceiling against vLLM/SGLang's own running + queued count, so a limit of 2 still means 2 with three proxies running. A backend with no /metrics, or one whose last reading is over ten seconds old, falls back to this replica's own count.

Which backends have metrics is detected, not configured. A backend that produces no reading is retried every five minutes and given up on after fifteen, at which point it is dropped from the scrape; one that answers with something that is not engine metrics settles on the second answer. Nothing needs to mark a provider as an engine, and nothing polls a cloud vendor in a loop. A backend given up on by deadline — nothing ever replied — gets one more look if it later passes a health probe, so an engine that was slow to load is not written off permanently.

Prefer the API over SQL

Changes made through /admin/* write an audit row and rebuild the snapshot. Direct psql writes do neither — the change still reaches proxies on their next snapshot poll, but nothing records who made it or why. If you must use SQL because no admin credential is available, say so explicitly in your report.

Verify

bash
# what the control plane now believes
curl -sk -b /tmp/ck https://192.168.10.129:4001/admin/frontend-models

# what a request would actually do
curl -sk -b /tmp/ck -X POST https://192.168.10.129:4001/admin/routing/dry-run \
  -H 'content-type: application/json' -d '{"model":"<virtual>","prompt_tokens":100}'

Changes reach the proxies on their snapshot poll, not instantly. A gateway 401 means the proxy is healthy and rejecting an unauthenticated request; a /health 503 means it has no usable snapshot yet.

© azrtydxb, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/fastllm-routing of azrtydxb/Fastllm-proxy.

Open the folder on GitHubat commit 5d53db8

Compare with similar skills

Fastllm Routing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fastllm Routing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fastllm Routing this skillazrtydxb/Fastllm-proxy108—~2.7kAutomated safety check: PassApache-2.0
Temps CLIgotempsh/temps833—~2kAutomated safety check: PassApache-2.0
Routing Architecturemajiayu000/litellm-rs118—~1.9kAutomated safety check: PassMIT
Supercheck Infrastructure Deploymentsupercheck-io/supercheck215—~1.4kAutomated safety check: NotesAGPL-3.0
Azure Resource Manager Mysql Dotnetmicrosoft/skills3.1k5 repos~3.5kAutomated safety check: PassMIT
Azure Resource Manager Postgresql Dotnetmicrosoft/skills3.1k5 repos~4kAutomated safety check: PassMIT

Similar skills

  • Temps CLI

    gotempsh/temps

    Operate Temps through the pinned @temps-sdk/cli package with bunx or npx.

    833 GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Routing Architecture

    majiayu000/litellm-rs

    LiteLLM-RS Routing Architecture. An agent skill from majiayu000/litellm-rs.

    118 GitHub stars~1.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Supercheck Infrastructure Deployment

    supercheck-io/supercheck

    Work on Supercheck Docker Compose, K3s, Kubernetes manifests, gVisor, OpenTofu/Hetzner, secrets, external services, autoscaling, backups, disaster recovery, DNS/TLS, or production deployment.

    215 GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Official

    Azure MySQL Flexible Server SDK for .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~3.5k tokens
    DevOps & CloudAuto-check passed
  • Azure PostgreSQL Flexible Server SDK for .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~4k tokens
    DevOps & CloudAuto-check passed
  • Preset

    microsoft/GitHub-Copilot-for-Azure

    Official

    Intelligently deploys Azure OpenAI models to optimal regions by analyzing capacity across all available regions.

    255 GitHub starsUsed in 1 repo~1.2k tokens
    DevOps & CloudAuto-check passed

More from azrtydxb/Fastllm-proxy

All 14 skills in this repo
  • Fastllm Agents

    azrtydxb/Fastllm-proxy

    Manage and invoke A2A agents behind FastLLM — register, patch, delete and list agents on the control plane, list them through the gateway, fetch an agent card, and invoke an agent by name.

    108 GitHub stars~512 tokensUpdated today
    Auto-check passed
  • Fastllm Backends

    azrtydxb/Fastllm-proxy

    Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

    108 GitHub stars~980 tokensUpdated today
    Auto-check passed
  • Fastllm Classifier

    azrtydxb/Fastllm-proxy

    Manage FastLLM prompt classes for semantic routing — create classes and their example prompts, list or delete them, and evaluate how a given prompt would be classified.

    108 GitHub stars~499 tokensUpdated today
    Auto-check passed
  • Fastllm Deployment

    azrtydxb/Fastllm-proxy

    Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and…

    108 GitHub stars~727 tokensUpdated today
    Auto-check passed
  • Fastllm Gateway

    azrtydxb/Fastllm-proxy

    Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…

    108 GitHub stars~926 tokensUpdated today
    Auto-check passed
  • Fastllm MCP

    azrtydxb/Fastllm-proxy

    Manage and use MCP servers behind FastLLM — register, patch, delete and list MCP servers on the control plane, and list or call their tools through the gateway.

    108 GitHub stars~515 tokensUpdated today
    Auto-check passed

Questions about Fastllm Routing

What does Fastllm Routing do?

Route requests across models in FastLLM — create, inspect or change frontend models, their default targets, weighted splits, and routing rules (by caller, prompt size, streaming, headers, budget…. Fastllm Routing is an agent skill from azrtydxb/Fastllm-proxy. Route requests across models in FastLLM — create, inspect or change frontend models, their default targets, weighted splits, and routing rules (by caller, prompt size, streaming, headers, budget, time of day, semantic class, or backend load for local/cloud spillover).

When should I use Fastllm Routing?

Fastllm Routing fits situations like: asked to expose a model under a client-facing name; decide which backend a request should hit; set up failover; explain why a request went where it did.

How do I install Fastllm Routing in Claude Code?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-routing -a claude-code`. Or copy the skill folder (.claude/skills/fastllm-routing in azrtydxb/Fastllm-proxy) into .claude/skills/fastllm-routing in your project. Claude Code loads it when a task matches its description.

How do I install Fastllm Routing in Codex?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-routing -a codex`. Or copy the skill folder (.claude/skills/fastllm-routing in azrtydxb/Fastllm-proxy) into .agents/skills/fastllm-routing in your project. Codex loads it when a task matches its description.

Can I use Fastllm Routing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-routing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fastllm-routing, .gemini/skills/fastllm-routing, .github/skills/fastllm-routing and .opencode/skills/fastllm-routing in your project.

What does Fastllm Routing need to run?

Going by SKILL.md and its folder, Fastllm Routing needs the command-line tools its instructions call (curl).

Does Fastllm Routing access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fastllm Routing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fastllm Routing use?

Fastllm Routing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fastllm Routing use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fastllm Routing?

Skills that share tags, products or a category with Fastllm Routing: Temps CLI (gotempsh/temps, 833 stars), Routing Architecture (majiayu000/litellm-rs, 118 stars), Supercheck Infrastructure Deployment (supercheck-io/supercheck, 215 stars) and Azure Resource Manager Mysql Dotnet (microsoft/skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fastllm Routing?

azrtydxb (a GitHub organization) maintains it in azrtydxb/Fastllm-proxy, which has 108 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 5, 2026.

Source: azrtydxb/Fastllm-proxy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.