Agent skill

Fastllm Models

by azrtydxb in azrtydxb/Fastllm-proxy

Register and maintain the models FastLLM can serve — create, patch or delete a model, attach backends to it, remove a backend, and set the deployment-wide fallback model.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Fastllm Models

skills CLI
$ npx skills add azrtydxb/Fastllm-proxy --skill fastllm-models -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install azrtydxb/Fastllm-proxy fastllm-models --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/azrtydxb/Fastllm-proxy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fastllm-models .claude/skills/fastllm-models && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fastllm-models
GitHub stars
108
Token cost
~1.8k tokens
SKILL.md length
697 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Register and maintain the models FastLLM can serve — create, patch or delete a model, attach backends to it, remove a backend, and set the deployment-wide fallback model.

  • Adding a new inference endpoint
  • SKILL.md covers Auth and Traps
  • Calls curl
  • Pointing a model at a different host

What it does

Fastllm Models is an agent skill from azrtydxb/Fastllm-proxy. Register and maintain the models FastLLM can serve — create, patch or delete a model, attach backends to it, remove a backend, and set the deployment-wide fallback model. Use when adding a new inference endpoint, pointing a model at a different host or port, retiring a backend, or choosing what catches a request when every other target fails. Not for choosing between models per request (fastllm-routing).

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model routing and gateways. The repository describes itself as: The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O… The licence is Apache-2.0.

When your agent uses it

  • Adding a new inference endpoint
  • Pointing a model at a different host
  • Retiring a backend
  • Choosing what catches a request when every other target fails

Example prompts

  • “/fastllm-models”

What it can do on your machine

Read from SKILL.md and the folder at commit 5d53db8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fastllm Models loads about 1.8k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 697 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from azrtydxb/Fastllm-proxy at commit 5d53db8, republished under its Apache-2.0 licence (© azrtydxb). 697 words, ~1,776 tokens.

Download SKILL.mdSave it as .claude/skills/fastllm-models/SKILL.md (or your agent's skills folder).
name
fastllm-models
description
Register and maintain the models FastLLM can serve — create, patch or delete a model, attach backends to it, remove a backend, and set the deployment-wide fallback model. Use when adding a new inference endpoint, pointing a model at a different host or port, retiring a backend, or choosing what catches a request when every other target fails. Not for choosing between models per request (fastllm-routing).

FastLLM models

Auth

Admin endpoints need a session cookie, not a bearer token — the gateway master key is not an admin credential.

bash
curl -sk -c /tmp/ck -X POST https://192.168.10.129:4001/login \
  -H 'content-type: application/json' -d '{"name":"<user>","password":"<pw>"}'
curl -sk -b /tmp/ck https://192.168.10.129:4001/admin/...
<!-- BEGIN GENERATED: endpoints -->
MethodPathSummaryBody fields
PATCH/admin/backends/{id}Change what one model costs, is called, and how it is protected at one provider. An explicit null clears a field; an absent field is left aloneupstream_model, input_price_per_mtok, output_price_per_mtok, default_max_tokens, upstream_timeout_seconds, admission_max_concurrent, options, admission_high_water, admission_max_queued, admission_max_wait_seconds*
DELETE/admin/backends/{id}Detach one model from one provider. The model, its usage history and the provider itself are left alone—
GET/admin/fallback-modelRead fallback-model—
PUT/admin/fallback-modelSet fallback-modelprovider_model_id*
GET/admin/provider-catalogueKnown providers and how to reach them—
GET/admin/provider-modelsRead provider models—
POST/admin/provider-modelsCreate provider modelsname, description, default, cache_ttl_seconds, context_length*
PATCH/admin/provider-models/{id}Correct a model in place. An explicit null clears a field; an absent field is left alonename, description, cache_ttl_seconds, context_length
DELETE/admin/provider-models/{id}Delete models id—
POST/admin/provider-models/{id}/backendsCreate models id backendsprovider_id, api_base, upstream_model, upstream_api_key, Authorization, protocol, auth_header, auth_scheme, default_max_tokens, input_price_per_mtok, output_price_per_mtok, credential_kind, extra_headers
GET/admin/providersRead providers—
POST/admin/providersAdd a provider: an endpoint and the credential that reaches itname, kind, catalogue_key, api_base, protocol, auth_header, auth_scheme, upstream_api_key, credential_kind, extra_headers
POST/admin/providers/registerRegister or refresh a provider's leaseapi_base, node, name, engine, ttl_seconds
PATCH/admin/providers/{id}Rename a provider, move it, or rotate its credential. An absent upstream_api_key leaves the stored one alone; "" clears itname, kind, api_base, protocol, auth_header, auth_scheme, upstream_api_key, credential_kind, extra_headers*
DELETE/admin/providers/{id}Delete a provider that serves no models—
GET/admin/providers/{id}/available-modelsWhat a provider is currently serving—
POST/admin/providers/{id}/oauth/callbackComplete an OAuth flow: exchanges the authorization code for tokens and stores them encrypted against the provider. Body: {state, code}—
POST/admin/providers/{id}/oauth/connectBegin an OAuth flow for the provider: generates the PKCE challenge and returns the authorization URL to visit—
POST/admin/providers/{id}/oauth/disconnectClear the provider's stored OAuth tokens—
GET/admin/providers/{id}/oauth/statusReport whether the provider holds live OAuth tokens and how long they remain valid—

* optional field

<!-- END GENERATED: endpoints -->
Show full SKILL.md (376 more words)Show less

Traps

A provider model and a frontend model may share a name, and normally do. The frontend model wins during resolution, and migration 0034 gives every provider model one of the same name so it stays callable. This used to be a 409 in both create paths; it no longer is.

The fallback model catches a frontend model whose chain ran out. It is the last resort when a rule author could not anticipate a failure mode; it is skipped when already present in the chain, so naming it explicitly does not double it.

A backend that fails health checks leaves rotation but is not dropped from the chain. When nothing is healthy the request still goes somewhere and the real upstream error reaches the client, which beats a synthetic 503.

A model may run at several providers, and they form one pool. POST /admin/provider-models/{id}/backends again with a different provider_id attaches it there too; router.rs then chooses between them per request (prefix-cache affinity, least-loaded, and so on). The same provider twice is a 409. Prices, upstream_model and default_max_tokens are on the attachment, not the model -- the same weights cost different amounts at different vendors -- and PATCH /admin/backends/{id} is what changes them.

A provider is created before its models, not by them. POST /admin/providers takes the endpoint and its credential — from a catalogue key for a cloud vendor, or a typed api_base for anything else. Attaching a model then only has to name it:

bash
# The endpoint and its key, once.
curl -sk -b /tmp/ck -X POST https://192.168.10.129:4001/admin/providers \
  -H 'content-type: application/json' \
  -d '{"catalogue_key":"anthropic","upstream_api_key":"sk-ant-..."}'

# What that provider is actually serving, before deciding what to register.
curl -sk -b /tmp/ck https://192.168.10.129:4001/admin/providers/7/available-models

# The model, on that provider.
curl -sk -b /tmp/ck -X POST \
  https://192.168.10.129:4001/admin/provider-models/42/backends \
  -H 'content-type: application/json' \
  -d '{"provider_id":7,"upstream_model":"claude-sonnet-4-5"}'

POST .../backends with an api_base instead still works and still find-or-creates a provider — that is how every backend was attached before providers were records, and every existing script does it that way.

provider_id and the fields describing an endpoint are mutually exclusive. Sending upstream_api_key alongside a provider_id is a 400, not a silent preference for one source: the caller would otherwise believe they had set a credential while the provider's is what actually gets sent. Change those with PATCH /admin/providers/{id}, which rotates the key for every model on it in one write.

A catalogue base_url can contain a <placeholder>. Bedrock and Vertex both encode a region, and Vertex a project. POST /admin/providers refuses an address that still has one in it rather than storing something that resolves nowhere and then reports itself unreachable.

© azrtydxb, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/fastllm-models of azrtydxb/Fastllm-proxy.

Open the folder on GitHubat commit 5d53db8

Compare with similar skills

Fastllm Models next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fastllm Models compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fastllm Models this skillazrtydxb/Fastllm-proxy108—~1.8kAutomated safety check: PassApache-2.0
Shogun Bloom Configyohey-w/multi-agent-shogun1.4k—~3.1kAutomated safety check: PassMIT
Codemie Analyticscodemie-ai/codemie-code294—~7.5kAutomated safety check: PassApache-2.0
Model Routernidhi-singh02/agent-router112—~1.2kAutomated safety check: PassMIT
Codex Model Routing Teamzjp1997720/codex-model-routing-team158—~736Automated safety check: PassMIT
Add Modelget-convex/convex-evals130—~1.5kAutomated safety check: NotesApache-2.0

Similar skills

  • Shogun Bloom Config

    yohey-w/multi-agent-shogun

    Interactive wizard: guided questions with multiple-choice options about subscriptions, then outputs a ready-to-paste capabilitytiers YAML + fixed agent model assignments.

    1.4k GitHub stars~3.1k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Codemie Analytics

    codemie-ai/codemie-code

    CodeMie Analytics expert — use this skill whenever the user asks about CodeMie usage data, AI adoption metrics, user leaderboards, CLI insights, spending, LiteLLM costs, token usage, or wants to…

    294 GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Model Router

    nidhi-singh02/agent-router

    A skill your agent uses when the user asks to pick a model, subscription, or reasoning effort, or to run router status, usage refresh, or resume a router session.

    112 GitHub stars~1.2k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Codex Model Routing Team

    zjp1997720/codex-model-routing-team

    在 Codex App 中为复杂、可并行的知识工作或编程任务自动创建多个可指定模型与推理强度的后台任务,由主 Agent 负责规划、分工、集成和验收。用于多来源调研、多章节内容、复杂 Skill/PPT、跨模块开发、独立验证或 2 个以上互不依赖工作流;也用于用户明确要求模型路由、后台 Worker、Agents Team…

    158 GitHub stars~736 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    get-convex/convex-evals

    Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

    130 GitHub stars~1.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • OmniRoute CLI Evals

    diegosouzapw/OmniRoute

    Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

    75k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from azrtydxb/Fastllm-proxy

All 14 skills in this repo
  • Fastllm Agents

    azrtydxb/Fastllm-proxy

    Manage and invoke A2A agents behind FastLLM — register, patch, delete and list agents on the control plane, list them through the gateway, fetch an agent card, and invoke an agent by name.

    108 GitHub stars~512 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Backends

    azrtydxb/Fastllm-proxy

    Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

    108 GitHub stars~980 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Classifier

    azrtydxb/Fastllm-proxy

    Manage FastLLM prompt classes for semantic routing — create classes and their example prompts, list or delete them, and evaluate how a given prompt would be classified.

    108 GitHub stars~499 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Deployment

    azrtydxb/Fastllm-proxy

    Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and…

    108 GitHub stars~727 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Gateway

    azrtydxb/Fastllm-proxy

    Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…

    108 GitHub stars~926 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm MCP

    azrtydxb/Fastllm-proxy

    Manage and use MCP servers behind FastLLM — register, patch, delete and list MCP servers on the control plane, and list or call their tools through the gateway.

    108 GitHub stars~515 tokensUpdated 4 days ago
    Auto-check passed

Questions about Fastllm Models

What does Fastllm Models do?

Register and maintain the models FastLLM can serve — create, patch or delete a model, attach backends to it, remove a backend, and set the deployment-wide fallback model. Fastllm Models is an agent skill from azrtydxb/Fastllm-proxy. Register and maintain the models FastLLM can serve — create, patch or delete a model, attach backends to it, remove a backend, and set the deployment-wide fallback model.

When should I use Fastllm Models?

Fastllm Models fits situations like: adding a new inference endpoint; pointing a model at a different host; retiring a backend; choosing what catches a request when every other target fails.

How do I install Fastllm Models in Claude Code?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-models -a claude-code`. Or copy the skill folder (.claude/skills/fastllm-models in azrtydxb/Fastllm-proxy) into .claude/skills/fastllm-models in your project. Claude Code loads it when a task matches its description.

How do I install Fastllm Models in Codex?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-models -a codex`. Or copy the skill folder (.claude/skills/fastllm-models in azrtydxb/Fastllm-proxy) into .agents/skills/fastllm-models in your project. Codex loads it when a task matches its description.

Can I use Fastllm Models in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-models -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fastllm-models, .gemini/skills/fastllm-models, .github/skills/fastllm-models and .opencode/skills/fastllm-models in your project.

What does Fastllm Models need to run?

Going by SKILL.md and its folder, Fastllm Models needs the command-line tools its instructions call (curl).

Does Fastllm Models access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fastllm Models safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fastllm Models use?

Fastllm Models is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fastllm Models use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fastllm Models?

Skills that share tags, products or a category with Fastllm Models: Shogun Bloom Config (yohey-w/multi-agent-shogun, 1.4k stars), Codemie Analytics (codemie-ai/codemie-code, 294 stars), Model Router (nidhi-singh02/agent-router, 112 stars) and Codex Model Routing Team (zjp1997720/codex-model-routing-team, 158 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fastllm Models?

azrtydxb (a GitHub organization) maintains it in azrtydxb/Fastllm-proxy, which has 108 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 5, 2026.

Source: azrtydxb/Fastllm-proxy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.