Agent skill

Fastllm Gateway

by azrtydxb in azrtydxb/Fastllm-proxy

Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Fastllm Gateway

skills CLI
$ npx skills add azrtydxb/Fastllm-proxy --skill fastllm-gateway -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install azrtydxb/Fastllm-proxy fastllm-gateway --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/azrtydxb/Fastllm-proxy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fastllm-gateway .claude/skills/fastllm-gateway && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fastllm-gateway
GitHub stars
108
Token cost
~926 tokens
SKILL.md length
383 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…

  • Calling a model through the proxy
  • SKILL.md covers Auth and Traps
  • Calls curl
  • Testing that a model

What it does

Fastllm Gateway is an agent skill from azrtydxb/Fastllm-proxy. Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image generation and edits, and listing available models. Use when calling a model through the proxy, testing that a model or frontend model actually serves, or debugging a 401, 404 or 503 from a client.

Its SKILL.md is about 930 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM API integration, Embeddings and Transcription. It works with OpenAI. The repository describes itself as: The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O… The licence is Apache-2.0.

When your agent uses it

  • Calling a model through the proxy
  • Testing that a model
  • Frontend model actually serves
  • Debugging a 401

Example prompts

  • “/fastllm-gateway”

What it can do on your machine

Read from SKILL.md and the folder at commit 59a47cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fastllm Gateway loads about 926 tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 383 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~926

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from azrtydxb/Fastllm-proxy at commit 59a47cf, republished under its Apache-2.0 licence (© azrtydxb). 383 words, ~926 tokens.

Download SKILL.mdSave it as .claude/skills/fastllm-gateway/SKILL.md (or your agent's skills folder).
name
fastllm-gateway
description
Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image generation and edits, and listing available models. Use when calling a model through the proxy, testing that a model or frontend model actually serves, or debugging a 401, 404 or 503 from a client.

FastLLM gateway

Auth

The gateway takes a principal API key as a bearer token — a different credential from the admin session. Keys are stored SHA-256 hashed and cannot be read back from the database; if you do not have one, you cannot call /v1/*.

bash
curl http://192.168.10.125/v1/chat/completions -H "Authorization: Bearer <key>" \
  -H 'content-type: application/json' -d '{"model":"<name>","messages":[...]}'
<!-- BEGIN GENERATED: endpoints -->
MethodPathSummaryBody fields
POST/v1/audio/speechProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/audio/transcriptionsProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/audio/translationsProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/chat/completionsProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/completionsProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/embeddingsProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/images/editsProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/images/generationsProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/messagesAnthropic Messages API. Translated to a chat completion and served by the ordinary request path, so routing, budgets, rate limits and RBAC apply unchanged—
POST/v1/messages/count_tokensBest-effort input token count for a Messages request. An estimate from the text, answered locally—
GET/v1/modelsModels this key may invoke. Filtered by the caller's grants. Anthropic-shaped when the request carries anthropic-version—
POST/v1/moderationsProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/rerankProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/responsesProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—
POST/v1/scoreProxied to the backend serving model. Forwarded byte-for-byte for an openai backend—

* optional field

<!-- END GENERATED: endpoints -->
Show full SKILL.md (103 more words)Show less

Traps

401 means the gateway is healthy. It reached the proxy and was rejected for credentials. A connection refused or 000 is the failure worth chasing.

404 model_not_found on a frontend model means no viable target, not an unknown name — check the chain resolves to something routable.

An unknown model name is a 404 regardless of permissions, deliberately, so "403 vs 404" cannot be used to probe which models exist.

Reasoning field names differ by backend engine. vLLM emits reasoning; SGLang emits reasoning_content. A client hardcoded to one shows blank reasoning against the other — check both before concluding a model is not thinking.

© azrtydxb, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/fastllm-gateway of azrtydxb/Fastllm-proxy.

Open the folder on GitHubat commit 59a47cf

Compare with similar skills

Fastllm Gateway next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fastllm Gateway compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fastllm Gateway this skillazrtydxb/Fastllm-proxy108—~926Automated safety check: PassApache-2.0
Azure AI Openai Dotnetmicrosoft/skills3.1k5 repos~3.4kAutomated safety check: PassMIT
Xsaimoeru-ai/airi50k1 repos~1.3kAutomated safety check: PassMIT
Local AI App Integrationamd/skills408—~6kAutomated safety check: PassMIT
AI Image GenZJU-REAL/Easel3.4k—~787Automated safety check: NotesApache-2.0
Unified LLM APIPrism-Shadow/penguin-harness2.5k—~6.7kAutomated safety check: PassApache-2.0

Similar skills

  • Azure AI Openai Dotnet

    microsoft/skills

    Official

    Azure OpenAI SDK for .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Xsai

    moeru-ai/airi

    A skill your agent uses when the user is building with xsai or any @xsai/ package, or is evaluating xsAI for a small OpenAI-compatible workflow with text generation, streaming, tool calling…

    50k GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Integrates local AI capabilities into applications using Embeddable Lemonade.

    408 GitHub stars~6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • AI Image Gen

    ZJU-REAL/Easel

    通用 AI 生图:文生图 / 图生图 / 图像变体。当用户说 AI 生图、AI 画图、文生图、图生图、生成图片、生成配图、图像生成、AI 出图、AI 作图、换图、改图、图像编辑、给我画一张、生成一张图 时使用。支持 OpenAI 兼容 API 与 apimart 异步 API,用户自备 API key。

    3.4k GitHub stars~787 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • Unified LLM API

    Prism-Shadow/penguin-harness

    Call model APIs through @prismshadow/mmsp (MMSP) — streaming text generation, image generation, speech synthesis, embeddings and the supported-model registry with one client.

    2.5k GitHub stars~6.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Talking Avatar Voice Chat App

    buildfastwithai/gen-ai-experiments

    Builds a realtime voice-chat app around a talking character portrait made from your photo or a text description, with mouth sprites driven by the audio.

    785 GitHub stars~1.7k tokensUpdated 18 days ago
    AI & LLM EngineeringAuto-check passed

More from azrtydxb/Fastllm-proxy

All 14 skills in this repo
  • Fastllm Agents

    azrtydxb/Fastllm-proxy

    Manage and invoke A2A agents behind FastLLM — register, patch, delete and list agents on the control plane, list them through the gateway, fetch an agent card, and invoke an agent by name.

    108 GitHub stars~512 tokensUpdated today
    Auto-check passed
  • Fastllm Backends

    azrtydxb/Fastllm-proxy

    Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

    108 GitHub stars~980 tokensUpdated today
    Auto-check passed
  • Fastllm Classifier

    azrtydxb/Fastllm-proxy

    Manage FastLLM prompt classes for semantic routing — create classes and their example prompts, list or delete them, and evaluate how a given prompt would be classified.

    108 GitHub stars~499 tokensUpdated today
    Auto-check passed
  • Fastllm Deployment

    azrtydxb/Fastllm-proxy

    Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and…

    108 GitHub stars~727 tokensUpdated today
    Auto-check passed
  • Fastllm MCP

    azrtydxb/Fastllm-proxy

    Manage and use MCP servers behind FastLLM — register, patch, delete and list MCP servers on the control plane, and list or call their tools through the gateway.

    108 GitHub stars~515 tokensUpdated today
    Auto-check passed
  • Fastllm Models

    azrtydxb/Fastllm-proxy

    Register and maintain the models FastLLM can serve — create, patch or delete a model, attach backends to it, remove a backend, and set the deployment-wide fallback model.

    108 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Works with

Questions about Fastllm Gateway

What does Fastllm Gateway do?

Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…. Fastllm Gateway is an agent skill from azrtydxb/Fastllm-proxy. Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image generation and edits, and listing available models.

When should I use Fastllm Gateway?

Fastllm Gateway fits situations like: calling a model through the proxy; testing that a model; frontend model actually serves; debugging a 401.

How do I install Fastllm Gateway in Claude Code?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-gateway -a claude-code`. Or copy the skill folder (.claude/skills/fastllm-gateway in azrtydxb/Fastllm-proxy) into .claude/skills/fastllm-gateway in your project. Claude Code loads it when a task matches its description.

How do I install Fastllm Gateway in Codex?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-gateway -a codex`. Or copy the skill folder (.claude/skills/fastllm-gateway in azrtydxb/Fastllm-proxy) into .agents/skills/fastllm-gateway in your project. Codex loads it when a task matches its description.

Can I use Fastllm Gateway in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-gateway -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fastllm-gateway, .gemini/skills/fastllm-gateway, .github/skills/fastllm-gateway and .opencode/skills/fastllm-gateway in your project.

What does Fastllm Gateway need to run?

Going by SKILL.md and its folder, Fastllm Gateway needs the command-line tools its instructions call (curl).

Does Fastllm Gateway access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fastllm Gateway safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fastllm Gateway use?

Fastllm Gateway is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fastllm Gateway use?

About 926 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fastllm Gateway?

Skills that share tags, products or a category with Fastllm Gateway: Azure AI Openai Dotnet (microsoft/skills, 3.1k stars), Xsai (moeru-ai/airi, 50k stars), Local AI App Integration (amd/skills, 408 stars) and AI Image Gen (ZJU-REAL/Easel, 3.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fastllm Gateway?

azrtydxb (a GitHub organization) maintains it in azrtydxb/Fastllm-proxy, which has 108 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 10, 2026.

Source: azrtydxb/Fastllm-proxy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.