Agent skill

Fastllm Deployment

by azrtydxb in azrtydxb/Fastllm-proxy

Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and…

Apache-2.0Auto-check passedDevOps & Cloud

Install Fastllm Deployment

skills CLI
$ npx skills add azrtydxb/Fastllm-proxy --skill fastllm-deployment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install azrtydxb/Fastllm-proxy fastllm-deployment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/azrtydxb/Fastllm-proxy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fastllm-deployment .claude/skills/fastllm-deployment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fastllm-deployment
GitHub stars
108
Token cost
~727 tokens
SKILL.md length
297 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and…

  • A configuration change has not taken effect
  • SKILL.md covers Auth and Traps
  • Calls curl
  • A proxy is serving stale policy

What it does

Fastllm Deployment is an agent skill from azrtydxb/Fastllm-proxy. Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and readiness. Use when a configuration change has not taken effect, when a proxy is serving stale policy, or to confirm what a running instance actually believes.

Its SKILL.md is about 730 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Deployment. It works with OpenAPI. The repository describes itself as: The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O… The licence is Apache-2.0.

When your agent uses it

  • A configuration change has not taken effect
  • A proxy is serving stale policy
  • Confirm what a running instance actually believes

Example prompts

  • “/fastllm-deployment”

What it can do on your machine

Read from SKILL.md and the folder at commit 5d53db8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fastllm Deployment loads about 727 tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 297 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~727

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from azrtydxb/Fastllm-proxy at commit 5d53db8, republished under its Apache-2.0 licence (© azrtydxb). 297 words, ~727 tokens.

Download SKILL.mdSave it as .claude/skills/fastllm-deployment/SKILL.md (or your agent's skills folder).
name
fastllm-deployment
description
Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and readiness. Use when a configuration change has not taken effect, when a proxy is serving stale policy, or to confirm what a running instance actually believes.

FastLLM deployment

Auth

Admin endpoints need a session cookie, not a bearer token — the gateway master key is not an admin credential.

bash
curl -sk -c /tmp/ck -X POST https://192.168.10.129:4001/login \
  -H 'content-type: application/json' -d '{"name":"<user>","password":"<pw>"}'
curl -sk -b /tmp/ck https://192.168.10.129:4001/admin/...
<!-- BEGIN GENERATED: endpoints -->
MethodPathSummaryBody fields
GET/admin/configWhat this process was started with, and who is asking—
GET/admin/deploymentThe FastllmProxy resource running this deployment, and its status—
PATCH/admin/deploymentChange the shape of this deployment: image, replicas, policy, autoscaling—
POST/admin/snapshot/rebuildRebuild and republish the snapshot now—
GET/docsSwagger UI over the spec. The page pulls its bundle from a CDN, so an air-gapped deployment gets an empty page while /openapi.json still works—
GET/healthPer-backend health, in-flight and error counts. Unauthenticated - exposes backend addresses—
POST/health-reportPer-replica backend health. Proxy token only—
GET/healthzLiveness. Unauthenticated and does no database work, so probing it often costs nothing—
GET/openapi.jsonThis document. Unauthenticated: a spec you need a session to read is one nobody generates a client from—
GET/snapshotThe flattened routing table a proxy replica polls. Proxy token only - returns decrypted upstream credentials—
POST/usageBatched usage reporting from a proxy replica. Proxy token onlyevents

* optional field

<!-- END GENERATED: endpoints -->

Traps

Config changes reach proxies on their snapshot poll, not instantly. If a change appears in /admin/* but not in behaviour, check the proxy actually refreshed before assuming the change was wrong.

/health 503 on the gateway means no usable snapshot, not a dead process. A 401 from the gateway means it is healthy and rejecting an unauthenticated request — that is a successful smoke test, not a failure.

AppState::apply_snapshot is the single write path, and it rebuilds the routing registry in the same call so the two cannot diverge. Never write the snapshot cell directly.

The request path performs no I/O. tests/no_io_on_hot_path.rs guards this and must be extended whenever new work lands there.

© azrtydxb, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/fastllm-deployment of azrtydxb/Fastllm-proxy.

Open the folder on GitHubat commit 5d53db8

Compare with similar skills

Fastllm Deployment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fastllm Deployment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fastllm Deployment this skillazrtydxb/Fastllm-proxy108—~727Automated safety check: PassApache-2.0
Tianji Worker Operationsmsgbyte/tianji3.1k—~1.1kAutomated safety check: PassApache-2.0
Databricks Model Servingdatabricks/databricks-agent-skills3451 repos~3.4kAutomated safety check: PassCustom licence
C4 Containeraiskillstore/marketplace4337 repos~1.4kAutomated safety check: PassNone
Azure AI Projects Pyaiskillstore/marketplace4334 repos~2.1kAutomated safety check: PassNone
Attio Upgrade Migrationjeremylongshore/tons-of-skills-marketplace2.8k—~1kAutomated safety check: PassMIT

Similar skills

  • Operates Tianji Workers: create, test, deploy, invoke, schedule, pause and roll back them, plus manage their environment variables and shared modules.

    3.1k GitHub stars~1.1k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Databricks Model Serving

    databricks/databricks-agent-skills

    Official

    Databricks Model Serving endpoint lifecycle and ops. An agent skill from databricks/databricks-agent-skills.

    345 GitHub starsUsed in 1 repo~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • C4 Container

    aiskillstore/marketplace

    Expert C4 Container-level documentation specialist. An agent skill from aiskillstore/marketplace.

    433 GitHub starsUsed in 7 repos~1.4k tokens
    Backend & APIsAuto-check passed
  • Azure AI Projects Py

    aiskillstore/marketplace

    Build AI applications on Microsoft Foundry using the azure-ai-projects SDK.

    433 GitHub starsUsed in 4 repos~2.1k tokens
    DevOps & CloudAuto-check passed
  • Attio Upgrade Migration

    jeremylongshore/tons-of-skills-marketplace

    Migrate an Attio integration through a contract-led inventory, OpenAPI and documentation diff, shadow validation, canary rollout, reconciliation, and tested rollback.

    2.8k GitHub stars~1k tokensUpdated today
    Backend & APIsAuto-check passed
  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed

More from azrtydxb/Fastllm-proxy

All 14 skills in this repo
  • Fastllm Agents

    azrtydxb/Fastllm-proxy

    Manage and invoke A2A agents behind FastLLM — register, patch, delete and list agents on the control plane, list them through the gateway, fetch an agent card, and invoke an agent by name.

    108 GitHub stars~512 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Backends

    azrtydxb/Fastllm-proxy

    Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

    108 GitHub stars~980 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Classifier

    azrtydxb/Fastllm-proxy

    Manage FastLLM prompt classes for semantic routing — create classes and their example prompts, list or delete them, and evaluate how a given prompt would be classified.

    108 GitHub stars~499 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Gateway

    azrtydxb/Fastllm-proxy

    Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…

    108 GitHub stars~926 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm MCP

    azrtydxb/Fastllm-proxy

    Manage and use MCP servers behind FastLLM — register, patch, delete and list MCP servers on the control plane, and list or call their tools through the gateway.

    108 GitHub stars~515 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Models

    azrtydxb/Fastllm-proxy

    Register and maintain the models FastLLM can serve — create, patch or delete a model, attach backends to it, remove a backend, and set the deployment-wide fallback model.

    108 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Fastllm Deployment

What does Fastllm Deployment do?

Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and…. Fastllm Deployment is an agent skill from azrtydxb/Fastllm-proxy. Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and readiness.

When should I use Fastllm Deployment?

Fastllm Deployment fits situations like: A configuration change has not taken effect; A proxy is serving stale policy; confirm what a running instance actually believes.

How do I install Fastllm Deployment in Claude Code?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-deployment -a claude-code`. Or copy the skill folder (.claude/skills/fastllm-deployment in azrtydxb/Fastllm-proxy) into .claude/skills/fastllm-deployment in your project. Claude Code loads it when a task matches its description.

How do I install Fastllm Deployment in Codex?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-deployment -a codex`. Or copy the skill folder (.claude/skills/fastllm-deployment in azrtydxb/Fastllm-proxy) into .agents/skills/fastllm-deployment in your project. Codex loads it when a task matches its description.

Can I use Fastllm Deployment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-deployment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fastllm-deployment, .gemini/skills/fastllm-deployment, .github/skills/fastllm-deployment and .opencode/skills/fastllm-deployment in your project.

What does Fastllm Deployment need to run?

Going by SKILL.md and its folder, Fastllm Deployment needs the command-line tools its instructions call (curl).

Does Fastllm Deployment access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fastllm Deployment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fastllm Deployment use?

Fastllm Deployment is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fastllm Deployment use?

About 727 tokens (SKILL.md is roughly 2.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fastllm Deployment?

Skills that share tags, products or a category with Fastllm Deployment: Tianji Worker Operations (msgbyte/tianji, 3.1k stars), Databricks Model Serving (databricks/databricks-agent-skills, 345 stars), C4 Container (aiskillstore/marketplace, 433 stars) and Azure AI Projects Py (aiskillstore/marketplace, 433 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fastllm Deployment?

azrtydxb (a GitHub organization) maintains it in azrtydxb/Fastllm-proxy, which has 108 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 5, 2026.

Source: azrtydxb/Fastllm-proxy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.