Agent skill

Fastllm Backends

by azrtydxb in azrtydxb/Fastllm-proxy

Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Fastllm Backends

skills CLI
$ npx skills add azrtydxb/Fastllm-proxy --skill fastllm-backends -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install azrtydxb/Fastllm-proxy fastllm-backends --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/azrtydxb/Fastllm-proxy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fastllm-backends .claude/skills/fastllm-backends && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fastllm-backends
GitHub stars
108
Token cost
~980 tokens
SKILL.md length
471 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

  • A Spark node is unresponsive
  • SKILL.md covers The memory rule, sparkrun, Diagnosing bad output and Speculative decoding
  • Calls curl
  • A model endpoint is down

What it does

Fastllm Backends is an agent skill from azrtydxb/Fastllm-proxy. Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding settings for GB10, and diagnosing a model that will not load, returns empty responses, or serves corrupted output. Use when a Spark node is unresponsive, a model endpoint is down, throughput looks wrong, or a container will not start. Not for FastLLM's own routing or API (fastllm-routing, fastllm-operate).

Its SKILL.md is about 980 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and SGLang. The repository describes itself as: The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O… The licence is Apache-2.0.

When your agent uses it

  • A Spark node is unresponsive
  • A model endpoint is down
  • Throughput looks wrong
  • A container will not start

Example prompts

  • “/fastllm-backends”

What it can do on your machine

Read from SKILL.md and the folder at commit 5d53db8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fastllm Backends loads about 980 tokens when it runs. Until then it costs about 129 tokens; SKILL.md has 471 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~129
When it runs · the whole SKILL.md, loaded when a task matches
~980

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from azrtydxb/Fastllm-proxy at commit 5d53db8, republished under its Apache-2.0 licence (© azrtydxb). 471 words, ~980 tokens.

Download SKILL.mdSave it as .claude/skills/fastllm-backends/SKILL.md (or your agent's skills folder).
name
fastllm-backends
description
Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding settings for GB10, and diagnosing a model that will not load, returns empty responses, or serves corrupted output. Use when a Spark node is unresponsive, a model endpoint is down, throughput looks wrong, or a container will not start. Not for FastLLM's own routing or API (fastllm-routing, fastllm-operate).

DGX Spark inference backends

Two nodes: dgx-spark (192.168.10.246) and dgx-spark2 (192.168.10.245), GB10 / SM121 / aarch64, 121 GB unified memory shared between CPU and GPU, joined by a direct ConnectX-7 link on 192.168.10.47.x (enp1s0f0np0; the f1 ports are down on this pair).

The memory rule

--gpu-memory-utilization 0.45 is load-bearing, not conservative. On unified memory the GPU budget and the OS share one pool. Raising it to 0.80 made a node answer ping while SSH and the API both timed out — the box thrashes and needs a power cycle. Measured: at 0.45, 72 GB used / 48 GB available while serving.

Symptom to recognise: ping works, SSH hangs at banner exchange. That is memory starvation, not a network fault.

sparkrun

  • sparkrun stop <id> without --cluster silently fails — it loses the cluster's SSH user, prints Permission denied and Workload stopped in the same output, and the workload keeps running. Always pass --cluster.
  • Upstream images need executor_config: {entrypoint: ""} in the recipe; sparkrun launches containers as <image> bash -c "<bootstrap>" and an image whose ENTRYPOINT is ["vllm","serve"] parses that as model arguments.
  • Prefix the serve command with cd /cache/huggingface && — the image's WORKDIR belongs to its own user and multiprocessing children die on os.chdir.
  • Redirect HOME, VLLM_CACHE_ROOT and TRITON_CACHE_DIR into the mounted HF cache, or compile caches are unwritable and are lost every start.
  • Node count derives from tp * pp, so a tp=1 job runs _solo on one host regardless of min_nodes. Data parallelism means launching one job per node plus a router.
  • sparkrun run can hang indefinitely on a half-closed Hugging Face connection, printing nothing under --no-follow. Kill and retry.
Show full SKILL.md (208 more words)Show less

Diagnosing bad output

Separate the model from the plumbing before blaming either. /v1/completions bypasses both the chat template and the reasoning parser:

bash
curl -s http://<node>:8000/v1/completions -H 'content-type: application/json' \
  -d '{"model":"<m>","prompt":"<|im_start|>user\nhi<|im_end|>\n<|im_start|>assistant\n<think>\n","max_tokens":200}'

If raw output is good but chat completions are empty, the parser or template is eating it — not the model. A chat response with finish_reason: length, zero content and zero reasoning while completion_tokens is at the cap means an unclosed thinking block.

Reasoning field names differ by engine: vLLM emits reasoning, SGLang emits reasoning_content. Reading the wrong key reports zero on a field that does not exist — check both before concluding a model is not thinking.

A chat template that pre-fills an open <think> tag must be paired with a parser that splits on the closing tag alone. Pairing it with one that expects an opening tag in the output discards everything.

Speculative decoding

Throughput is accept_length × steps/s. Both matter, and acceptance is dominated by content, not by configuration — measured 0.19 on prose with reasoning versus 0.95 on code, on the same server and config. Quote a tok/s figure only alongside the workload it was measured on.

Draft depth is a property of the drafter, not a tuning knob: a block drafter trained at K=7 loses accuracy when run at K=5, rather than simply drafting less.

© azrtydxb, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/fastllm-backends of azrtydxb/Fastllm-proxy.

Open the folder on GitHubat commit 5d53db8

Compare with similar skills

Fastllm Backends next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fastllm Backends compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fastllm Backends this skillazrtydxb/Fastllm-proxy108—~980Automated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Debug InferenceNVIDIA/OpenShell16k—~1.9kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
One EvalOpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Debug Inference

    NVIDIA/OpenShell

    Official

    Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

    16k GitHub stars~1.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • One Eval

    OpenDCAI/One-Eval

    驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

    165 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

More from azrtydxb/Fastllm-proxy

All 14 skills in this repo
  • Fastllm Agents

    azrtydxb/Fastllm-proxy

    Manage and invoke A2A agents behind FastLLM — register, patch, delete and list agents on the control plane, list them through the gateway, fetch an agent card, and invoke an agent by name.

    108 GitHub stars~512 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Classifier

    azrtydxb/Fastllm-proxy

    Manage FastLLM prompt classes for semantic routing — create classes and their example prompts, list or delete them, and evaluate how a given prompt would be classified.

    108 GitHub stars~499 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Deployment

    azrtydxb/Fastllm-proxy

    Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and…

    108 GitHub stars~727 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Gateway

    azrtydxb/Fastllm-proxy

    Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…

    108 GitHub stars~926 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm MCP

    azrtydxb/Fastllm-proxy

    Manage and use MCP servers behind FastLLM — register, patch, delete and list MCP servers on the control plane, and list or call their tools through the gateway.

    108 GitHub stars~515 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Models

    azrtydxb/Fastllm-proxy

    Register and maintain the models FastLLM can serve — create, patch or delete a model, attach backends to it, remove a backend, and set the deployment-wide fallback model.

    108 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Fastllm Backends

What does Fastllm Backends do?

Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…. Fastllm Backends is an agent skill from azrtydxb/Fastllm-proxy. Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding settings for GB10, and diagnosing a model that will not load, returns empty responses, or serves corrupted output.

When should I use Fastllm Backends?

Fastllm Backends fits situations like: A Spark node is unresponsive; A model endpoint is down; throughput looks wrong; A container will not start.

How do I install Fastllm Backends in Claude Code?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-backends -a claude-code`. Or copy the skill folder (.claude/skills/fastllm-backends in azrtydxb/Fastllm-proxy) into .claude/skills/fastllm-backends in your project. Claude Code loads it when a task matches its description.

How do I install Fastllm Backends in Codex?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-backends -a codex`. Or copy the skill folder (.claude/skills/fastllm-backends in azrtydxb/Fastllm-proxy) into .agents/skills/fastllm-backends in your project. Codex loads it when a task matches its description.

Can I use Fastllm Backends in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-backends -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fastllm-backends, .gemini/skills/fastllm-backends, .github/skills/fastllm-backends and .opencode/skills/fastllm-backends in your project.

What does Fastllm Backends need to run?

Going by SKILL.md and its folder, Fastllm Backends needs the command-line tools its instructions call (curl).

Does Fastllm Backends access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fastllm Backends safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fastllm Backends use?

Fastllm Backends is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fastllm Backends use?

About 980 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fastllm Backends?

Skills that share tags, products or a category with Fastllm Backends: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Debug Inference (NVIDIA/OpenShell, 16k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fastllm Backends?

azrtydxb (a GitHub organization) maintains it in azrtydxb/Fastllm-proxy, which has 108 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 5, 2026.

Source: azrtydxb/Fastllm-proxy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.