Agent skill

Fastllm Observability

by azrtydxb in azrtydxb/Fastllm-proxy

Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status.

Apache-2.0Auto-check passedDevOps & Cloud

Install Fastllm Observability

skills CLI
$ npx skills add azrtydxb/Fastllm-proxy --skill fastllm-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install azrtydxb/Fastllm-proxy fastllm-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/azrtydxb/Fastllm-proxy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/fastllm-observability .claude/skills/fastllm-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fastllm-observability
GitHub stars
108
Token cost
~916 tokens
SKILL.md length
442 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status.

  • Asked how much a caller spent
  • SKILL.md covers Auth and Facts worth knowing
  • Calls curl
  • What changed and who changed it

What it does

Fastllm Observability is an agent skill from azrtydxb/Fastllm-proxy. Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status. Use when asked how much a caller spent, what changed and who changed it, whether a backend is healthy, or to investigate an error-rate or latency question.

Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Observability, Monitoring and alerting and Forecasting and time series. It works with Prometheus. The repository describes itself as: The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O… The licence is Apache-2.0.

When your agent uses it

  • Asked how much a caller spent
  • What changed and who changed it
  • Whether a backend is healthy
  • Investigate an error-rate

Example prompts

  • “/fastllm-observability”

What it can do on your machine

Read from SKILL.md and the folder at commit 5d53db8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fastllm Observability loads about 916 tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 442 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~916

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from azrtydxb/Fastllm-proxy at commit 5d53db8, republished under its Apache-2.0 licence (© azrtydxb). 442 words, ~916 tokens.

Download SKILL.mdSave it as .claude/skills/fastllm-observability/SKILL.md (or your agent's skills folder).
name
fastllm-observability
description
Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status. Use when asked how much a caller spent, what changed and who changed it, whether a backend is healthy, or to investigate an error-rate or latency question.

FastLLM observability

Auth

Admin endpoints need a session cookie, not a bearer token — the gateway master key is not an admin credential.

bash
curl -sk -c /tmp/ck -X POST https://192.168.10.129:4001/login \
  -H 'content-type: application/json' -d '{"name":"<user>","password":"<pw>"}'
curl -sk -b /tmp/ck https://192.168.10.129:4001/admin/...
<!-- BEGIN GENERATED: endpoints -->
MethodPathSummaryBody fields
GET/admin/auditThe change log, newest first, keyset-paginated—
GET/admin/fleetWhat each proxy replica reports, kept per replica and never merged—
GET/admin/healthRead health—
GET/admin/nodesThe hosts registering their own endpoints, rolled up per node. An agent is not a row: it is a node several dynamic providers share, and its lease is what says it is alive—
GET/admin/timeseriesBucketed traffic, latency and spend. Empty buckets come back as explicit zeros; latency is null where there was nothing to measure—
GET/admin/usageAggregate usage and spend, grouped by model, principal, frontend model or day—
GET/metricsPrometheus text. Unauthenticated—

* optional field

<!-- END GENERATED: endpoints -->

Facts worth knowing

The audit trail is middleware, not hand-wired. Everything that is not a GET under /admin/* passes through it, so a newly added endpoint is audited before it is written. GETs are deliberately not audited — auditing reads would bury the changes in noise.

Direct database writes produce no audit row. If a change was made with psql because no admin credential was available, the audit trail will not show it; say so explicitly rather than letting the absence imply nothing happened.

A usage row exists for every attributable request, including ones whose response carried no token counts. usage_reported distinguishes "consumed nothing" from "counts unknown" — treating them alike understates consumption.

/admin/fleet never averages replicas together. Every replica losing a backend is a dead backend; one replica losing it is a partition.

Show full SKILL.md (179 more words)Show less

snapshot_version spread is not a measure of staleness. A version is the microsecond the control plane built that snapshot, and it republishes only when the content changed — so the gap between two consecutive versions is the time between two real config changes, not any replica's lag. A replica one version behind can show a gap of a second or of a minute at identical health.

Judge it by time instead. Proxies poll every config_poll_seconds and report health every health_report_interval_seconds (both in GET /admin/config), so a replica is stuck only if the newest snapshot has been available for longer than their sum — or if it has been behind across several samples spanning that long. The second test is the one that works on a busy gateway, where Budget.tokens_used being part of the snapshot means traffic alone republishes it every few seconds and nothing is ever old. One GET /admin/fleet cannot distinguish a stuck replica from a converging one there; take a few, spaced.

Reading any spread at all as a fault reports a healthy fleet as split after every change.

© azrtydxb, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/fastllm-observability of azrtydxb/Fastllm-proxy.

Open the folder on GitHubat commit 5d53db8

Compare with similar skills

Fastllm Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fastllm Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fastllm Observability this skillazrtydxb/Fastllm-proxy108—~916Automated safety check: PassApache-2.0
Happy Infra Metrics and Grafanaslopus/happy24k—~2kAutomated safety check: NotesMIT
WizTelemetry Platform Servicekubesphere/kubesphere17k—~1.8kAutomated safety check: PassCustom licence
Redis Observabilityredis/agent-skills1662 repos~911Automated safety check: PassMIT
Developing Funboost Mixinydf0509/funboost895—~2.1kAutomated safety check: PassNone
Prometheus System Health Checkprometheus/prometheus-mcp121—~584Automated safety check: PassApache-2.0

Similar skills

  • Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API.

    24k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • WizTelemetry Platform Service

    kubesphere/kubesphere

    Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions.

    17k GitHub stars~1.8k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Redis Observability

    redis/agent-skills

    Official

    Redis observability guidance — which metrics to monitor (memory, connections, hit ratio, ops/sec, rejected connections), which built-in commands to reach for during incident triage (SLOWLOG, INFO…

    166 GitHub starsUsed in 2 repos~911 tokens
    DevOps & CloudAuto-check passed
  • 当需要为 funboost 创建 Consumer 或 Publisher 的 Mixin 扩展类时使用。触发场景:添加监控、熔断、限流、链路追踪等横切关注点,编写自定义前置/后置处理钩子。关键词:mixin, consumeroverridecls, publisheroverridecls, ConsumerMixin, 自定义消费者, hook, 拦截器, 熔断器, 监控…

    895 GitHub stars~2.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Prometheus System Health Check

    prometheus/prometheus-mcp

    Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load.

    121 GitHub stars~584 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed

More from azrtydxb/Fastllm-proxy

All 14 skills in this repo
  • Fastllm Agents

    azrtydxb/Fastllm-proxy

    Manage and invoke A2A agents behind FastLLM — register, patch, delete and list agents on the control plane, list them through the gateway, fetch an agent card, and invoke an agent by name.

    108 GitHub stars~512 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Backends

    azrtydxb/Fastllm-proxy

    Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

    108 GitHub stars~980 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Classifier

    azrtydxb/Fastllm-proxy

    Manage FastLLM prompt classes for semantic routing — create classes and their example prompts, list or delete them, and evaluate how a given prompt would be classified.

    108 GitHub stars~499 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Deployment

    azrtydxb/Fastllm-proxy

    Inspect and control the running FastLLM deployment — read effective configuration and deployment settings, force a snapshot rebuild, fetch the snapshot the proxies consume, and check liveness and…

    108 GitHub stars~727 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm Gateway

    azrtydxb/Fastllm-proxy

    Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…

    108 GitHub stars~926 tokensUpdated 4 days ago
    Auto-check passed
  • Fastllm MCP

    azrtydxb/Fastllm-proxy

    Manage and use MCP servers behind FastLLM — register, patch, delete and list MCP servers on the control plane, and list or call their tools through the gateway.

    108 GitHub stars~515 tokensUpdated 4 days ago
    Auto-check passed

Works with

Categories

Questions about Fastllm Observability

What does Fastllm Observability do?

Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status. Fastllm Observability is an agent skill from azrtydxb/Fastllm-proxy. Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status.

When should I use Fastllm Observability?

Fastllm Observability fits situations like: asked how much a caller spent; what changed and who changed it; whether a backend is healthy; investigate an error-rate.

How do I install Fastllm Observability in Claude Code?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-observability -a claude-code`. Or copy the skill folder (.claude/skills/fastllm-observability in azrtydxb/Fastllm-proxy) into .claude/skills/fastllm-observability in your project. Claude Code loads it when a task matches its description.

How do I install Fastllm Observability in Codex?

Run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-observability -a codex`. Or copy the skill folder (.claude/skills/fastllm-observability in azrtydxb/Fastllm-proxy) into .agents/skills/fastllm-observability in your project. Codex loads it when a task matches its description.

Can I use Fastllm Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add azrtydxb/Fastllm-proxy --skill fastllm-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fastllm-observability, .gemini/skills/fastllm-observability, .github/skills/fastllm-observability and .opencode/skills/fastllm-observability in your project.

What does Fastllm Observability need to run?

Going by SKILL.md and its folder, Fastllm Observability needs the command-line tools its instructions call (curl).

Does Fastllm Observability access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Fastllm Observability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fastllm Observability use?

Fastllm Observability is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fastllm Observability use?

About 916 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fastllm Observability?

Skills that share tags, products or a category with Fastllm Observability: Happy Infra Metrics and Grafana (slopus/happy, 24k stars), WizTelemetry Platform Service (kubesphere/kubesphere, 17k stars), Redis Observability (redis/agent-skills, 166 stars) and Developing Funboost Mixin (ydf0509/funboost, 895 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fastllm Observability?

azrtydxb (a GitHub organization) maintains it in azrtydxb/Fastllm-proxy, which has 108 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 5, 2026.

Source: azrtydxb/Fastllm-proxy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.