Release It
wondelai/skills
Build production-ready systems with stability patterns: circuit breakers, bulkheads, timeouts, and retry logic.
Agent skill
by rednote-machine-learning in rednote-machine-learning/Inference-autopilot
Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.
$ npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install rednote-machine-learning/Inference-autopilot inference-autopilot --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "inference-autopilot" agent skill from https://github.com/rednote-machine-learning/Inference-autopilot/tree/main into .claude/skills/inference-autopilot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "inference-autopilot", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install rednote-machine-learning/Inference-autopilot inference-autopilot --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "inference-autopilot" agent skill from https://github.com/rednote-machine-learning/Inference-autopilot/tree/main into .agents/skills/inference-autopilot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "inference-autopilot", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install rednote-machine-learning/Inference-autopilot inference-autopilot --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "inference-autopilot" agent skill from https://github.com/rednote-machine-learning/Inference-autopilot/tree/main into .cursor/skills/inference-autopilot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "inference-autopilot", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install rednote-machine-learning/Inference-autopilot inference-autopilot --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "inference-autopilot" agent skill from https://github.com/rednote-machine-learning/Inference-autopilot/tree/main into .gemini/skills/inference-autopilot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "inference-autopilot", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install rednote-machine-learning/Inference-autopilot inference-autopilotInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "inference-autopilot" agent skill from https://github.com/rednote-machine-learning/Inference-autopilot/tree/main into .github/skills/inference-autopilot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "inference-autopilot", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install rednote-machine-learning/Inference-autopilot inference-autopilot --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "inference-autopilot" agent skill from https://github.com/rednote-machine-learning/Inference-autopilot/tree/main into .opencode/skills/inference-autopilot/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "inference-autopilot", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
inference-autopilotAnalyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.
Inference Autopilot is an agent skill from rednote-machine-learning/Inference-autopilot. Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs. Use when Codex needs to tune SGLang launch parameters, run bounded single-host GPU experiments, plan deployment topology, inspect GPU or CPU profiles, identify scheduler/KV/communication/kernel bottlenecks, propose operator optimizations, or validate that a candidate configuration improves performance without correctness or SLO regressions.
Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 51 other files, including scripts, reference files and assets (for example `.github/workflows/sglang-parameter-compat.yml`, `.github/workflows/version-consistency.yml` and `PARAMETER_EVOLUTION.md`).
It sits in DevOps & Cloud, covering Site reliability engineering and Deployment. It works with SGLang. The repository describes itself as: Evidence-driven SGLang inference deployment optimization. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 61eb1c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Inference Autopilot loads about 4.5k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 132 tokens; SKILL.md has 1,968 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from rednote-machine-learning/Inference-autopilot at commit 61eb1c0, republished under its Apache-2.0 licence (© rednote-machine-learning). 1,968 words, ~4,539 tokens.
.claude/skills/inference-autopilot/SKILL.md (or your agent's skills folder). This skill also uses 47 other files; get the full folder from GitHub.Run an evidence-driven optimization loop. Treat this skill as the control plane and the bundled scripts as deterministic utilities. Default to local, private, dry-run behavior.
For a new private single-host deployment, start with the one-shot interface. The
user supplies only local paths, workload, SLO, objective, budget, and optional
GPU visibility. The script discovers NVIDIA or AMD GPU memory and topology,
reads the local model config and weight sizes, regenerates the current SGLang
parameter contract from server_args.py and sglang.launch_server --help, captures a bounded Nsight Systems baseline, routes trace and workload
evidence to parameter families, runs bounded screening, applies mode-specific
confirmation, and emits a deployable command only after the confirmation gates
pass:
inferopt init --output /absolute/private/task.json
inferopt doctor --task /absolute/private/task.json --output /absolute/private/doctor.json
inferopt plan --task /absolute/private/task.json --output /absolute/private/plan.json
inferopt run --task /absolute/private/task.json --yes --output /absolute/private/final.json
inferopt report --result /absolute/private/final.json --output /absolute/private/report.mdinferopt run only detects and reports missing fused MoE tuning configs. Never
run the high-cost kernel search as part of the normal workflow. If the user
explicitly requests it after reviewing the report, use the standalone command
emitted by the report:
inferopt tune-moe --task RUN_DIR/task.json --profile RUN_DIR/profile/nsys-diagnosis.json --result RUN_DIR/final.json --output-dir RUN_DIR/optional-fused-moe-tuning --yes --output RUN_DIR/optional-fused-moe-tuning.jsonTreat generated configs as non-deployable until a separate end-to-end A/B
benchmark passes the original workload, SLO, noise, and confirmation gates.
Generated files must remain under configs/triton_<version>/ and use the
installed SGLang naming contract (E, per-TP N, runtime GPU name, dtype,
optional block shape/per-channel marker) with a JSON mapping from token batch
size M to kernel tile parameters. If the log requests an _down.json file,
the normal tuner is only partial: use SGLang's separate top-k capture/tuner
workflow and pass --topk-ids-dir. Inference Autopilot refuses to authorize an
up-only result in that case; both matching up and _down files are required.
For direct generation of paired files without end-to-end validation, use:
inferopt generate-moe-config \
--repository /path/to/sglang \
--model-path /path/to/model \
--tp-size 4 --dtype fp8_w8a8 \
--topk-ids-dir /path/to/topk_ids \
--output-dir /path/to/moe-config-output --yesThis invokes the official separate tuner and refuses to create a fake _down.json without routing captures. Use --mode standard only when SGLang did not request a down-kernel config.
The standalone inferopt command does not require Codex. doctor never starts
a server: it validates the local model, GPU/runtime, current SGLang CLI, profiler
availability, and a transparent single-GPU memory estimate. It may recommend a
quantized checkpoint class when the current model cannot fit. It also reports
legal TP sizes across the complete selected single-host GPU set and treats a
missing nsys installation as a blocking preflight error. It never downloads,
switches, or approves a model variant without an explicit quality evaluation.
The underlying scripts remain available for controlled integration:
cp assets/task.autopilot.example.json /absolute/private/task.json
python3 scripts/autopilot.py validate --task /absolute/private/task.json
python3 scripts/autopilot.py plan --task /absolute/private/task.json --output /absolute/private/plan.json
python3 scripts/autopilot.py run --task /absolute/private/task.json --yes --output /absolute/private/final-pointer.jsonRead references/hardware-profiles.json for the official-source GPU capability catalog. Runtime discovery overrides catalog values. Preserve the checked-out SGLang version's automatic defaults as the baseline; never transplant a winning parameter from one GPU or workload to another without measurement.
scripts/inferopt.py validate --spec <task.json>.scripts/inferopt.py inventory --output <run-dir>/inventory.json on every target host when access is available.For authorized single-host execution, copy assets/task.execute.example.json, read references/execution-schema.md, validate it, render the exact command plan, obtain explicit approval, and only then run:
python3 scripts/autotune.py validate --spec task.json
python3 scripts/autotune.py plan --spec task.json --output plan.json
python3 scripts/autotune.py run --spec task.json --yes --output final-pointer.json
python3 scripts/autotune.py report --run-dir /absolute/completed-run --output decision.jsonIf the framework is SGLang, read references/sglang-adapter.md. For another engine, discover its actual launch, benchmark, metrics, and profiling interfaces instead of assuming SGLang flags.
dry_run. Require the user or task specification to opt into execution.--yes as confirmation that the generated command plan was reviewed; never add it before approval.Read references/safety-policy.md before executing benchmarks, profilers, or code changes.
Set deployment_mode in the autopilot task. Use online_latency for an
interactive service: maximize the selected objective while every declared E2E,
TTFT, TPOT, or ITL gate passes. When no latency SLO is declared, the configured
secondary-regression limit still protects observed latency. Use
offline_throughput for batch inference: maximize sustained
aggregate throughput at calibrated batch pressure; latency remains recorded
but is only a gate when the task explicitly declares it.
For offline workloads without SLOs, skip capacity calibration, remove the
client concurrency cap, and let the resolved SGLang admission policy determine
sustained pressure. Start the baseline service exactly twice: one bounded nsys
capture and one unprofiled benchmark. Preserve that benchmark as an immutable
reference for all candidates; never search or override max_running_requests.
For an SLO-constrained workload, control load with the benchmark client's
max_concurrency; treat the server's resolved max_running_requests as
diagnostic evidence rather than a tuning parameter. Persist every benchmark
command and resolved server value.
For a model with a verified official cookbook, first compare complete, locally
valid capability bundles: include the relevant model feature (for example MTP),
prefix/KV cache policy, scheduler/admission, memory pool, CUDA Graph, and MoE
backend variants when the workload and hardware make them applicable. Profile
the fastest SLO-valid initial configuration, even if its single screening
sample has not yet cleared the final improvement threshold. Then use Nsight
evidence to select and refine parameter families, and screen and confirm every
candidate against the original target workload. Never label a result as global
best: report it as the best configuration within the checked-out SGLang
parameter contract, tested search space, hardware, workload, and SLO gates.
Use search_depth: thorough (the default) to add a one-factor sensitivity
screen for every high-impact compatible family even when a short trace lacks a
single hotspot. Use evidence_guided only when experiment budget is tight.
Every run emits search-plan.json.parameter_audit, which accounts for each
CLI-visible, non-deprecated ServerArgs as selected, excluded, or
inapplicable with a concrete reason. It prevents a short candidate list from
being mistaken for the whole startup-parameter surface.
Render per-stage trial progress as an elapsed-time ASCII bar with completed and
planned counts. Count completed, failed, and capability-skipped trials as
processed. Do not invent a percentage for Nsight capture or model startup when
their duration is not knowable in advance, and keep redirected logs free of
terminal control characters.
After the one-factor screen, the optimizer uses the remaining trial budget to
test explicit combinations of independent screened winners. Every candidate
that clears the configured improvement threshold is considered: pairs with
the strongest candidate run first, followed by other pairs and larger
combinations while budget remains. Parameter/environment conflicts and exact
duplicates are excluded. Positive sub-threshold candidates may use leftover
slots but never displace an above-threshold combination. Offline no-SLO runs
reuse the preserved baseline and rerun only the selected candidate once;
SLO-constrained runs retain repeated A/B confirmation. It does not perform an
unbounded Cartesian product.
Extract:
Do not optimize an unspecified scalar. Convert multiple SLOs into hard constraints plus one primary objective, for example: maximize request goodput subject to P99 TTFT and P99 TPOT.
Use scripts/inferopt.py analyze for SGLang-compatible benchmark JSONL and scripts/inferopt.py compare to gate candidates. Use scripts/autotune.py only for an isolated, authorized, single-host SGLang experiment.
Treat a one-run screening_winner as a hypothesis. For SLO-constrained modes require repeated A/B measurements, acceptable objective CV, and every repetition passing SLO. For offline no-SLO mode, capture both the short screening baseline and a cache-flushed confirmation-reference window while the same baseline service is loaded, then rerun only the selected candidate with the exact reference dataset, request count, and duration contract. Never compare the longer final candidate with the short screening baseline.
Use recommended_configuration for the deployment decision. Preserve a confirmed baseline when recommendation_status=retain_confirmed_baseline; an empty winner can mean candidates were correctly rejected, not that the experiment failed.
Read references/diagnosis-playbook.md. First classify the dominant bottleneck:
Do not start kernel work while queueing, cache misses, bad topology, or an unsuitable backend dominates end-to-end time.
Search in this order unless evidence justifies a different order:
Change one conceptual factor per ablation. Use a coarse family screen before
fine tuning. Enumerate every current server_args.py parameter into a family,
and record every compatible parameter as selected, excluded with its reason,
or inapplicable to the current model/topology/deployment mode. Reuse prior trials
and reject infeasible configurations before launching a server.
Use the least expensive tool that can answer the current question. Read references/profiler-routing.md.
Keep profiling windows short and representative. Do not compare profiled latency directly with unprofiled production latency.
Before editing a kernel, record:
Build an isolated correctness and performance benchmark. Cover boundary shapes and realistic distributions, not only one favorable shape. Integrate only after isolated and end-to-end gates pass.
Require:
Deliver:
python3 scripts/inferopt.py validate --spec task.json
python3 scripts/inferopt.py inventory --output runs/inventory.json
python3 scripts/inferopt.py plan --spec task.json --output runs/plan.json
python3 scripts/inferopt.py analyze --input runs/baseline.jsonl --spec task.json --output runs/baseline-summary.json
python3 scripts/inferopt.py compare --baseline runs/baseline-summary.json --candidate runs/trial-001-summary.json --spec task.json
python3 scripts/autotune.py validate --spec task.json
python3 scripts/autotune.py plan --spec task.json --output plan.json
python3 scripts/autotune.py run --spec task.json --yes
python3 scripts/autopilot.py validate --task task.json
python3 scripts/autopilot.py plan --task task.json --output plan.json
python3 scripts/autopilot.py run --task task.json --yes --output final-pointer.jsoninferopt.py never starts a server. autopilot.py starts only the local SGLang
process groups it creates, regenerates an installed-version parameter contract for every run, performs a bounded Nsight Systems CUDA-range capture,
then generates and executes two-stage autotune.py runs using structured
allowlisted parameters. It cannot execute shell snippets, remote SSH, Slurm,
Kubernetes, arbitrary modules, production endpoints, package installation,
source edits, or kernel changes. It records a shape-matched Nsight Compute
follow-up only when one kernel is trace-proven hot; it never edits a kernel.
© rednote-machine-learning, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 47 other files (scripts, references, assets) in the repository root of rednote-machine-learning/Inference-autopilot.
Open the folder on GitHubat commit 61eb1c0
Inference Autopilot next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Inference Autopilot this skillrednote-machine-learning/Inference-autopilot | 142 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | |
| Release Itwondelai/skills | 2.4k | — | ~4k | Automated safety check: Pass | MIT | |
| Testing In Productionpetrkindlmann/qa-skills | 163 | — | ~5.3k | Automated safety check: Pass | MIT | |
| Delivery Managerborghei/Claude-Skills | 874 | — | ~2.2k | Automated safety check: Pass | MIT | |
| Release Engineeringmagnus919/agent-skills | 111 | — | ~3.9k | Automated safety check: Pass | MIT | |
| Platform Operationsrsmdt/the-startup | 536 | — | ~991 | Automated safety check: Pass | MIT |
wondelai/skills
Build production-ready systems with stability patterns: circuit breakers, bulkheads, timeouts, and retry logic.
petrkindlmann/qa-skills
Safe-release techniques DURING rollout: feature flags, progressive rollouts, canary analysis, guardrail metrics, production smoke tests, and synthetic users.
borghei/Claude-Skills
Expert delivery management for release planning, deployment strategy, incident response, change management, SLA/error-budget tracking, and DORA metrics across continuous delivery pipelines.
magnus919/agent-skills
Design, automate, and operate end-to-end software releases: release process models and pipelines (trunk-based development, CD stages, release trains), progressive delivery and feature flags…
rsmdt/the-startup
Unified platform operations guidance for CI/CD pipeline design, deployment strategies, observability, SLI/SLOs, and incident-ready rollouts.
MicrosoftDocs/Agent-Skills
Expert knowledge for Azure Reliability development including best practices, decision making, architecture & design patterns, limits & quotas, and deployment.
Works with
Categories
Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs. Inference Autopilot is an agent skill from rednote-machine-learning/Inference-autopilot. Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.
Inference Autopilot fits situations like: Codex needs to tune SGLang launch parameters; run bounded single-host GPU experiments; plan deployment topology; identify scheduler/KV/communication/kernel bottlenecks.
Run `npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a claude-code`. Or copy the skill folder (the rednote-machine-learning/Inference-autopilot repository) into .claude/skills/inference-autopilot in your project. Claude Code loads it when a task matches its description.
Run `npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a codex`. Or copy the skill folder (the rednote-machine-learning/Inference-autopilot repository) into .agents/skills/inference-autopilot in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/inference-autopilot, .gemini/skills/inference-autopilot, .github/skills/inference-autopilot and .opencode/skills/inference-autopilot in your project.
Going by SKILL.md and its folder, Inference Autopilot needs the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Inference Autopilot is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Inference Autopilot: Release It (wondelai/skills, 2.4k stars), Testing In Production (petrkindlmann/qa-skills, 163 stars), Delivery Manager (borghei/Claude-Skills, 874 stars) and Release Engineering (magnus919/agent-skills, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
rednote-machine-learning (a GitHub organization) maintains it in rednote-machine-learning/Inference-autopilot, which has 142 GitHub stars. The repository was last updated on September 7, 2026.
Source: rednote-machine-learning/Inference-autopilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.