Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.

Apache-2.0Auto-check passedDevOps & Cloud

Install Inference Autopilot

skills CLI
$ npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rednote-machine-learning/Inference-autopilot inference-autopilot --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
inference-autopilot
GitHub stars
142
Token cost
~4.5k tokens
SKILL.md length
1,968 words
Files
48 (incl. scripts, references, assets)
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.

  • Works in 7 steps: Normalize The Objective → Establish A Reproducible Baseline → Classify Before Tuning → …
  • Codex needs to tune SGLang launch parameters
  • SKILL.md covers Start Here, Safety Contract, Optimization Loop and Bundled Commands
  • Calls python3

What it does

Inference Autopilot is an agent skill from rednote-machine-learning/Inference-autopilot. Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs. Use when Codex needs to tune SGLang launch parameters, run bounded single-host GPU experiments, plan deployment topology, inspect GPU or CPU profiles, identify scheduler/KV/communication/kernel bottlenecks, propose operator optimizations, or validate that a candidate configuration improves performance without correctness or SLO regressions.

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 51 other files, including scripts, reference files and assets (for example `.github/workflows/sglang-parameter-compat.yml`, `.github/workflows/version-consistency.yml` and `PARAMETER_EVOLUTION.md`).

It sits in DevOps & Cloud, covering Site reliability engineering and Deployment. It works with SGLang. The repository describes itself as: Evidence-driven SGLang inference deployment optimization. The licence is Apache-2.0.

When your agent uses it

  • Codex needs to tune SGLang launch parameters
  • Run bounded single-host GPU experiments
  • Plan deployment topology
  • Identify scheduler/KV/communication/kernel bottlenecks

Example prompts

  • “/inference-autopilot”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Normalize The Objective
  2. Establish A Reproducible Baseline
  3. Classify Before Tuning
  4. Search From Coarse To Fine
  5. Escalate Profiling By Evidence
  6. Optimize Operators Only When Justified
  7. Validate And Report

What it can do on your machine

Read from SKILL.md and the folder at commit 61eb1c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Inference Autopilot loads about 4.5k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 132 tokens; SKILL.md has 1,968 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~132
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from rednote-machine-learning/Inference-autopilot at commit 61eb1c0, republished under its Apache-2.0 licence (© rednote-machine-learning). 1,968 words, ~4,539 tokens.

Download SKILL.mdSave it as .claude/skills/inference-autopilot/SKILL.md (or your agent's skills folder). This skill also uses 47 other files; get the full folder from GitHub.
name
inference-autopilot
description
Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs. Use when Codex needs to tune SGLang launch parameters, run bounded single-host GPU experiments, plan deployment topology, inspect GPU or CPU profiles, identify scheduler/KV/communication/kernel bottlenecks, propose operator optimizations, or validate that a candidate configuration improves performance without correctness or SLO regressions.

Inference Autopilot

Run an evidence-driven optimization loop. Treat this skill as the control plane and the bundled scripts as deterministic utilities. Default to local, private, dry-run behavior.

Start Here

For a new private single-host deployment, start with the one-shot interface. The user supplies only local paths, workload, SLO, objective, budget, and optional GPU visibility. The script discovers NVIDIA or AMD GPU memory and topology, reads the local model config and weight sizes, regenerates the current SGLang parameter contract from server_args.py and sglang.launch_server --help, captures a bounded Nsight Systems baseline, routes trace and workload evidence to parameter families, runs bounded screening, applies mode-specific confirmation, and emits a deployable command only after the confirmation gates pass:

bash
inferopt init --output /absolute/private/task.json
inferopt doctor --task /absolute/private/task.json --output /absolute/private/doctor.json
inferopt plan --task /absolute/private/task.json --output /absolute/private/plan.json
inferopt run --task /absolute/private/task.json --yes --output /absolute/private/final.json
inferopt report --result /absolute/private/final.json --output /absolute/private/report.md

inferopt run only detects and reports missing fused MoE tuning configs. Never run the high-cost kernel search as part of the normal workflow. If the user explicitly requests it after reviewing the report, use the standalone command emitted by the report:

bash
inferopt tune-moe --task RUN_DIR/task.json --profile RUN_DIR/profile/nsys-diagnosis.json --result RUN_DIR/final.json --output-dir RUN_DIR/optional-fused-moe-tuning --yes --output RUN_DIR/optional-fused-moe-tuning.json

Treat generated configs as non-deployable until a separate end-to-end A/B benchmark passes the original workload, SLO, noise, and confirmation gates. Generated files must remain under configs/triton_<version>/ and use the installed SGLang naming contract (E, per-TP N, runtime GPU name, dtype, optional block shape/per-channel marker) with a JSON mapping from token batch size M to kernel tile parameters. If the log requests an _down.json file, the normal tuner is only partial: use SGLang's separate top-k capture/tuner workflow and pass --topk-ids-dir. Inference Autopilot refuses to authorize an up-only result in that case; both matching up and _down files are required.

For direct generation of paired files without end-to-end validation, use:

bash
inferopt generate-moe-config \
  --repository /path/to/sglang \
  --model-path /path/to/model \
  --tp-size 4 --dtype fp8_w8a8 \
  --topk-ids-dir /path/to/topk_ids \
  --output-dir /path/to/moe-config-output --yes

This invokes the official separate tuner and refuses to create a fake _down.json without routing captures. Use --mode standard only when SGLang did not request a down-kernel config.

The standalone inferopt command does not require Codex. doctor never starts a server: it validates the local model, GPU/runtime, current SGLang CLI, profiler availability, and a transparent single-GPU memory estimate. It may recommend a quantized checkpoint class when the current model cannot fit. It also reports legal TP sizes across the complete selected single-host GPU set and treats a missing nsys installation as a blocking preflight error. It never downloads, switches, or approves a model variant without an explicit quality evaluation.

The underlying scripts remain available for controlled integration:

bash
cp assets/task.autopilot.example.json /absolute/private/task.json
python3 scripts/autopilot.py validate --task /absolute/private/task.json
python3 scripts/autopilot.py plan --task /absolute/private/task.json --output /absolute/private/plan.json
python3 scripts/autopilot.py run --task /absolute/private/task.json --yes --output /absolute/private/final-pointer.json

Read references/hardware-profiles.json for the official-source GPU capability catalog. Runtime discovery overrides catalog values. Preserve the checked-out SGLang version's automatic defaults as the baseline; never transplant a winning parameter from one GPU or workload to another without measurement.

  1. Locate the inference repository and read its local instructions.
  2. Create a task specification using references/input-schema.md.
  3. Run scripts/inferopt.py validate --spec <task.json>.
  4. Run scripts/inferopt.py inventory --output <run-dir>/inventory.json on every target host when access is available.
  5. Inspect the generated inventory and task constraints before launching anything.
  6. Follow the optimization loop below.

For authorized single-host execution, copy assets/task.execute.example.json, read references/execution-schema.md, validate it, render the exact command plan, obtain explicit approval, and only then run:

bash
python3 scripts/autotune.py validate --spec task.json
python3 scripts/autotune.py plan --spec task.json --output plan.json
python3 scripts/autotune.py run --spec task.json --yes --output final-pointer.json
python3 scripts/autotune.py report --run-dir /absolute/completed-run --output decision.json

If the framework is SGLang, read references/sglang-adapter.md. For another engine, discover its actual launch, benchmark, metrics, and profiling interfaces instead of assuming SGLang flags.

Safety Contract

  • Default to dry_run. Require the user or task specification to opt into execution.
  • Run only on hosts, endpoints, repositories, models, datasets, and output directories in scope.
  • Never modify a production deployment, autoscaler, traffic router, cloud resource, driver, system service, or shared cluster without explicit authorization.
  • Never download a model or dependency, reserve paid hardware, publish results, or send traces externally without explicit authorization.
  • Never kill processes not started by the current experiment. Record owned PIDs and terminate only those PIDs.
  • Never use broad process-kill commands as cleanup.
  • Redact prompts, API keys, tokens, user identifiers, and proprietary model paths from reports.
  • Enforce trial, wall-time, GPU-hour, and failure budgets. Stop immediately on correctness failure, repeated OOM, thermal/power anomaly, or error-rate violation.
  • Keep baseline artifacts immutable. Store every candidate in a distinct trial directory.
  • Do not accept a faster candidate unless correctness, stability, and SLO gates pass.
  • Treat --yes as confirmation that the generated command plan was reviewed; never add it before approval.

Read references/safety-policy.md before executing benchmarks, profilers, or code changes.

Optimization Loop

Deployment Modes

Set deployment_mode in the autopilot task. Use online_latency for an interactive service: maximize the selected objective while every declared E2E, TTFT, TPOT, or ITL gate passes. When no latency SLO is declared, the configured secondary-regression limit still protects observed latency. Use offline_throughput for batch inference: maximize sustained aggregate throughput at calibrated batch pressure; latency remains recorded but is only a gate when the task explicitly declares it.

For offline workloads without SLOs, skip capacity calibration, remove the client concurrency cap, and let the resolved SGLang admission policy determine sustained pressure. Start the baseline service exactly twice: one bounded nsys capture and one unprofiled benchmark. Preserve that benchmark as an immutable reference for all candidates; never search or override max_running_requests. For an SLO-constrained workload, control load with the benchmark client's max_concurrency; treat the server's resolved max_running_requests as diagnostic evidence rather than a tuning parameter. Persist every benchmark command and resolved server value. For a model with a verified official cookbook, first compare complete, locally valid capability bundles: include the relevant model feature (for example MTP), prefix/KV cache policy, scheduler/admission, memory pool, CUDA Graph, and MoE backend variants when the workload and hardware make them applicable. Profile the fastest SLO-valid initial configuration, even if its single screening sample has not yet cleared the final improvement threshold. Then use Nsight evidence to select and refine parameter families, and screen and confirm every candidate against the original target workload. Never label a result as global best: report it as the best configuration within the checked-out SGLang parameter contract, tested search space, hardware, workload, and SLO gates. Use search_depth: thorough (the default) to add a one-factor sensitivity screen for every high-impact compatible family even when a short trace lacks a single hotspot. Use evidence_guided only when experiment budget is tight. Every run emits search-plan.json.parameter_audit, which accounts for each CLI-visible, non-deprecated ServerArgs as selected, excluded, or inapplicable with a concrete reason. It prevents a short candidate list from being mistaken for the whole startup-parameter surface. Render per-stage trial progress as an elapsed-time ASCII bar with completed and planned counts. Count completed, failed, and capability-skipped trials as processed. Do not invent a percentage for Nsight capture or model startup when their duration is not knowable in advance, and keep redirected logs free of terminal control characters. After the one-factor screen, the optimizer uses the remaining trial budget to test explicit combinations of independent screened winners. Every candidate that clears the configured improvement threshold is considered: pairs with the strongest candidate run first, followed by other pairs and larger combinations while budget remains. Parameter/environment conflicts and exact duplicates are excluded. Positive sub-threshold candidates may use leftover slots but never displace an above-threshold combination. Offline no-SLO runs reuse the preserved baseline and rerun only the selected candidate once; SLO-constrained runs retain repeated A/B confirmation. It does not perform an unbounded Cartesian product.

Show full SKILL.md (803 more words)Show less
1. Normalize The Objective

Extract:

  • hardware and interconnect topology;
  • model architecture, dtype, quantization, context limit, and draft model;
  • workload arrival process, input/output distributions, prefix reuse, modalities, structured outputs, and concurrency;
  • TTFT, TPOT/ITL, end-to-end latency, throughput, goodput, cost, power, and correctness constraints;
  • experiment budget and allowed mutation scope.

Do not optimize an unspecified scalar. Convert multiple SLOs into hard constraints plus one primary objective, for example: maximize request goodput subject to P99 TTFT and P99 TPOT.

2. Establish A Reproducible Baseline
  • Pin repository commit, container/image, model revision, driver/runtime, environment variables, and exact launch command.
  • Warm up before measurement.
  • Warm up before every measured window. Require at least 30 seconds of completed-request measurement by default; when a run is shorter, automatically increase request count and repeat it, preserving the short raw sample as an audit artifact.
  • Use at least five times the maximum concurrency in steady-state serving tests when affordable.
  • Repeat noisy trials; preserve raw JSONL, logs, metrics, and profiler traces.
  • Record idle and loaded GPU memory, utilization, power, clocks, host CPU, network, cache hit rate, request failures, accepted draft tokens, and queueing metrics when available.
  • Validate output correctness before trusting performance numbers.

Use scripts/inferopt.py analyze for SGLang-compatible benchmark JSONL and scripts/inferopt.py compare to gate candidates. Use scripts/autotune.py only for an isolated, authorized, single-host SGLang experiment.

Treat a one-run screening_winner as a hypothesis. For SLO-constrained modes require repeated A/B measurements, acceptable objective CV, and every repetition passing SLO. For offline no-SLO mode, capture both the short screening baseline and a cache-flushed confirmation-reference window while the same baseline service is loaded, then rerun only the selected candidate with the exact reference dataset, request count, and duration contract. Never compare the longer final candidate with the short screening baseline.

Use recommended_configuration for the deployment decision. Preserve a confirmed baseline when recommendation_status=retain_confirmed_baseline; an empty winner can mean candidates were correctly rejected, not that the experiment failed.

3. Classify Before Tuning

Read references/diagnosis-playbook.md. First classify the dominant bottleneck:

  • admission or queueing;
  • prefill compute;
  • decode memory bandwidth;
  • KV capacity, reuse, transfer, or host IO;
  • CPU scheduler/tokenizer/frontend;
  • TP/EP/PP collective communication;
  • MoE imbalance or All-to-All;
  • speculative draft/verify waste;
  • multimodal preprocessing;
  • one or more GPU operators.

Do not start kernel work while queueing, cache misses, bad topology, or an unsuitable backend dominates end-to-end time.

4. Search From Coarse To Fine

Search in this order unless evidence justifies a different order:

  1. deployment topology and parallelism;
  2. memory feasibility and KV allocation;
  3. scheduler, batching, and chunked prefill;
  4. attention, GEMM, MoE, and communication backends;
  5. prefix cache, hierarchical cache, sparse attention, or PD transfer;
  6. speculative decoding;
  7. operator implementation and kernel parameters.

Change one conceptual factor per ablation. Use a coarse family screen before fine tuning. Enumerate every current server_args.py parameter into a family, and record every compatible parameter as selected, excluded with its reason, or inapplicable to the current model/topology/deployment mode. Reuse prior trials and reject infeasible configurations before launching a server.

5. Escalate Profiling By Evidence

Use the least expensive tool that can answer the current question. Read references/profiler-routing.md.

  • Start with serving metrics and logs.
  • Use framework/PyTorch traces for stage and operator attribution.
  • Use Nsight Systems or equivalent for CPU/GPU overlap and communication timelines.
  • Use Nsight Compute or a kernel benchmark only after isolating a specific hot operator and representative shapes.

Keep profiling windows short and representative. Do not compare profiled latency directly with unprofiled production latency.

6. Optimize Operators Only When Justified

Before editing a kernel, record:

  • operator share of GPU-active execution time, and independently measured request-level attribution before claiming an end-to-end bound;
  • exact shapes, dtypes, layouts, architecture, batch and sequence regimes;
  • current backend and competing implementations;
  • numerical tolerance and determinism requirements;
  • expected upper-bound end-to-end gain from Amdahl's law.

Build an isolated correctness and performance benchmark. Cover boundary shapes and realistic distributions, not only one favorable shape. Integrate only after isolated and end-to-end gates pass.

7. Validate And Report

Require:

  • correctness equivalence or an explicitly approved quality tradeoff;
  • all hard SLOs pass at the target load;
  • improvement exceeds the configured minimum and noise margin;
  • no material error-rate, memory, stability, fairness, or cost regression;
  • repeated or confidence-supported results.

Deliver:

  • recommended launch/deployment configuration;
  • baseline versus winner and confidence/noise statement;
  • diagnosed bottleneck with supporting evidence;
  • rejected trials and why they failed;
  • raw artifact locations and reproduction commands;
  • remaining risks and next experiments;
  • code changes only when they produced validated gains.

Bundled Commands

bash
python3 scripts/inferopt.py validate --spec task.json
python3 scripts/inferopt.py inventory --output runs/inventory.json
python3 scripts/inferopt.py plan --spec task.json --output runs/plan.json
python3 scripts/inferopt.py analyze --input runs/baseline.jsonl --spec task.json --output runs/baseline-summary.json
python3 scripts/inferopt.py compare --baseline runs/baseline-summary.json --candidate runs/trial-001-summary.json --spec task.json
python3 scripts/autotune.py validate --spec task.json
python3 scripts/autotune.py plan --spec task.json --output plan.json
python3 scripts/autotune.py run --spec task.json --yes
python3 scripts/autopilot.py validate --task task.json
python3 scripts/autopilot.py plan --task task.json --output plan.json
python3 scripts/autopilot.py run --task task.json --yes --output final-pointer.json

inferopt.py never starts a server. autopilot.py starts only the local SGLang process groups it creates, regenerates an installed-version parameter contract for every run, performs a bounded Nsight Systems CUDA-range capture, then generates and executes two-stage autotune.py runs using structured allowlisted parameters. It cannot execute shell snippets, remote SSH, Slurm, Kubernetes, arbitrary modules, production endpoints, package installation, source edits, or kernel changes. It records a shape-matched Nsight Compute follow-up only when one kernel is trace-proven hot; it never edits a kernel.

© rednote-machine-learning, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 47 other files (scripts, references, assets) in the repository root of rednote-machine-learning/Inference-autopilot.

  • SKILL.md
  • .github/workflows/sglang-parameter-compat.yml
  • .github/workflows/version-consistency.yml
  • .gitignore
  • LICENSE
  • MANIFEST.in
  • PARAMETER_EVOLUTION.md
  • PROGRESS.md
  • README.md
  • README_zh.md
  • VERSION
  • agents/openai.yaml
  • assets/inference-autopilot-logo.svg
  • assets/task.autopilot.example.json
  • assets/task.example.json
  • assets/task.execute.example.json
  • pyproject.toml
  • … and 31 more

Open the folder on GitHubat commit 61eb1c0

Compare with similar skills

Inference Autopilot next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Inference Autopilot compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Inference Autopilot this skillrednote-machine-learning/Inference-autopilot142—~4.5kAutomated safety check: PassApache-2.0
Release Itwondelai/skills2.4k—~4kAutomated safety check: PassMIT
Testing In Productionpetrkindlmann/qa-skills163—~5.3kAutomated safety check: PassMIT
Delivery Managerborghei/Claude-Skills874—~2.2kAutomated safety check: PassMIT
Release Engineeringmagnus919/agent-skills111—~3.9kAutomated safety check: PassMIT
Platform Operationsrsmdt/the-startup536—~991Automated safety check: PassMIT

Similar skills

  • Release It

    wondelai/skills

    Build production-ready systems with stability patterns: circuit breakers, bulkheads, timeouts, and retry logic.

    2.4k GitHub stars~4k tokensUpdated 27 days ago
    DevOps & CloudAuto-check passed
  • Testing In Production

    petrkindlmann/qa-skills

    Safe-release techniques DURING rollout: feature flags, progressive rollouts, canary analysis, guardrail metrics, production smoke tests, and synthetic users.

    163 GitHub stars~5.3k tokensUpdated 3 mo ago
    DevOps & CloudAuto-check passed
  • Delivery Manager

    borghei/Claude-Skills

    Expert delivery management for release planning, deployment strategy, incident response, change management, SLA/error-budget tracking, and DORA metrics across continuous delivery pipelines.

    874 GitHub stars~2.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Release Engineering

    magnus919/agent-skills

    Design, automate, and operate end-to-end software releases: release process models and pipelines (trunk-based development, CD stages, release trains), progressive delivery and feature flags…

    111 GitHub stars~3.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Platform Operations

    rsmdt/the-startup

    Unified platform operations guidance for CI/CD pipeline design, deployment strategies, observability, SLI/SLOs, and incident-ready rollouts.

    536 GitHub stars~991 tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Azure Reliability

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Reliability development including best practices, decision making, architecture & design patterns, limits & quotas, and deployment.

    775 GitHub stars~2.2k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed

Works with

Categories

Questions about Inference Autopilot

What does Inference Autopilot do?

Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs. Inference Autopilot is an agent skill from rednote-machine-learning/Inference-autopilot. Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.

When should I use Inference Autopilot?

Inference Autopilot fits situations like: Codex needs to tune SGLang launch parameters; run bounded single-host GPU experiments; plan deployment topology; identify scheduler/KV/communication/kernel bottlenecks.

How do I install Inference Autopilot in Claude Code?

Run `npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a claude-code`. Or copy the skill folder (the rednote-machine-learning/Inference-autopilot repository) into .claude/skills/inference-autopilot in your project. Claude Code loads it when a task matches its description.

How do I install Inference Autopilot in Codex?

Run `npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a codex`. Or copy the skill folder (the rednote-machine-learning/Inference-autopilot repository) into .agents/skills/inference-autopilot in your project. Codex loads it when a task matches its description.

Can I use Inference Autopilot in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rednote-machine-learning/Inference-autopilot --skill inference-autopilot -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/inference-autopilot, .gemini/skills/inference-autopilot, .github/skills/inference-autopilot and .opencode/skills/inference-autopilot in your project.

What does Inference Autopilot need to run?

Going by SKILL.md and its folder, Inference Autopilot needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Inference Autopilot access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Inference Autopilot safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Inference Autopilot use?

Inference Autopilot is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Inference Autopilot use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.1k tokens, read only when the agent opens those files.

What are the alternatives to Inference Autopilot?

Skills that share tags, products or a category with Inference Autopilot: Release It (wondelai/skills, 2.4k stars), Testing In Production (petrkindlmann/qa-skills, 163 stars), Delivery Manager (borghei/Claude-Skills, 874 stars) and Release Engineering (magnus919/agent-skills, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Inference Autopilot?

rednote-machine-learning (a GitHub organization) maintains it in rednote-machine-learning/Inference-autopilot, which has 142 GitHub stars. The repository was last updated on September 7, 2026.

Source: rednote-machine-learning/Inference-autopilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.