Official agent skill

Aks GPU Inference

by microsoft in microsoft/GitHub-Copilot-for-Azure

Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence.

OfficialMITAuto-check passedAI & LLM Engineering

Install Aks GPU Inference

skills CLI
$ npx skills add microsoft/GitHub-Copilot-for-Azure --skill aks-gpu-inference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/GitHub-Copilot-for-Azure aks-gpu-inference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/GitHub-Copilot-for-Azure.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/aks-skills/skills/aks-gpu-inference .claude/skills/aks-gpu-inference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aks-gpu-inference
GitHub stars
255
Token cost
~764 tokens
SKILL.md length
304 words
Files
6 (incl. references)
Skills in repo
61
Repo updated
First seen
Licence
MIT

At a glance

Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence.

  • Works in 6 steps: Bind subscription, cluster, kube… → Capture exact events/status and observed… → Route to scheduling, → …
  • : setup (airunway-aks-setup)
  • SKILL.md covers Quick Reference, When to Use This Skill, MCP Tools and Host Capability Gate, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Aks GPU Inference is an agent skill from microsoft/GitHub-Copilot-for-Azure, published by the product's own GitHub organization. Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence. WHEN: 'Insufficient nvidia.com/gpu', GPU pod Pending, model-load OOM, DCGM/VRAM, KAITO Workspace not ready, or GPU autoscaling. DO NOT USE FOR: setup (airunway-aks-setup), non-GPU incidents (aks-troubleshooting), standalone VM quota (azure-quotas), or generic cost (cost-analysis or cost-optimization from the optional azure-cost plugin).

Its SKILL.md is about 760 tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/gpu-cost-and-scaling.md`, `references/gpu-observability.md` and `references/gpu-scheduling.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with Microsoft Azure and NVIDIA AI Platform. The repository describes itself as: GitHub Copilot for Azure. The licence is MIT.

When your agent uses it

  • : setup (airunway-aks-setup)
  • Non-GPU incidents (aks-troubleshooting)
  • Standalone VM quota (azure-quotas)
  • Generic cost (cost-analysis

Example prompts

  • “Insufficient nvidia.com/gpu”
  • “/aks-gpu-inference”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Bind subscription, cluster, kube context, pool, and affected resource.
  2. Capture exact events/status and observed gpuProfile; state which reads ran.
  3. Route to scheduling,
  4. Separate container/host OOM from device-allocation failures. For OOMKilled,
  5. For KAITO not-ready, warn that Workspace deletion leaves its GPU pools;
  6. Report evidence, confidence, missing evidence, and owner handoff.

What it can do on your machine

Read from SKILL.md and the folder at commit ce94fce. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Aks GPU Inference loads about 764 tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 304 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~764
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/GitHub-Copilot-for-Azure at commit ce94fce, republished under its MIT licence (© microsoft). 304 words, ~764 tokens.

Download SKILL.mdSave it as .claude/skills/aks-gpu-inference/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
aks-gpu-inference
description
Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence. WHEN: 'Insufficient nvidia.com/gpu', GPU pod Pending, model-load OOM, DCGM/VRAM, KAITO Workspace not ready, or GPU autoscaling. DO NOT USE FOR: setup (airunway-aks-setup), non-GPU incidents (aks-troubleshooting), standalone VM quota (azure-quotas), or generic cost (cost-analysis or cost-optimization from the optional azure-cost plugin).
license
MIT
metadata.author
Microsoft
metadata.version
0.0.0-placeholder

AKS GPU Inference Day-2

Quick Reference

PropertyValue
Best forAKS GPU/KAITO failures
EvidenceBound target, events/status, gpuProfile, DCGM

When to Use This Skill

Use for scheduling, VRAM/OOM, KAITO readiness, and workload scaling. Route exclusions as described above.

MCP Tools

Use a fitting host-advertised Azure read or the references' read-only queries.

Host Capability Gate

Before executing any referenced pipeline, verify that the host authorizes the required shell and kubectl/az commands against the bound target. File-backed collection also requires approved artifact storage and access to any bundled files it uses. A governed Azure CLI tool alone does not establish these capabilities. Equivalent host reads may replace commands only where their advertised schemas provide the required evidence.

If execution is unavailable or prohibited, say so and analyze supplied or redacted events, status, logs, and metrics, or give the operator a scoped collection plan. State which reads did not run and leave conclusions requiring missing evidence unconfirmed. Never route kubectl through Azure MCP or bypass host policy. The authorization requirements below still apply.

Workflow

  1. Bind subscription, cluster, kube context, pool, and affected resource.
  2. Capture exact events/status and observed gpuProfile; state which reads ran.
  3. Route to scheduling, observability, KAITO, or scaling.
  4. Separate container/host OOM from device-allocation failures. For OOMKilled, correlate container memory limits/usage and node conditions. For Managed+Install device-memory evidence, use the exporter on port 19400 and incident-window FB_USED/FB_FREE; sizing tables do not establish a cause.
  5. For KAITO not-ready, warn that Workspace deletion leaves its GPU pools; cleanup is separate and needs explicit authorization.
  6. Report evidence, confidence, missing evidence, and owner handoff.

Require explicit authorization before scaling, cordon/drain, deletion, add-on enablement, or monitoring mutation.

Error Handling

ConditionResponse
Missing/contradictory evidenceMark unconfirmed; request the owner/read
Unknown profilePreserve evidence; do not prescribe stack repair
Mutation requiredPropose separately and wait for authorization

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in plugins/aks-skills/skills/aks-gpu-inference of microsoft/GitHub-Copilot-for-Azure.

  • SKILL.md
  • references/gpu-cost-and-scaling.md
  • references/gpu-observability.md
  • references/gpu-scheduling.md
  • references/kaito-workspaces.md
  • version.json

Open the folder on GitHubat commit ce94fce

Compare with similar skills

Aks GPU Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Aks GPU Inference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Aks GPU Inference this skillmicrosoft/GitHub-Copilot-for-Azure255—~764Automated safety check: PassMIT
TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT
Vllm Deploy Dockervllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNone

Similar skills

  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 4 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Integrate Model

    tryonlabs/opentryon

    Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner).

    551 GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from microsoft/GitHub-Copilot-for-Azure

All 61 skills in this repo
  • Capacity

    microsoft/GitHub-Copilot-for-Azure

    Official

    Discovers available Azure OpenAI model capacity across regions and projects.

    255 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Deploy Model

    microsoft/GitHub-Copilot-for-Azure

    Official

    Unified Azure OpenAI model deployment skill with intelligent intent-based routing.

    255 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Azure Storage

    microsoft/GitHub-Copilot-for-Azure

    Official

    Azure Storage Services including Blob Storage, File Shares, Queue Storage, Table Storage, and Data Lake.

    255 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Microsoft Foundry

    microsoft/GitHub-Copilot-for-Azure

    Official

    Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end.

    255 GitHub starsUsed in 1 repo~6.7k tokens
    Auto-check passed
  • Entra Agent Id

    microsoft/GitHub-Copilot-for-Azure

    Official

    Provision Microsoft Entra Agent Identity Blueprints, BlueprintPrincipals, and per-instance Agent Identities via Microsoft Graph, and configure OAuth 2.0 token exchange (fmipath, OBO, cross-tenant)…

    255 GitHub starsUsed in 2 repos~4k tokens
    Auto-check passed
  • Azure Kubernetes Automatic Readiness

    microsoft/GitHub-Copilot-for-Azure

    Official

    Assess Kubernetes workloads and cluster configuration for AKS Automatic compatibility.

    255 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed

Questions about Aks GPU Inference

What does Aks GPU Inference do?

Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence. Aks GPU Inference is an agent skill from microsoft/GitHub-Copilot-for-Azure, published by the product's own GitHub organization. Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence.

When should I use Aks GPU Inference?

Aks GPU Inference fits situations like: : setup (airunway-aks-setup); non-GPU incidents (aks-troubleshooting); standalone VM quota (azure-quotas); generic cost (cost-analysis.

How do I install Aks GPU Inference in Claude Code?

Run `npx skills add microsoft/GitHub-Copilot-for-Azure --skill aks-gpu-inference -a claude-code`. Or copy the skill folder (plugins/aks-skills/skills/aks-gpu-inference in microsoft/GitHub-Copilot-for-Azure) into .claude/skills/aks-gpu-inference in your project. Claude Code loads it when a task matches its description.

How do I install Aks GPU Inference in Codex?

Run `npx skills add microsoft/GitHub-Copilot-for-Azure --skill aks-gpu-inference -a codex`. Or copy the skill folder (plugins/aks-skills/skills/aks-gpu-inference in microsoft/GitHub-Copilot-for-Azure) into .agents/skills/aks-gpu-inference in your project. Codex loads it when a task matches its description.

Can I use Aks GPU Inference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/GitHub-Copilot-for-Azure --skill aks-gpu-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aks-gpu-inference, .gemini/skills/aks-gpu-inference, .github/skills/aks-gpu-inference and .opencode/skills/aks-gpu-inference in your project.

What does Aks GPU Inference need to run?

SKILL.md names no scripts, command-line tools or credentials: Aks GPU Inference is instructions for the agent only.

Does Aks GPU Inference access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Aks GPU Inference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Aks GPU Inference use?

Aks GPU Inference is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Aks GPU Inference use?

About 764 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to Aks GPU Inference?

Skills that share tags, products or a category with Aks GPU Inference: TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars), Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Aks GPU Inference?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/GitHub-Copilot-for-Azure, which has 255 GitHub stars. The repository holds 61 skills in this directory. The repository was last updated on October 9, 2026.

Source: microsoft/GitHub-Copilot-for-Azure on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.