Official agent skill

Infrastructure

by microsoft in microsoft/physical-ai-toolchain

Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology

OfficialMITAuto-check passedDevOps & Cloud

Install Infrastructure

skills CLI
$ npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/physical-ai-toolchain infrastructure --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/infrastructure .claude/skills/infrastructure && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
infrastructure
GitHub stars
122
Token cost
~1.6k tokens
SKILL.md length
376 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology

  • Works in 6 steps: Initialize Azure subscription → Configure Terraform variables → Provision infrastructure → …
  • Tasks that involve Infrastructure as code
  • SKILL.md covers Prerequisites, Deployment Workflow, Network Mode Selection and Common Operations, plus 3 more sections
  • Calls terraform, az and kubectl

What it does

Infrastructure is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization. Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Infrastructure as code and Container orchestration. It works with Microsoft Azure, Terraform, Kubernetes and Azure Kubernetes Service. The licence is MIT.

When your agent uses it

  • Tasks that involve Infrastructure as code
  • Tasks that involve Container orchestration

Example prompts

  • “/infrastructure”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Initialize Azure subscription
  2. Configure Terraform variables
  3. Provision infrastructure
  4. Deploy VPN (private clusters only)
  5. Connect to cluster
  6. Run setup scripts

What it can do on your machine

Read from SKILL.md and the folder at commit 5d38197. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • terraform
    • az
    • kubectl
    • shellcheck

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use az and kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Infrastructure loads about 1.6k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 376 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/physical-ai-toolchain at commit 5d38197, republished under its MIT licence (© microsoft). 376 words, ~1,584 tokens.

Download SKILL.mdSave it as .claude/skills/infrastructure/SKILL.md (or your agent's skills folder).
name
infrastructure
description
Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology

Infrastructure Skill

Deploy and manage Azure cloud infrastructure for the Physical AI Toolchain — Terraform IaC, AKS cluster configuration, GPU node pools, and network topology.

Prerequisites

ToolRequirement
Azure CLIaz login authenticated
Terraform1.5+
kubectlMatching cluster version
Helm4.2+
shellcheckFor script validation

Deployment Workflow

Follow these steps in order for a complete deployment.

Step 1 — Initialize Azure subscription
bash
source infrastructure/terraform/prerequisites/az-sub-init.sh

Exports ARM_SUBSCRIPTION_ID and validates Azure CLI authentication.

Step 2 — Configure Terraform variables
bash
cd infrastructure/terraform
cp terraform.tfvars.example terraform.tfvars

Edit terraform.tfvars with environment-specific values. Example configurations are in infrastructure/examples/:

FileScenario
terraform.tfvars.devSingle spot GPU pool, public networking
terraform.tfvars.prodMultiple GPU pools, full private networking, HA
terraform.tfvars.hybridPrivate data services, public AKS API server
Step 3 — Provision infrastructure
bash
terraform init
terraform plan -var-file=terraform.tfvars
terraform apply -var-file=terraform.tfvars
Step 4 — Deploy VPN (private clusters only)

Required when should_enable_private_aks_cluster = true:

bash
cd infrastructure/terraform/vpn
terraform init && terraform apply
Step 5 — Connect to cluster
bash
az aks get-credentials --resource-group <rg> --name <aks>
kubectl cluster-info
Step 6 — Run setup scripts
bash
cd infrastructure/setup
./01-deploy-robotics-charts.sh
./02-deploy-azureml-extension.sh
./03-deploy-osmo.sh --private-service-ip <unused-aks-subnet-ip>

Scripts must run in numeric order. Each supports --config-preview for dry-run output. On a new cluster, 03-deploy-osmo.sh needs --private-service-ip with a free address in the AKS subnet for the internal load balancer in front of OSMO; later runs reuse that address.

Network Mode Selection

Three network modes control connectivity and security:

Modeshould_enable_private_endpointshould_enable_private_aks_clusterVPN Required
Full PrivatetruetrueYes
HybridtruefalseNo
Full PublicfalsefalseNo

Full Private is the default and recommended for production. Hybrid mode allows kubectl access without VPN while keeping data services private.

Common Operations

Show full SKILL.md (151 more words)Show less
Plan changes
bash
cd infrastructure/terraform
terraform plan -var-file=terraform.tfvars
Apply changes
bash
terraform apply -var-file=terraform.tfvars
Destroy infrastructure
bash
terraform destroy -var-file=terraform.tfvars
VPN setup
bash
cd infrastructure/terraform/vpn
terraform init && terraform apply
DNS configuration
bash
cd infrastructure/terraform/dns
terraform init && terraform apply
Validate setup scripts
bash
shellcheck infrastructure/setup/01-deploy-robotics-charts.sh
infrastructure/setup/01-deploy-robotics-charts.sh --config-preview
Check Terraform formatting
bash
terraform fmt -check -recursive infrastructure/terraform/

Directory Structure

text
infrastructure/
├── terraform/                         # Infrastructure as Code
│   ├── main.tf                        # Module composition
│   ├── variables.tf                   # Input variables
│   ├── outputs.tf                     # Output values
│   ├── versions.tf                    # Provider requirements
│   ├── terraform.tfvars.example       # Example configuration
│   ├── prerequisites/                 # Azure subscription setup
│   ├── modules/                       # Terraform modules
│   ├── vpn/                           # Standalone VPN deployment
│   ├── automation/                    # Standalone automation deployment
│   └── dns/                           # Standalone DNS deployment
├── setup/                             # Post-deploy cluster configuration
│   ├── 01-deploy-robotics-charts.sh   # GPU Operator, KAI Scheduler
│   ├── 02-deploy-azureml-extension.sh # AzureML K8s extension
│   ├── 03-deploy-osmo.sh             # OSMO control plane and backend
│   ├── defaults.conf                  # Central version and namespace config
│   └── lib/                           # Shared shell libraries
├── specifications/                    # Domain specification documents
└── examples/                          # Example tfvars configurations

GPU Configuration Reference

GPUVM SKUDriver Sourcegpu_driverMIG Strategy
A10Standard_NV36ads_A10_v5AKS-managedInstallN/A
RTX PRO 6000Standard_NC144ds_xl_RTXPRO6000BSE_v6 (1 GPU, 96 GB)AKS-managed GRID driverInstallsingle
H100Standard_NC40ads_H100_v5GPU OperatorNoneDisabled

Only RTX PRO 6000 pools created with gpu_driver = "None" need the nvidia.com/gpu.deploy.driver=false label, which hands them to the fallback GRID driver DaemonSet. Preview RTX sizes (128, 256, or 320 vCPUs) no longer deploy. Each NC144ds_xl node needs 144 vCPUs of RTX PRO 6000 quota, plus one more node's worth for an upgrade surge; park a pool with autoscaling off and node_count = 0 until quota exists.

Documentation

GuideDescription
Infrastructure READMEDomain overview and quick start
Terraform READMETerraform configuration reference
Setup READMESetup script reference
Infrastructure DeploymentFull deployment walkthrough
GPU ConfigurationDetailed GPU driver and operator reference

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/infrastructure of microsoft/physical-ai-toolchain.

Open the folder on GitHubat commit 5d38197

Compare with similar skills

Infrastructure next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Infrastructure compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Infrastructure this skillmicrosoft/physical-ai-toolchain122—~1.6kAutomated safety check: PassMIT
Cloud Devopsdavila7/claude-code-templates32k4 repos~1.4kAutomated safety check: PassMIT
Iac Securityhardw00t/ai-security-arsenal104—~2.4kAutomated safety check: PassNone
Agent Bom Scan InfraLeoYeAI/openclaw-master-skills2.2k—~1.5kAutomated safety check: PassApache-2.0
Aks Deployment Skilltimothywarner/chatgptclass143—~916Automated safety check: PassCustom licence
Asdfjjmartres/opencode133—~2.1kAutomated safety check: NotesMIT

Similar skills

  • Cloud Devops

    davila7/claude-code-templates

    Cloud infrastructure and DevOps workflow covering AWS, Azure, GCP, Kubernetes, Terraform, CI/CD, monitoring, and cloud-native development.

    32k GitHub starsUsed in 4 repos~1.4k tokens
    DevOps & CloudAuto-check passed
  • Iac Security

    hardw00t/ai-security-arsenal

    Infrastructure-as-Code security scanning router for Terraform, CloudFormation, Kubernetes manifests, Helm, ARM/Bicep.

    104 GitHub stars~2.4k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check passed
  • Agent Bom Scan Infra

    LeoYeAI/openclaw-master-skills

    Scan infrastructure-as-code, cloud configurations, and find secrets.

    2.2k GitHub stars~1.5k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Aks Deployment Skill

    timothywarner/chatgptclass

    Deploy and operate workloads on Azure Kubernetes Service (AKS) the safe way.

    143 GitHub stars~916 tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed
  • Asdf

    jjmartres/opencode

    A skill your agent uses whenever the user wants to install, configure, or use asdf (asdf-vm), the universal version manager.

    133 GitHub stars~2.1k tokensUpdated 5 mo ago
    DevOps & CloudAuto-check: notes
  • Supercheck Infrastructure Deployment

    supercheck-io/supercheck

    Work on Supercheck Docker Compose, K3s, Kubernetes manifests, gVisor, OpenTofu/Hetzner, secrets, external services, autoscaling, backups, disaster recovery, DNS/TLS, or production deployment.

    215 GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from microsoft/physical-ai-toolchain

  • Osmo Lerobot Training

    microsoft/physical-ai-toolchain

    Official

    Submit, monitor, analyze, and evaluate LeRobot imitation learning training jobs on OSMO with Azure ML MLflow integration and inference evaluation - Brought to you by microsoft/physical-ai-toolchain

    122 GitHub stars~3.8k tokensUpdated today
    Auto-check: notes
  • Azureml K3s Compute Target Setup

    microsoft/physical-ai-toolchain

    Official

    Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.

    122 GitHub stars~5.7k tokensUpdated today
    Auto-check: notes
  • Environment Deployment

    microsoft/physical-ai-toolchain

    Official

    Generate, transfer, and consume environment-specific Azure, AKS, OSMO, ACR, and Azure ML deployment bundles.

    122 GitHub stars~5.8k tokensUpdated today
    Auto-check passed
  • Fleet Deployment

    microsoft/physical-ai-toolchain

    Official

    Deploy trained robot policies to edge fleets via FluxCD GitOps, image automation, and deployment gating

    122 GitHub stars~518 tokensUpdated today
    Auto-check passed
  • Fleet Intelligence

    microsoft/physical-ai-toolchain

    Official

    Monitor robot fleet telemetry via Azure IoT Operations, drift detection, Grafana dashboards, and Fabric analytics

    122 GitHub stars~598 tokensUpdated today
    Auto-check passed
  • Synthetic Data

    microsoft/physical-ai-toolchain

    Official

    Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines

    122 GitHub stars~469 tokensUpdated today
    Auto-check passed

Categories

Questions about Infrastructure

What does Infrastructure do?

Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology. Infrastructure is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization.

When should I use Infrastructure?

Infrastructure fits situations like: tasks that involve Infrastructure as code; tasks that involve Container orchestration.

How do I install Infrastructure in Claude Code?

Run `npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a claude-code`. Or copy the skill folder (.github/skills/infrastructure in microsoft/physical-ai-toolchain) into .claude/skills/infrastructure in your project. Claude Code loads it when a task matches its description.

How do I install Infrastructure in Codex?

Run `npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a codex`. Or copy the skill folder (.github/skills/infrastructure in microsoft/physical-ai-toolchain) into .agents/skills/infrastructure in your project. Codex loads it when a task matches its description.

Can I use Infrastructure in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/infrastructure, .gemini/skills/infrastructure, .github/skills/infrastructure and .opencode/skills/infrastructure in your project.

What does Infrastructure need to run?

Going by SKILL.md and its folder, Infrastructure needs the command-line tools its instructions call (terraform, az, kubectl and shellcheck).

Does Infrastructure access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Infrastructure safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Infrastructure use?

Infrastructure is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Infrastructure use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Infrastructure?

Skills that share tags, products or a category with Infrastructure: Cloud Devops (davila7/claude-code-templates, 32k stars), Iac Security (hardw00t/ai-security-arsenal, 104 stars), Agent Bom Scan Infra (LeoYeAI/openclaw-master-skills, 2.2k stars) and Aks Deployment Skill (timothywarner/chatgptclass, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Infrastructure?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/physical-ai-toolchain, which has 122 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.

Source: microsoft/physical-ai-toolchain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.