Cloud Devops
davila7/claude-code-templates
Cloud infrastructure and DevOps workflow covering AWS, Azure, GCP, Kubernetes, Terraform, CI/CD, monitoring, and cloud-native development.
Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology
$ npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/physical-ai-toolchain infrastructure --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/infrastructure .claude/skills/infrastructure && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "infrastructure" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/infrastructure into .claude/skills/infrastructure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "infrastructure", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/infrastructureType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/physical-ai-toolchain infrastructure --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/infrastructure .agents/skills/infrastructure && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "infrastructure" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/infrastructure into .agents/skills/infrastructure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "infrastructure", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/physical-ai-toolchain infrastructure --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/infrastructure .cursor/skills/infrastructure && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "infrastructure" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/infrastructure into .cursor/skills/infrastructure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "infrastructure", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/physical-ai-toolchain.git --path .github/skills/infrastructure--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/physical-ai-toolchain infrastructure --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/infrastructure .gemini/skills/infrastructure && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "infrastructure" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/infrastructure into .gemini/skills/infrastructure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "infrastructure", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/physical-ai-toolchain infrastructureInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/infrastructure .github/skills/infrastructure && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "infrastructure" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/infrastructure into .github/skills/infrastructure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "infrastructure", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/physical-ai-toolchain infrastructure --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/infrastructure .opencode/skills/infrastructure && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "infrastructure" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/infrastructure into .opencode/skills/infrastructure/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "infrastructure", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
infrastructureDeploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology
Infrastructure is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization. Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Infrastructure as code and Container orchestration. It works with Microsoft Azure, Terraform, Kubernetes and Azure Kubernetes Service. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 5d38197. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
terraformazkubectlshellcheckFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use az and kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Infrastructure loads about 1.6k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 376 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/physical-ai-toolchain at commit 5d38197, republished under its MIT licence (© microsoft). 376 words, ~1,584 tokens.
.claude/skills/infrastructure/SKILL.md (or your agent's skills folder).Deploy and manage Azure cloud infrastructure for the Physical AI Toolchain — Terraform IaC, AKS cluster configuration, GPU node pools, and network topology.
| Tool | Requirement |
|---|---|
| Azure CLI | az login authenticated |
| Terraform | 1.5+ |
| kubectl | Matching cluster version |
| Helm | 4.2+ |
| shellcheck | For script validation |
Follow these steps in order for a complete deployment.
source infrastructure/terraform/prerequisites/az-sub-init.shExports ARM_SUBSCRIPTION_ID and validates Azure CLI authentication.
cd infrastructure/terraform
cp terraform.tfvars.example terraform.tfvarsEdit terraform.tfvars with environment-specific values. Example configurations are in infrastructure/examples/:
| File | Scenario |
|---|---|
terraform.tfvars.dev | Single spot GPU pool, public networking |
terraform.tfvars.prod | Multiple GPU pools, full private networking, HA |
terraform.tfvars.hybrid | Private data services, public AKS API server |
terraform init
terraform plan -var-file=terraform.tfvars
terraform apply -var-file=terraform.tfvarsRequired when should_enable_private_aks_cluster = true:
cd infrastructure/terraform/vpn
terraform init && terraform applyaz aks get-credentials --resource-group <rg> --name <aks>
kubectl cluster-infocd infrastructure/setup
./01-deploy-robotics-charts.sh
./02-deploy-azureml-extension.sh
./03-deploy-osmo.sh --private-service-ip <unused-aks-subnet-ip>Scripts must run in numeric order. Each supports --config-preview for dry-run output. On a new cluster, 03-deploy-osmo.sh needs --private-service-ip with a free address in the AKS subnet for the internal load balancer in front of OSMO; later runs reuse that address.
Three network modes control connectivity and security:
| Mode | should_enable_private_endpoint | should_enable_private_aks_cluster | VPN Required |
|---|---|---|---|
| Full Private | true | true | Yes |
| Hybrid | true | false | No |
| Full Public | false | false | No |
Full Private is the default and recommended for production. Hybrid mode allows kubectl access without VPN while keeping data services private.
cd infrastructure/terraform
terraform plan -var-file=terraform.tfvarsterraform apply -var-file=terraform.tfvarsterraform destroy -var-file=terraform.tfvarscd infrastructure/terraform/vpn
terraform init && terraform applycd infrastructure/terraform/dns
terraform init && terraform applyshellcheck infrastructure/setup/01-deploy-robotics-charts.sh
infrastructure/setup/01-deploy-robotics-charts.sh --config-previewterraform fmt -check -recursive infrastructure/terraform/infrastructure/
├── terraform/ # Infrastructure as Code
│ ├── main.tf # Module composition
│ ├── variables.tf # Input variables
│ ├── outputs.tf # Output values
│ ├── versions.tf # Provider requirements
│ ├── terraform.tfvars.example # Example configuration
│ ├── prerequisites/ # Azure subscription setup
│ ├── modules/ # Terraform modules
│ ├── vpn/ # Standalone VPN deployment
│ ├── automation/ # Standalone automation deployment
│ └── dns/ # Standalone DNS deployment
├── setup/ # Post-deploy cluster configuration
│ ├── 01-deploy-robotics-charts.sh # GPU Operator, KAI Scheduler
│ ├── 02-deploy-azureml-extension.sh # AzureML K8s extension
│ ├── 03-deploy-osmo.sh # OSMO control plane and backend
│ ├── defaults.conf # Central version and namespace config
│ └── lib/ # Shared shell libraries
├── specifications/ # Domain specification documents
└── examples/ # Example tfvars configurations| GPU | VM SKU | Driver Source | gpu_driver | MIG Strategy |
|---|---|---|---|---|
| A10 | Standard_NV36ads_A10_v5 | AKS-managed | Install | N/A |
| RTX PRO 6000 | Standard_NC144ds_xl_RTXPRO6000BSE_v6 (1 GPU, 96 GB) | AKS-managed GRID driver | Install | single |
| H100 | Standard_NC40ads_H100_v5 | GPU Operator | None | Disabled |
Only RTX PRO 6000 pools created with gpu_driver = "None" need the nvidia.com/gpu.deploy.driver=false label, which hands them to the fallback GRID driver DaemonSet. Preview RTX sizes (128, 256, or 320 vCPUs) no longer deploy. Each NC144ds_xl node needs 144 vCPUs of RTX PRO 6000 quota, plus one more node's worth for an upgrade surge; park a pool with autoscaling off and node_count = 0 until quota exists.
| Guide | Description |
|---|---|
| Infrastructure README | Domain overview and quick start |
| Terraform README | Terraform configuration reference |
| Setup README | Setup script reference |
| Infrastructure Deployment | Full deployment walkthrough |
| GPU Configuration | Detailed GPU driver and operator reference |
© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .github/skills/infrastructure of microsoft/physical-ai-toolchain.
Open the folder on GitHubat commit 5d38197
Infrastructure next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Infrastructure this skillmicrosoft/physical-ai-toolchain | 122 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Cloud Devopsdavila7/claude-code-templates | 32k | 4 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Iac Securityhardw00t/ai-security-arsenal | 104 | — | ~2.4k | Automated safety check: Pass | None | |
| Agent Bom Scan InfraLeoYeAI/openclaw-master-skills | 2.2k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Aks Deployment Skilltimothywarner/chatgptclass | 143 | — | ~916 | Automated safety check: Pass | Custom licence | |
| Asdfjjmartres/opencode | 133 | — | ~2.1k | Automated safety check: Notes | MIT |
davila7/claude-code-templates
Cloud infrastructure and DevOps workflow covering AWS, Azure, GCP, Kubernetes, Terraform, CI/CD, monitoring, and cloud-native development.
hardw00t/ai-security-arsenal
Infrastructure-as-Code security scanning router for Terraform, CloudFormation, Kubernetes manifests, Helm, ARM/Bicep.
LeoYeAI/openclaw-master-skills
Scan infrastructure-as-code, cloud configurations, and find secrets.
timothywarner/chatgptclass
Deploy and operate workloads on Azure Kubernetes Service (AKS) the safe way.
jjmartres/opencode
A skill your agent uses whenever the user wants to install, configure, or use asdf (asdf-vm), the universal version manager.
supercheck-io/supercheck
Work on Supercheck Docker Compose, K3s, Kubernetes manifests, gVisor, OpenTofu/Hetzner, secrets, external services, autoscaling, backups, disaster recovery, DNS/TLS, or production deployment.
microsoft/physical-ai-toolchain
Submit, monitor, analyze, and evaluate LeRobot imitation learning training jobs on OSMO with Azure ML MLflow integration and inference evaluation - Brought to you by microsoft/physical-ai-toolchain
microsoft/physical-ai-toolchain
Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.
microsoft/physical-ai-toolchain
Generate, transfer, and consume environment-specific Azure, AKS, OSMO, ACR, and Azure ML deployment bundles.
microsoft/physical-ai-toolchain
Deploy trained robot policies to edge fleets via FluxCD GitOps, image automation, and deployment gating
microsoft/physical-ai-toolchain
Monitor robot fleet telemetry via Azure IoT Operations, drift detection, Grafana dashboards, and Fabric analytics
microsoft/physical-ai-toolchain
Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines
Categories
Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology. Infrastructure is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization.
Infrastructure fits situations like: tasks that involve Infrastructure as code; tasks that involve Container orchestration.
Run `npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a claude-code`. Or copy the skill folder (.github/skills/infrastructure in microsoft/physical-ai-toolchain) into .claude/skills/infrastructure in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a codex`. Or copy the skill folder (.github/skills/infrastructure in microsoft/physical-ai-toolchain) into .agents/skills/infrastructure in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/physical-ai-toolchain --skill infrastructure -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/infrastructure, .gemini/skills/infrastructure, .github/skills/infrastructure and .opencode/skills/infrastructure in your project.
Going by SKILL.md and its folder, Infrastructure needs the command-line tools its instructions call (terraform, az, kubectl and shellcheck).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Infrastructure is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Infrastructure: Cloud Devops (davila7/claude-code-templates, 32k stars), Iac Security (hardw00t/ai-security-arsenal, 104 stars), Agent Bom Scan Infra (LeoYeAI/openclaw-master-skills, 2.2k stars) and Aks Deployment Skill (timothywarner/chatgptclass, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/physical-ai-toolchain, which has 122 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.
Source: microsoft/physical-ai-toolchain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.