Official agent skill

Azureml K3s Compute Target Setup

by microsoft in microsoft/physical-ai-toolchain

Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.

OfficialMITAuto-check: notesDevOps & Cloud

Install Azureml K3s Compute Target Setup

skills CLI
$ npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/physical-ai-toolchain azureml-k3s-compute-target-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/azureml-k3s-compute-target-setup .claude/skills/azureml-k3s-compute-target-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
azureml-k3s-compute-target-setup
GitHub stars
123
Token cost
~5.7k tokens
SKILL.md length
1,941 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.

  • Works in 4 steps: Installs a training-only Azure ML… → Waits for the InstanceType CRD and… → Attaches the cluster as a Kubernetes… → …
  • Tasks that involve Container orchestration
  • SKILL.md covers Prerequisites, Choose Where Each Command Runs, Connect the Host to Arc and Install K3s and Configure the…, plus 6 more sections
  • Calls az, kubectl and sh; needs HF_TOKEN

What it does

Azureml K3s Compute Target Setup is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization. Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target. Includes GPU smoke-test and validation instructions for Azure ML jobs on the Arc-connected cluster - Brought to you by microsoft/physical-ai-toolchain

Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Container orchestration and QA and bug reports. It works with Microsoft Azure, Azure Machine Learning, Kubernetes and NVIDIA AI Platform. The licence is MIT.

When your agent uses it

  • Tasks that involve Container orchestration
  • Tasks that involve QA and bug reports

Example prompts

  • “/azureml-k3s-compute-target-setup”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Installs a training-only Azure ML extension from azureml-arc-config.template.json, with the extension's own device plugin and DCGM…
  2. Waits for the InstanceType CRD and applies azureml-instance-types-hil.yaml through Arc cluster connect: defaultinstancetype for CPU jobs…
  3. Attaches the cluster as a Kubernetes compute with a system-assigned identity in the azureml namespace.
  4. Grants that identity AzureML Data Scientist on the workspace, Storage Blob Data Contributor on its storage account, and AcrPull on its…

What it can do on your machine

Read from SKILL.md and the folder at commit 0b12fe8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • az
    • kubectl
    • sh
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use az and kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Azureml K3s Compute Target Setup loads about 5.7k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 1,941 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:171
    TOKEN` in the untracked repository-root `.env.local`, never as a CLI argument or in chat. The submission script loads `.
  • NoteMentions a .env fileSKILL.md:184
    GROUP`, and `AZUREML_WORKSPACE_NAME` in `.env.local` or pass `--subscription-id`, `--resource-group`, and `--workspace-n
  • NoteMentions a .env fileSKILL.md:256
    on the model page, refresh the token in `.env.local` if needed, then resubmit |

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/physical-ai-toolchain at commit 0b12fe8, republished under its MIT licence (© microsoft). 1,941 words, ~5,703 tokens.

Download SKILL.mdSave it as .claude/skills/azureml-k3s-compute-target-setup/SKILL.md (or your agent's skills folder).
name
azureml-k3s-compute-target-setup
description
Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target. Includes GPU smoke-test and validation instructions for Azure ML jobs on the Arc-connected cluster - Brought to you by microsoft/physical-ai-toolchain

Azure ML K3s Compute Target Setup

Set up K3s on an Ubuntu host with an NVIDIA GPU, connect the host and cluster to Azure Arc, and attach the cluster to Azure ML as a Kubernetes compute target. Then prove real CUDA execution with a bounded smoke job before submitting a long-running training job. A GPU reservation in the Azure ML job spec is not proof the container can see the device; validate at the container level every time the runtime stack changes.

Prerequisites

RequirementPurpose
Ubuntu host with an NVIDIA GPU and driver installedK3s workload target
Existing Arc resource group, subscription ID, tenant IDazcmagent and connectedk8s registration targets
Azure ML workspace deployed by this repository's Terraform, with its outputs on the workstation05-attach-hil-azureml-compute.sh reads the workspace from those outputs
Contributor on the Arc cluster's resource group, and rights to assign roles on the workspace, its storage, and its container registryThe attach creates an Azure Relay and grants the compute identity access
az CLI with connectedk8s, k8s-extension, ml, and ssh extensionsArc and Azure ML operations, and remote access to the host
NVIDIA Container Toolkit installed on the hostProvides the nvidia-container-runtime binary K3s detects
HuggingFace account with access to any gated base modelRequired only when warm-starting from a gated repository (for example google/paligemma-3b-pt-224)

Choose Where Each Command Runs

Host setup and the kubectl checks in this skill run on the GPU host, because the K3s kubeconfig points at the host-local API server. The attach script reaches the cluster through Arc cluster connect, so it runs from a workstation, as do job submission commands.

StepRuns on
Connect the host to Arc, install K3s, enable the GPU, connect the cluster to ArcGPU host, from a clone of this repository
kubectl checks against the host's K3sGPU host
05-attach-hil-azureml-compute.sh: extension, InstanceTypes, compute attach, role assignmentsWorkstation with this repository's Terraform outputs
Job submission, az ml job show, and az ml job streamAny workstation

Before running a host step, confirm you are on the target host. Run these checks and compare the hostname with the intended host:

bash
hostname
nvidia-smi -L
systemctl is-active k3s

If nvidia-smi is missing or the hostname does not match, you are not on the GPU host. Connect to it before continuing:

  • Before the host is Arc-connected, use regular SSH or console access to reach it, then clone this repository on the host.
  • After Connect the Host to Arc completes, use az ssh arc. It needs no public IP address or inbound port, and requires Owner or Contributor on the Arc-enabled server and the Microsoft.HybridConnectivity resource provider.

Enable SSH on the Arc-enabled server once, from a workstation:

bash
az extension add --name ssh
az provider register -n Microsoft.HybridConnectivity

machine_id=/subscriptions/<subscription-id>/resourceGroups/<arc-resource-group>/providers/Microsoft.HybridCompute/machines/<arc-server-name>
az rest --method put \
  --uri "https://management.azure.com${machine_id}/providers/Microsoft.HybridConnectivity/endpoints/default?api-version=2023-03-15" \
  --body '{"properties": {"type": "default"}}'
az rest --method put \
  --uri "https://management.azure.com${machine_id}/providers/Microsoft.HybridConnectivity/endpoints/default/serviceconfigurations/SSH?api-version=2023-03-15" \
  --body '{"properties": {"serviceName": "SSH", "port": 22}}'

Open a session on the host, then run the host steps inside it:

bash
az ssh arc --resource-group <arc-resource-group> --name <arc-server-name> --local-user <host-user>

Inside the session, point kubectl at the protected kubeconfig written by the K3s installer:

bash
export KUBECONFIG="$HOME/.local/share/physical-ai-toolchain/hil/kubeconfig.yaml"

Connect the Host to Arc

Run on the GPU host. Preview, then connect the Ubuntu server:

bash
data-pipeline/setup/edge/03-connect-arc-server.sh \
  --subscription-id <subscription-id> --tenant-id <tenant-id> \
  --resource-group <arc-resource-group> --location <location> \
  --server-name <arc-server-name> --config-preview

data-pipeline/setup/edge/03-connect-arc-server.sh \
  --subscription-id <subscription-id> --tenant-id <tenant-id> \
  --resource-group <arc-resource-group> --location <location> \
  --server-name <arc-server-name>

Install K3s and Configure the GPU Default Runtime

Run on the GPU host. Install the pinned, owned K3s compute plane, passing --default-runtime nvidia on any host where GPU-backed Azure ML jobs will run:

bash
data-pipeline/setup/hil/01-install-k3s.sh --default-runtime nvidia --config-preview
data-pipeline/setup/hil/01-install-k3s.sh --default-runtime nvidia

[!IMPORTANT] Azure ML InstanceType limits.nvidia.com/gpu only reserves the device through the Kubernetes device plugin. It does not set a pod runtimeClassName, so on a K3s host whose default runtime is not NVIDIA, the container starts with no /dev/nvidia* devices and torch.cuda.device_count() returns 0 even though the job reports a reserved GPU. Setting default-runtime: nvidia in the K3s config is required whenever the Azure ML Kubernetes extension does not expose per-pod runtime-class configuration.

Then run 05-enable-k3s-gpu.sh on the host. Its preview checks the driver, the NVIDIA Container Toolkit, the K3s service and API, and the nvidia RuntimeClass, and says whether the run will restart K3s:

bash
data-pipeline/setup/hil/05-enable-k3s-gpu.sh --config-preview
data-pipeline/setup/hil/05-enable-k3s-gpu.sh

The script installs the digest-pinned NVIDIA device plugin so the node advertises nvidia.com/gpu. When nvidia is already the default runtime, as after --default-runtime nvidia, it leaves K3s running. On hosts installed without that option, it writes the K3s default-runtime: nvidia drop-in and restarts K3s once, so run it when no job containers are running.

Verify the NVIDIA device plugin is healthy and the node advertises allocatable GPUs before attaching Azure ML:

bash
kubectl get pods -n kube-system -l name=nvidia-device-plugin-ds
kubectl get node <node-name> -o jsonpath='{.status.allocatable.nvidia\.com/gpu}'

Connect the K3s Cluster to Arc

Run on the GPU host. Connect the cluster with the OIDC issuer and workload identity enabled, and grant yourself cluster-admin so Arc cluster connect accepts your identity:

bash
data-pipeline/setup/edge/05-connect-arc-kubernetes.sh \
  --subscription-id <subscription-id> --tenant-id <tenant-id> \
  --resource-group <arc-resource-group> --location <location> \
  --cluster-name <arc-cluster-name> --kubeconfig <kubeconfig-path> \
  --enable-workload-identity --cluster-admin-signed-in-user --config-preview

data-pipeline/setup/edge/05-connect-arc-kubernetes.sh \
  --subscription-id <subscription-id> --tenant-id <tenant-id> \
  --resource-group <arc-resource-group> --location <location> \
  --cluster-name <arc-cluster-name> --kubeconfig <kubeconfig-path> \
  --enable-workload-identity --cluster-admin-signed-in-user

The attach script applies InstanceTypes through Arc cluster connect, which checks K3s RBAC for the identity that runs it. When someone else runs the attach, grant them with --cluster-admin-object-id <object-id> instead.

Attach the Cluster to Azure ML

infrastructure/setup/02-deploy-azureml-extension.sh targets the Terraform-managed AKS cluster. For an Arc-connected K3s cluster, run 05-attach-hil-azureml-compute.sh from a workstation that has this repository's Terraform outputs. Preview, then attach:

bash
infrastructure/setup/05-attach-hil-azureml-compute.sh \
  --arc-cluster-resource-id /subscriptions/<subscription-id>/resourceGroups/<arc-resource-group>/providers/Microsoft.Kubernetes/connectedClusters/<arc-cluster-name> \
  --compute-name <compute-name> --config-preview

infrastructure/setup/05-attach-hil-azureml-compute.sh \
  --arc-cluster-resource-id /subscriptions/<subscription-id>/resourceGroups/<arc-resource-group>/providers/Microsoft.Kubernetes/connectedClusters/<arc-cluster-name> \
  --compute-name <compute-name> --require-gpu

The script:

  1. Installs a training-only Azure ML extension from azureml-arc-config.template.json, with the extension's own device plugin and DCGM exporter off, or updates only the settings that differ on an existing extension.
  2. Waits for the InstanceType CRD and applies azureml-instance-types-hil.yaml through Arc cluster connect: defaultinstancetype for CPU jobs and gpu for one GPU. The gpu type has no node selector, so the node needs no label.
  3. Attaches the cluster as a Kubernetes compute with a system-assigned identity in the azureml namespace.
  4. Grants that identity AzureML Data Scientist on the workspace, Storage Blob Data Contributor on its storage account, and AcrPull on its container registry, so jobs can pull images from it.

Compute names are 16 characters at most, and the default, k8s-<cluster>, is truncated, so pass --compute-name. --require-gpu stops before any change unless a node reports allocatable nvidia.com/gpu. For InstanceTypes that request more GPUs, pass your own manifest with --instance-types-manifest, and request only what the node advertises.

The extension creates an Azure Relay namespace and hybrid connection in the Arc cluster's resource group. Don't modify them, because the compute depends on them.

Azure ML Credentials and Dataset Inputs

RequirementWhere it applies
az login session with rights to the target subscription and workspaceAll az ml submission commands
AzureML Data Scientist on the workspace and Storage Blob Data Contributor on its storage account for the compute identityGranted by 05-attach-hil-azureml-compute.sh; required for data asset mounts, outputs, and MLflow
HF_TOKEN with access to the gated base repositoryOnly when --policy-repo-id or the HuggingFace dataset path resolves to a gated repository
Datastore-backed Azure ML data asset, referenced with an explicit numeric version (azureml:NAME:VERSION)--dataset-asset; shorthand references without a version are rejected to keep runs reproducible

Store HF_TOKEN in the untracked repository-root .env.local, never as a CLI argument or in chat. The submission script loads .env.local and forwards HF_TOKEN to the job, so --hf-token is not needed. Tokens passed as CLI arguments are visible to any process inspecting the host, so rotate a token immediately if it was ever exposed that way. .amlignore already excludes .env and .env.* from the Azure ML code snapshot.

Data asset mount failures during job start often mean the registered asset version does not resolve against its backing datastore. Register a new datastore-backed version and reference that explicit version rather than reusing a broken one:

bash
az ml data create --name <dataset-name> --version <next-version> \
  --type uri_folder --path azureml://datastores/<datastore>/paths/<path>
Show full SKILL.md (768 more words)Show less

Submit a Bounded GPU Smoke Test

To check the GPU and the services training depends on before any model runs, submit training/smoke/scripts/submit-azureml-gpu-smoke.sh --compute <compute-name> --instance-type gpu --stream first. See Smoke-Test a GPU Target.

Submit a short run (10 to 20 steps) before committing to a full training job. Keep --save-freq at or below the step count so at least one checkpoint round-trips. Pass --compute with the attached compute name. When Terraform outputs are unavailable, set AZURE_SUBSCRIPTION_ID, AZURE_RESOURCE_GROUP, and AZUREML_WORKSPACE_NAME in .env.local or pass --subscription-id, --resource-group, and --workspace-name:

bash
training/vla/scripts/submit-azureml-vla-pi0-training.sh \
  -d <dataset-repo-id> --dataset-asset azureml:<dataset-name>:<version> \
  -p pi05 --policy-repo-id lerobot/pi05_base \
  --training-steps 10 --batch-size 16 --save-freq 10 --log-freq 1 \
  --train-expert-only --mixed-precision bf16 \
  --compute <compute-name> --instance-type gpu \
  -j <job-name> --config-preview

training/vla/scripts/submit-azureml-vla-pi0-training.sh \
  -d <dataset-repo-id> --dataset-asset azureml:<dataset-name>:<version> \
  -p pi05 --policy-repo-id lerobot/pi05_base \
  --training-steps 10 --batch-size 16 --save-freq 10 --log-freq 1 \
  --train-expert-only --mixed-precision bf16 \
  --compute <compute-name> --instance-type gpu \
  -j <job-name>

Validate GPU Usage Without Interrupting the Job

Never cancel a running job to validate it. Use read-only checks against the live pod and a bounded log stream instead.

The az ml commands run from any workstation. The kubectl commands run on the GPU host; when you submitted the job from another machine, open an az ssh arc session on the host as described in Choose Where Each Command Runs and set KUBECONFIG there first.

Confirm the job status from the workstation, then find its pod on the host:

bash
az ml job show --name <job-name> --query '{status:status}' -o json
bash
kubectl get pods -n azureml -l azureml.job.name=<job-name>

Confirm NVIDIA devices and live utilization inside the execution-wrapper container, on the host:

bash
kubectl exec -n azureml <pod-name> -c <pod-name-execution-wrapper> -- \
  sh -c 'ls /dev/nvidia*; nvidia-smi --query-gpu=name,memory.total,memory.used,utilization.gpu --format=csv,noheader'

Confirm PyTorch itself reports the device, using the job's installed virtual environment rather than a system Python:

bash
kubectl exec -n azureml <pod-name> -c <pod-name-execution-wrapper> -- \
  /opt/lerobot-venv/bin/python -c "import torch; print(torch.cuda.is_available(), torch.cuda.device_count())"

Stream a bounded window of user logs from the workstation and look for the training entrypoint's own detection line rather than relying on nvidia-smi alone, because it confirms the training process itself, not just the container, sees CUDA:

bash
timeout --signal=INT 20s az ml job stream --name <job-name>

The training entrypoint logs [GPU-DETECT] torch.cuda.device_count()=<n>, CUDA_VISIBLE_DEVICES=<value> from train.py before invoking lerobot-train. A count of 0 on a job with a reserved GPU means the runtime injection is broken, not that the GPU is unavailable. timeout only stops the local log viewer; it does not cancel the Azure ML job.

[GPU-DETECT] only proves the container sees a device, not that training is using it. Confirm real GPU execution from the lerobot-train step log lines themselves: a rising mem_gb value across steps, a step:<n> counter advancing toward --steps, and a Checkpoint policy after step <n> line once the run finishes.

A smoke run reaching all configured steps within tens of seconds with mem_gb in the low double digits confirms the GPU did the work. A run stuck at step 0 or taking hours per step, even with device_count=1, still indicates the GPU is not actually being used.

An Azure ML job can stay in Running after the training loop finishes while large checkpoints upload; a pretrained_model payload of several gigabytes can take many minutes after the last step:<n> log line. Do not treat this as a hang.

A [MLflow] Failed to log artifacts for <step> message citing a task-queue flush timeout is a transient MLflow tracking-API timeout, not a training or upload failure, as long as the raw upload progress bar that follows it reaches 100%.

Troubleshooting

SymptomLikely CauseResolution
Job reserves one GPU but [GPU-DETECT] torch.cuda.device_count()=0Pod has no runtimeClassName and K3s default runtime is not NVIDIARun data-pipeline/setup/hil/05-enable-k3s-gpu.sh on the host once no job containers are active
nvidia-smi missing or /dev/nvidia* absent inside the containerSame root cause as aboveConfirm with the kubectl exec device checks in the validation section, then apply the K3s default runtime fix
403 fetching a gated HuggingFace repositoryAccount lacks gated-repo access, or HF_TOKEN is staleSign in to huggingface.co as the account that owns HF_TOKEN, request access on the model page, refresh the token in .env.local if needed, then resubmit
Data asset mount fails at job startAsset version not backed by a resolvable datastore pathRegister a new datastore-backed asset version and reference it explicitly
05-attach-hil-azureml-compute.sh stops at Arc cluster connectYour identity has no K3s RBAC on the cluster, or the proxy port is in useGrant access with 05-connect-arc-kubernetes.sh --cluster-admin-signed-in-user or --cluster-admin-object-id, or pass --proxy-port
Job stays Queued with the gpu instance typeThe node reports no allocatable nvidia.com/gpu, usually because the device plugin isn't runningRun 05-enable-k3s-gpu.sh on the host, then check kubectl get node <node-name> -o jsonpath='{.status.allocatable.nvidia\.com/gpu}'
Training runs entirely on CPU with no errorSame GPU runtime-injection root cause; PyTorch silently falls backApply the K3s default runtime fix before assuming a code-level bug
Job stays Running well after the last step:<n> log lineLarge checkpoint still uploading to blob storageCheck for an active upload progress bar in the log before assuming a hang
[MLflow] Failed to log artifacts for <step>: ... Failed to flush task queue within 300.0 secondsTransient MLflow tracking-API timeout, unrelated to the checkpoint data itselfConfirm the raw upload progress bar immediately after it reaches 100%; no data is lost

Brought to you by microsoft/physical-ai-toolchain

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/azureml-k3s-compute-target-setup of microsoft/physical-ai-toolchain.

Open the folder on GitHubat commit 0b12fe8

Compare with similar skills

Azureml K3s Compute Target Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Azureml K3s Compute Target Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Azureml K3s Compute Target Setup this skillmicrosoft/physical-ai-toolchain123—~5.7kAutomated safety check: NotesMIT
Devops Pipelineluongnv89/skills131—~4.8kAutomated safety check: PassMIT
Azure Diagnosticsmicrosoft/azure-skills1.5k1 repos~1.6kAutomated safety check: PassMIT
Nim Operator InstallNVIDIA/k8s-nim-operator159—~4.7kAutomated safety check: PassApache-2.0
Provider Bug Reviewmondoohq/mql411—~2.9kAutomated safety check: PassCustom licence
Nim Operator UninstallNVIDIA/k8s-nim-operator159—~3.6kAutomated safety check: PassApache-2.0

Similar skills

  • Devops Pipeline

    luongnv89/skills

    Configure pre-commit hooks and lean GitHub Actions for shift-left quality assurance.

    131 GitHub stars~4.8k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Azure Diagnostics

    microsoft/azure-skills

    Official

    Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage.

    1.5k GitHub starsUsed in 1 repo~1.6k tokens
    DevOps & CloudAuto-check passed
  • Nim Operator Install

    NVIDIA/k8s-nim-operator

    Official

    Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional…

    159 GitHub stars~4.7k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Deep static code review of an mql provider for logic errors, nil-handling bugs, pagination truncation, caching/id collisions, and other defects that silently give users wrong data.

    411 GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Nim Operator Uninstall

    NVIDIA/k8s-nim-operator

    Official

    Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and…

    159 GitHub stars~3.6k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Official

    A skill your agent uses when validating DCGM Exporter in a local GPU-backed k3d/Kubernetes environment.

    1.9k GitHub stars~116 tokensUpdated 19 days ago
    DevOps & CloudAuto-check passed

More from microsoft/physical-ai-toolchain

  • Osmo Lerobot Training

    microsoft/physical-ai-toolchain

    Official

    Submit, monitor, analyze, and evaluate LeRobot imitation learning training jobs on OSMO with Azure ML MLflow integration and inference evaluation - Brought to you by microsoft/physical-ai-toolchain

    123 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check: notes
  • Environment Deployment

    microsoft/physical-ai-toolchain

    Official

    Generate, transfer, and consume environment-specific Azure, AKS, OSMO, ACR, and Azure ML deployment bundles.

    123 GitHub stars~5.8k tokensUpdated yesterday
    Auto-check passed
  • Fleet Deployment

    microsoft/physical-ai-toolchain

    Official

    Deploy trained robot policies to edge fleets via FluxCD GitOps, image automation, and deployment gating

    123 GitHub stars~518 tokensUpdated yesterday
    Auto-check passed
  • Fleet Intelligence

    microsoft/physical-ai-toolchain

    Official

    Monitor robot fleet telemetry via Azure IoT Operations, drift detection, Grafana dashboards, and Fabric analytics

    123 GitHub stars~598 tokensUpdated yesterday
    Auto-check passed
  • Infrastructure

    microsoft/physical-ai-toolchain

    Official

    Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology

    123 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Synthetic Data

    microsoft/physical-ai-toolchain

    Official

    Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines

    123 GitHub stars~469 tokensUpdated yesterday
    Auto-check passed

Questions about Azureml K3s Compute Target Setup

What does Azureml K3s Compute Target Setup do?

Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target. Azureml K3s Compute Target Setup is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization. Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.

When should I use Azureml K3s Compute Target Setup?

Azureml K3s Compute Target Setup fits situations like: tasks that involve Container orchestration; tasks that involve QA and bug reports.

How do I install Azureml K3s Compute Target Setup in Claude Code?

Run `npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a claude-code`. Or copy the skill folder (.github/skills/azureml-k3s-compute-target-setup in microsoft/physical-ai-toolchain) into .claude/skills/azureml-k3s-compute-target-setup in your project. Claude Code loads it when a task matches its description.

How do I install Azureml K3s Compute Target Setup in Codex?

Run `npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a codex`. Or copy the skill folder (.github/skills/azureml-k3s-compute-target-setup in microsoft/physical-ai-toolchain) into .agents/skills/azureml-k3s-compute-target-setup in your project. Codex loads it when a task matches its description.

Can I use Azureml K3s Compute Target Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/azureml-k3s-compute-target-setup, .gemini/skills/azureml-k3s-compute-target-setup, .github/skills/azureml-k3s-compute-target-setup and .opencode/skills/azureml-k3s-compute-target-setup in your project.

What does Azureml K3s Compute Target Setup need to run?

Going by SKILL.md and its folder, Azureml K3s Compute Target Setup needs the command-line tools its instructions call (az, kubectl, sh and python) and credentials named HF_TOKEN.

Does Azureml K3s Compute Target Setup access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Azureml K3s Compute Target Setup safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Azureml K3s Compute Target Setup use?

Azureml K3s Compute Target Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Azureml K3s Compute Target Setup use?

About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Azureml K3s Compute Target Setup?

Skills that share tags, products or a category with Azureml K3s Compute Target Setup: Devops Pipeline (luongnv89/skills, 131 stars), Azure Diagnostics (microsoft/azure-skills, 1.5k stars), Nim Operator Install (NVIDIA/k8s-nim-operator, 159 stars) and Provider Bug Review (mondoohq/mql, 411 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Azureml K3s Compute Target Setup?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/physical-ai-toolchain, which has 123 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 7, 2026.

Source: microsoft/physical-ai-toolchain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.