Devops Pipeline
luongnv89/skills
Configure pre-commit hooks and lean GitHub Actions for shift-left quality assurance.
Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.
$ npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/physical-ai-toolchain azureml-k3s-compute-target-setup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/azureml-k3s-compute-target-setup .claude/skills/azureml-k3s-compute-target-setup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "azureml-k3s-compute-target-setup" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/azureml-k3s-compute-target-setup into .claude/skills/azureml-k3s-compute-target-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "azureml-k3s-compute-target-setup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/azureml-k3s-compute-target-setupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/physical-ai-toolchain azureml-k3s-compute-target-setup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/azureml-k3s-compute-target-setup .agents/skills/azureml-k3s-compute-target-setup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "azureml-k3s-compute-target-setup" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/azureml-k3s-compute-target-setup into .agents/skills/azureml-k3s-compute-target-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "azureml-k3s-compute-target-setup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/physical-ai-toolchain azureml-k3s-compute-target-setup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/azureml-k3s-compute-target-setup .cursor/skills/azureml-k3s-compute-target-setup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "azureml-k3s-compute-target-setup" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/azureml-k3s-compute-target-setup into .cursor/skills/azureml-k3s-compute-target-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "azureml-k3s-compute-target-setup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/physical-ai-toolchain.git --path .github/skills/azureml-k3s-compute-target-setup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/physical-ai-toolchain azureml-k3s-compute-target-setup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/azureml-k3s-compute-target-setup .gemini/skills/azureml-k3s-compute-target-setup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "azureml-k3s-compute-target-setup" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/azureml-k3s-compute-target-setup into .gemini/skills/azureml-k3s-compute-target-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "azureml-k3s-compute-target-setup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/physical-ai-toolchain azureml-k3s-compute-target-setupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/azureml-k3s-compute-target-setup .github/skills/azureml-k3s-compute-target-setup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "azureml-k3s-compute-target-setup" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/azureml-k3s-compute-target-setup into .github/skills/azureml-k3s-compute-target-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "azureml-k3s-compute-target-setup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/physical-ai-toolchain azureml-k3s-compute-target-setup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/azureml-k3s-compute-target-setup .opencode/skills/azureml-k3s-compute-target-setup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "azureml-k3s-compute-target-setup" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/azureml-k3s-compute-target-setup into .opencode/skills/azureml-k3s-compute-target-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "azureml-k3s-compute-target-setup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
azureml-k3s-compute-target-setupSet up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.
Azureml K3s Compute Target Setup is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization. Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target. Includes GPU smoke-test and validation instructions for Azure ML jobs on the Arc-connected cluster - Brought to you by microsoft/physical-ai-toolchain
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Container orchestration and QA and bug reports. It works with Microsoft Azure, Azure Machine Learning, Kubernetes and NVIDIA AI Platform. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 0b12fe8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
azkubectlshpythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use az and kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Azureml K3s Compute Target Setup loads about 5.7k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 1,941 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
TOKEN` in the untracked repository-root `.env.local`, never as a CLI argument or in chat. The submission script loads `.GROUP`, and `AZUREML_WORKSPACE_NAME` in `.env.local` or pass `--subscription-id`, `--resource-group`, and `--workspace-non the model page, refresh the token in `.env.local` if needed, then resubmit |Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/physical-ai-toolchain at commit 0b12fe8, republished under its MIT licence (© microsoft). 1,941 words, ~5,703 tokens.
.claude/skills/azureml-k3s-compute-target-setup/SKILL.md (or your agent's skills folder).Set up K3s on an Ubuntu host with an NVIDIA GPU, connect the host and cluster to Azure Arc, and attach the cluster to Azure ML as a Kubernetes compute target. Then prove real CUDA execution with a bounded smoke job before submitting a long-running training job. A GPU reservation in the Azure ML job spec is not proof the container can see the device; validate at the container level every time the runtime stack changes.
| Requirement | Purpose |
|---|---|
| Ubuntu host with an NVIDIA GPU and driver installed | K3s workload target |
| Existing Arc resource group, subscription ID, tenant ID | azcmagent and connectedk8s registration targets |
| Azure ML workspace deployed by this repository's Terraform, with its outputs on the workstation | 05-attach-hil-azureml-compute.sh reads the workspace from those outputs |
| Contributor on the Arc cluster's resource group, and rights to assign roles on the workspace, its storage, and its container registry | The attach creates an Azure Relay and grants the compute identity access |
az CLI with connectedk8s, k8s-extension, ml, and ssh extensions | Arc and Azure ML operations, and remote access to the host |
| NVIDIA Container Toolkit installed on the host | Provides the nvidia-container-runtime binary K3s detects |
| HuggingFace account with access to any gated base model | Required only when warm-starting from a gated repository (for example google/paligemma-3b-pt-224) |
Host setup and the kubectl checks in this skill run on the GPU host, because the K3s kubeconfig points at the host-local API server. The attach script reaches the cluster through Arc cluster connect, so it runs from a workstation, as do job submission commands.
| Step | Runs on |
|---|---|
| Connect the host to Arc, install K3s, enable the GPU, connect the cluster to Arc | GPU host, from a clone of this repository |
kubectl checks against the host's K3s | GPU host |
05-attach-hil-azureml-compute.sh: extension, InstanceTypes, compute attach, role assignments | Workstation with this repository's Terraform outputs |
Job submission, az ml job show, and az ml job stream | Any workstation |
Before running a host step, confirm you are on the target host. Run these checks and compare the hostname with the intended host:
hostname
nvidia-smi -L
systemctl is-active k3sIf nvidia-smi is missing or the hostname does not match, you are not on the GPU host. Connect to it before continuing:
az ssh arc. It needs no public IP address or inbound port, and requires Owner or Contributor on the Arc-enabled server and the Microsoft.HybridConnectivity resource provider.Enable SSH on the Arc-enabled server once, from a workstation:
az extension add --name ssh
az provider register -n Microsoft.HybridConnectivity
machine_id=/subscriptions/<subscription-id>/resourceGroups/<arc-resource-group>/providers/Microsoft.HybridCompute/machines/<arc-server-name>
az rest --method put \
--uri "https://management.azure.com${machine_id}/providers/Microsoft.HybridConnectivity/endpoints/default?api-version=2023-03-15" \
--body '{"properties": {"type": "default"}}'
az rest --method put \
--uri "https://management.azure.com${machine_id}/providers/Microsoft.HybridConnectivity/endpoints/default/serviceconfigurations/SSH?api-version=2023-03-15" \
--body '{"properties": {"serviceName": "SSH", "port": 22}}'Open a session on the host, then run the host steps inside it:
az ssh arc --resource-group <arc-resource-group> --name <arc-server-name> --local-user <host-user>Inside the session, point kubectl at the protected kubeconfig written by the K3s installer:
export KUBECONFIG="$HOME/.local/share/physical-ai-toolchain/hil/kubeconfig.yaml"Run on the GPU host. Preview, then connect the Ubuntu server:
data-pipeline/setup/edge/03-connect-arc-server.sh \
--subscription-id <subscription-id> --tenant-id <tenant-id> \
--resource-group <arc-resource-group> --location <location> \
--server-name <arc-server-name> --config-preview
data-pipeline/setup/edge/03-connect-arc-server.sh \
--subscription-id <subscription-id> --tenant-id <tenant-id> \
--resource-group <arc-resource-group> --location <location> \
--server-name <arc-server-name>Run on the GPU host. Install the pinned, owned K3s compute plane, passing --default-runtime nvidia on any host where GPU-backed Azure ML jobs will run:
data-pipeline/setup/hil/01-install-k3s.sh --default-runtime nvidia --config-preview
data-pipeline/setup/hil/01-install-k3s.sh --default-runtime nvidia[!IMPORTANT] Azure ML InstanceType
limits.nvidia.com/gpuonly reserves the device through the Kubernetes device plugin. It does not set a podruntimeClassName, so on a K3s host whose default runtime is not NVIDIA, the container starts with no/dev/nvidia*devices andtorch.cuda.device_count()returns0even though the job reports a reserved GPU. Settingdefault-runtime: nvidiain the K3s config is required whenever the Azure ML Kubernetes extension does not expose per-pod runtime-class configuration.
Then run 05-enable-k3s-gpu.sh on the host. Its preview checks the driver, the NVIDIA Container Toolkit, the K3s service and API, and the nvidia RuntimeClass, and says whether the run will restart K3s:
data-pipeline/setup/hil/05-enable-k3s-gpu.sh --config-preview
data-pipeline/setup/hil/05-enable-k3s-gpu.shThe script installs the digest-pinned NVIDIA device plugin so the node advertises nvidia.com/gpu. When nvidia is already the default runtime, as after --default-runtime nvidia, it leaves K3s running. On hosts installed without that option, it writes the K3s default-runtime: nvidia drop-in and restarts K3s once, so run it when no job containers are running.
Verify the NVIDIA device plugin is healthy and the node advertises allocatable GPUs before attaching Azure ML:
kubectl get pods -n kube-system -l name=nvidia-device-plugin-ds
kubectl get node <node-name> -o jsonpath='{.status.allocatable.nvidia\.com/gpu}'Run on the GPU host. Connect the cluster with the OIDC issuer and workload identity enabled, and grant yourself cluster-admin so Arc cluster connect accepts your identity:
data-pipeline/setup/edge/05-connect-arc-kubernetes.sh \
--subscription-id <subscription-id> --tenant-id <tenant-id> \
--resource-group <arc-resource-group> --location <location> \
--cluster-name <arc-cluster-name> --kubeconfig <kubeconfig-path> \
--enable-workload-identity --cluster-admin-signed-in-user --config-preview
data-pipeline/setup/edge/05-connect-arc-kubernetes.sh \
--subscription-id <subscription-id> --tenant-id <tenant-id> \
--resource-group <arc-resource-group> --location <location> \
--cluster-name <arc-cluster-name> --kubeconfig <kubeconfig-path> \
--enable-workload-identity --cluster-admin-signed-in-userThe attach script applies InstanceTypes through Arc cluster connect, which checks K3s RBAC for the identity that runs it. When someone else runs the attach, grant them with --cluster-admin-object-id <object-id> instead.
infrastructure/setup/02-deploy-azureml-extension.sh targets the Terraform-managed AKS cluster. For an Arc-connected K3s cluster, run 05-attach-hil-azureml-compute.sh from a workstation that has this repository's Terraform outputs. Preview, then attach:
infrastructure/setup/05-attach-hil-azureml-compute.sh \
--arc-cluster-resource-id /subscriptions/<subscription-id>/resourceGroups/<arc-resource-group>/providers/Microsoft.Kubernetes/connectedClusters/<arc-cluster-name> \
--compute-name <compute-name> --config-preview
infrastructure/setup/05-attach-hil-azureml-compute.sh \
--arc-cluster-resource-id /subscriptions/<subscription-id>/resourceGroups/<arc-resource-group>/providers/Microsoft.Kubernetes/connectedClusters/<arc-cluster-name> \
--compute-name <compute-name> --require-gpuThe script:
InstanceType CRD and applies azureml-instance-types-hil.yaml through Arc cluster connect: defaultinstancetype for CPU jobs and gpu for one GPU. The gpu type has no node selector, so the node needs no label.azureml namespace.Compute names are 16 characters at most, and the default, k8s-<cluster>, is truncated, so pass --compute-name. --require-gpu stops before any change unless a node reports allocatable nvidia.com/gpu. For InstanceTypes that request more GPUs, pass your own manifest with --instance-types-manifest, and request only what the node advertises.
The extension creates an Azure Relay namespace and hybrid connection in the Arc cluster's resource group. Don't modify them, because the compute depends on them.
| Requirement | Where it applies |
|---|---|
az login session with rights to the target subscription and workspace | All az ml submission commands |
AzureML Data Scientist on the workspace and Storage Blob Data Contributor on its storage account for the compute identity | Granted by 05-attach-hil-azureml-compute.sh; required for data asset mounts, outputs, and MLflow |
HF_TOKEN with access to the gated base repository | Only when --policy-repo-id or the HuggingFace dataset path resolves to a gated repository |
Datastore-backed Azure ML data asset, referenced with an explicit numeric version (azureml:NAME:VERSION) | --dataset-asset; shorthand references without a version are rejected to keep runs reproducible |
Store HF_TOKEN in the untracked repository-root .env.local, never as a CLI argument or in chat. The submission script loads .env.local and forwards HF_TOKEN to the job, so --hf-token is not needed. Tokens passed as CLI arguments are visible to any process inspecting the host, so rotate a token immediately if it was ever exposed that way. .amlignore already excludes .env and .env.* from the Azure ML code snapshot.
Data asset mount failures during job start often mean the registered asset version does not resolve against its backing datastore. Register a new datastore-backed version and reference that explicit version rather than reusing a broken one:
az ml data create --name <dataset-name> --version <next-version> \
--type uri_folder --path azureml://datastores/<datastore>/paths/<path>To check the GPU and the services training depends on before any model runs, submit training/smoke/scripts/submit-azureml-gpu-smoke.sh --compute <compute-name> --instance-type gpu --stream first. See Smoke-Test a GPU Target.
Submit a short run (10 to 20 steps) before committing to a full training job. Keep --save-freq at or below the step count so at least one checkpoint round-trips. Pass --compute with the attached compute name. When Terraform outputs are unavailable, set AZURE_SUBSCRIPTION_ID, AZURE_RESOURCE_GROUP, and AZUREML_WORKSPACE_NAME in .env.local or pass --subscription-id, --resource-group, and --workspace-name:
training/vla/scripts/submit-azureml-vla-pi0-training.sh \
-d <dataset-repo-id> --dataset-asset azureml:<dataset-name>:<version> \
-p pi05 --policy-repo-id lerobot/pi05_base \
--training-steps 10 --batch-size 16 --save-freq 10 --log-freq 1 \
--train-expert-only --mixed-precision bf16 \
--compute <compute-name> --instance-type gpu \
-j <job-name> --config-preview
training/vla/scripts/submit-azureml-vla-pi0-training.sh \
-d <dataset-repo-id> --dataset-asset azureml:<dataset-name>:<version> \
-p pi05 --policy-repo-id lerobot/pi05_base \
--training-steps 10 --batch-size 16 --save-freq 10 --log-freq 1 \
--train-expert-only --mixed-precision bf16 \
--compute <compute-name> --instance-type gpu \
-j <job-name>Never cancel a running job to validate it. Use read-only checks against the live pod and a bounded log stream instead.
The az ml commands run from any workstation. The kubectl commands run on the GPU host; when you submitted the job from another machine, open an az ssh arc session on the host as described in Choose Where Each Command Runs and set KUBECONFIG there first.
Confirm the job status from the workstation, then find its pod on the host:
az ml job show --name <job-name> --query '{status:status}' -o jsonkubectl get pods -n azureml -l azureml.job.name=<job-name>Confirm NVIDIA devices and live utilization inside the execution-wrapper container, on the host:
kubectl exec -n azureml <pod-name> -c <pod-name-execution-wrapper> -- \
sh -c 'ls /dev/nvidia*; nvidia-smi --query-gpu=name,memory.total,memory.used,utilization.gpu --format=csv,noheader'Confirm PyTorch itself reports the device, using the job's installed virtual environment rather than a system Python:
kubectl exec -n azureml <pod-name> -c <pod-name-execution-wrapper> -- \
/opt/lerobot-venv/bin/python -c "import torch; print(torch.cuda.is_available(), torch.cuda.device_count())"Stream a bounded window of user logs from the workstation and look for the training entrypoint's own detection line rather than relying on nvidia-smi alone, because it confirms the training process itself, not just the container, sees CUDA:
timeout --signal=INT 20s az ml job stream --name <job-name>The training entrypoint logs [GPU-DETECT] torch.cuda.device_count()=<n>, CUDA_VISIBLE_DEVICES=<value> from train.py before invoking lerobot-train. A count of 0 on a job with a reserved GPU means the runtime injection is broken, not that the GPU is unavailable. timeout only stops the local log viewer; it does not cancel the Azure ML job.
[GPU-DETECT] only proves the container sees a device, not that training is using it. Confirm real GPU execution from the lerobot-train step log lines themselves: a rising mem_gb value across steps, a step:<n> counter advancing toward --steps, and a Checkpoint policy after step <n> line once the run finishes.
A smoke run reaching all configured steps within tens of seconds with mem_gb in the low double digits confirms the GPU did the work. A run stuck at step 0 or taking hours per step, even with device_count=1, still indicates the GPU is not actually being used.
An Azure ML job can stay in Running after the training loop finishes while large checkpoints upload; a pretrained_model payload of several gigabytes can take many minutes after the last step:<n> log line. Do not treat this as a hang.
A [MLflow] Failed to log artifacts for <step> message citing a task-queue flush timeout is a transient MLflow tracking-API timeout, not a training or upload failure, as long as the raw upload progress bar that follows it reaches 100%.
| Symptom | Likely Cause | Resolution |
|---|---|---|
Job reserves one GPU but [GPU-DETECT] torch.cuda.device_count()=0 | Pod has no runtimeClassName and K3s default runtime is not NVIDIA | Run data-pipeline/setup/hil/05-enable-k3s-gpu.sh on the host once no job containers are active |
nvidia-smi missing or /dev/nvidia* absent inside the container | Same root cause as above | Confirm with the kubectl exec device checks in the validation section, then apply the K3s default runtime fix |
403 fetching a gated HuggingFace repository | Account lacks gated-repo access, or HF_TOKEN is stale | Sign in to huggingface.co as the account that owns HF_TOKEN, request access on the model page, refresh the token in .env.local if needed, then resubmit |
| Data asset mount fails at job start | Asset version not backed by a resolvable datastore path | Register a new datastore-backed asset version and reference it explicitly |
05-attach-hil-azureml-compute.sh stops at Arc cluster connect | Your identity has no K3s RBAC on the cluster, or the proxy port is in use | Grant access with 05-connect-arc-kubernetes.sh --cluster-admin-signed-in-user or --cluster-admin-object-id, or pass --proxy-port |
Job stays Queued with the gpu instance type | The node reports no allocatable nvidia.com/gpu, usually because the device plugin isn't running | Run 05-enable-k3s-gpu.sh on the host, then check kubectl get node <node-name> -o jsonpath='{.status.allocatable.nvidia\.com/gpu}' |
| Training runs entirely on CPU with no error | Same GPU runtime-injection root cause; PyTorch silently falls back | Apply the K3s default runtime fix before assuming a code-level bug |
Job stays Running well after the last step:<n> log line | Large checkpoint still uploading to blob storage | Check for an active upload progress bar in the log before assuming a hang |
[MLflow] Failed to log artifacts for <step>: ... Failed to flush task queue within 300.0 seconds | Transient MLflow tracking-API timeout, unrelated to the checkpoint data itself | Confirm the raw upload progress bar immediately after it reaches 100%; no data is lost |
Brought to you by microsoft/physical-ai-toolchain
© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .github/skills/azureml-k3s-compute-target-setup of microsoft/physical-ai-toolchain.
Open the folder on GitHubat commit 0b12fe8
Azureml K3s Compute Target Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Azureml K3s Compute Target Setup this skillmicrosoft/physical-ai-toolchain | 123 | — | ~5.7k | Automated safety check: Notes | MIT | |
| Devops Pipelineluongnv89/skills | 131 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Azure Diagnosticsmicrosoft/azure-skills | 1.5k | 1 repos | ~1.6k | Automated safety check: Pass | MIT | |
| Nim Operator InstallNVIDIA/k8s-nim-operator | 159 | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | |
| Provider Bug Reviewmondoohq/mql | 411 | — | ~2.9k | Automated safety check: Pass | Custom licence | |
| Nim Operator UninstallNVIDIA/k8s-nim-operator | 159 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 |
luongnv89/skills
Configure pre-commit hooks and lean GitHub Actions for shift-left quality assurance.
microsoft/azure-skills
Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage.
NVIDIA/k8s-nim-operator
Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional…
mondoohq/mql
Deep static code review of an mql provider for logic errors, nil-handling bugs, pagination truncation, caching/id collisions, and other defects that silently give users wrong data.
NVIDIA/k8s-nim-operator
Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and…
NVIDIA/dcgm-exporter
A skill your agent uses when validating DCGM Exporter in a local GPU-backed k3d/Kubernetes environment.
microsoft/physical-ai-toolchain
Submit, monitor, analyze, and evaluate LeRobot imitation learning training jobs on OSMO with Azure ML MLflow integration and inference evaluation - Brought to you by microsoft/physical-ai-toolchain
microsoft/physical-ai-toolchain
Generate, transfer, and consume environment-specific Azure, AKS, OSMO, ACR, and Azure ML deployment bundles.
microsoft/physical-ai-toolchain
Deploy trained robot policies to edge fleets via FluxCD GitOps, image automation, and deployment gating
microsoft/physical-ai-toolchain
Monitor robot fleet telemetry via Azure IoT Operations, drift detection, Grafana dashboards, and Fabric analytics
microsoft/physical-ai-toolchain
Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology
microsoft/physical-ai-toolchain
Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines
Categories
Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target. Azureml K3s Compute Target Setup is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization. Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.
Azureml K3s Compute Target Setup fits situations like: tasks that involve Container orchestration; tasks that involve QA and bug reports.
Run `npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a claude-code`. Or copy the skill folder (.github/skills/azureml-k3s-compute-target-setup in microsoft/physical-ai-toolchain) into .claude/skills/azureml-k3s-compute-target-setup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a codex`. Or copy the skill folder (.github/skills/azureml-k3s-compute-target-setup in microsoft/physical-ai-toolchain) into .agents/skills/azureml-k3s-compute-target-setup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/physical-ai-toolchain --skill azureml-k3s-compute-target-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/azureml-k3s-compute-target-setup, .gemini/skills/azureml-k3s-compute-target-setup, .github/skills/azureml-k3s-compute-target-setup and .opencode/skills/azureml-k3s-compute-target-setup in your project.
Going by SKILL.md and its folder, Azureml K3s Compute Target Setup needs the command-line tools its instructions call (az, kubectl, sh and python) and credentials named HF_TOKEN.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Azureml K3s Compute Target Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Azureml K3s Compute Target Setup: Devops Pipeline (luongnv89/skills, 131 stars), Azure Diagnostics (microsoft/azure-skills, 1.5k stars), Nim Operator Install (NVIDIA/k8s-nim-operator, 159 stars) and Provider Bug Review (mondoohq/mql, 411 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/physical-ai-toolchain, which has 123 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 7, 2026.
Source: microsoft/physical-ai-toolchain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.