Agent skill

Optimize Slurm Topology

by NVlabs in NVlabs/alpasim

Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.

Apache-2.0Auto-check passedDevOps & Cloud

Install Optimize Slurm Topology

skills CLI
$ npx skills add NVlabs/alpasim --skill optimize-slurm-topology -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVlabs/alpasim optimize-slurm-topology --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVlabs/alpasim.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/optimize-slurm-topology .claude/skills/optimize-slurm-topology && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
optimize-slurm-topology
GitHub stars
1.3k
Token cost
~1.6k tokens
SKILL.md length
812 words
Files
5 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.

  • Works in 4 steps: Start local telemetry once with → The argument is provided by the… → Keep the local telemetry stack running… → …
  • Tuning service GPU placement
  • SKILL.md covers Inputs, Telemetry and Memory Setup, Experiment Loop and Supporting documents, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Optimize Slurm Topology is an agent skill from NVlabs/alpasim. Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts. Use when tuning service GPU placement, replicaspercontainer, runtime.nrworkers, endpoint nconcurrentrollouts, NRE/physics cache sizes, or Slurm experiment batches for full-duration rollout throughput.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `agents/openai.yaml`, `references/experiment-record.md` and `references/metrics.md`).

It sits in DevOps & Cloud, covering Monitoring and alerting. It works with Prometheus, Grafana and NVIDIA AI Platform. The repository describes itself as: AlpaSim is an open-source autonomous vehicle simulation platform designed for development and testing of end-to-end AV policies. The licence is Apache-2.0.

When your agent uses it

  • Tuning service GPU placement
  • Replicaspercontainer
  • Runtime.nrworkers
  • Endpoint nconcurrentrollouts

Example prompts

  • “/optimize-slurm-topology”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Start local telemetry once with
  2. The argument is provided by the experiment logs.
  3. Keep the local telemetry stack running until all experiment candidates have
  4. Create a repo-local experiment record from references/experiment-record.md.

What it can do on your machine

Read from SKILL.md and the folder at commit affc2ea. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Optimize Slurm Topology loads about 1.6k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 812 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVlabs/alpasim at commit affc2ea, republished under its Apache-2.0 licence (© NVlabs). 812 words, ~1,608 tokens.

Download SKILL.mdSave it as .claude/skills/optimize-slurm-topology/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
optimize-slurm-topology
description
Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts. Use when tuning service GPU placement, replicas_per_container, runtime.nr_workers, endpoint n_concurrent_rollouts, NRE/physics cache sizes, or Slurm experiment batches for full-duration rollout throughput.

Optimize Slurm Topology

Use this workflow to iteratively improve AlpaSim rollout throughput on Slurm. Optimize for full 20s rollout throughput, not startup-only behavior.

Inputs

When this skill is invoked, the user must specify a full run command as starting point for the optimization. This command provides a base topology as starting point, as well as the target configuration (e.g. driver, sceneset, n_rollouts, cluster, and any other parameters such as number of cameras, simulation frequency, etc.).

If the user doesn't specify a full run command, ask them!

Telemetry and Memory Setup

Use one persistent local Prometheus/Grafana instance for the whole experiment.

  1. Start local telemetry once with src/tools/scripts/start-prometheus-grafana.sh <file-sd-dir-or-ssh-path> --grafana-port 3003 --prometheus-port 9093. Use the non-default ports to avoid conflicts with user-started telemetry stacks.
  2. The <file-sd-dir-or-ssh-path> argument is provided by the experiment logs. For example, the default value on IAD is <iad-ssh-alias>:/lustre/fsw/portfolios/av/projects/av_alpamayo_reasoning/data/av_alpamayo_sim/.cache/prometheus/file-sd
  3. Keep the local telemetry stack running until all experiment candidates have been evaluated. Then stop it with src/tools/scripts/start-prometheus-grafana.sh stop
  4. Create a repo-local experiment record from references/experiment-record.md. Default path: docs/experiments/topology-opt-<driver>-<cluster>-<YYYYMMDD>.md. This file will be your experiment log and memory. It should contain all necessary information to understand why a topology was tried and what the results were.

Experiment Loop

  1. Start from a known topology and run a baseline experiment.
  2. If you already have a baseline, start one or multiple candidate experiments in parallel (at most 3).
  3. The experiments have an initial startup time of a couple of about 5 min before they appear in Prometheus. After that, they start producing rollouts. However, because all rollouts are initially started simultenously, there's significant congestion in the first 10-15 minutes. Wait until you can see this congestion has cleared and the system has reached a steady state. Use the 5m seconds_per_rollout only as an early diagnostic. Once the full-run seconds_per_rollout has stabilized, typically after 30-45 minutes, use it as the primary optimization target.
  4. Reject candidates that OOM or crash. Analyze the reason for failure and avoid repeating the same mistake. A typical reason is insufficient GPU memory.
  5. Once a candidate reaches steady state, analyze it carefully and document (see references/metrics.md and references/topology-knobs.md):
  • Its stabilized full-run seconds_per_rollout, using the 5m value only to diagnose recent behavior and confirm that the run remains healthy.
  • Its bottlenecks, using alpasim:rpc_queue_depth_at_start_latest:max as the primary bottleneck signal and alpasim:rpc_queue_depth_at_start_latest:min to detect workers that are starving or receiving uneven load.
  • Its used and available resources, including per-GPU utilization, memory consumption, memory pressure, and memory headroom.
  • Opportunities for improvement.
  1. Skill improvement reflections: Did you learn something new about the system that was not yet covered in the skill? This can include, for example:
  • How to run experiments or query results.
  • Which Prometheus queries are useful.
  • How the topology knobs affect throughput and memory.
  • Are there additional metrics that we should introduce to better understand the system?
  • Any additional scripts that you wrote to help with the experiments that would be useful to add to the skill.
  • Or anything else that you think is useful to remember for future experiments.
Show full SKILL.md (305 more words)Show less
  1. Keep two record sections current (see references/experiment-record.md). These records should include the output of both step 4 and step 5, and should be updated after every candidate experiment:
    • a short Markdown progress document (including table) for humans and quick parsing;
    • a more detailed JSON memory block with fixed inputs, runs, metrics, GPU utilization, GPU memory consumption and headroom, decisions, and links to run artifacts.
  2. Decide on the next topology changes and go back to step 1.
  3. You should stop running experiments as soon as you have enough data to support your reasoning and decision. It is not required to let them run until the end. However, do not compare or stop a healthy candidate before its full-run seconds_per_rollout has stabilized, normally 30-45 minutes after rollout production begins. Note that "steady state" can still contain cyclic behavior.
  4. Stop when you can't make progress over multiple iterations or when you don't believe there is more enough free resources to improve throughput.

Supporting documents

  • Read references/experiment-record.md for how to keep a record of the experiment and its candidates.
  • Read references/metrics.md for current metric names, PromQL queries, and interpretation.
  • Read references/topology-knobs.md for guidelines on how to change topology and what to expect from each change.

Final Reporting

At the end of the optimization, re-read the experiment record and summarize the results in a DETAILED final report, including:

  1. Baseline topology and initial speed (primary stabilized full-run seconds_per_rollout, plus 5m seconds_per_rollout for recent-behavior context), including per-GPU utilization and memory consumption/headroom.
  2. All tried topologies, their reasoning, expected effect, measured result, resource usage, and decision.
  3. Best topology found, why it won, remaining bottlenecks or constraints, and how much GPU utilization and memory headroom remains for further tuning.
  4. Any advice on how the skill or the instructions can be improved.
  5. Link the repo-local experiment record.

© NVlabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in .agents/skills/optimize-slurm-topology of NVlabs/alpasim.

  • SKILL.md
  • agents/openai.yaml
  • references/experiment-record.md
  • references/metrics.md
  • references/topology-knobs.md

Open the folder on GitHubat commit affc2ea

Compare with similar skills

Optimize Slurm Topology next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Optimize Slurm Topology compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Optimize Slurm Topology this skillNVlabs/alpasim1.3k—~1.6kAutomated safety check: PassApache-2.0
Syncmetapawurb/hotpath-rs1.9k—~1.2kAutomated safety check: NotesMIT
Dashboard Previewm4r1k/Eneru149—~1.4kAutomated safety check: PassMIT
Graftm4r1k/Eneru1491 repos~2.3kAutomated safety check: PassMIT
Release Reviewm4r1k/Eneru149—~1.9kAutomated safety check: PassMIT
Archestra Dev Observabilityarchestra-ai/archestra4.4k—~1.2kAutomated safety check: PassCustom licence

Similar skills

  • Syncmeta

    pawurb/hotpath-rs

    Sync changes from the hotpath, hotpath-macros and hotpath-drain crates to their meta counterparts (hotpath-meta, hotpath-macros-meta and hotpath-drain-meta).

    1.9k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Visually verify Eneru browser-dashboard changes against a live daemon or audit an exact deployment.

    149 GitHub stars~1.4k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Graft

    m4r1k/Eneru

    This repo is indexed by graft/. An agent skill from m4r1k/Eneru.

    149 GitHub starsUsed in 1 repo~2.3k tokens
    DevOps & CloudAuto-check passed
  • Release Review

    m4r1k/Eneru

    Mandatory pre-release deep review for minor/major releases (X.Y.0 / X.0.0).

    149 GitHub stars~1.9k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Archestra Dev Observability

    archestra-ai/archestra

    A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup.

    4.4k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Frontmcp Observability

    agentfront/frontmcp

    A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server.

    146 GitHub stars~4.6k tokensUpdated today
    DevOps & CloudAuto-check passed

More from NVlabs/alpasim

  • Doc Reviewer

    NVlabs/alpasim

    Reviews recent code changes and checks if documentation needs updates.

    1.3k GitHub stars~1.2k tokensUpdated 23 days ago
    Auto-check passed

Categories

Questions about Optimize Slurm Topology

What does Optimize Slurm Topology do?

Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts. Optimize Slurm Topology is an agent skill from NVlabs/alpasim. Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts.

When should I use Optimize Slurm Topology?

Optimize Slurm Topology fits situations like: tuning service GPU placement; replicaspercontainer; runtime.nrworkers; endpoint nconcurrentrollouts.

How do I install Optimize Slurm Topology in Claude Code?

Run `npx skills add NVlabs/alpasim --skill optimize-slurm-topology -a claude-code`. Or copy the skill folder (.agents/skills/optimize-slurm-topology in NVlabs/alpasim) into .claude/skills/optimize-slurm-topology in your project. Claude Code loads it when a task matches its description.

How do I install Optimize Slurm Topology in Codex?

Run `npx skills add NVlabs/alpasim --skill optimize-slurm-topology -a codex`. Or copy the skill folder (.agents/skills/optimize-slurm-topology in NVlabs/alpasim) into .agents/skills/optimize-slurm-topology in your project. Codex loads it when a task matches its description.

Can I use Optimize Slurm Topology in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVlabs/alpasim --skill optimize-slurm-topology -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimize-slurm-topology, .gemini/skills/optimize-slurm-topology, .github/skills/optimize-slurm-topology and .opencode/skills/optimize-slurm-topology in your project.

What does Optimize Slurm Topology need to run?

SKILL.md names no scripts, command-line tools or credentials: Optimize Slurm Topology is instructions for the agent only.

Does Optimize Slurm Topology access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Optimize Slurm Topology safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Optimize Slurm Topology use?

Optimize Slurm Topology is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Optimize Slurm Topology use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.8k tokens, read only when the agent opens those files.

What are the alternatives to Optimize Slurm Topology?

Skills that share tags, products or a category with Optimize Slurm Topology: Syncmeta (pawurb/hotpath-rs, 1.9k stars), Dashboard Preview (m4r1k/Eneru, 149 stars), Graft (m4r1k/Eneru, 149 stars) and Release Review (m4r1k/Eneru, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Optimize Slurm Topology?

NVlabs (a GitHub organization) maintains it in NVlabs/alpasim, which has 1,259 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 18, 2026.

Source: NVlabs/alpasim on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.