Agent skill

Model Compute Simulator

by BBuf in BBuf/AI-Infra-Auto-Driven-SKILLS

Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

No licenceAuto-check passedAI & LLM Engineering

Install Model Compute Simulator

skills CLI
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS model-compute-simulation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/model-compute-simulation .claude/skills/model-compute-simulation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-compute-simulation
GitHub stars
900
Token cost
~4.5k tokens
SKILL.md length
1,936 words
Files
6 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
None found

At a glance

Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

  • Works in 4 steps: Load model config → Generate execution flow and tensor… → Estimate MFU with measured latency → …
  • Listing tensor shapes and per-op FLOPs for a model's decode or prefill step
  • SKILL.md covers Overview, Inputs, Shape and timing limits for… and Workflow, plus 5 more sections
  • Runs Python scripts from its folder; calls python3; reaches huggingface.co

What it does

The simulator loads a model config, builds the representative operator sequence, prints tensor shapes and per-op FLOPs, and can estimate MFU from a measured latency. The agent first gathers inputs: the model name, resolved through model-config-index.json, the GPU type for peak FLOPS from gpu-specs.json, dtype (bf16 by default, with fp8 doubling the peak), batch size and sequence length (decode at batch 1 by default), TP, DP and EP settings, and a per-GPU forward-pass latency, without which no MFU is given.

If the model is not indexed, the agent asks for a config.json path or adds an indexed config, and a raw nested Hugging Face config can be passed with --config. The bundled scripts are model_compute_simulator.py, config_normalization.py and extract_compute_flow_from_trace.py, which maps profiler traces to operators. For speculative MoE the real target and draft row counts come from the trace rather than from acceptance length, and an unverified template divided by TP or EP is not to be reported as measured FLOPs.

When your agent uses it

  • Listing tensor shapes and per-op FLOPs for a model's decode or prefill step
  • Computing MFU from a measured latency on a specific GPU
  • Mapping kernels in a profiler trace back to model operators
  • Comparing how different TP, DP or EP settings change compute per GPU

Example prompts

  • “Estimate per-op FLOPs for decode at batch size 1 on my GPU in bf16, using the model's config.json.”
  • “Map the kernels in this profiler trace to model ops and report MFU.”
  • “What changes in FLOPs per GPU if I move this model from TP 4 to TP 8?”

Requirements

  • Python
  • The model's config.json, or a model listed in the bundled config index
  • A measured per-GPU latency, only if you want MFU

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Load model config
  2. Generate execution flow and tensor dimensions
  3. Estimate MFU with measured latency
  4. Per-operator MFU with kernel-level latency

What it can do on your machine

Read from SKILL.md and the folder at commit 6dc9c66. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co

    Also links to:

    • github.com
    • nvidia.com
    • amd.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Compute Simulator loads about 4.5k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 57 tokens; SKILL.md has 1,936 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~19k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,936 words (~4,522 tokens).

“Use this when the question is about operator order, tensor dimensions, FLOPs, MFU, or parallelism checks. The simulator loads a model config, builds the representative operator sequence, prints tensor shapes and FLOPs, and can estimate MFU from measured latency.”

— opening of SKILL.md by BBuf
name
model-compute-simulation

Read the full SKILL.md on GitHub

Files

SKILL.md and 5 other files (scripts, references) in skills/model-compute-simulation of BBuf/AI-Infra-Auto-Driven-SKILLS.

  • SKILL.md
  • references/gpu-specs.json
  • references/model-config-index.json
  • scripts/config_normalization.py
  • scripts/extract_compute_flow_from_trace.py
  • scripts/model_compute_simulator.py

Open the folder on GitHubat commit 6dc9c66

Compare with similar skills

Model Compute Simulator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Compute Simulator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Compute Simulator this skillBBuf/AI-Infra-Auto-Driven-SKILLS900—~4.5kAutomated safety check: PassNone
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Ascend Model Adapter for vLLMvllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs13k5 repos~1.3kAutomated safety check: PassMIT
bitsandbytes Model QuantizationOrchestra-Research/AI-Research-SKILLs13k3 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Ascend Model Adapter for vLLM

    vllm-project/vllm-ascend

    Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

    2.9k GitHub stars~2.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 5 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • bitsandbytes Model Quantization

    Orchestra-Research/AI-Research-SKILLs

    Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

    13k GitHub starsUsed in 3 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Speculative Decoding

    Orchestra-Research/AI-Research-SKILLs

    Explains three ways to speed up LLM inference: draft-model speculative decoding, Medusa heads and lookahead decoding with Jacobi iteration, and when each one fits.

    13k GitHub starsUsed in 3 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed

More from BBuf/AI-Infra-Auto-Driven-SKILLS

All 10 skills in this repo
  • SGLang Model Day-0 Support

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

    900 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • LLM Serving Capacity Planner

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

    900 GitHub stars~2.5k tokensUpdated 2 days ago
    Auto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    900 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed
  • Model Architecture Diagram Finder

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.

    900 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • SGLang Maintainer-Style Review

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reviews SGLang changes the way its maintainers do, drawing on a bundled corpus of public PR review threads and a flowchart of how the diff runs.

    900 GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed
  • Torch Profiler Layer Track

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

    900 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Model Compute Simulator

What does Model Compute Simulator do?

Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks. The simulator loads a model config, builds the representative operator sequence, prints tensor shapes and per-op FLOPs, and can estimate MFU from a measured latency.json, dtype (bf16 by default, with fp8 doubling the peak), batch size and sequence length (decode at batch 1 by default), TP, DP and EP settings, and a per-GPU forward-pass latency, without which no MFU is given.

When should I use Model Compute Simulator?

Model Compute Simulator fits situations like: listing tensor shapes and per-op FLOPs for a model's decode or prefill step; computing MFU from a measured latency on a specific GPU; mapping kernels in a profiler trace back to model operators; comparing how different TP, DP or EP settings change compute per GPU.

How do I install Model Compute Simulator in Claude Code?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a claude-code`. Or copy the skill folder (skills/model-compute-simulation in BBuf/AI-Infra-Auto-Driven-SKILLS) into .claude/skills/model-compute-simulation in your project. Claude Code loads it when a task matches its description.

How do I install Model Compute Simulator in Codex?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a codex`. Or copy the skill folder (skills/model-compute-simulation in BBuf/AI-Infra-Auto-Driven-SKILLS) into .agents/skills/model-compute-simulation in your project. Codex loads it when a task matches its description.

Can I use Model Compute Simulator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-compute-simulation, .gemini/skills/model-compute-simulation, .github/skills/model-compute-simulation and .opencode/skills/model-compute-simulation in your project.

What does Model Compute Simulator need to run?

Going by SKILL.md and its folder, Model Compute Simulator needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python; The model's config.json, or a model listed in the bundled config index; A measured per-GPU latency, only if you want MFU.

Does Model Compute Simulator access the network?

SKILL.md names 4 domains. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. As links in the text: github.com, nvidia.com and amd.com. This is read from the text; nothing was executed.

Is Model Compute Simulator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Model Compute Simulator use?

No licence was found for Model Compute Simulator or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Model Compute Simulator use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.

What are the alternatives to Model Compute Simulator?

Skills that share tags, products or a category with Model Compute Simulator: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Ascend Model Adapter for vLLM (vllm-project/vllm-ascend, 2.9k stars), Graphsignal (graphsignal/graphsignal, 257 stars) and TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Compute Simulator?

BBuf (a GitHub user) maintains it in BBuf/AI-Infra-Auto-Driven-SKILLS, which has 900 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 5, 2026.

Source: BBuf/AI-Infra-Auto-Driven-SKILLS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.