Hugging Face Local Model Evals
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS model-compute-simulation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/model-compute-simulation .claude/skills/model-compute-simulation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "model-compute-simulation" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/model-compute-simulation into .claude/skills/model-compute-simulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-compute-simulation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/model-compute-simulationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS model-compute-simulation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/model-compute-simulation .agents/skills/model-compute-simulation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "model-compute-simulation" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/model-compute-simulation into .agents/skills/model-compute-simulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-compute-simulation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS model-compute-simulation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/model-compute-simulation .cursor/skills/model-compute-simulation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "model-compute-simulation" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/model-compute-simulation into .cursor/skills/model-compute-simulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-compute-simulation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git --path skills/model-compute-simulation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS model-compute-simulation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/model-compute-simulation .gemini/skills/model-compute-simulation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "model-compute-simulation" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/model-compute-simulation into .gemini/skills/model-compute-simulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-compute-simulation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS model-compute-simulationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/model-compute-simulation .github/skills/model-compute-simulation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "model-compute-simulation" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/model-compute-simulation into .github/skills/model-compute-simulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-compute-simulation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS model-compute-simulation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/model-compute-simulation .opencode/skills/model-compute-simulation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "model-compute-simulation" agent skill from https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS/tree/main/skills/model-compute-simulation into .opencode/skills/model-compute-simulation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "model-compute-simulation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
model-compute-simulationBuilds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.
The simulator loads a model config, builds the representative operator sequence, prints tensor shapes and per-op FLOPs, and can estimate MFU from a measured latency. The agent first gathers inputs: the model name, resolved through model-config-index.json, the GPU type for peak FLOPS from gpu-specs.json, dtype (bf16 by default, with fp8 doubling the peak), batch size and sequence length (decode at batch 1 by default), TP, DP and EP settings, and a per-GPU forward-pass latency, without which no MFU is given.
If the model is not indexed, the agent asks for a config.json path or adds an indexed config, and a raw nested Hugging Face config can be passed with --config. The bundled scripts are model_compute_simulator.py, config_normalization.py and extract_compute_flow_from_trace.py, which maps profiler traces to operators. For speculative MoE the real target and draft row counts come from the trace rather than from acceptance length, and an unverified template divided by TP or EP is not to be reported as measured FLOPs.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6dc9c66. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.coAlso links to:
github.comnvidia.comamd.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Model Compute Simulator loads about 4.5k tokens when it runs, and up to ~19k if it reads all its reference files. Until then it costs about 57 tokens; SKILL.md has 1,936 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Without a licence we can't republish the file, so here is its outline and opening line. It has 1,936 words (~4,522 tokens).
“Use this when the question is about operator order, tensor dimensions, FLOPs, MFU, or parallelism checks. The simulator loads a model config, builds the representative operator sequence, prints tensor shapes and FLOPs, and can estimate MFU from measured latency.”
SKILL.md and 5 other files (scripts, references) in skills/model-compute-simulation of BBuf/AI-Infra-Auto-Driven-SKILLS.
Open the folder on GitHubat commit 6dc9c66
Model Compute Simulator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Model Compute Simulator this skillBBuf/AI-Infra-Auto-Driven-SKILLS | 900 | — | ~4.5k | Automated safety check: Pass | None | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Ascend Model Adapter for vLLMvllm-project/vllm-ascend | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | |
| bitsandbytes Model QuantizationOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~2.5k | Automated safety check: Pass | MIT |
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
vllm-project/vllm-ascend
Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Orchestra-Research/AI-Research-SKILLs
Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.
Orchestra-Research/AI-Research-SKILLs
Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.
Orchestra-Research/AI-Research-SKILLs
Explains three ways to speed up LLM inference: draft-model speculative decoding, Medusa heads and lookahead decoding with Jacobi iteration, and when each one fits.
BBuf/AI-Infra-Auto-Driven-SKILLS
Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.
BBuf/AI-Infra-Auto-Driven-SKILLS
Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
BBuf/AI-Infra-Auto-Driven-SKILLS
Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.
BBuf/AI-Infra-Auto-Driven-SKILLS
Reviews SGLang changes the way its maintainers do, drawing on a bundled corpus of public PR review threads and a flowchart of how the diff runs.
BBuf/AI-Infra-Auto-Driven-SKILLS
Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.
Works with
Categories
Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks. The simulator loads a model config, builds the representative operator sequence, prints tensor shapes and per-op FLOPs, and can estimate MFU from a measured latency.json, dtype (bf16 by default, with fp8 doubling the peak), batch size and sequence length (decode at batch 1 by default), TP, DP and EP settings, and a per-GPU forward-pass latency, without which no MFU is given.
Model Compute Simulator fits situations like: listing tensor shapes and per-op FLOPs for a model's decode or prefill step; computing MFU from a measured latency on a specific GPU; mapping kernels in a profiler trace back to model operators; comparing how different TP, DP or EP settings change compute per GPU.
Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a claude-code`. Or copy the skill folder (skills/model-compute-simulation in BBuf/AI-Infra-Auto-Driven-SKILLS) into .claude/skills/model-compute-simulation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a codex`. Or copy the skill folder (skills/model-compute-simulation in BBuf/AI-Infra-Auto-Driven-SKILLS) into .agents/skills/model-compute-simulation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-compute-simulation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-compute-simulation, .gemini/skills/model-compute-simulation, .github/skills/model-compute-simulation and .opencode/skills/model-compute-simulation in your project.
Going by SKILL.md and its folder, Model Compute Simulator needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python; The model's config.json, or a model listed in the bundled config index; A measured per-GPU latency, only if you want MFU.
SKILL.md names 4 domains. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. As links in the text: github.com, nvidia.com and amd.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
No licence was found for Model Compute Simulator or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.
About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Model Compute Simulator: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Ascend Model Adapter for vLLM (vllm-project/vllm-ascend, 2.9k stars), Graphsignal (graphsignal/graphsignal, 257 stars) and TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
BBuf (a GitHub user) maintains it in BBuf/AI-Infra-Auto-Driven-SKILLS, which has 900 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 5, 2026.
Source: BBuf/AI-Infra-Auto-Driven-SKILLS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.