Agent skill

Model PR History Knowledge

by BBuf in BBuf/AI-Infra-Auto-Driven-SKILLS

A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.

No licenceAuto-check passedAI & LLM Engineering

Install Model PR History Knowledge

skills CLI
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-pr-history-knowledge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS model-pr-history-knowledge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/model-pr-optimization-history .claude/skills/model-pr-history-knowledge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-pr-history-knowledge
GitHub stars
911
Token cost
~1.5k tokens
SKILL.md length
631 words
Files
360 (incl. scripts)
Skills in repo
10
Repo updated
First seen
Licence
None found

At a glance

A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.

  • Works in 5 steps: Infer the model-family slug from the… → Read the matching SGLang history first… → Extract only actionable evidence → …
  • TokenSpeed serving/model optimization task needs prior model-family PR evidence
  • SKILL.md covers Query, Workflow, Model Slugs and Optimization Workflow Contract
  • Runs Python scripts from its folder; calls python3

What it does

Model PR History Knowledge is an agent skill from BBuf/AI-Infra-Auto-Driven-SKILLS. Use when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence. Query and read the PR-driven history docs under model-pr-optimization-history before choosing source paths, fast paths, kernel/fusion ideas, regression risks, or validation lanes.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 366 other files, including scripts (for example `README.md`, `open-pr-watch.md` and `scripts/query.py`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with SGLang, NVIDIA AI Platform and vLLM.

When your agent uses it

  • TokenSpeed serving/model optimization task needs prior model-family PR evidence
  • Tasks that involve LLM inference and serving

Example prompts

  • “/model-pr-history-knowledge”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Infer the model-family slug from the user's model id, checkpoint path, or
  2. Read the matching SGLang history first for SGLang patch work. Read competitor
  3. Extract only actionable evidence
  4. Save a short note in the active run artifacts, for example
  5. Do not copy long PR cards into the final answer. Cite paths and summarize the

What it can do on your machine

Read from SKILL.md and the folder at commit 6dc9c66. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model PR History Knowledge loads about 1.5k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 631 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 631 words (~1,461 tokens).

“This is a PR-driven knowledge base for model optimization history. It is not a set of per-model skills. Each model family keeps bilingual docs with inspected PR diffs, implementation file coverage, timelines, changed files, code excerpts, and validation/risk notes.”

— opening of SKILL.md by BBuf
name
model-pr-history-knowledge

Read the full SKILL.md on GitHub

Files

SKILL.md and 359 other files (scripts) in model-pr-optimization-history of BBuf/AI-Infra-Auto-Driven-SKILLS.

  • SKILL.md
  • README.md
  • open-pr-watch.md
  • scripts/query.py
  • sglang/README.md
  • sglang/deepseek-ocr-2/README.en.md
  • sglang/deepseek-ocr-2/README.zh.md
  • sglang/deepseek-ocr/README.en.md
  • sglang/deepseek-ocr/README.zh.md
  • sglang/deepseek-v3-r1/README.en.md
  • sglang/deepseek-v3-r1/README.zh.md
  • sglang/deepseek-v31/README.en.md
  • sglang/deepseek-v31/README.zh.md
  • sglang/deepseek-v32/README.en.md
  • … and 346 more

Open the folder on GitHubat commit 6dc9c66

Compare with similar skills

Model PR History Knowledge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model PR History Knowledge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model PR History Knowledge this skillBBuf/AI-Infra-Auto-Driven-SKILLS911—~1.5kAutomated safety check: PassNone
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT
Jetson LLM ServeNVIDIA/skills3.5k2 repos~3kAutomated safety check: NotesApache-2.0
Jetson Inference Mem TuneNVIDIA/skills3.5k1 repos~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 3 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Jetson LLM Serve

    NVIDIA/skills

    Official

    Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

    3.5k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check: notes
  • Official

    Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

    3.5k GitHub starsUsed in 1 repo~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Agentsop LLM Engine Selection

    agentsope/SkillAlchemy

    Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

    459 GitHub stars~6.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from BBuf/AI-Infra-Auto-Driven-SKILLS

All 10 skills in this repo
  • SGLang Model Day-0 Support

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

    911 GitHub stars~2.3k tokensUpdated 3 days ago
    Auto-check passed
  • LLM Serving Capacity Planner

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

    911 GitHub stars~2.5k tokensUpdated 3 days ago
    Auto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    Auto-check passed
  • Model Architecture Diagram Finder

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.

    911 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Model Compute Simulator

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

    911 GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • SGLang Maintainer-Style Review

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reviews SGLang changes the way its maintainers do, drawing on a bundled corpus of public PR review threads and a flowchart of how the diff runs.

    911 GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed

Questions about Model PR History Knowledge

What does Model PR History Knowledge do?

A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence. Model PR History Knowledge is an agent skill from BBuf/AI-Infra-Auto-Driven-SKILLS. Use when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.

When should I use Model PR History Knowledge?

Model PR History Knowledge fits situations like: tokenSpeed serving/model optimization task needs prior model-family PR evidence; tasks that involve LLM inference and serving.

How do I install Model PR History Knowledge in Claude Code?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-pr-history-knowledge -a claude-code`. Or copy the skill folder (model-pr-optimization-history in BBuf/AI-Infra-Auto-Driven-SKILLS) into .claude/skills/model-pr-history-knowledge in your project. Claude Code loads it when a task matches its description.

How do I install Model PR History Knowledge in Codex?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-pr-history-knowledge -a codex`. Or copy the skill folder (model-pr-optimization-history in BBuf/AI-Infra-Auto-Driven-SKILLS) into .agents/skills/model-pr-history-knowledge in your project. Codex loads it when a task matches its description.

Can I use Model PR History Knowledge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill model-pr-history-knowledge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-pr-history-knowledge, .gemini/skills/model-pr-history-knowledge, .github/skills/model-pr-history-knowledge and .opencode/skills/model-pr-history-knowledge in your project.

What does Model PR History Knowledge need to run?

Going by SKILL.md and its folder, Model PR History Knowledge needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Model PR History Knowledge access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Model PR History Knowledge safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Model PR History Knowledge use?

No licence was found for Model PR History Knowledge or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Model PR History Knowledge use?

About 1.5k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model PR History Knowledge?

Skills that share tags, products or a category with Model PR History Knowledge: Dstack Prototyping (dstackai/dstack, 2.3k stars), Graphsignal (graphsignal/graphsignal, 257 stars), SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Jetson LLM Serve (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model PR History Knowledge?

BBuf (a GitHub user) maintains it in BBuf/AI-Infra-Auto-Driven-SKILLS, which has 911 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 5, 2026.

Source: BBuf/AI-Infra-Auto-Driven-SKILLS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.