Agent skill

Sglang Bisect CI Regression

by sgl-project in sgl-project/sglang

Investigate consistently failing SGLang CI tests by extracting the failure signature from scheduled or rerun workflows, bisecting the passing/failing commit window, checking runner or hardware…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Sglang Bisect CI Regression

skills CLI
$ npx skills add sgl-project/sglang --skill sglang-bisect-ci-regression -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang sglang-bisect-ci-regression --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/sglang-bisect-ci-regression .claude/skills/sglang-bisect-ci-regression && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sglang-bisect-ci-regression
GitHub stars
37k
Used in
2 other repos
Token cost
~2.5k tokens
SKILL.md length
782 words
Files
1
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Investigate consistently failing SGLang CI tests by extracting the failure signature from scheduled or rerun workflows, bisecting the passing/failing commit window, checking runner or hardware…

  • Works in 6 steps: Extract the Failure Signature → Temporal Bisection → Runner/Hardware Analysis → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers Slash Command, When to Use This Skill, Arguments and Background: Scheduled CI Runs, plus 3 more sections
  • Calls gh, ssh and git

What it does

Sglang Bisect CI Regression is an agent skill from sgl-project/sglang. Investigate consistently failing SGLang CI tests by extracting the failure signature from scheduled or rerun workflows, bisecting the passing/failing commit window, checking runner or hardware specificity, and optionally reproducing on a remote GPU host.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with SGLang and Docker. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/sglang-bisect-ci-regression”

Requirements

  • Python 3
  • Docker

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Extract the Failure Signature
  2. Temporal Bisection
  3. Runner/Hardware Analysis
  4. Code Analysis
  5. Remote Reproduction (Optional)
  6. Report

What it can do on your machine

Read from SKILL.md and the folder at commit dab108b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • ssh
    • git
    • pip
    • curl
    • scp

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, ssh, git, pip, curl and scp, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sglang Bisect CI Regression loads about 2.5k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 782 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit dab108b, republished under its Apache-2.0 licence (© sgl-project). 782 words, ~2,499 tokens.

Download SKILL.mdSave it as .claude/skills/sglang-bisect-ci-regression/SKILL.md (or your agent's skills folder).
name
sglang-bisect-ci-regression
description
Investigate consistently failing SGLang CI tests by extracting the failure signature from scheduled or rerun workflows, bisecting the passing/failing commit window, checking runner or hardware specificity, and optionally reproducing on a remote GPU host.

SGLang Bisect CI Regression

Investigate a consistently failing CI test to find the root cause - whether it's a code regression from a specific PR, a hardware/runner-specific issue, or an environment change. Optionally reproduce the failure on a remote GPU server.

Slash Command

/sglang-bisect-ci-regression <test_name_or_ci_url> [ssh_target] [docker_container]

When to Use This Skill

  • A CI test is failing consistently on main (scheduled runs)
  • You need to find which PR introduced a regression
  • You suspect a runner-specific or GPU-specific issue
  • You want to reproduce a CI failure on a remote server

Arguments

  • First argument (required): Test file name (e.g. test_lora_tp.py) or a GitHub Actions job URL
  • Second argument (optional): SSH target for remote reproduction (e.g. user@host)
  • Third argument (optional): Docker container name on the SSH target (e.g. sglang_dev)

If SSH target and docker container are not provided, the skill will only perform the CI log analysis and bisection, without remote reproduction. Ask the user for these if reproduction is needed and they weren't provided.

Background: Scheduled CI Runs

SGLang uses the pr-test.yml workflow with scheduled runs (cron-triggered) to periodically test the main branch. These runs are the primary data source for detecting regressions:

Always use these scheduled runs (not PR-triggered runs) when bisecting regressions on main. The --event schedule filter in gh run list ensures you only see these periodic main-branch runs.

Workflow

Phase 1: Extract the Failure Signature
  1. Get the failing test details from CI logs. If given a URL, fetch logs directly. If given a test name, find recent scheduled runs of pr-test.yml on main that failed:
bash
# List recent scheduled runs targeting main (the primary source of truth for regressions)
# These are cron-triggered runs visible at:
# https://github.com/sgl-project/sglang/actions/workflows/pr-test.yml?query=event%3Aschedule
gh run list --repo sgl-project/sglang --workflow="pr-test.yml" --event schedule --branch main --limit 20 --json databaseId,conclusion,createdAt,headSha

# Find the job containing the test
gh run view {RUN_ID} --repo sgl-project/sglang --json jobs --jq '.jobs[] | select(.conclusion == "failure") | {name, conclusion, databaseId}'

# Get the failure details
gh run view {RUN_ID} --repo sgl-project/sglang --job {JOB_ID} --log 2>&1 | grep -E -B 5 -A 30 "AssertionError|FAIL|Error|{TEST_NAME}"
  1. Record the failure signature:
    • Exact error message and assertion
    • Affected test method name
    • Model/config involved
    • Numeric values (e.g., tolerance diffs, scores)
    • Whether the failure is deterministic (same values across runs)
Phase 2: Temporal Bisection
  1. Find the boundary between passing and failing runs. Walk through the scheduled run history (from the pr-test.yml schedule runs on main) to identify:
    • Last known PASSING run (sha + date)
    • First known FAILING run (sha + date)
bash
# For each scheduled run, check the specific partition/job status
gh run view {RUN_ID} --repo sgl-project/sglang --json jobs --jq '.jobs[] | select(.name == "{JOB_NAME}") | {conclusion, databaseId}'

# Verify a specific test passed or failed in a run
gh run view {RUN_ID} --repo sgl-project/sglang --job {JOB_ID} --log 2>&1 | grep -E "{TEST_NAME}|PASSED|FAILED|logprobs mismatch" | head -10
  1. List commits between the boundary:
bash
git log --oneline {LAST_PASS_SHA}..{FIRST_FAIL_SHA}
  1. Filter for relevant commits that touch files related to the failing test (model layers, kernels, test utilities, etc.):
bash
git log --oneline {LAST_PASS_SHA}..{FIRST_FAIL_SHA} -- {relevant_paths}
Phase 3: Runner/Hardware Analysis
  1. Check if the failure is runner-specific. Extract the runner identity from each failing and passing run:
bash
# Get runner name and machine
gh run view {RUN_ID} --repo sgl-project/sglang --job {JOB_ID} --log 2>&1 | grep -E "Runner name|Machine name" | head -5

# Get GPU/driver info
gh run view {RUN_ID} --repo sgl-project/sglang --job {JOB_ID} --log 2>&1 | grep -i -E "NVIDIA-SMI|Driver Version|CUDA Version" | head -5

# Get package versions
gh run view {RUN_ID} --repo sgl-project/sglang --job {JOB_ID} --log 2>&1 | grep -E "sgl.kernel.*==|flashinfer.*==" | head -5
  1. Correlate runners with pass/fail outcomes. Build a table:
Run IDDateRunnerGPU TypeDriverResult

If all failures map to a specific runner type/GPU and all passes map to another, the issue is hardware-specific, not a code regression.

Show full SKILL.md (328 more words)Show less
Phase 4: Code Analysis
  1. If a code regression is suspected (failures not runner-specific), examine the candidate commits:

    • Read the changed files
    • Understand how the changes could affect the failing test
    • Look for prefill-vs-decode differences, TP-specific paths, kernel changes
  2. If a hardware issue is suspected, analyze:

    • Kernel compatibility (CUDA compute capability)
    • Driver version differences
    • All-reduce / NCCL behavior differences
    • CUDA graph capture differences across GPU architectures
Phase 5: Remote Reproduction (Optional)

Only if SSH target and docker container were provided.

  1. Verify the remote environment:
bash
ssh {SSH_TARGET} "docker exec {CONTAINER} nvidia-smi --query-gpu=name,driver_version --format=csv"
ssh {SSH_TARGET} "docker exec {CONTAINER} pip show sgl-kernel sglang flashinfer-python 2>&1 | grep -E 'Name:|Version:'"
  1. Ensure latest code is installed. If the container is stale, update:
bash
# Try fetching latest main
ssh {SSH_TARGET} "docker exec {CONTAINER} bash -c 'cd /path/to/sglang && git fetch origin main && git checkout origin/main'"
# Or download and install from tarball if git auth fails
ssh {SSH_TARGET} "docker exec {CONTAINER} bash -c 'cd /tmp && curl -L https://github.com/sgl-project/sglang/archive/refs/heads/main.tar.gz | tar xz && cd sglang-main && pip install -e \"python[all]\"'"
# Reinstall (after git fetch)
ssh {SSH_TARGET} "docker exec {CONTAINER} bash -c 'cd /path/to/sglang && pip install -e \"python[all]\"'"
# Install test dependencies if needed
ssh {SSH_TARGET} "docker exec {CONTAINER} pip install peft rouge-score"
  1. Create a minimal reproduction script that:

    • Uses if __name__ == '__main__' with mp.set_start_method("spawn")
    • Runs the specific failing test configuration
    • Prints key metrics (diffs, scores, outputs)
    • Exits with code 1 on failure
  2. Copy and run the reproduction script:

bash
scp /tmp/repro_script.py {SSH_TARGET}:/tmp/
ssh {SSH_TARGET} "docker cp /tmp/repro_script.py {CONTAINER}:/tmp/"
ssh {SSH_TARGET} "docker exec -e CUDA_VISIBLE_DEVICES=0,1 {CONTAINER} python3 /tmp/repro_script.py"
  1. Run control experiments to isolate the variable:
    • If suspecting TP issue: run with TP=1 as control
    • If suspecting GPU issue: compare same code on different GPU
    • If suspecting a specific commit: test before/after that commit
Phase 6: Report
  1. Produce a structured report:
markdown
## CI Regression Bisection Report

### Failure Signature
- **Test**: {test_file}::{test_method}
- **Error**: {exact error message}
- **Key metrics**: {numeric values}
- **Deterministic**: Yes/No

### Root Cause Classification
One of:
- **Code Regression**: PR #{number} introduced the bug
- **Hardware-Specific**: Fails on {GPU_TYPE}, passes on others
- **Environment Change**: New runner/driver/package version
- **Pre-existing Flakiness**: Intermittent, not a new regression

### Evidence
| Condition | Result |
|-----------|--------|
| {condition1} | PASS/FAIL |
| {condition2} | PASS/FAIL |

### Timeline
- {date}: Last known pass ({sha}, {runner})
- {date}: First known fail ({sha}, {runner})
- {date}: Confirmed reproduction on {server}

### Recommended Fix
- **Short-term**: {workaround}
- **Long-term**: {proper fix}

Key Patterns to Recognize

PatternDiagnosis
Same SHA passes on runner A, fails on runner BHardware/runner-specific
All runners fail after commit XCode regression from commit X
Intermittent - same runner sometimes passes/failsFlaky test or race condition
Prefill OK but decode failsTP/all-reduce issue in decode path
Works with TP=1, fails with TP>1Tensor parallelism bug
Exact same numeric diff every timeDeterministic bug, not flakiness

Important Notes

  • Always check runner identity before concluding it's a code regression. Many "consistent" failures are actually runner-specific.
  • Test partition assignments change over time as tests are added/removed. A test may move between partitions, landing on different runner types.
  • H200 runners use /root/actions-runner/ path and machine names like gpu-h200-worker-*. Non-H200 runners use /public_sglang_ci/runner-* paths.
  • When running remote reproduction, use run_in_background for long-running tests and check output with TaskOutput.
  • Container environments may be stale - always verify package versions match CI before drawing conclusions.

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/sglang-bisect-ci-regression of sgl-project/sglang.

Open the folder on GitHubat commit dab108b

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Sglang Bisect CI Regression next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sglang Bisect CI Regression compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sglang Bisect CI Regression this skillsgl-project/sglang37k2 repos~2.5kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Upgrade Depsareal-project/AReaL5.8k—~6kAutomated safety check: PassApache-2.0
Hyperloom SetupAMD-AGI/Hyperloom219—~7.1kAutomated safety check: NotesCustom licence
Quark Torch LLM Evalamd/Quark182—~6.2kAutomated safety check: PassMIT
Install Miles Diffusionradixark/miles_diffusion110—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Upgrade Deps

    areal-project/AReaL

    Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~6k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Hyperloom Setup

    AMD-AGI/Hyperloom

    Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

    219 GitHub stars~7.1k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • End-to-end LLM accuracy evaluation on AMD ROCm (ROCm-only) — container setup, vLLM/SGLang/ATOM serving, lm-eval / lighteval / evalscope benchmarks.

    182 GitHub stars~6.2k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Install Miles Diffusion

    radixark/miles_diffusion

    Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them.

    110 GitHub stars~1.6k tokensUpdated today
    DevOps & CloudAuto-check passed
  • slime RL Post-Training

    Orchestra-Research/AI-Research-SKILLs

    Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.

    13k GitHub starsUsed in 4 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed

More from sgl-project/sglang

All 32 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Works with

Questions about Sglang Bisect CI Regression

What does Sglang Bisect CI Regression do?

Investigate consistently failing SGLang CI tests by extracting the failure signature from scheduled or rerun workflows, bisecting the passing/failing commit window, checking runner or hardware…. Sglang Bisect CI Regression is an agent skill from sgl-project/sglang. Investigate consistently failing SGLang CI tests by extracting the failure signature from scheduled or rerun workflows, bisecting the passing/failing commit window, checking runner or hardware specificity, and optionally reproducing on a remote GPU host.

When should I use Sglang Bisect CI Regression?

Sglang Bisect CI Regression fits situations like: AI & LLM Engineering work in your project.

How do I install Sglang Bisect CI Regression in Claude Code?

Run `npx skills add sgl-project/sglang --skill sglang-bisect-ci-regression -a claude-code`. Or copy the skill folder (.agents/skills/sglang-bisect-ci-regression in sgl-project/sglang) into .claude/skills/sglang-bisect-ci-regression in your project. Claude Code loads it when a task matches its description.

How do I install Sglang Bisect CI Regression in Codex?

Run `npx skills add sgl-project/sglang --skill sglang-bisect-ci-regression -a codex`. Or copy the skill folder (.agents/skills/sglang-bisect-ci-regression in sgl-project/sglang) into .agents/skills/sglang-bisect-ci-regression in your project. Codex loads it when a task matches its description.

Can I use Sglang Bisect CI Regression in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill sglang-bisect-ci-regression -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sglang-bisect-ci-regression, .gemini/skills/sglang-bisect-ci-regression, .github/skills/sglang-bisect-ci-regression and .opencode/skills/sglang-bisect-ci-regression in your project.

What does Sglang Bisect CI Regression need to run?

Going by SKILL.md and its folder, Sglang Bisect CI Regression needs the command-line tools its instructions call (gh, ssh, git, pip, curl and scp). Our summary lists: Python 3; Docker.

Does Sglang Bisect CI Regression access the network?

SKILL.md contains no URLs. Its commands use gh, ssh, git, pip and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Sglang Bisect CI Regression safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sglang Bisect CI Regression use?

Sglang Bisect CI Regression is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sglang Bisect CI Regression use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sglang Bisect CI Regression?

Skills that share tags, products or a category with Sglang Bisect CI Regression: Dstack Prototyping (dstackai/dstack, 2.3k stars), Upgrade Deps (areal-project/AReaL, 5.8k stars), Hyperloom Setup (AMD-AGI/Hyperloom, 219 stars) and Quark Torch LLM Eval (amd/Quark, 182 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sglang Bisect CI Regression?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,973 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 11, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.