CI/CD Failure Troubleshooting
ruby-git/ruby-git
Diagnoses and fixes failing GitHub Actions runs by identifying the failure, fetching only the relevant logs, finding the root cause and reproducing it locally.
Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures.
$ npx skills add sgl-project/sglang --skill ci-workflow-guide -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sgl-project/sglang ci-workflow-guide --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ci-workflow-guide .claude/skills/ci-workflow-guide && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ci-workflow-guide" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/ci-workflow-guide into .claude/skills/ci-workflow-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ci-workflow-guide", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sgl-project/sglang/tree/main/.agents/skills/ci-workflow-guideType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sgl-project/sglang --skill ci-workflow-guide -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sgl-project/sglang ci-workflow-guide --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/ci-workflow-guide .agents/skills/ci-workflow-guide && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ci-workflow-guide" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/ci-workflow-guide into .agents/skills/ci-workflow-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ci-workflow-guide", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill ci-workflow-guide -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sgl-project/sglang ci-workflow-guide --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/ci-workflow-guide .cursor/skills/ci-workflow-guide && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ci-workflow-guide" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/ci-workflow-guide into .cursor/skills/ci-workflow-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ci-workflow-guide", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sgl-project/sglang.git --path .agents/skills/ci-workflow-guide--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sgl-project/sglang --skill ci-workflow-guide -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sgl-project/sglang ci-workflow-guide --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/ci-workflow-guide .gemini/skills/ci-workflow-guide && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ci-workflow-guide" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/ci-workflow-guide into .gemini/skills/ci-workflow-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ci-workflow-guide", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sgl-project/sglang ci-workflow-guideInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sgl-project/sglang --skill ci-workflow-guide -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/ci-workflow-guide .github/skills/ci-workflow-guide && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ci-workflow-guide" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/ci-workflow-guide into .github/skills/ci-workflow-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ci-workflow-guide", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill ci-workflow-guide -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sgl-project/sglang ci-workflow-guide --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/ci-workflow-guide .opencode/skills/ci-workflow-guide && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ci-workflow-guide" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/ci-workflow-guide into .opencode/skills/ci-workflow-guide/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ci-workflow-guide", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ci-workflow-guideGuide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures.
CI Workflow Guide is an agent skill from sgl-project/sglang. Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures. Use when modifying CI workflows, adding stages, debugging CI pipeline issues, or understanding how tests are dispatched and gated across stages.
Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering CI/CD, Failing and flaky tests and Debugging. It works with SGLang and GitHub Actions. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1c42ad3. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.sglang.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
CI Workflow Guide loads about 5.5k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 1,833 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sgl-project/sglang at commit 1c42ad3, republished under its Apache-2.0 licence (© sgl-project). 1,833 words, ~5,456 tokens.
.claude/skills/ci-workflow-guide/SKILL.md (or your agent's skills folder).This skill covers the CI infrastructure layer — how tests are dispatched, gated, and aborted on failure across stages. For test authoring (templates, fixtures, registration, model selection), see the write-sglang-test skill.
base-{a,b,c}-test-{gpu_count}-gpu-{hardware} (e.g., base-b-test-1-gpu-small)test/registered/ (e.g., hicache maps to test/registered/hicache/test_*.py){gpu_count}-gpu-{hardware} (e.g., 1-gpu-5090, 4-gpu-h100, 8-gpu-h200)| File | Role |
|---|---|
.github/workflows/pr-test.yml | Main workflow — all stages, jobs, conditions, matrix definitions |
.github/workflows/pr-test-extra.yml | Extra workflow — gated by BOTH run-ci and run-ci-extra labels |
.github/workflows/pr-gate.yml | PR gating: draft check, run-ci label, per-user rate limiting |
.github/actions/check-pr-test-health/action.yml | Cross-job fail-fast: queries API for any failed job |
.github/actions/wait-for-jobs/action.yml | Stage gating: polls API until stage jobs complete |
.github/actions/check-maintenance/action.yml | Maintenance mode check |
.github/scripts/ci-labels.cjs | Resolves the four CI control labels into dispatch axes |
test/run_suite.py | Suite runner: collects, filters, partitions, executes tests |
python/sglang/test/ci/ci_register.py | Test registration (AST-parsed markers), LPT auto-partition |
python/sglang/test/ci/ci_utils.py | run_unittest_files(): execution, retry, continue-on-error |
scripts/ci/utils/slash_command_handler.py | Handles slash commands from PR comments |
┌──────────────┐
│ build kernel │
└──────┬───────┘
│
├─ check-changes ──── detects which packages changed
│ (main_package, sgl_kernel, jit_kernel, multimodal_gen)
│
├─ call-gate ──────── pr-gate.yml (draft? label? rate limit?)
│
├─────────────────────────────────────────────────────┐
│ │
▼ │
┌─────────────────────────────────────┐ │
│ Base A (~3 min) │ │
│ pre-flight check │ │
│ │ │
│ ┌─────────────────────────────┐ │ │
│ │ base-a-test-1-gpu-small │ │ │
│ │ (small GPUs) │ │ │
│ └─────────────────────────────┘ │ │
│ ┌─────────────────────────────┐ │ │
│ │ base-a-test-cpu │ │ │
│ │ (CPU) │ │ │
│ └─────────────────────────────┘ │ │
└──────┬──────────────────────────────┘ │
│ │
▼ ▼
┌─────────────────────────────────────┐ ┌──────────────────────────┐
│ Base B (~30 min) │ │ kernel test │
│ base tests │ └──────────────────────────┘
│ │ ┌──────────────────────────┐
│ ┌─────────────────────────────┐ │ │ multimodal gen test │
│ │ base-b-test-1-gpu-small │ │ └──────────────────────────┘
│ │ (small GPUs, e.g. 5090) │ │
│ └─────────────────────────────┘ │
│ ┌─────────────────────────────┐ │
│ │ base-b-test-1-gpu-large │ │
│ │ (large GPUs, e.g. H100) │ │
│ └─────────────────────────────┘ │
│ ┌─────────────────────────────┐ │
│ │ base-b-test-2-gpu-large │ │
│ │ (large GPUs, e.g. H100) │ │
│ └─────────────────────────────┘ │
└──────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────┐
│ Base C (~30 min) │
│ advanced tests │
│ │
│ ┌─────────────────────────────┐ │
│ │ base-c-test-4-gpu-h100 │ │
│ │ (H100 GPUs) │ │
│ └─────────────────────────────┘ │
│ ┌─────────────────────────────┐ │
│ │ base-c-test-8-gpu-h200 │ │
│ │ (8 x H200 GPUs) │ │
│ └─────────────────────────────┘ │
│ ┌─────────────────────────────┐ │
│ │ base-c-test-4-gpu-b200 │ │
│ │ (4 x B200 GPUs) │ │
│ └─────────────────────────────┘ │
│ ┌─────────────────────────────┐ │
│ │ Other advanced tests │ │
│ │ (DeepEP, PD Disagg, GB300) │ │
│ └─────────────────────────────┘ │
└──────┬──────────────────────────────┘
│
▼
┌─────────────────────────────────────┐
│ pr-test-finish │
│ aggregates all results, fails if │
│ any job failed/cancelled │
└─────────────────────────────────────┘Every stage test job includes a check-pr-test-health step after checkout — if any job in the run has already failed, the job fails fast (red X) with a root cause annotation.
Scheduled runs skip wait-for-base-* jobs, running all stages in parallel. Fail-fast is also disabled.
4 layers of fail-fast, from fine to coarse:
| Layer | Mechanism | Granularity | Disabled on schedule? |
|---|---|---|---|
| 1. Test method → file | unittest -f (failfast) | One test method fails → entire test file stops immediately | Yes |
| 2. File → suite | run_unittest_files() default | One test file fails → entire suite stops (--continue-on-error off) | Yes |
| 3. Job → job (same stage) | check-pr-test-health action | One job fails → other waiting jobs in same stage fail-fast (red X) | Yes |
| 4. Stage → stage (cross-stage) | wait-for-base-* + needs | Base A fails → base B/C jobs skip entirely (never get a runner) | Yes (wait jobs skipped) |
-f flag appended to all python3 -m pytest / unittest invocations in ci_utils.py--continue-on-error flag in run_suite.py — off for PRs, on for scheduled runscheck-pr-test-health auto-detects schedule event and skips; filters out cascade failures to show only root cause jobswait-for-base-* jobs are conditioned on github.event_name == 'pull_request' — skipped for scheduled runs| Aspect | PR (pull_request) | Scheduled (cron, every 6h) | Manual dispatch (workflow_dispatch) |
|---|---|---|---|
| Stage ordering | Sequential: A → B → C via wait-for-base-* | Parallel (all at once) | Single target stage only |
| Cross-job fail-fast | Yes (check-pr-test-health) | Yes | Yes |
| continue-on-error | No (stop at first failure within suite) | Yes (run all tests) | No |
| Retry | Enabled | Enabled | Enabled |
| max_parallel | 3 (default), 14 if max-concurrency label | 14 | 3 (default), 14 if max-concurrency |
| PR gate | Yes (draft, label, rate limit) | Skipped | Skipped |
| Concurrency | cancel-in-progress: true per PR | Queue (no cancel) | Isolated per stage+SHA |
Four labels relax these limits for one PR: bypass-fail-fast, parallel-stages,
max-concurrency, and highest-priority (all three). .github/scripts/ci-labels.cjs
resolves them; the contribution guide
describes what each one does.
wait-for-jobs action)wait-for-base-a and wait-for-base-b are lightweight ubuntu-latest jobs that poll the GitHub Actions API.
How it works:
listJobsForWorkflowRun to list all jobs in the current runbase-b-test-1-gpu-small (3))conclusion === 'failure' → fail immediately (fail-fast)expected_count → successpoll-interval-seconds (default: 60s) and retrymax-wait-minutes (240 min for base-a, 480 min for base-b)Job specs example (base-b):
[
{"prefix": "base-b-test-1-gpu-small", "expected_count": 8},
{"prefix": "base-b-test-1-gpu-large", "expected_count": 14},
{"prefix": "base-b-test-2-gpu-large", "expected_count": 4},
{"prefix": "base-b-test-4-gpu-b200", "expected_count": 1}
]Critical:
expected_countmust match the matrix size. If you add/remove matrix entries, update the wait job's spec accordingly.
PR only: Condition github.event_name == 'pull_request' && !inputs.target_stage — scheduled runs and manual dispatches skip these entirely, allowing parallel execution.
check-pr-test-health action)Composite action called after checkout in every stage test job (21 jobs total across pr-test.yml, pr-test-multimodal-gen.yml, pr-test-sgl-kernel.yml, pr-test-jit-kernel.yml).
How it works:
listJobsForWorkflowRun for the current workflow runconclusion === 'failure' whose failing step is NOT check-pr-test-health (excludes cascade failures)core.setFailed() with the list of root cause job namesCascade filtering: When job A fails fast due to the health check, it also has conclusion: failure. Without filtering, job B would list both the original failure AND job A's fail-fast. The filter checks each failed job's steps array — if the failing step name contains check-pr-test-health or Check PR test health, it's excluded from the root cause list.
Usage pattern:
steps:
- name: Checkout code
uses: actions/checkout@v4
...
- uses: ./.github/actions/check-pr-test-health
id: pr-test-health
- name: Install dependencies # skipped automatically if health check failed
... # (default if: success() is false)
- name: Run test # also skipped
...Visual effect: Job shows red X (failure) with error annotation showing root cause job names. Subsequent steps are naturally skipped (default if: success() is false after a failed step). No per-step if guards needed.
No stage filtering: Checks ALL jobs in the run, not just the current stage. Any failure anywhere triggers fail-fast.
Error message example:
Fail-fast: skipping — root cause job(s): base-b-test-1-gpu-small (0), base-b-test-1-gpu-small (1)Controlled by run_unittest_files() in python/sglang/test/ci/ci_utils.py.
| Flag | PR default | Scheduled default | Effect |
|---|---|---|---|
--continue-on-error | Off | On | Off: stop at first failure. On: run all files, report all failures at end |
--enable-retry | On | On | Retry retriable failures (accuracy/perf assertions) |
--max-attempts | 2 | 2 | Max attempts per file including initial run |
When a test fails and retry is enabled, the output is classified:
Non-retriable (checked first — real code errors):
SyntaxError, ImportError, ModuleNotFoundError, NameError, TypeError, AttributeError, RuntimeError, CUDA out of memory, OOM, Segmentation fault, core dumped, ConnectionRefusedError, FileNotFoundError
Retriable (accuracy/performance):
AssertionError with comparison patterns (not greater than, not less than, not equal to), accuracy, score, latency, throughput, timeout
Default: Unknown AssertionError → retriable. Other unknown failures → not retriable.
continue_on_error is setIn pr-test.yml's check-changes job:
schedule runs or run_all_tests flag → continue_on_error = 'true'continue_on_error = 'false'Each test job propagates via:
env:
CONTINUE_ON_ERROR_FLAG: ${{ needs.check-changes.outputs.continue_on_error == 'true' && '--continue-on-error' || '' }}
run: |
python3 run_suite.py --hw cuda --suite <name> $CONTINUE_ON_ERROR_FLAGLarge suites are split across matrix jobs using the LPT (Longest Processing Time) heuristic in ci_register.py:auto_partition():
est_time descending, filename as tie-breaker (deterministic)Partition table (CUDA per-commit suites):
| Suite | Partitions | Runner | max_parallel |
|---|---|---|---|
base-a-test-1-gpu-small | 1 (no matrix) | 1-gpu-5090 | — |
base-a-test-cpu | 4 | ubuntu-latest | — |
base-b-test-1-gpu-small | 8 | 1-gpu-5090 | 8 |
base-b-test-1-gpu-large | 14 | 1-gpu-h100 | dynamic (3 or 14) |
base-b-test-2-gpu-large | 4 | 2-gpu-h100 | — |
base-b-test-4-gpu-b200 | 1 (no matrix) | 4-gpu-b200 | — |
base-b-kernel-unit-test-1-gpu-large | 1 (no matrix) | 1-gpu-h100 | — |
base-b-kernel-unit-test-4-gpu-b200 | 1 (no matrix) | 4-gpu-b200 | — |
base-b-kernel-unit-test-8-gpu-h200 | 1 (no matrix) | 8-gpu-h200 | — |
base-b-kernel-benchmark-test-1-gpu-large | 1 (no matrix) | 1-gpu-h100 | — |
base-c-test-4-gpu-h100 | 3 | 4-gpu-h100 | — |
base-c-test-8-gpu-h200 | 4 | 8-gpu-h200 | — |
base-c-test-8-gpu-h20 | 2 | 8-gpu-h20 | — |
base-c-test-deepep-4-gpu-h100 | 1 (no matrix) | 4-gpu-h100 | — |
base-c-test-4-gpu-b200 | 3 | 4-gpu-b200 | — |
base-c-test-4-gpu-b200-small | 3 | 4-gpu-b200-low-disk | — |
base-c-test-8-gpu-b200 | registered only | 8-gpu-b200 | — |
base-c-test-4-gpu-gb200 | registered only | 4-gpu-gb200 | — |
Suite names are generated, not hand-written: each comes from a test's
register_*_ci(stage=..., runner_config=...)as{stage}-test-{runner_config}, andrunner_configmaps to theRunnercolumn viascripts/ci/runner_configs.yml.Note: Kernel suites (
base-b-kernel-*) run viapr-test-jit-kernel.ymlandpr-test-sgl-kernel.yml, not the mainpr-test.yml.base-c-test-8-gpu-b200is registered intest/run_suite.pybut not wired to PR CI. The GB200 job is currently commented out inpr-test.ymluntil a company-owned runner is provisioned. Multimodal diffusion usespython/sglang/multimodal_gen/test/run_suite.py, nottest/run_suite.py.
Workflow usage:
strategy:
matrix:
partition: [0, 1, 2, 3, 4, 5, 6, 7]
steps:
- run: python3 run_suite.py --hw cuda --suite base-b-test-1-gpu-small \
--auto-partition-id ${{ matrix.partition }} --auto-partition-size 8Determines which test suites to run based on file changes.
| Trigger | Method | Details |
|---|---|---|
pull_request | dorny/paths-filter | Detects changes via GitHub diff |
workflow_dispatch (with pr_head_sha) | GitHub API | repos/{repo}/compare/main...{sha} |
schedule / run_all_tests | Force all true | Runs everything |
| Output | Triggers |
|---|---|
main_package | Base A/B/C test suites |
sgl_kernel | Kernel wheel builds + kernel test suites; also switches B200 jobs to kernel-build runner labels outside target_stage mode |
jit_kernel | JIT kernel test workflow |
multimodal_gen | Multimodal-gen test workflow |
Note: In
target_stagemode,sgl_kernelis only active wheninclude_wheel_build=true. Without that opt-in, kernel-change reruns fail validation instead of running a target stage without freshly built wheels. Outsidetarget_stage,sgl_kernel=trueswitches B200 jobs from4-gpu-b200/4-gpu-b200-low-diskto4-gpu-b200-kernel/4-gpu-b200-kernel-low-disk.
group: pr-test-{event_name}-{branch}-{pr_sha}-{stage}| Segment | Source | Purpose |
|---|---|---|
event_name | github.event_name | Prevents scheduled runs colliding with fork PRs named main |
branch | github.head_ref || github.ref_name | Per-branch isolation |
pr_sha | inputs.pr_head_sha || 'current' | Isolates manual dispatches from main runs |
stage | inputs.target_stage || 'all' | Allows parallel stage dispatches |
cancel-in-progress: true for pull_request events (new push cancels old run), false for workflow_call.
pr-test.yml with needs: [check-changes, call-gate, wait-for-base-X, ...]if: condition pattern from an existing same-stage job (handles target_stage, schedule, main_package)checkout stepcheck-pr-test-health step (after checkout) — if any prior job failed, core.setFailed() fires and all subsequent steps auto-skip via default if: success()check-maintenance stepdownload-artifact step if sgl_kernel changedinstall dependencies steprun test step with $CONTINUE_ON_ERROR_FLAGupload-cuda-coredumps step with if: always()PER_COMMIT_SUITES in test/run_suite.py--auto-partition-id and --auto-partition-size to the run commandwait-for-base-X job spec with the new job name and expected_count (if matrix)pr-test-finish.needs list| Symptom | Likely cause | What to check |
|---|---|---|
| All stage-B/C jobs green but steps skipped | Earlier job failed, check-pr-test-health triggered | Find the actual failed job (red X) |
wait-for-base-b timeout | expected_count doesn't match matrix size | Verify job spec counts match matrix: array length |
pr-test-finish fails but all jobs green | A job was cancelled (counts as failure in finish) | Check concurrency cancellation |
| Tests pass locally but fail in CI | Partition assignment, runner GPU type, or est_time inaccuracy | Check which partition the test lands in; verify runner label |
| Flaky test retried and passed | Retriable failure (accuracy/perf) | Check [CI Retry] markers in job logs |
| Flaky test NOT retried | Matched non-retriable pattern | Check if error matches NON_RETRIABLE_PATTERNS in ci_utils.py |
| Command | Effect |
|---|---|
/tag-run-ci-label | Adds run-ci label to PR |
/tag-run-ci-label extra | Adds both run-ci and run-ci-extra labels |
/rerun-failed-ci | Reruns failed jobs in the latest workflow run |
/tag-and-rerun-ci | Adds run-ci label + reruns failed |
/tag-and-rerun-ci extra | Adds both run-ci and run-ci-extra labels + reruns failed |
/run-full-ci | Short form of /tag-and-rerun-ci extra (baseline + extra). Alias: /rerun-full-ci |
/run-extra-ci | Adds both labels + reruns only PR Test Extra, leaving baseline runs alone. Alias: /rerun-extra-ci |
/rerun-test <test-file> [<test-file> ...] | Reruns specific test file(s) via rerun-test.yml. A file arg containing a glob metacharacter (*, ?, [...]) expands against test/registered/ and the multimodal test dir to every matching test_*.py (e.g. /rerun-test test_*backend*.py — wrap in backticks so GitHub doesn't italicize the *); matches are deduped, grouped by dispatch shape, and can't carry a ::test selector. No match → single ⛔ reply, nothing dispatched. |
/rerun-group <group> [<group> ...] | Expands registered test groups, then reuses /rerun-test |
Handled by scripts/ci/utils/slash_command_handler.py → .github/workflows/slash-command-handler.yml.
pr-test.yml and pr-test-extra.yml both listen for pull_request.labeled (in addition to opened/synchronize/reopened). The check-changes.if gate has two clauses:
labeled events: the just-added label must be one of the gating labels (run-ci for pr-test, run-ci or run-ci-extra for pr-test-extra) — otherwise every unrelated label addition would dispatch a full CI run.This is what lets /tag-run-ci-label (and the extra variant) trigger a fresh CI run without an extra push.
Caveat — skipped runs cannot be un-skipped by run.rerun(): GitHub's rerun API reuses the original event payload, so rerunning a pull_request-event run that was skipped because of missing labels will skip again (label set in the frozen payload doesn't update). The only way to recover a label-skipped run is to add the missing label, which fires a fresh labeled event with the current label set. handle_rerun_failed_ci in the slash handler is for rerunning failed/non-label-skipped runs; it cannot revive label-skipped ones.
© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/ci-workflow-guide of sgl-project/sglang.
Open the folder on GitHubat commit 1c42ad3
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.
CI Workflow Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| CI Workflow Guide this skillsgl-project/sglang | 37k | 2 repos | ~5.5k | Automated safety check: Pass | Apache-2.0 | |
| CI/CD Failure Troubleshootingruby-git/ruby-git | 1.8k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Debugging CI FailuresPostHog/posthog-foss | 721 | — | ~5.7k | Automated safety check: Pass | MIT | |
| CI Debuggingcitypaul/.dotfiles | 739 | — | ~1.5k | Automated safety check: Notes | Custom licence | |
| Authoring CI WorkflowsPostHog/posthog-foss | 721 | — | ~11k | Automated safety check: Pass | MIT | |
| ReleaseSma1lboy/rove | 146 | — | ~3.1k | Automated safety check: Warn | MIT |
ruby-git/ruby-git
Diagnoses and fixes failing GitHub Actions runs by identifying the failure, fetching only the relevant logs, finding the root cause and reproducing it locally.
PostHog/posthog-foss
Debugs failing GitHub Actions CI runs for PostHog PRs, commits, and branches, and answers broad CI-health questions ("is CI red?", "is master green today?", "what's broken right now?").
citypaul/.dotfiles
Systematic CI/CD failure diagnosis using hypothesis-first investigation, local reproduction, and environment delta analysis.
PostHog/posthog-foss
A skill your agent uses when adding or editing a GitHub Actions workflow, composite action, or reusable workflow under .github/ — new CI jobs, triggers, matrices, checkout/clone tuning, action…
Sma1lboy/rove
Autonomously cut a Rove (@sma1lboy/rove) release end-to-end — detect the semver bump from pending changesets (flagging an upstream minor you didn't intend), run the release gates, dispatch the…
manor-os/manor-ai
A skill your agent uses when Manor GitHub Actions, .github/workflows/ci.yml, OSS smoke/regression jobs, web source smoke, frontend build, lint, or public CI failure logs need diagnosis or repair.
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
sgl-project/sglang
Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.
sgl-project/sglang
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
sgl-project/sglang
Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.
Works with
Categories
Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures. CI Workflow Guide is an agent skill from sgl-project/sglang. Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures.
CI Workflow Guide fits situations like: modifying CI workflows; debugging CI pipeline issues; understanding how tests are dispatched and gated across stages.
Run `npx skills add sgl-project/sglang --skill ci-workflow-guide -a claude-code`. Or copy the skill folder (.agents/skills/ci-workflow-guide in sgl-project/sglang) into .claude/skills/ci-workflow-guide in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sgl-project/sglang --skill ci-workflow-guide -a codex`. Or copy the skill folder (.agents/skills/ci-workflow-guide in sgl-project/sglang) into .agents/skills/ci-workflow-guide in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill ci-workflow-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ci-workflow-guide, .gemini/skills/ci-workflow-guide, .github/skills/ci-workflow-guide and .opencode/skills/ci-workflow-guide in your project.
Going by SKILL.md and its folder, CI Workflow Guide needs the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: docs.sglang.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
CI Workflow Guide is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with CI Workflow Guide: CI/CD Failure Troubleshooting (ruby-git/ruby-git, 1.8k stars), Debugging CI Failures (PostHog/posthog-foss, 721 stars), CI Debugging (citypaul/.dotfiles, 739 stars) and Authoring CI Workflows (PostHog/posthog-foss, 721 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,829 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.
Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.