Agent skill

CI Workflow Guide

by sgl-project in sgl-project/sglang

Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures.

Apache-2.0Auto-check passedDevOps & Cloud

Install CI Workflow Guide

skills CLI
$ npx skills add sgl-project/sglang --skill ci-workflow-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang ci-workflow-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ci-workflow-guide .claude/skills/ci-workflow-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ci-workflow-guide
GitHub stars
37k
Used in
2 other repos
Token cost
~5.5k tokens
SKILL.md length
1,833 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures.

  • Works in 6 steps: Calls listJobsForWorkflowRun to list all… → Matches jobs by exact name or prefix… → If any matched job has conclusion ===… → …
  • Modifying CI workflows
  • SKILL.md covers Naming Conventions, Key Files, Architecture Overview and Fail-Fast Layers, plus 10 more sections
  • Calls python3

What it does

CI Workflow Guide is an agent skill from sgl-project/sglang. Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures. Use when modifying CI workflows, adding stages, debugging CI pipeline issues, or understanding how tests are dispatched and gated across stages.

Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering CI/CD, Failing and flaky tests and Debugging. It works with SGLang and GitHub Actions. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • Modifying CI workflows
  • Debugging CI pipeline issues
  • Understanding how tests are dispatched and gated across stages

Example prompts

  • “/ci-workflow-guide”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Calls listJobsForWorkflowRun to list all jobs in the current run
  2. Matches jobs by exact name or prefix (for matrix jobs, e.g., base-b-test-1-gpu-small (3))
  3. If any matched job has conclusion === 'failure' → fail immediately (fail-fast)
  4. If all matched jobs are completed and count matches expected_count → success
  5. Otherwise → sleep poll-interval-seconds (default: 60s) and retry
  6. Timeout after max-wait-minutes (240 min for base-a, 480 min for base-b)

What it can do on your machine

Read from SKILL.md and the folder at commit 1c42ad3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.sglang.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CI Workflow Guide loads about 5.5k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 1,833 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~5.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit 1c42ad3, republished under its Apache-2.0 licence (© sgl-project). 1,833 words, ~5,456 tokens.

Download SKILL.mdSave it as .claude/skills/ci-workflow-guide/SKILL.md (or your agent's skills folder).
name
ci-workflow-guide
description
Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures. Use when modifying CI workflows, adding stages, debugging CI pipeline issues, or understanding how tests are dispatched and gated across stages.

SGLang CI Workflow Orchestration Guide

This skill covers the CI infrastructure layer — how tests are dispatched, gated, and aborted on failure across stages. For test authoring (templates, fixtures, registration, model selection), see the write-sglang-test skill.


Naming Conventions

  • Suite: base-{a,b,c}-test-{gpu_count}-gpu-{hardware} (e.g., base-b-test-1-gpu-small)
  • Test group: Directory-level registered test group under test/registered/ (e.g., hicache maps to test/registered/hicache/test_*.py)
  • CI runner: {gpu_count}-gpu-{hardware} (e.g., 1-gpu-5090, 4-gpu-h100, 8-gpu-h200)

Key Files

FileRole
.github/workflows/pr-test.ymlMain workflow — all stages, jobs, conditions, matrix definitions
.github/workflows/pr-test-extra.ymlExtra workflow — gated by BOTH run-ci and run-ci-extra labels
.github/workflows/pr-gate.ymlPR gating: draft check, run-ci label, per-user rate limiting
.github/actions/check-pr-test-health/action.ymlCross-job fail-fast: queries API for any failed job
.github/actions/wait-for-jobs/action.ymlStage gating: polls API until stage jobs complete
.github/actions/check-maintenance/action.ymlMaintenance mode check
.github/scripts/ci-labels.cjsResolves the four CI control labels into dispatch axes
test/run_suite.pySuite runner: collects, filters, partitions, executes tests
python/sglang/test/ci/ci_register.pyTest registration (AST-parsed markers), LPT auto-partition
python/sglang/test/ci/ci_utils.pyrun_unittest_files(): execution, retry, continue-on-error
scripts/ci/utils/slash_command_handler.pyHandles slash commands from PR comments

Architecture Overview

 ┌──────────────┐
 │ build kernel │
 └──────┬───────┘
        │
        ├─ check-changes ──── detects which packages changed
        │                      (main_package, sgl_kernel, jit_kernel, multimodal_gen)
        │
        ├─ call-gate ──────── pr-gate.yml (draft? label? rate limit?)
        │
        ├─────────────────────────────────────────────────────┐
        │                                                     │
        ▼                                                     │
 ┌─────────────────────────────────────┐                      │
 │          Base A (~3 min)            │                      │
 │         pre-flight check            │                      │
 │                                     │                      │
 │  ┌─────────────────────────────┐    │                      │
 │  │ base-a-test-1-gpu-small    │    │                      │
 │  │ (small GPUs)                │    │                      │
 │  └─────────────────────────────┘    │                      │
 │  ┌─────────────────────────────┐    │                      │
 │  │ base-a-test-cpu            │    │                      │
 │  │ (CPU)                       │    │                      │
 │  └─────────────────────────────┘    │                      │
 └──────┬──────────────────────────────┘                      │
        │                                                     │
        ▼                                                     ▼
 ┌─────────────────────────────────────┐          ┌──────────────────────────┐
 │          Base B (~30 min)           │          │      kernel test         │
 │            base tests               │          └──────────────────────────┘
 │                                     │          ┌──────────────────────────┐
 │  ┌─────────────────────────────┐    │          │   multimodal gen test    │
 │  │ base-b-test-1-gpu-small    │    │          └──────────────────────────┘
 │  │ (small GPUs, e.g. 5090)     │    │
 │  └─────────────────────────────┘    │
 │  ┌─────────────────────────────┐    │
 │  │ base-b-test-1-gpu-large    │    │
 │  │ (large GPUs, e.g. H100)     │    │
 │  └─────────────────────────────┘    │
 │  ┌─────────────────────────────┐    │
 │  │ base-b-test-2-gpu-large    │    │
 │  │ (large GPUs, e.g. H100)     │    │
 │  └─────────────────────────────┘    │
 └──────┬──────────────────────────────┘
        │
        ▼
 ┌─────────────────────────────────────┐
 │          Base C (~30 min)           │
 │          advanced tests             │
 │                                     │
 │  ┌─────────────────────────────┐    │
 │  │ base-c-test-4-gpu-h100     │    │
 │  │ (H100 GPUs)                 │    │
 │  └─────────────────────────────┘    │
 │  ┌─────────────────────────────┐    │
 │  │ base-c-test-8-gpu-h200     │    │
 │  │ (8 x H200 GPUs)             │    │
 │  └─────────────────────────────┘    │
 │  ┌─────────────────────────────┐    │
 │  │ base-c-test-4-gpu-b200     │    │
 │  │ (4 x B200 GPUs)             │    │
 │  └─────────────────────────────┘    │
 │  ┌─────────────────────────────┐    │
 │  │ Other advanced tests        │    │
 │  │ (DeepEP, PD Disagg, GB300)  │    │
 │  └─────────────────────────────┘    │
 └──────┬──────────────────────────────┘
        │
        ▼
 ┌─────────────────────────────────────┐
 │         pr-test-finish              │
 │  aggregates all results, fails if   │
 │  any job failed/cancelled           │
 └─────────────────────────────────────┘

Every stage test job includes a check-pr-test-health step after checkout — if any job in the run has already failed, the job fails fast (red X) with a root cause annotation.

Scheduled runs skip wait-for-base-* jobs, running all stages in parallel. Fail-fast is also disabled.


Fail-Fast Layers

4 layers of fail-fast, from fine to coarse:

LayerMechanismGranularityDisabled on schedule?
1. Test method → fileunittest -f (failfast)One test method fails → entire test file stops immediatelyYes
2. File → suiterun_unittest_files() defaultOne test file fails → entire suite stops (--continue-on-error off)Yes
3. Job → job (same stage)check-pr-test-health actionOne job fails → other waiting jobs in same stage fail-fast (red X)Yes
4. Stage → stage (cross-stage)wait-for-base-* + needsBase A fails → base B/C jobs skip entirely (never get a runner)Yes (wait jobs skipped)
  • Layer 1: -f flag appended to all python3 -m pytest / unittest invocations in ci_utils.py
  • Layer 2: --continue-on-error flag in run_suite.py — off for PRs, on for scheduled runs
  • Layer 3: check-pr-test-health auto-detects schedule event and skips; filters out cascade failures to show only root cause jobs
  • Layer 4: wait-for-base-* jobs are conditioned on github.event_name == 'pull_request' — skipped for scheduled runs

Execution Modes

AspectPR (pull_request)Scheduled (cron, every 6h)Manual dispatch (workflow_dispatch)
Stage orderingSequential: A → B → C via wait-for-base-*Parallel (all at once)Single target stage only
Cross-job fail-fastYes (check-pr-test-health)YesYes
continue-on-errorNo (stop at first failure within suite)Yes (run all tests)No
RetryEnabledEnabledEnabled
max_parallel3 (default), 14 if max-concurrency label143 (default), 14 if max-concurrency
PR gateYes (draft, label, rate limit)SkippedSkipped
Concurrencycancel-in-progress: true per PRQueue (no cancel)Isolated per stage+SHA

Four labels relax these limits for one PR: bypass-fail-fast, parallel-stages, max-concurrency, and highest-priority (all three). .github/scripts/ci-labels.cjs resolves them; the contribution guide describes what each one does.


Stage Gating (wait-for-jobs action)

wait-for-base-a and wait-for-base-b are lightweight ubuntu-latest jobs that poll the GitHub Actions API.

How it works:

  1. Calls listJobsForWorkflowRun to list all jobs in the current run
  2. Matches jobs by exact name or prefix (for matrix jobs, e.g., base-b-test-1-gpu-small (3))
  3. If any matched job has conclusion === 'failure' → fail immediately (fail-fast)
  4. If all matched jobs are completed and count matches expected_count → success
  5. Otherwise → sleep poll-interval-seconds (default: 60s) and retry
  6. Timeout after max-wait-minutes (240 min for base-a, 480 min for base-b)

Job specs example (base-b):

json
[
  {"prefix": "base-b-test-1-gpu-small", "expected_count": 8},
  {"prefix": "base-b-test-1-gpu-large", "expected_count": 14},
  {"prefix": "base-b-test-2-gpu-large", "expected_count": 4},
  {"prefix": "base-b-test-4-gpu-b200", "expected_count": 1}
]

Critical: expected_count must match the matrix size. If you add/remove matrix entries, update the wait job's spec accordingly.

PR only: Condition github.event_name == 'pull_request' && !inputs.target_stage — scheduled runs and manual dispatches skip these entirely, allowing parallel execution.


Cross-Job Fail-Fast (check-pr-test-health action)

Composite action called after checkout in every stage test job (21 jobs total across pr-test.yml, pr-test-multimodal-gen.yml, pr-test-sgl-kernel.yml, pr-test-jit-kernel.yml).

How it works:

  1. Queries listJobsForWorkflowRun for the current workflow run
  2. Filters for root cause failures only — jobs with conclusion === 'failure' whose failing step is NOT check-pr-test-health (excludes cascade failures)
  3. If root cause failures found → calls core.setFailed() with the list of root cause job names
  4. If none → does nothing (step succeeds)

Cascade filtering: When job A fails fast due to the health check, it also has conclusion: failure. Without filtering, job B would list both the original failure AND job A's fail-fast. The filter checks each failed job's steps array — if the failing step name contains check-pr-test-health or Check PR test health, it's excluded from the root cause list.

Usage pattern:

yaml
steps:
  - name: Checkout code
    uses: actions/checkout@v4
    ...

  - uses: ./.github/actions/check-pr-test-health
    id: pr-test-health

  - name: Install dependencies        # skipped automatically if health check failed
    ...                                # (default if: success() is false)

  - name: Run test                     # also skipped
    ...

Visual effect: Job shows red X (failure) with error annotation showing root cause job names. Subsequent steps are naturally skipped (default if: success() is false after a failed step). No per-step if guards needed.

No stage filtering: Checks ALL jobs in the run, not just the current stage. Any failure anywhere triggers fail-fast.

Error message example:

Fail-fast: skipping — root cause job(s): base-b-test-1-gpu-small (0), base-b-test-1-gpu-small (1)

Within-Suite Failure Handling

Controlled by run_unittest_files() in python/sglang/test/ci/ci_utils.py.

Flags
FlagPR defaultScheduled defaultEffect
--continue-on-errorOffOnOff: stop at first failure. On: run all files, report all failures at end
--enable-retryOnOnRetry retriable failures (accuracy/perf assertions)
--max-attempts22Max attempts per file including initial run
Retry Classification

When a test fails and retry is enabled, the output is classified:

Non-retriable (checked first — real code errors): SyntaxError, ImportError, ModuleNotFoundError, NameError, TypeError, AttributeError, RuntimeError, CUDA out of memory, OOM, Segmentation fault, core dumped, ConnectionRefusedError, FileNotFoundError

Retriable (accuracy/performance): AssertionError with comparison patterns (not greater than, not less than, not equal to), accuracy, score, latency, throughput, timeout

Default: Unknown AssertionError → retriable. Other unknown failures → not retriable.

How continue_on_error is set

In pr-test.yml's check-changes job:

  • schedule runs or run_all_tests flag → continue_on_error = 'true'
  • PR runs → continue_on_error = 'false'

Each test job propagates via:

yaml
env:
  CONTINUE_ON_ERROR_FLAG: ${{ needs.check-changes.outputs.continue_on_error == 'true' && '--continue-on-error' || '' }}
run: |
  python3 run_suite.py --hw cuda --suite <name> $CONTINUE_ON_ERROR_FLAG

Test Partitioning

Large suites are split across matrix jobs using the LPT (Longest Processing Time) heuristic in ci_register.py:auto_partition():

  1. Sort tests by est_time descending, filename as tie-breaker (deterministic)
  2. Greedily assign each test to the partition with smallest cumulative time
  3. Result: roughly equal total time per partition

Partition table (CUDA per-commit suites):

SuitePartitionsRunnermax_parallel
base-a-test-1-gpu-small1 (no matrix)1-gpu-5090—
base-a-test-cpu4ubuntu-latest—
base-b-test-1-gpu-small81-gpu-50908
base-b-test-1-gpu-large141-gpu-h100dynamic (3 or 14)
base-b-test-2-gpu-large42-gpu-h100—
base-b-test-4-gpu-b2001 (no matrix)4-gpu-b200—
base-b-kernel-unit-test-1-gpu-large1 (no matrix)1-gpu-h100—
base-b-kernel-unit-test-4-gpu-b2001 (no matrix)4-gpu-b200—
base-b-kernel-unit-test-8-gpu-h2001 (no matrix)8-gpu-h200—
base-b-kernel-benchmark-test-1-gpu-large1 (no matrix)1-gpu-h100—
base-c-test-4-gpu-h10034-gpu-h100—
base-c-test-8-gpu-h20048-gpu-h200—
base-c-test-8-gpu-h2028-gpu-h20—
base-c-test-deepep-4-gpu-h1001 (no matrix)4-gpu-h100—
base-c-test-4-gpu-b20034-gpu-b200—
base-c-test-4-gpu-b200-small34-gpu-b200-low-disk—
base-c-test-8-gpu-b200registered only8-gpu-b200—
base-c-test-4-gpu-gb200registered only4-gpu-gb200—

Suite names are generated, not hand-written: each comes from a test's register_*_ci(stage=..., runner_config=...) as {stage}-test-{runner_config}, and runner_config maps to the Runner column via scripts/ci/runner_configs.yml.

Note: Kernel suites (base-b-kernel-*) run via pr-test-jit-kernel.yml and pr-test-sgl-kernel.yml, not the main pr-test.yml. base-c-test-8-gpu-b200 is registered in test/run_suite.py but not wired to PR CI. The GB200 job is currently commented out in pr-test.yml until a company-owned runner is provisioned. Multimodal diffusion uses python/sglang/multimodal_gen/test/run_suite.py, not test/run_suite.py.

Workflow usage:

yaml
strategy:
  matrix:
    partition: [0, 1, 2, 3, 4, 5, 6, 7]
steps:
  - run: python3 run_suite.py --hw cuda --suite base-b-test-1-gpu-small \
           --auto-partition-id ${{ matrix.partition }} --auto-partition-size 8

Show full SKILL.md (713 more words)Show less

check-changes Job

Determines which test suites to run based on file changes.

Detection Methods
TriggerMethodDetails
pull_requestdorny/paths-filterDetects changes via GitHub diff
workflow_dispatch (with pr_head_sha)GitHub APIrepos/{repo}/compare/main...{sha}
schedule / run_all_testsForce all trueRuns everything
Output Flags
OutputTriggers
main_packageBase A/B/C test suites
sgl_kernelKernel wheel builds + kernel test suites; also switches B200 jobs to kernel-build runner labels outside target_stage mode
jit_kernelJIT kernel test workflow
multimodal_genMultimodal-gen test workflow

Note: In target_stage mode, sgl_kernel is only active when include_wheel_build=true. Without that opt-in, kernel-change reruns fail validation instead of running a target stage without freshly built wheels. Outside target_stage, sgl_kernel=true switches B200 jobs from 4-gpu-b200 / 4-gpu-b200-low-disk to 4-gpu-b200-kernel / 4-gpu-b200-kernel-low-disk.


Concurrency Control

group: pr-test-{event_name}-{branch}-{pr_sha}-{stage}
SegmentSourcePurpose
event_namegithub.event_namePrevents scheduled runs colliding with fork PRs named main
branchgithub.head_ref || github.ref_namePer-branch isolation
pr_shainputs.pr_head_sha || 'current'Isolates manual dispatches from main runs
stageinputs.target_stage || 'all'Allows parallel stage dispatches

cancel-in-progress: true for pull_request events (new push cancels old run), false for workflow_call.


How To: Add a New Stage Job

  1. Define the job in pr-test.yml with needs: [check-changes, call-gate, wait-for-base-X, ...]
  2. Copy the if: condition pattern from an existing same-stage job (handles target_stage, schedule, main_package)
  3. Add checkout step
  4. Add check-pr-test-health step (after checkout) — if any prior job failed, core.setFailed() fires and all subsequent steps auto-skip via default if: success()
  5. Add check-maintenance step
  6. Add download-artifact step if sgl_kernel changed
  7. Add install dependencies step
  8. Add run test step with $CONTINUE_ON_ERROR_FLAG
  9. Add upload-cuda-coredumps step with if: always()
  10. Register the suite name in PER_COMMIT_SUITES in test/run_suite.py
  11. If using matrix, add --auto-partition-id and --auto-partition-size to the run command
  12. Update wait-for-base-X job spec with the new job name and expected_count (if matrix)
  13. Add the job to pr-test-finish.needs list

How To: Debug CI Failures

SymptomLikely causeWhat to check
All stage-B/C jobs green but steps skippedEarlier job failed, check-pr-test-health triggeredFind the actual failed job (red X)
wait-for-base-b timeoutexpected_count doesn't match matrix sizeVerify job spec counts match matrix: array length
pr-test-finish fails but all jobs greenA job was cancelled (counts as failure in finish)Check concurrency cancellation
Tests pass locally but fail in CIPartition assignment, runner GPU type, or est_time inaccuracyCheck which partition the test lands in; verify runner label
Flaky test retried and passedRetriable failure (accuracy/perf)Check [CI Retry] markers in job logs
Flaky test NOT retriedMatched non-retriable patternCheck if error matches NON_RETRIABLE_PATTERNS in ci_utils.py

Slash Commands

CommandEffect
/tag-run-ci-labelAdds run-ci label to PR
/tag-run-ci-label extraAdds both run-ci and run-ci-extra labels
/rerun-failed-ciReruns failed jobs in the latest workflow run
/tag-and-rerun-ciAdds run-ci label + reruns failed
/tag-and-rerun-ci extraAdds both run-ci and run-ci-extra labels + reruns failed
/run-full-ciShort form of /tag-and-rerun-ci extra (baseline + extra). Alias: /rerun-full-ci
/run-extra-ciAdds both labels + reruns only PR Test Extra, leaving baseline runs alone. Alias: /rerun-extra-ci
/rerun-test <test-file> [<test-file> ...]Reruns specific test file(s) via rerun-test.yml. A file arg containing a glob metacharacter (*, ?, [...]) expands against test/registered/ and the multimodal test dir to every matching test_*.py (e.g. /rerun-test test_*backend*.py — wrap in backticks so GitHub doesn't italicize the *); matches are deduped, grouped by dispatch shape, and can't carry a ::test selector. No match → single ⛔ reply, nothing dispatched.
/rerun-group <group> [<group> ...]Expands registered test groups, then reuses /rerun-test

Handled by scripts/ci/utils/slash_command_handler.py → .github/workflows/slash-command-handler.yml.

Label-gated workflow dispatch (pr-test, pr-test-extra)

pr-test.yml and pr-test-extra.yml both listen for pull_request.labeled (in addition to opened/synchronize/reopened). The check-changes.if gate has two clauses:

  1. For labeled events: the just-added label must be one of the gating labels (run-ci for pr-test, run-ci or run-ci-extra for pr-test-extra) — otherwise every unrelated label addition would dispatch a full CI run.
  2. All events: the PR must currently carry the required labels.

This is what lets /tag-run-ci-label (and the extra variant) trigger a fresh CI run without an extra push.

Caveat — skipped runs cannot be un-skipped by run.rerun(): GitHub's rerun API reuses the original event payload, so rerunning a pull_request-event run that was skipped because of missing labels will skip again (label set in the frozen payload doesn't update). The only way to recover a label-skipped run is to add the missing label, which fires a fresh labeled event with the current label set. handle_rerun_failed_ci in the slash handler is for rerunning failed/non-label-skipped runs; it cannot revive label-skipped ones.

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/ci-workflow-guide of sgl-project/sglang.

Open the folder on GitHubat commit 1c42ad3

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

CI Workflow Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CI Workflow Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CI Workflow Guide this skillsgl-project/sglang37k2 repos~5.5kAutomated safety check: PassApache-2.0
CI/CD Failure Troubleshootingruby-git/ruby-git1.8k—~1.9kAutomated safety check: PassMIT
Debugging CI FailuresPostHog/posthog-foss721—~5.7kAutomated safety check: PassMIT
CI Debuggingcitypaul/.dotfiles739—~1.5kAutomated safety check: NotesCustom licence
Authoring CI WorkflowsPostHog/posthog-foss721—~11kAutomated safety check: PassMIT
ReleaseSma1lboy/rove146—~3.1kAutomated safety check: WarnMIT

Similar skills

  • Diagnoses and fixes failing GitHub Actions runs by identifying the failure, fetching only the relevant logs, finding the root cause and reproducing it locally.

    1.8k GitHub stars~1.9k tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed
  • Debugging CI Failures

    PostHog/posthog-foss

    Official

    Debugs failing GitHub Actions CI runs for PostHog PRs, commits, and branches, and answers broad CI-health questions ("is CI red?", "is master green today?", "what's broken right now?").

    721 GitHub stars~5.7k tokensUpdated today
    Testing & QAAuto-check passed
  • CI Debugging

    citypaul/.dotfiles

    Systematic CI/CD failure diagnosis using hypothesis-first investigation, local reproduction, and environment delta analysis.

    739 GitHub stars~1.5k tokensUpdated 4 days ago
    DevOps & CloudAuto-check: notes
  • Authoring CI Workflows

    PostHog/posthog-foss

    Official

    A skill your agent uses when adding or editing a GitHub Actions workflow, composite action, or reusable workflow under .github/ — new CI jobs, triggers, matrices, checkout/clone tuning, action…

    721 GitHub stars~11k tokensUpdated today
    DevelopmentAuto-check passed
  • Release

    Sma1lboy/rove

    Autonomously cut a Rove (@sma1lboy/rove) release end-to-end — detect the semver bump from pending changesets (flagging an upstream minor you didn't intend), run the release gates, dispatch the…

    146 GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check: warnings
  • Manor CI Triage

    manor-os/manor-ai

    A skill your agent uses when Manor GitHub Actions, .github/workflows/ci.yml, OSS smoke/regression jobs, web source smoke, frontend build, lint, or public CI failure logs need diagnosis or repair.

    161 GitHub stars~533 tokensUpdated 29 days ago
    DevOps & CloudAuto-check passed

More from sgl-project/sglang

All 31 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Questions about CI Workflow Guide

What does CI Workflow Guide do?

Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures. CI Workflow Guide is an agent skill from sgl-project/sglang. Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures.

When should I use CI Workflow Guide?

CI Workflow Guide fits situations like: modifying CI workflows; debugging CI pipeline issues; understanding how tests are dispatched and gated across stages.

How do I install CI Workflow Guide in Claude Code?

Run `npx skills add sgl-project/sglang --skill ci-workflow-guide -a claude-code`. Or copy the skill folder (.agents/skills/ci-workflow-guide in sgl-project/sglang) into .claude/skills/ci-workflow-guide in your project. Claude Code loads it when a task matches its description.

How do I install CI Workflow Guide in Codex?

Run `npx skills add sgl-project/sglang --skill ci-workflow-guide -a codex`. Or copy the skill folder (.agents/skills/ci-workflow-guide in sgl-project/sglang) into .agents/skills/ci-workflow-guide in your project. Codex loads it when a task matches its description.

Can I use CI Workflow Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill ci-workflow-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ci-workflow-guide, .gemini/skills/ci-workflow-guide, .github/skills/ci-workflow-guide and .opencode/skills/ci-workflow-guide in your project.

What does CI Workflow Guide need to run?

Going by SKILL.md and its folder, CI Workflow Guide needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does CI Workflow Guide access the network?

SKILL.md names 1 domain. As links in the text: docs.sglang.io. This is read from the text; nothing was executed.

Is CI Workflow Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does CI Workflow Guide use?

CI Workflow Guide is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CI Workflow Guide use?

About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to CI Workflow Guide?

Skills that share tags, products or a category with CI Workflow Guide: CI/CD Failure Troubleshooting (ruby-git/ruby-git, 1.8k stars), Debugging CI Failures (PostHog/posthog-foss, 721 stars), CI Debugging (citypaul/.dotfiles, 739 stars) and Authoring CI Workflows (PostHog/posthog-foss, 721 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CI Workflow Guide?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,829 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.