Agent skill

Generate Profile

by sgl-project in sgl-project/sglang

Generate an e2e profiling trace of an SGLang server run. An agent skill from sgl-project/sglang.

Apache-2.0Auto-check passedTesting & QA

Install Generate Profile

skills CLI
$ npx skills add sgl-project/sglang --skill generate-profile -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang generate-profile --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/generate-profile .claude/skills/generate-profile && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
generate-profile
GitHub stars
37k
Used in
2 other repos
Token cost
~1.1k tokens
SKILL.md length
394 words
Files
1
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generate an e2e profiling trace of an SGLang server run. An agent skill from sgl-project/sglang.

  • Works in 6 steps: Launch the server → Wait for server readiness → Validate accuracy (sanity check) → …
  • Tasks that involve End-to-end testing
  • SKILL.md covers Prerequisites, Step-by-step Workflow, Example Full Run and Customization, plus 1 more section
  • Calls python3, curl and pip

What it does

Generate Profile is an agent skill from sgl-project/sglang. Generate an e2e profiling trace of an SGLang server run. Launches a server, validates accuracy, captures a Chrome-compatible trace, and returns the profile path.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering End-to-end testing. It works with SGLang and Qwen. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve End-to-end testing

Example prompts

  • “/generate-profile”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Launch the server
  2. Wait for server readiness
  3. Validate accuracy (sanity check)
  4. Generate the profile
  5. Kill the server
  6. Report the profile path

What it can do on your machine

Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • curl
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ui.perfetto.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Generate Profile loads about 1.1k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 394 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 394 words, ~1,103 tokens.

Download SKILL.mdSave it as .claude/skills/generate-profile/SKILL.md (or your agent's skills folder).
name
generate-profile
description
Generate an e2e profiling trace of an SGLang server run. Launches a server, validates accuracy, captures a Chrome-compatible trace, and returns the profile path.

Generate an E2E Profile of an SGLang Server Run

This skill launches an SGLang server, validates it with a quick accuracy test, generates a profiling trace, and returns the profile file path.

Prerequisites

  • A working SGLang installation (pip install -e . or equivalent)
  • At least one available CUDA GPU

Step-by-step Workflow

Step 1: Launch the server
bash
CUDA_VISIBLE_DEVICES=<gpu_id> sglang serve --model-path <model> --port <port> &
  • Default model: Qwen/Qwen3-8B (good balance of speed and quality)
  • Default port: 30000
  • The server runs in the background. Save the PID for cleanup.
  • Use the GPU specified by the user's preferences (check memory files for GPU preferences).
Step 2: Wait for server readiness

Poll the health endpoint until the server is ready:

bash
for i in $(seq 1 120); do
  if curl -s http://127.0.0.1:<port>/health 2>/dev/null | grep -q "ok\|healthy"; then
    echo "Server ready"
    break
  fi
  sleep 5
done

The server prints "The server is fired up and ready to roll!" to stdout when ready. The health endpoint returns 200 once the server can accept requests.

Typical startup time: 30-90 seconds depending on model size and whether CUDA graphs are being compiled.

Step 3: Validate accuracy (sanity check)
bash
sgl-eval run gsm8k --base-url http://127.0.0.1:<port>/v1 --num-examples 20
  • Expected accuracy: > 0.8 for capable models (Qwen3-8B, Llama-3.1-8B-Instruct, etc.)
  • This is a quick sanity check, not a rigorous benchmark.
  • If accuracy is unexpectedly low, something is wrong — do not proceed to profiling.
Step 4: Generate the profile
bash
python3 -m sglang.test.send_one --profile

This command:

  1. Sends a request to the server
  2. Triggers the profiler for 5 steps (default)
  3. Generates a trace file under /tmp/<timestamp>/
  4. The trace directory contains:
    • <timestamp>-TP-0.trace.json.gz — Chrome trace format (open in chrome://tracing or Perfetto)
    • server_args.json — the server configuration used

Output format:

Dump profiling traces to /tmp/<timestamp>

The profile path is printed to stdout. Parse it from the output.

Optional flags:

  • --profile-steps N — number of profiling steps (default: 5)
  • --profile-by-stage — profile by stage (prefill/decode separately)
  • --profile-prefix <path> — custom output prefix
Show full SKILL.md (123 more words)Show less
Step 5: Kill the server
bash
pkill -9 -f "sglang.launch_server\|sglang serve\|sglang.srt"

Wait a moment and verify no sglang processes remain:

bash
sleep 2 && pgrep -af "sglang serve" || echo "Server killed"
Step 6: Report the profile path

Return the profile directory path (e.g., /tmp/1773999986.4769795) and list its contents so the user knows what files were generated.

Example Full Run

bash
# 1. Launch server
source cleanup/bin/activate
CUDA_VISIBLE_DEVICES=1 sglang serve --model-path Qwen/Qwen3-8B --port 30000 &

# 2. Wait for ready
for i in $(seq 1 120); do
  curl -s http://127.0.0.1:30000/health | grep -q "ok" && break
  sleep 5
done

# 3. Accuracy check
python3 -m sglang.test.run_eval --host 127.0.0.1 --port 30000 --eval-name gsm8k --num-examples 20
# Expected: Accuracy > 0.8

# 4. Profile
python3 -m sglang.test.send_one --profile
# Output: "Dump profiling traces to /tmp/1773999986.4769795"

# 5. Cleanup
pkill -9 -f "sglang.launch_server\|sglang serve\|sglang.srt"
sleep 2

# 6. Check output
ls -la /tmp/1773999986.4769795/
# 1773999986.4851577-TP-0.trace.json.gz  (Chrome trace)
# server_args.json                        (server config)

Customization

  • Different port: Pass --port <port> and use --host 127.0.0.1 --port <port> for test commands
  • Multi-GPU: Use --tp <N> for tensor parallelism; trace files will be generated per TP rank
  • Longer profile: Use --profile-steps 10 for more steps in the trace
  • Stage profiling: Use --profile-by-stage to separate prefill and decode phases

Viewing the Profile

Open the .trace.json.gz file in:

Both support the gzipped Chrome trace format natively.

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/generate-profile of sgl-project/sglang.

Open the folder on GitHubat commit f620d73

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Generate Profile next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Generate Profile compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Generate Profile this skillsgl-project/sglang37k2 repos~1.1kAutomated safety check: PassApache-2.0
Qwen Code E2E TestingQwenLM/qwen-code28k—~2.1kAutomated safety check: PassApache-2.0
tmux Real User TestingQwenLM/qwen-code28k—~2.3kAutomated safety check: PassApache-2.0
Wally Device E2ERunanywhereAI/wally1.6k—~845Automated safety check: PassMIT
Wally E2ERunanywhereAI/wally1.6k—~3.8kAutomated safety check: PassMIT
Run Agmente E2Erebornix/Agmente545—~414Automated safety check: PassMIT

Similar skills

  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Testing & QAAuto-check passed
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Wally Device E2E

    RunanywhereAI/wally

    Run wally's LLM e2e on Apple Neural Engine (NeuRT) and Snapdragon Hexagon NPU (QHexRT) devices.

    1.6k GitHub stars~845 tokensUpdated today
    Testing & QAAuto-check passed
  • Wally E2E

    RunanywhereAI/wally

    Verify a built wally binary against a pinned C++ desktop kit on macOS and Windows.

    1.6k GitHub stars~3.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Run Agmente E2E

    rebornix/Agmente

    Run Agmente iOS end-to-end tests against a local ACP agent (Gemini, Claude, Qwen, or Vibe), validate core RPC flow, and perform mandatory cleanup.

    545 GitHub stars~414 tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • LLM Pipeline Profiler Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

    925 GitHub stars~3.9k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

More from sgl-project/sglang

All 32 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Works with

Categories

Questions about Generate Profile

What does Generate Profile do?

Generate an e2e profiling trace of an SGLang server run. An agent skill from sgl-project/sglang. Generate Profile is an agent skill from sgl-project/sglang. Generate an e2e profiling trace of an SGLang server run.

When should I use Generate Profile?

Generate Profile fits situations like: tasks that involve End-to-end testing.

How do I install Generate Profile in Claude Code?

Run `npx skills add sgl-project/sglang --skill generate-profile -a claude-code`. Or copy the skill folder (.agents/skills/generate-profile in sgl-project/sglang) into .claude/skills/generate-profile in your project. Claude Code loads it when a task matches its description.

How do I install Generate Profile in Codex?

Run `npx skills add sgl-project/sglang --skill generate-profile -a codex`. Or copy the skill folder (.agents/skills/generate-profile in sgl-project/sglang) into .agents/skills/generate-profile in your project. Codex loads it when a task matches its description.

Can I use Generate Profile in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill generate-profile -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generate-profile, .gemini/skills/generate-profile, .github/skills/generate-profile and .opencode/skills/generate-profile in your project.

What does Generate Profile need to run?

Going by SKILL.md and its folder, Generate Profile needs the command-line tools its instructions call (python3, curl and pip). Our summary lists: Python 3.

Does Generate Profile access the network?

SKILL.md names 1 domain. As links in the text: ui.perfetto.dev. This is read from the text; nothing was executed.

Is Generate Profile safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Generate Profile use?

Generate Profile is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Generate Profile use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Generate Profile?

Skills that share tags, products or a category with Generate Profile: Qwen Code E2E Testing (QwenLM/qwen-code, 28k stars), tmux Real User Testing (QwenLM/qwen-code, 28k stars), Wally Device E2E (RunanywhereAI/wally, 1.6k stars) and Wally E2E (RunanywhereAI/wally, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Generate Profile?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.