Agent skill

Engine Sglang

by autonomous-ai in autonomous-ai/openharness

Serve a Hugging Face model with SGLang on a Linux machine with an NVIDIA or AMD GPU, configured from the model's SGLang cookbook page — or, when it has none, from the model's own files — and join it…

MITAuto-check passedAI & LLM Engineering

Install Engine Sglang

skills CLI
$ npx skills add autonomous-ai/openharness --skill engine-sglang -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/openharness engine-sglang --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-sglang .claude/skills/engine-sglang && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
engine-sglang
GitHub stars
1.1k
Token cost
~1.5k tokens
SKILL.md length
672 words
Files
1
Skills in repo
100
Repo updated
First seen
Licence
MIT

At a glance

Serve a Hugging Face model with SGLang on a Linux machine with an NVIDIA or AMD GPU, configured from the model's SGLang cookbook page — or, when it has none, from the model's own files — and join it…

  • Works in 2 steps: The cookbook page — "$GRID_FLEET" recipe… → No cookbook page — read the model:…
  • Tasks that involve Model hubs and datasets
  • SKILL.md covers When to use it, 1. The cookbook page —…, 2. No cookbook page — read the… and Install and version, plus 4 more sections
  • Calls uv, python3 and docker; reaches github.com and raw.githubusercontent.com

What it does

Engine Sglang is an agent skill from autonomous-ai/openharness. Serve a Hugging Face model with SGLang on a Linux machine with an NVIDIA or AMD GPU, configured from the model's SGLang cookbook page — or, when it has none, from the model's own files — and join it to the person's fleet. Load before installing, configuring, starting or stopping SGLang.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets. It works with SGLang, Linux, NVIDIA AI Platform and Hugging Face. The repository describes itself as: The ultimate harness for coding agents and beyond. All your agents. All your machines. One command center. Start with code, then follow your curiosity and build across… The licence is MIT.

When your agent uses it

  • Tasks that involve Model hubs and datasets

Example prompts

  • “s SGLang cookbook page — or, when it has none, from the model”
  • “/engine-sglang”

Requirements

  • Python 3
  • Docker

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. The cookbook page — "$GRID_FLEET" recipe sglang ORG/NAME
  2. No cookbook page — read the model: "$GRID_FLEET" model-facts ORG/NAME

What it can do on your machine

Read from SKILL.md and the folder at commit 50da5db. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • python3
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • raw.githubusercontent.com

    Also links to:

    • docs.sglang.io
    • sglang.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Engine Sglang loads about 1.5k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 672 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/openharness at commit 50da5db, republished under its MIT licence (© autonomous-ai). 672 words, ~1,511 tokens.

Download SKILL.mdSave it as .claude/skills/engine-sglang/SKILL.md (or your agent's skills folder).
name
engine-sglang
description
Serve a Hugging Face model with SGLang on a Linux machine with an NVIDIA or AMD GPU, configured from the model's SGLang cookbook page — or, when it has none, from the model's own files — and join it to the person's fleet. Load before installing, configuring, starting or stopping SGLang.

SGLang (Linux + NVIDIA CUDA or AMD ROCm)

Official sources, read 2026-09-29. Agent index: docs.sglang.io/llms.txt (every page is available as .md). Cookbook · Models · Quickstart · Server arguments · Tool parser · Separate reasoning · Native API · Transformers fallback · source: github.com/sgl-project/sglang Tested: not runnable on macOS (below); no GPU run yet. Tags: [doc] official source, [run] seen on a machine, [?] unverified.

When to use it

  • Linux with an NVIDIA GPU of compute capability 8.0 or newer (A10, A100, L4, L40S, H100), Python ≥ 3.10 [doc]; AMD GPUs, Intel Xeon CPUs and TPUs have their own guides [doc]. Hugging Face models, many requests at once, several GPUs.
  • Never on macOS. The current release ships Linux wheels only and requires CUDA builds of its kernels; on a Mac the installer falls back to an old release whose server does not import [run].
  • A GGUF file for one person: Grid's own engine serves it with nothing to install.

1. The cookbook page — "$GRID_FLEET" recipe sglang ORG/NAME

Finds the model family's page through llms.txt (the most specific family name the model name starts with) and prints: source (the page), installation (its version requirement — some families need a release or the main branch — and install commands), serveCommands (the page's own launch commands per GPU type), toolCallParsers, reasoningParsers, and configurationTips. Read the page itself for anything else. Exit code 3 means no page: go to step 2.

2. No cookbook page — read the model: "$GRID_FLEET" model-facts ORG/NAME

  1. support.sglang.listed: true means SGLang's model registry has the architecture; null means this check cannot tell. SGLang can run most decoder models through --model-impl transformers [doc] — a test, not a promise.
  2. modelCard.serveCommands and parsers: the model authors' own SGLang command. Use it when present.
  3. Otherwise --tool-call-parser auto and --reasoning-parser auto: SGLang detects both from the model's chat template [doc]. Then the tool-call acceptance check decides; if it fails, say so and stop — do not cycle through parser names.
  4. contextLength must be ≥ 65536.

Install and version

  • Slow step, ask first. uv pip install --prerelease=allow sglang [doc]; a family whose cookbook page asks for the main branch: uv pip install --prerelease=allow 'git+https://github.com/sgl-project/sglang.git#subdirectory=python' [doc]. OSError: CUDA_HOME environment variable is not set → export CUDA_HOME=/usr/local/cuda-<version> [doc].
  • Docker: lmsysorg/sglang:latest (NVIDIA), lmsysorg/sglang-rocm:<tag> (AMD, tag per GPU generation on the cookbook page) [doc], run with --gpus all --shm-size 32g --ipc=host -p 127.0.0.1:P:30000 -v ~/.cache/huggingface:/root/.cache/huggingface [doc].
Show full SKILL.md (268 more words)Show less

Common configs (the cookbook's flags win)

Each adds the parsers from step 1 or 2, --host 127.0.0.1 --port P and --context-length of at least 65536 (the default is the model's own maximum [doc]).

CaseAdd
One GPU, one agent--max-running-requests 4
One GPU, many peopleleave --max-running-requests to SGLang
Several GPUs, one machine--tp-size <GPU count> [doc]
Out of memorylower --mem-fraction-static (weights plus KV pool share) [doc]; --kv-cache-dtype fp8_e4m3 [doc]

Start, ready, join

"$GRID_FLEET" serve sglang-P --env HF_HUB_OFFLINE=1 -- ~/.grid/envs/sglang/bin/sglang serve \
  --model-path <model id or snapshot dir> --served-model-name <id> --host 127.0.0.1 --port P <flags>

(python3 -m sglang.launch_server takes the same arguments [doc].) Defaults are 127.0.0.1 and port 30000 [doc]. Install into ~/.grid/envs/sglang so fleet models finds it; fleet serve keeps it alive after your shell returns (log run/sglang-P.log).

  • Verify: "$GRID_FLEET" verify --at http://127.0.0.1:P/v1 --model <id> --kind sglang (bounded, narrated).
  • Ready: the log says The server is fired up and ready to roll! [doc]; GET /health → 200; GET /health_generate generates one token [doc]; GET /v1/models lists <id>; one bounded answer; one tool call.
  • Join: "$GRID_FLEET" run -- join GRID --at http://127.0.0.1:P/v1 -m <id> --advertise-as ALIAS. Grid's detector has no SGLang probe and would mislabel it on 8000 or 8080 [run], so always --at, always with /v1.
  • Thinking off per request: "chat_template_kwargs": {"enable_thinking": false} where the cookbook page shows it [doc].

Stop

"$GRID_FLEET" stop sglang-P (or docker stop your container); confirm the port is free.

Known failures → what to do

SignDo
out of memory while servinglower --mem-fraction-static [doc]
no tool_callsthe cookbook's --tool-call-parser, else auto [doc]; still none → report
a CUDA crash, a multi-GPU hang, a production incidentread SGLang's own maintainer skills: .agents/skills/debug-cuda-crash, debug-distributed-hang, sglang-prod-incident-triage (SKILL.md in each) at https://raw.githubusercontent.com/sgl-project/sglang/main/
404 on /modelsthe --at URL lacks /v1

© autonomous-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in store/agents/autonomous-grid/skills/engine-sglang of autonomous-ai/openharness.

Open the folder on GitHubat commit 50da5db

Compare with similar skills

Engine Sglang next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Engine Sglang compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Engine Sglang this skillautonomous-ai/openharness1.1k—~1.5kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Megatron Memory Estimatoryzlnew/infra-skills149—~2.2kAutomated safety check: PassNone
Convergence TestAMD-AGI/Primus131—~2.1kAutomated safety check: PassCustom licence
Kermt Continue PretrainNVIDIA/skills3.5k1 repos~4.1kAutomated safety check: PassApache-2.0
Kermt EmbedNVIDIA/skills3.5k1 repos~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Megatron Memory Estimator

    yzlnew/infra-skills

    Estimate GPU memory usage for Megatron-based MoE (Mixture of Experts) and dense models.

    149 GitHub stars~2.2k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Convergence Test

    AMD-AGI/Primus

    Run, monitor, stop and report Primus convergence tests -- training a model on a real corpus and checking that the loss curve is healthy -- from a plain-language request such as "run convergence test…

    131 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Continue KERMT pretraining on a custom SMILES corpus with a groverbase, cmim, or hybrid checkpoint.

    3.5k GitHub starsUsed in 1 repo~4.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Kermt Embed

    NVIDIA/skills

    Official

    Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint.

    3.5k GitHub starsUsed in 1 repo~1.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches.

    3.5k GitHub stars~4.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes

More from autonomous-ai/openharness

All 100 skills in this repo
  • G-code Slicer Tool

    autonomous-ai/openharness

    Slices 3D mesh files into printer-profiled plain G-code through real slicer CLIs, with backend discovery, input inspection, dry runs and static validation.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Home Assistant Automation Builder

    autonomous-ai/openharness

    Turns a home-automation request into standard, testable automations.yaml, run against Home Assistant Core's real triggers and verified with its own trace tool.

    1.1k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Score Music Composer

    autonomous-ai/openharness

    Turns a musical brief into LilyPond concert-pitch music, checked parts for each instrument and a playable practice pack.

    1.1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • OrcaSlicer 3MF and G-code Workflow

    autonomous-ai/openharness

    Turns an STL and explicit printer and material requirements into compared OrcaSlicer plans, an editable 3MF project, checked G-code and a portable handoff.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Sheets and Docs Report Builder

    autonomous-ai/openharness

    Builds an editable DOCX report, a formula-driven XLSX workbook and a fresh LibreOffice PDF preview from one structured source file, then checks them together.

    1.1k GitHub stars~708 tokensUpdated today
    Auto-check passed
  • Bambu Labs

    autonomous-ai/openharness

    Dry-run, upload, and cautiously initiate local Bambu Lab print jobs from validated plain .gcode, using Bambu LAN FTPS/MQTT handoffs.

    1.1k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check: warnings

Questions about Engine Sglang

What does Engine Sglang do?

Serve a Hugging Face model with SGLang on a Linux machine with an NVIDIA or AMD GPU, configured from the model's SGLang cookbook page — or, when it has none, from the model's own files — and join it…. Engine Sglang is an agent skill from autonomous-ai/openharness. Serve a Hugging Face model with SGLang on a Linux machine with an NVIDIA or AMD GPU, configured from the model's SGLang cookbook page — or, when it has none, from the model's own files — and join it to the person's fleet.

When should I use Engine Sglang?

Engine Sglang fits situations like: tasks that involve Model hubs and datasets.

How do I install Engine Sglang in Claude Code?

Run `npx skills add autonomous-ai/openharness --skill engine-sglang -a claude-code`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-sglang in autonomous-ai/openharness) into .claude/skills/engine-sglang in your project. Claude Code loads it when a task matches its description.

How do I install Engine Sglang in Codex?

Run `npx skills add autonomous-ai/openharness --skill engine-sglang -a codex`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-sglang in autonomous-ai/openharness) into .agents/skills/engine-sglang in your project. Codex loads it when a task matches its description.

Can I use Engine Sglang in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/openharness --skill engine-sglang -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/engine-sglang, .gemini/skills/engine-sglang, .github/skills/engine-sglang and .opencode/skills/engine-sglang in your project.

What does Engine Sglang need to run?

Going by SKILL.md and its folder, Engine Sglang needs the command-line tools its instructions call (uv, python3 and docker). Our summary lists: Python 3; Docker.

Does Engine Sglang access the network?

SKILL.md names 4 domains. In commands or code: github.com and raw.githubusercontent.com; the agent is likely to contact these when it follows the instructions. As links in the text: docs.sglang.io and sglang.io. This is read from the text; nothing was executed.

Is Engine Sglang safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Engine Sglang use?

Engine Sglang is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Engine Sglang use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Engine Sglang?

Skills that share tags, products or a category with Engine Sglang: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Megatron Memory Estimator (yzlnew/infra-skills, 149 stars), Convergence Test (AMD-AGI/Primus, 131 stars) and Kermt Continue Pretrain (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Engine Sglang?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/openharness, which has 1,149 GitHub stars. The repository holds 100 skills in this directory. The repository was last updated on October 8, 2026.

Source: autonomous-ai/openharness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.