Search

LLM inference and serving

372 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

sgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0today
2

Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

huggingface/skills11k1 repo~4.6kAutomated safety check: PassApache-2.0yesterday
3

Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

amElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT2 days ago
4

Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

huggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0yesterday
5

Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

vllm-project/vllm-omni7.1k—~7.5kAutomated safety check: PassApache-2.0today
6

Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

vllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0today
7

Autonomous performance optimization: research, PoC, benchmark, implement, review, PR

raullenchai/Rapid-MLX3.9k—~1.6kAutomated safety check: NotesUnknowntoday
8

Bump or upgrade the pinned versions of Helmor's bundled agent CLIs, SDKs, and supporting binaries — Claude Code + claude-agent-sdk (lockstep), Codex, Cursor SDK, OpenCode, Kimi, Pi, and gh / glab /…

dohooo/helmor1.3k—~2.1kAutomated safety check: PassApache-2.01 mo ago
9

Run the pre-release end-to-end sweep of every user-facing surface — the 33 samples under samples/ (booted from their packaged artifacts and driven in a real browser via chrome-devtools MCP), the…

Atmosphere/atmosphere3.8k—~4.2kAutomated safety check: PassApache-2.0today
10

Guides building on the 0G Compute Network, a decentralized GPU marketplace for AI inference and fine-tuning, with SDK patterns and CLI commands.

internet-court/internet-court-skill6.5k1 repo~1.9kAutomated safety check: PassUnknown1 mo ago
11

Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

R6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT3 mo ago
12

Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

guqiong96/Lvllm4652 repos~349Automated safety check: PassApache-2.018 days ago
13

Create a new skill in the current repository. An agent skill from Hyk260/PureChat.

Hyk260/PureChat5461 repo~823Automated safety check: PassMIT23 days ago
14

Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

BBuf/AI-Infra-Auto-Driven-SKILLS925—~2.3kAutomated safety check: PassNo licence4 days ago
15

Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

huggingface/skills11k1 repo~6.9kAutomated safety check: PassApache-2.0yesterday
16

Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

vllm-project/vllm-omni7.1k—~1.4kAutomated safety check: PassApache-2.0today
17

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0today
18

Benchmark and compare LLM models for SUB/WAVE's on-air calls — track picks, segments, listener requests, DJ scripts, banter, and programme beats — in both candidate-pool and agent modes, using…

perminder-klair/subwave1.4k—~2.4kAutomated safety check: NotesMITyesterday
19

Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

intel/auto-round1.6k—~2.8kAutomated safety check: PassApache-2.0today
20

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

guqiong96/Lvllm4652 repos~1.5kAutomated safety check: PassApache-2.018 days ago
21
21.Debug InferenceOfficial

Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

NVIDIA/OpenShell16k—~1.9kAutomated safety check: PassApache-2.0today
22

Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

BBuf/AI-Infra-Auto-Driven-SKILLS925—~2.5kAutomated safety check: PassNo licence4 days ago
23

Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

vllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0today
24

Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.

vllm-project/vllm-omni7.1k—~3.8kAutomated safety check: PassApache-2.0today
25

Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

Blackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMITyesterday
26

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.011 days ago
27

A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

Mesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0today
28
28.Adapt New LLMOfficial

Adapt AutoRound to support a new LLM architecture that doesn't work out-of-the-box.

intel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0today
29

驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

OpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.01 mo ago
30

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

huggingface/skills11k3 repos~945Automated safety check: PassApache-2.0yesterday
31

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNo licence4 days ago
32

Processes sensitive local documents through PII Guard and a local Ollama model into a reversible redacted copy, without letting the main agent read the original or restored contents.

danyuchn/pii-guard249—~2.7kAutomated safety check: PassMIT7 days ago
33

Inspects and tunes the shared-vs-dedicated memory split on AMD Ryzen APUs with unified memory (UMA) so larger LLMs and image-gen models fit on the iGPU, or so reserved GPU memory is returned to the…

amd/skills406—~2.6kAutomated safety check: PassMITtoday
34

A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform…

warpfront/hipfire655—~2.6kAutomated safety check: PassUnknowntoday
35

Use iDeer as a daily paper-reading workflow for chatbot-first users such as Codex, Gemini, or ChatGPT.

AI45Lab/iDeer416—~3kAutomated safety check: NotesAGPL-3.02 mo ago
36

Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

ModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassUnknownyesterday
37

Run wally's LLM e2e on Apple Neural Engine (NeuRT) and Snapdragon Hexagon NPU (QHexRT) devices.

RunanywhereAI/wally1.6k—~845Automated safety check: PassMITtoday
38

Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).

intel/auto-round1.6k—~1.9kAutomated safety check: PassApache-2.0today
39

Adapt and port new LLM model architectures to this xinfer project.

guoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT1 mo ago
40

Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner).

tryonlabs/opentryon551—~1.1kAutomated safety check: PassUnknownyesterday
41

Check if package.json files are in sync with pnpm-lock.yaml.

Hyk260/PureChat546—~594Automated safety check: PassMIT23 days ago
42

Query images with a local Ollama vision model without loading the image into the main agent context.

gridaco/grida2.7k—~1.5kAutomated safety check: PassApache-2.0yesterday
43

Build, extend, and debug AI agents in Egroo using the Microsoft Agent Framework (C .NET).

jihadkhawaja/Egroo178—~1.9kAutomated safety check: PassApache-2.06 mo ago
44
44.Bank

Answer questions about the user's own bank accounts and money using the bank MCP tools (listaccounts, getbalances, gettransactions, watches).

noskillish/bankmcp277—~1.1kAutomated safety check: PassMIT11 days ago
45

Creates and validates dubbo-go-pixiu LLM gateway conf.yaml. An agent skill from apache/dubbo-go-pixiu.

apache/dubbo-go-pixiu568—~2.5kAutomated safety check: PassApache-2.06 days ago
46

Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

vllm-project/vllm-omni7.1k—~7kAutomated safety check: PassApache-2.0today
47

Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

EGalahad/sim2real145—~1.1kAutomated safety check: PassNo licence11 days ago
48

A skill your agent uses when a user asks how many tokens (or how much context/cost) ContextPilot has saved, or wants a ContextPilot savings status/summary inside Hermes Agent — e.g.

EfficientContext/ContextPilot140—~1.4kAutomated safety check: PassMITyesterday