Search

LLM inference and serving

373 skills found, page 3.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97

Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes.

NVIDIA-NeMo/Nemotron2.1k—~2.4kAutomated safety check: PassApache-2.05 days ago
98
98.Add Vlm ModelOfficial

Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

intel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0yesterday
99

GGUF format and llama.cpp quantization for efficient CPU/GPU inference.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.6kAutomated safety check: PassMIT3 mo ago
100

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.

Orchestra-Research/AI-Research-SKILLs13k3 repos~1.5kAutomated safety check: PassMIT3 mo ago
101

A template skill demonstrating the PAI SKILL.md format. An agent skill from nirholas/PAI.

nirholas/PAI113—~919Automated safety check: PassUnknowntoday
102

QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.

qualcomm/qai-appbuilder247—~4.1kAutomated safety check: PassBSD-3-Clauseyesterday
103

Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

agentsope/SkillAlchemy436—~3kAutomated safety check: PassMIT2 days ago
104

Databricks Model Serving endpoint lifecycle and ops. An agent skill from databricks/databricks-agent-skills.

databricks/databricks-agent-skills3451 repo~3.4kAutomated safety check: PassUnknownyesterday
105

Find evidence-backed simplification candidates in vLLM-Omni and, when requested, turn them into focused proposals or code changes.

vllm-project/vllm-omni7.1k—~2.7kAutomated safety check: PassApache-2.0today
106

Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.

sohu-mptc/FlashRec107—~915Automated safety check: PassApache-2.0yesterday
107

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

vllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.06 mo ago
108

Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed.

scragnog/HOT-Step-CPP174—~4.9kAutomated safety check: PassMIT2 days ago
109

Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

amd/Quark182—~3kAutomated safety check: PassMIT13 days ago
110

Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

oracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.01 mo ago
111

Guide a tester through hipfire bring-up, serve smoke, claim-scoped harnesses, and benchmark reporting on AMD RDNA/CDNA GPUs.

warpfront/hipfire658—~1.5kAutomated safety check: PassUnknownyesterday
112

Bump Heron version via the VERSION-file SSOT. An agent skill from Netis/heron.

Netis/heron102—~983Automated safety check: PassApache-2.07 days ago
113

A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.4kAutomated safety check: PassApache-2.0yesterday
114
114.Review PROfficial

Review or prepare a pull request for the AutoRound repository — checks registration points for new data types/backends/VLMs, validates Chinese translation parity for modified markdown files…

intel/auto-round1.6k—~1.7kAutomated safety check: PassApache-2.0yesterday
115

Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.1kAutomated safety check: PassMIT3 mo ago
116
116.Gptq

Post-training 4-bit quantization for LLMs with minimal accuracy loss.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
117

Half-Quadratic Quantization for LLMs without calibration data.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
118

Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.2kAutomated safety check: PassMIT3 mo ago
119

Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace.

Orchestra-Research/AI-Research-SKILLs13k2 repos~3.7kAutomated safety check: PassMIT3 mo ago
120

High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.1kAutomated safety check: NotesMIT3 mo ago
121

Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.5kAutomated safety check: PassMIT3 mo ago
122

Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
123

Explains three ways to speed up LLM inference: draft-model speculative decoding, Medusa heads and lookahead decoding with Jacobi iteration, and when each one fits.

Orchestra-Research/AI-Research-SKILLs13k2 repos~3.5kAutomated safety check: PassMIT3 mo ago
124

Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs.

MetaX-MACA/vLLM-metax180—~2.3kAutomated safety check: PassApache-2.0yesterday
125

Summarize a document at a requested length. An agent skill from nirholas/PAI.

nirholas/PAI113—~992Automated safety check: PassUnknowntoday
126

This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms.

GoogleCloudPlatform/accelerated-platforms106—~1.2kAutomated safety check: PassApache-2.0yesterday
127

Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.

JakeATX/llamAmpere166—~5.6kAutomated safety check: PassMITyesterday
128

Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

azrtydxb/Fastllm-proxy108—~980Automated safety check: PassApache-2.0today
129

Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.

vllm-project/vllm-omni7.1k—~5.5kAutomated safety check: PassApache-2.0today
130

Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

amd/skills408—~5.7kAutomated safety check: NotesMIT2 days ago
131

Analyze and optimize vllm-rlt inference performance using reproducible unprofiled benchmarks, paired ops-only/full profiles, source-level attribution, and correctness checks.

ThinkFlowLab/vllm-rlt149—~4.6kAutomated safety check: PassApache-2.0today
132

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

vllm-project/vllm-skills102—~2kAutomated safety check: PassApache-2.06 mo ago
133

Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

Krusty84/triton-ascend-agent-dev-kit106—~597Automated safety check: PassApache-2.01 mo ago
134

Routes a NanoClaw agent group to a local Ollama model instead of the Anthropic API, using environment overrides and an optional block on Anthropic hosts.

nanocoai/nanoclaw31k—~1.5kAutomated safety check: PassMITyesterday
135

Sync documentation with project state using ICAV workflow. An agent skill from Netis/heron.

Netis/heron102—~1.4kAutomated safety check: PassApache-2.07 days ago
136

Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

joshzyj/open-scholar-skill168—~7.4kAutomated safety check: NotesUnknown23 days ago
137

Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~3.9kAutomated safety check: PassNo licence6 days ago
138

Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision.

MetaX-MACA/vLLM-metax180—~2.8kAutomated safety check: PassApache-2.0yesterday
139

Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting.

Orchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT3 mo ago
140

Generate and run tests for vllm-project/vllm-omni with CI-aligned levels and markers; wire new tests into Buildkite (test-ready.yml for L1/L2, test-merge.yml for L3, test-nightly.yml for L4).

vllm-project/vllm-omni7.1k—~17kAutomated safety check: PassApache-2.0today
141

Entry point for hosting a model on Amazon SageMaker: asks a few questions, picks a deployment pathway and hands off to the specialist skills.

huggingface/skills11k1 repo~2.1kAutomated safety check: PassApache-2.03 days ago
142

Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers.

amd/Quark182—~2.9kAutomated safety check: PassMIT13 days ago
143

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

vllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.06 mo ago
144

Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.

amd/skills408—~1.7kAutomated safety check: NotesMIT2 days ago