Search

vLLM

173 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Produces ranked, evidence-grounded next steps when an ML experiment has stalled, drawing on a Leeroopedia knowledge base or on fetched docs and issues.

Leeroo-AI/superml195—~4.8kAutomated safety check: PassApache-2.06 mo ago
50

Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

amd/skills408—~4kAutomated safety check: NotesMITyesterday
51
51.Aqua DeploymentOfficial

Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.

oracle/accelerated-data-science125—~2.4kAutomated safety check: PassUPL-1.01 mo ago
52

Build vLLM (CPU) end-to-end with the oneDNN + ZenDNN (zen64) backend via direct integration (NO zentorch).

amd/ZenDNN158—~5.2kAutomated safety check: PassUnknown4 days ago
53

Review PRs and local changes for hsliuustc0106/vllm-rlt: Ouro engine and KV correctness, serving behavior, and BF16 accuracy/speed evidence.

ThinkFlowLab/vllm-rlt149—~1.1kAutomated safety check: PassApache-2.0today
54

Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

agentsope/SkillAlchemy436—~3kAutomated safety check: PassMIT2 days ago
55

Find evidence-backed simplification candidates in vLLM-Omni and, when requested, turn them into focused proposals or code changes.

vllm-project/vllm-omni7.1k—~2.7kAutomated safety check: PassApache-2.0today
56

Run RecIF beam×concurrency eval and FlashRec vs SGLang/vLLM/TRT-LLM baselines.

sohu-mptc/FlashRec107—~915Automated safety check: PassApache-2.0yesterday
57

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

vllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.06 mo ago
58

Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.

Leeroo-AI/superml195—~3.8kAutomated safety check: PassApache-2.06 mo ago
59

Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

amd/Quark182—~3kAutomated safety check: PassMIT13 days ago
60

verl-omni commit message + PR conventions and the mandatory contribution policy.

verl-project/verl-omni1.2k—~1.4kAutomated safety check: PassApache-2.0yesterday
61

Register, list, get, and manage LLM models in OCI AI Quick Actions (AQUA) using the ADS SDK.

oracle/accelerated-data-science125—~1.4kAutomated safety check: PassUPL-1.01 mo ago
62

Bump Heron version via the VERSION-file SSOT. An agent skill from Netis/heron.

Netis/heron102—~983Automated safety check: PassApache-2.06 days ago
63

A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.4kAutomated safety check: PassApache-2.0yesterday
64

Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.1kAutomated safety check: PassMIT3 mo ago
65

Half-Quadratic Quantization for LLMs without calibration data.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
66

High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.1kAutomated safety check: NotesMIT3 mo ago
67

Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
68

Guides reinforcement-learning research with torchforge, Meta's PyTorch-native library that keeps RL algorithms apart from infrastructure, including GRPO math-reasoning runs.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.5kAutomated safety check: PassMIT3 mo ago
69

Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT3 mo ago
70

Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs.

MetaX-MACA/vLLM-metax180—~2.3kAutomated safety check: PassApache-2.0yesterday
71

This is an experimental Skill. An agent skill from GoogleCloudPlatform/accelerated-platforms.

GoogleCloudPlatform/accelerated-platforms106—~1.2kAutomated safety check: PassApache-2.0yesterday
72

Run and troubleshoot the inference backends on the DGX Spark pair that FastLLM proxies to — starting or stopping models with vLLM, SGLang or sparkrun, choosing memory and speculative-decoding…

azrtydxb/Fastllm-proxy108—~980Automated safety check: PassApache-2.0today
73

Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works.

vllm-project/vllm-omni7.1k—~5.5kAutomated safety check: PassApache-2.0today
74

Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

amd/skills408—~5.7kAutomated safety check: NotesMITyesterday
75

Analyze and optimize vllm-rlt inference performance using reproducible unprofiled benchmarks, paired ops-only/full profiles, source-level attribution, and correctness checks.

ThinkFlowLab/vllm-rlt149—~4.6kAutomated safety check: PassApache-2.0today
76

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

vllm-project/vllm-skills102—~2kAutomated safety check: PassApache-2.06 mo ago
77

Sync documentation with project state using ICAV workflow. An agent skill from Netis/heron.

Netis/heron102—~1.4kAutomated safety check: PassApache-2.06 days ago
78

Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~3.9kAutomated safety check: PassNo licence6 days ago
79

Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision.

MetaX-MACA/vLLM-metax180—~2.8kAutomated safety check: PassApache-2.0yesterday
80

Generate and run tests for vllm-project/vllm-omni with CI-aligned levels and markers; wire new tests into Buildkite (test-ready.yml for L1/L2, test-merge.yml for L3, test-nightly.yml for L4).

vllm-project/vllm-omni7.1k—~17kAutomated safety check: PassApache-2.0today
81

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

vllm-project/vllm-skills102—~1.6kAutomated safety check: PassApache-2.06 mo ago
82

Autonomously optimizes end-to-end LLM inference throughput on AMD Instinct GPUs and reports a validated gain, using the Hyperloom multi-agent optimizer.

amd/skills408—~1.7kAutomated safety check: NotesMITyesterday
83

Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.3kAutomated safety check: PassMIT3 mo ago
84

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~7.5kAutomated safety check: PassNo licence6 days ago
85

Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades.

MetaX-MACA/vLLM-metax180—~1.1kAutomated safety check: PassApache-2.0yesterday
86

Deploys and manages VSS through setup.sh and its Docker Compose overlays.

open-edge-platform/edge-ai-libraries171—~4.1kAutomated safety check: PassApache-2.0yesterday
87

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

vllm-project/vllm-skills102—~1.4kAutomated safety check: PassApache-2.06 mo ago
88

Optimize Ollama configuration for the current machine's hardware.

luongnv89/skills131—~4.1kAutomated safety check: NotesMITtoday
89

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMITyesterday
90

A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~1.5kAutomated safety check: PassNo licence6 days ago
91

Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL.

areal-project/AReaL5.8k—~6kAutomated safety check: PassApache-2.0yesterday
92

A skill your agent uses whenever a developer needs to deploy VSS to Kubernetes, helm install VSS, configure values.yaml for VSS, or run VSS on k8s with GPU/vLLM for the…

open-edge-platform/edge-ai-libraries171—~3.8kAutomated safety check: PassApache-2.0yesterday
93
93.Aqua MetricsOfficial

Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.

oracle/accelerated-data-science125—~1.5kAutomated safety check: PassUPL-1.01 mo ago
94

DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

open-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.012 days ago
95

Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval.

henryalouf/ruflow157—~1.6kAutomated safety check: PassMIT4 mo ago
96

Trim large MetaX model directories for dummy smoke tests or real-checkpoint loading on limited GPUs.

MetaX-MACA/vLLM-metax180—~2.4kAutomated safety check: PassApache-2.0yesterday