Topic · AI & LLM Engineering

Best LLM inference and serving skills, page 6

Skills #241–288 of 372, ranked by score.

LLM inference and serving skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM inference and serving skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
241

Delegate coding to OpenCode CLI (multi-LLM open-source coding agent)

taracodlabs/aiden851—~917Automated safety check: NotesApache-2.025 days ago
242

Extend, test, debug, or integrate the Model Download microservice codebase.

open-edge-platform/edge-ai-libraries169—~2.8kAutomated safety check: PassApache-2.0today
243

Download and convert AI models using the Model Download microservice.

open-edge-platform/edge-ai-libraries169—~3.8kAutomated safety check: PassApache-2.0today
244

Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP.

scragnog/HOT-Step-CPP173—~6.6kAutomated safety check: PassMIT2 days ago
245

End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output.

amd/Quark181—~4.5kAutomated safety check: PassMIT11 days ago
246

Automatically fetch InferenceX benchmark data and generate daily performance reports for LLM inference on various hardware (NVIDIA, AMD, etc.).

ascend-ai-coding/awesome-ascend-skills174—~1.1kAutomated safety check: PassNo licencetoday
247

[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the…

rlaope/oh-my-hermes3.2k—~2.2kAutomated safety check: PassMITtoday
248
248.Jetson LLM ServeOfficial

Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

NVIDIA/skills3.5k1 repo~3kAutomated safety check: NotesApache-2.0today
249

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

NVIDIA/skills3.5k1 repo~1.2kAutomated safety check: PassApache-2.0today
250

Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.

open-thoughts/OpenThoughts-Agent301—~947Automated safety check: PassApache-2.010 days ago
251

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

open-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.010 days ago
252

Deeper backend map of overbae — module layout, celery queue topology, the span-only tracing model and its API surface, capabilities and toml sync, behaviour-keyed scoring, auth and guests, model…

overmind-core/overmind603—~10kAutomated safety check: PassAGPL-3.0today
253

Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills.

databricks/databricks-agent-skills345—~4.6kAutomated safety check: PassUnknowntoday
254

Interactively guide setup for connecting the Claude Code CLI to a local Ollama LLM on Mac.

receptron/mulmoclaude368—~1.9kAutomated safety check: NotesMITtoday
255

Canonical tempo-sync/rate/quantize divisions from src/audio/MusicTime.h (4 bars to 1/32 with dotted and triplets, plus 1/64).

n1m21n/Infinite268—~1.5kAutomated safety check: PassUnknowntoday
256
256.Deepstream DevOfficial

NVIDIA DeepStream SDK development with Python pyservicemaker API.

NVIDIA/skills3.5k—~3.3kAutomated safety check: PassApache-2.0today
257
257.Deepstream SopOfficial

A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

NVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0today
258

Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines.

NVIDIA/skills3.5k—~1.5kAutomated safety check: PassApache-2.0today
259

A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill.

NVIDIA/skills3.5k—~2.1kAutomated safety check: PassApache-2.0today
260
260.Tao Finetune ClipOfficial

CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

NVIDIA/skills3.5k—~4kAutomated safety check: NotesApache-2.0today
261

InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment.

NVIDIA/skills3.5k—~3.5kAutomated safety check: NotesApache-2.0today
262

The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action.

NVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0today
263

PyTorch-based TAO image classification. An agent skill from NVIDIA/skills.

NVIDIA/skills3.5k—~3.6kAutomated safety check: NotesApache-2.0today
264
264.Tao Train RtdetrOfficial

RT-DETR (Real-Time DEtection TRansformer) for 2D object detection.

NVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0today
265

Run coding tasks via opencode using free cloud models. An agent skill from luongnv89/skills.

luongnv89/skills131—~4.5kAutomated safety check: NotesMITtoday
266

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

agentsope/SkillAlchemy466—~6.1kAutomated safety check: PassMITtoday
267

Project-kickoff rubric for the self-host vs managed-cloud decision — when is running your own inference engine / LLM platform worth the ops cost vs paying per-token for a managed API?

agentsope/SkillAlchemy466—~6.3kAutomated safety check: PassMITtoday
268

Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy466—~6.1kAutomated safety check: PassMITtoday
269

Converts one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model through llama.cpp, automatically preferring an available Vulkan GPU and falling back…

godot-fun/gai182—~813Automated safety check: PassMITtoday
270

Redact, anonymize, sanitize, or remove PII locally with Distil-PII and llama.cpp; keep personal data and secret values out of model context, logs, and chat.

HybridAIOne/hybridclaw158—~1kAutomated safety check: PassMITtoday
271

Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama.

autonomous-ai/openharness1.2k—~2.2kAutomated safety check: NotesMITtoday
272

Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…

autonomous-ai/openharness1.2k—~1.9kAutomated safety check: PassMITtoday
273

A Principal ML Engineer interviewer that simulates a FAANG-style ML system design interview covering the full lifecycle from data to production.

PrepLabsAI/InterviewMentor112—~4.2kAutomated safety check: PassMIT2 days ago
274

Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama.

glebis/claude-skills390—~1.4kAutomated safety check: PassMITyesterday
275

Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.

open-thoughts/OpenThoughts-Agent301—~5.8kAutomated safety check: PassApache-2.010 days ago
276
276.Aks GPU InferenceOfficial

Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence.

microsoft/GitHub-Copilot-for-Azure255—~764Automated safety check: PassMITtoday
277

Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.

amd/Quark181—~4.8kAutomated safety check: PassMIT11 days ago
278

Validate Quark ONNX quantization output using four lightweight checks: auxiliary file copy alignment, expected non-quantized initializer MD5 byte-identity (inline rawdata + external-data byte…

amd/Quark181—~2.5kAutomated safety check: PassMIT11 days ago
279

Route Quark ONNX user goals to the correct atomic skill. An agent skill from amd/Quark.

amd/Quark181—~2.7kAutomated safety check: PassMIT11 days ago
280

Apply existing ShapeShifter graph passes to an .onnx model via the quark-cli shapeshifter CLI or a ShapeShifter YAML.

amd/Quark181—~1.9kAutomated safety check: PassMIT11 days ago
281

Detect upstream Quark ONNX changes that affect the ONNX skill family and classify required updates.

amd/Quark181—~3.3kAutomated safety check: PassMIT11 days ago
282

Partition an ONNX model graph into named functional subgraphs and emit a subgraphpartition.json file.

amd/Quark181—~2.7kAutomated safety check: PassMIT11 days ago
283

Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.

amd/Quark181—~1.9kAutomated safety check: NotesMIT11 days ago
284

Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run.

amd/Quark181—~1.5kAutomated safety check: PassMIT11 days ago
285

Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole.

amd/Quark181—~2.3kAutomated safety check: PassMIT11 days ago
286

L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate.

amd/Quark181—~2.6kAutomated safety check: PassMIT11 days ago
287

Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output.

amd/Quark181—~2.8kAutomated safety check: PassMIT11 days ago
288

Inspect a target model and prepare metadata for Quark PTQ planning.

amd/Quark181—~2kAutomated safety check: PassMIT11 days ago