Topic · AI & LLM Engineering
Best LLM inference and serving skills, page 6
LLM inference and serving skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 241 | 241.Opencode Delegate coding to OpenCode CLI (multi-LLM open-source coding agent) | taracodlabs/ | 851 | — | ~917 | Automated safety check: Notes | Apache-2.0 | 25 days ago |
| 242 | Extend, test, debug, or integrate the Model Download microservice codebase. | open-edge-platform/ | 169 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 243 | Download and convert AI models using the Model Download microservice. | open-edge-platform/ | 169 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 244 | 244.Model Management Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP. | scragnog/ | 173 | — | ~6.6k | Automated safety check: Pass | MIT | 2 days ago |
| 245 | End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output. | amd/ | 181 | — | ~4.5k | Automated safety check: Pass | MIT | 11 days ago |
| 246 | Automatically fetch InferenceX benchmark data and generate daily performance reports for LLM inference on various hardware (NVIDIA, AMD, etc.). | ascend-ai-coding/ | 174 | — | ~1.1k | Automated safety check: Pass | No licence | today |
| 247 | [omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the… | rlaope/ | 3.2k | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 248 | Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin. | NVIDIA/ | 3.5k | 1 repo | ~3k | Automated safety check: Notes | Apache-2.0 | today |
| 249 | Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck. | NVIDIA/ | 3.5k | 1 repo | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 250 | Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes. | open-thoughts/ | 301 | — | ~947 | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 251 | Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy… | open-thoughts/ | 301 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 252 | Deeper backend map of overbae — module layout, celery queue topology, the span-only tracing model and its API surface, capabilities and toml sync, behaviour-keyed scoring, auth and guests, model… | overmind-core/ | 603 | — | ~10k | Automated safety check: Pass | AGPL-3.0 | today |
| 253 | Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills. | databricks/ | 345 | — | ~4.6k | Automated safety check: Pass | Unknown | today |
| 254 | Interactively guide setup for connecting the Claude Code CLI to a local Ollama LLM on Mac. | receptron/ | 368 | — | ~1.9k | Automated safety check: Notes | MIT | today |
| 255 | Canonical tempo-sync/rate/quantize divisions from src/audio/MusicTime.h (4 bars to 1/32 with dotted and triplets, plus 1/64). | n1m21n/ | 268 | — | ~1.5k | Automated safety check: Pass | Unknown | today |
| 256 | NVIDIA DeepStream SDK development with Python pyservicemaker API. | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 257 | A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether… | NVIDIA/ | 3.5k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | today |
| 258 | Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines. | NVIDIA/ | 3.5k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 259 | A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. | NVIDIA/ | 3.5k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 260 | CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment. | NVIDIA/ | 3.5k | — | ~4k | Automated safety check: Notes | Apache-2.0 | today |
| 261 | InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment. | NVIDIA/ | 3.5k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | today |
| 262 | The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action. | NVIDIA/ | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | today |
| 263 | PyTorch-based TAO image classification. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | today |
| 264 | RT-DETR (Real-Time DEtection TRansformer) for 2D object detection. | NVIDIA/ | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | today |
| 265 | 265.Opencode Runner Run coding tasks via opencode using free cloud models. An agent skill from luongnv89/skills. | luongnv89/ | 131 | — | ~4.5k | Automated safety check: Notes | MIT | today |
| 266 | Cross-engine decision rubric for self-hosting or recommending an LLM serving stack. | agentsope/ | 466 | — | ~6.1k | Automated safety check: Pass | MIT | today |
| 267 | Project-kickoff rubric for the self-host vs managed-cloud decision — when is running your own inference engine / LLM platform worth the ops cost vs paying per-token for a managed API? | agentsope/ | 466 | — | ~6.3k | Automated safety check: Pass | MIT | today |
| 268 | 268.Agentsop Vllm Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy. | agentsope/ | 466 | — | ~6.1k | Automated safety check: Pass | MIT | today |
| 269 | 269.Image To Text Converts one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model through llama.cpp, automatically preferring an available Vulkan GPU and falling back… | godot-fun/ | 182 | — | ~813 | Automated safety check: Pass | MIT | today |
| 270 | Redact, anonymize, sanitize, or remove PII locally with Distil-PII and llama.cpp; keep personal data and secret values out of model context, logs, and chat. | HybridAIOne/ | 158 | — | ~1k | Automated safety check: Pass | MIT | today |
| 271 | 271.Engine Ollama Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama. | autonomous-ai/ | 1.2k | — | ~2.2k | Automated safety check: Notes | MIT | today |
| 272 | 272.Engine Vllm Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it… | autonomous-ai/ | 1.2k | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 273 | A Principal ML Engineer interviewer that simulates a FAANG-style ML system design interview covering the full lifecycle from data to production. | PrepLabsAI/ | 112 | — | ~4.2k | Automated safety check: Pass | MIT | 2 days ago |
| 274 | 274.Local Models Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama. | glebis/ | 390 | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 275 | Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent. | open-thoughts/ | 301 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 276 | Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence. | microsoft/ | 255 | — | ~764 | Automated safety check: Pass | MIT | today |
| 277 | Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent. | amd/ | 181 | — | ~4.8k | Automated safety check: Pass | MIT | 11 days ago |
| 278 | Validate Quark ONNX quantization output using four lightweight checks: auxiliary file copy alignment, expected non-quantized initializer MD5 byte-identity (inline rawdata + external-data byte… | amd/ | 181 | — | ~2.5k | Automated safety check: Pass | MIT | 11 days ago |
| 279 | Route Quark ONNX user goals to the correct atomic skill. An agent skill from amd/Quark. | amd/ | 181 | — | ~2.7k | Automated safety check: Pass | MIT | 11 days ago |
| 280 | Apply existing ShapeShifter graph passes to an .onnx model via the quark-cli shapeshifter CLI or a ShapeShifter YAML. | amd/ | 181 | — | ~1.9k | Automated safety check: Pass | MIT | 11 days ago |
| 281 | Detect upstream Quark ONNX changes that affect the ONNX skill family and classify required updates. | amd/ | 181 | — | ~3.3k | Automated safety check: Pass | MIT | 11 days ago |
| 282 | Partition an ONNX model graph into named functional subgraphs and emit a subgraphpartition.json file. | amd/ | 181 | — | ~2.7k | Automated safety check: Pass | MIT | 11 days ago |
| 283 | Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. | amd/ | 181 | — | ~1.9k | Automated safety check: Notes | MIT | 11 days ago |
| 284 | Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run. | amd/ | 181 | — | ~1.5k | Automated safety check: Pass | MIT | 11 days ago |
| 285 | Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole. | amd/ | 181 | — | ~2.3k | Automated safety check: Pass | MIT | 11 days ago |
| 286 | L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate. | amd/ | 181 | — | ~2.6k | Automated safety check: Pass | MIT | 11 days ago |
| 287 | Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output. | amd/ | 181 | — | ~2.8k | Automated safety check: Pass | MIT | 11 days ago |
| 288 | Inspect a target model and prepare metadata for Quark PTQ planning. | amd/ | 181 | — | ~2k | Automated safety check: Pass | MIT | 11 days ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents546
- Deep learning408
- Embeddings373
- Prompt engineering348
- Retrieval-augmented generation348
- LLM evaluation308
- Fine-tuning308
- Speech recognition and synthesis305
- Structured output and tool calling268
- Model routing and gateways257
- LLM cost and token optimization254
- LLM API integration250
- LLM observability242
- LLM guardrails217
- Computer vision195
- Model hubs and datasets179
- GPU and accelerator computing172
- Diffusion and image models165
- Natural language processing134
- Reinforcement learning66
- AI interpretability23