Search

LLM inference and serving

373 skills found, page 6.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
241
241.Airunway Aks SetupOfficial

Set up AI Runway on AKS — from bare cluster to running model.

microsoft/GitHub-Copilot-for-Azure2551 repo~1.1kAutomated safety check: PassMITyesterday
242

Debug live Azure Kubernetes Service (AKS) incidents with a read-only, evidence-first investigation.

microsoft/GitHub-Copilot-for-Azure255—~3.2kAutomated safety check: PassMITyesterday
243

Delegate coding to OpenCode CLI (multi-LLM open-source coding agent)

taracodlabs/aiden852—~917Automated safety check: NotesApache-2.028 days ago
244

Extend, test, debug, or integrate the Model Download microservice codebase.

open-edge-platform/edge-ai-libraries171—~2.8kAutomated safety check: PassApache-2.0yesterday
245

Download and convert AI models using the Model Download microservice.

open-edge-platform/edge-ai-libraries171—~3.8kAutomated safety check: PassApache-2.0yesterday
246

Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP.

scragnog/HOT-Step-CPP174—~6.6kAutomated safety check: PassMIT2 days ago
247

End-to-end ONNX PTQ workflow for AMD Quark — from a .onnx file (and calibration data) to a quantized .onnx output.

amd/Quark182—~4.5kAutomated safety check: PassMIT13 days ago
248

[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the…

rlaope/oh-my-hermes3.2k—~2.2kAutomated safety check: PassMITyesterday
249

Automatically fetch InferenceX benchmark data and generate daily performance reports for LLM inference on various hardware (NVIDIA, AMD, etc.).

ascend-ai-coding/awesome-ascend-skills174—~1.1kAutomated safety check: PassNo licenceyesterday
250
250.Jetson LLM ServeOfficial

Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

NVIDIA/skills3.6k1 repo~3kAutomated safety check: NotesApache-2.02 days ago
251

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

NVIDIA/skills3.6k1 repo~1.2kAutomated safety check: PassApache-2.02 days ago
252

Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.

open-thoughts/OpenThoughts-Agent301—~947Automated safety check: PassApache-2.013 days ago
253

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

open-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.013 days ago
254

Deeper backend map of overbae — module layout, celery queue topology, the span-only tracing model and its API surface, capabilities and toml sync, behaviour-keyed scoring, auth and guests, model…

overmind-core/overmind608—~10kAutomated safety check: PassAGPL-3.02 days ago
255

Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills.

databricks/databricks-agent-skills345—~4.6kAutomated safety check: PassUnknownyesterday
256

Interactively guide setup for connecting the Claude Code CLI to a local Ollama LLM on Mac.

receptron/mulmoclaude371—~1.9kAutomated safety check: NotesMITyesterday
257

Canonical tempo-sync/rate/quantize divisions from src/audio/MusicTime.h (4 bars to 1/32 with dotted and triplets, plus 1/64).

n1m21n/Infinite271—~1.5kAutomated safety check: PassUnknownyesterday
258
258.Deepstream DevOfficial

NVIDIA DeepStream SDK development with Python pyservicemaker API.

NVIDIA/skills3.6k—~3.3kAutomated safety check: PassApache-2.02 days ago
259
259.Deepstream SopOfficial

A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

NVIDIA/skills3.6k—~4.7kAutomated safety check: NotesApache-2.02 days ago
260

Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines.

NVIDIA/skills3.6k—~1.5kAutomated safety check: PassApache-2.02 days ago
261

A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill.

NVIDIA/skills3.6k—~2.1kAutomated safety check: PassApache-2.02 days ago
262
262.Tao Finetune ClipOfficial

CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

NVIDIA/skills3.6k—~4kAutomated safety check: NotesApache-2.02 days ago
263

InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment.

NVIDIA/skills3.6k—~3.5kAutomated safety check: NotesApache-2.02 days ago
264

The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action.

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.02 days ago
265

PyTorch-based TAO image classification. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~3.6kAutomated safety check: NotesApache-2.02 days ago
266
266.Tao Train RtdetrOfficial

RT-DETR (Real-Time DEtection TRansformer) for 2D object detection.

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.02 days ago
267

Run coding tasks via opencode using free cloud models. An agent skill from luongnv89/skills.

luongnv89/skills131—~4.5kAutomated safety check: NotesMIT2 days ago
268

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago
269

Project-kickoff rubric for the self-host vs managed-cloud decision — when is running your own inference engine / LLM platform worth the ops cost vs paying per-token for a managed API?

agentsope/SkillAlchemy436—~6.3kAutomated safety check: PassMIT2 days ago
270

Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago
271

Converts one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model through llama.cpp, automatically preferring an available Vulkan GPU and falling back…

godot-fun/gai183—~813Automated safety check: PassMITyesterday
272

Redact, anonymize, sanitize, or remove PII locally with Distil-PII and llama.cpp; keep personal data and secret values out of model context, logs, and chat.

HybridAIOne/hybridclaw159—~1kAutomated safety check: PassMITyesterday
273

Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama.

autonomous-ai/openharness1.2k—~2.2kAutomated safety check: NotesMITyesterday
274

Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…

autonomous-ai/openharness1.2k—~1.9kAutomated safety check: PassMITyesterday
275

A Principal ML Engineer interviewer that simulates a FAANG-style ML system design interview covering the full lifecycle from data to production.

PrepLabsAI/InterviewMentor112—~4.2kAutomated safety check: PassMIT4 days ago
276

Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama.

glebis/claude-skills391—~1.4kAutomated safety check: PassMIT3 days ago
277

Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.

open-thoughts/OpenThoughts-Agent301—~5.8kAutomated safety check: PassApache-2.013 days ago
278
278.Aks GPU InferenceOfficial

Diagnose Day-2 AKS GPU and KAITO incidents using profile-aware, read-only evidence.

microsoft/GitHub-Copilot-for-Azure255—~764Automated safety check: PassMITyesterday
279

Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.

amd/Quark182—~4.8kAutomated safety check: PassMIT13 days ago
280

Validate Quark ONNX quantization output using four lightweight checks: auxiliary file copy alignment, expected non-quantized initializer MD5 byte-identity (inline rawdata + external-data byte…

amd/Quark182—~2.5kAutomated safety check: PassMIT13 days ago
281

Route Quark ONNX user goals to the correct atomic skill. An agent skill from amd/Quark.

amd/Quark182—~2.7kAutomated safety check: PassMIT13 days ago
282

Apply existing ShapeShifter graph passes to an .onnx model via the quark-cli shapeshifter CLI or a ShapeShifter YAML.

amd/Quark182—~1.9kAutomated safety check: PassMIT13 days ago
283

Detect upstream Quark ONNX changes that affect the ONNX skill family and classify required updates.

amd/Quark182—~3.3kAutomated safety check: PassMIT13 days ago
284

Partition an ONNX model graph into named functional subgraphs and emit a subgraphpartition.json file.

amd/Quark182—~2.7kAutomated safety check: PassMIT13 days ago
285

Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.

amd/Quark182—~1.9kAutomated safety check: NotesMIT13 days ago
286

Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run.

amd/Quark182—~1.5kAutomated safety check: PassMIT13 days ago
287

Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole.

amd/Quark182—~2.3kAutomated safety check: PassMIT13 days ago
288

L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate.

amd/Quark182—~2.6kAutomated safety check: PassMIT13 days ago