Search

LLM inference and serving

373 skills found, page 4.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
145

Úsalo cuando el usuario pida redactar una factura, una cuenta de cobro o una nota de cobro.

gustavoeenriquez/MakerAi212—~127Automated safety check: PassMIT2 days ago
146

Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.3kAutomated safety check: PassMIT3 mo ago
147

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~7.5kAutomated safety check: PassNo licence6 days ago
148

Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades.

MetaX-MACA/vLLM-metax180—~1.1kAutomated safety check: PassApache-2.0yesterday
149

Deploys and manages VSS through setup.sh and its Docker Compose overlays.

open-edge-platform/edge-ai-libraries171—~4.1kAutomated safety check: PassApache-2.0yesterday
150

Installation and configuration skill for Agent Brain document search system.

SpillwaveSolutions/agent-brain119—~7kAutomated safety check: NotesMIT21 days ago
151

Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

amd/Quark182—~1.4kAutomated safety check: PassMIT13 days ago
152

A skill your agent uses when a new Ollama Cloud model is announced or available (e.g.

heypinchy/pinchy182—~3.9kAutomated safety check: NotesAGPL-3.020 days ago
153

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns.

vllm-project/vllm-skills102—~1.4kAutomated safety check: PassApache-2.06 mo ago
154

Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks.

davila7/claude-code-templates33k9 repos~2.3kAutomated safety check: PassMITtoday
155

Optimize Ollama configuration for the current machine's hardware.

luongnv89/skills131—~4.1kAutomated safety check: NotesMITtoday
156

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMIT2 days ago
157

Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference.

NVIDIA-NeMo/Nemotron2.1k—~1.8kAutomated safety check: PassApache-2.05 days ago
158

A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~1.5kAutomated safety check: PassNo licence6 days ago
159
159.Aqua MetricsOfficial

Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI.

oracle/accelerated-data-science125—~1.5kAutomated safety check: PassUPL-1.01 mo ago
160

DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

open-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.013 days ago
161

Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

Jeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT8 days ago
162

Create or debug an AReno multi-turn agentic dataset, runagent implementation, tool schemas, tool execution, reward, loss masks, interactive TUI game with OpenAI-compatible LLM inference, or agentic…

inclusionAI/AReno323—~435Automated safety check: PassApache-2.0yesterday
163

A skill your agent uses when learning, configuring, or troubleshooting Prime Agent (PrimeIntellect-ai/prime-agent), including installation, providers, custom OpenAI-compatible models, local…

wcygan/dotfiles194—~1.8kAutomated safety check: PassNo licenceyesterday
164

Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama.

artokun/comfyui-mcp803—~897Automated safety check: NotesMIT6 days ago
165

A skill your agent uses when running quantization of a BF16/FP16 GGUF repo and Skippy layer-package creation as one local or Hugging Face Jobs workflow, publishing both artifacts to Hugging Face.

Mesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0today
166

Adds an MCP server so the NanoClaw container agent can send prompts to local Ollama models, with optional tools to manage the model library.

nanocoai/nanoclaw31k—~3kAutomated safety check: NotesMITyesterday
167

Organizes research with the self-hosted Open Notebook alternative to NotebookLM.

K-Dense-AI/scientific-agent-skills48k1 repo~2.8kAutomated safety check: PassMIT6 days ago
168

3DGS compression-to-deployment pipeline: quantization (scalar/VQ/mixed-precision), pruning (coreset/adaptive/variational/merge/Bayesian), progressive streaming & LoD, Web/WebGPU/mobile deployment…

jaccen/Awesome-Gaussian-Skills161—~5.5kAutomated safety check: PassApache-2.0today
169

Trim large MetaX model directories for dummy smoke tests or real-checkpoint loading on limited GPUs.

MetaX-MACA/vLLM-metax180—~2.4kAutomated safety check: PassApache-2.0yesterday
170

llama.cpp local GGUF inference + HF Hub model discovery. An agent skill from Tommy-yw/RunbookHermes.

Tommy-yw/RunbookHermes5464 repos~2.2kAutomated safety check: PassMIT4 mo ago
171
171.App

Opinionated app components building on top of ./ui primitives

JakeATX/llamAmpere166—~146Automated safety check: PassMITyesterday
172

Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags.

waybarrios/opencode-power-pack534—~4.6kAutomated safety check: PassApache-2.05 days ago
173

L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script…

amd/Quark182—~3.4kAutomated safety check: PassMIT13 days ago
174

Diagnose and fix OCI AI Quick Actions (AQUA) issues including deployment failures, OOM errors, authorization problems, capacity issues, container errors, and policy misconfigurations.

oracle/accelerated-data-science125—~1.8kAutomated safety check: PassUPL-1.01 mo ago
175

Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…

Tommy-yw/RunbookHermes5463 repos~2.3kAutomated safety check: NotesMIT4 mo ago
176

Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

vllm-project/vllm-omni7.1k—~3kAutomated safety check: PassApache-2.0today
177

Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…

ai-dynamo/dynamo8.3k—~4.9kAutomated safety check: PassApache-2.0today
178

Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces.

darknecrocities/DomoDomo---All-in-one-Tool239—~17kAutomated safety check: PassNo licence5 days ago
179

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills102—~2.5kAutomated safety check: NotesApache-2.06 mo ago
180

Build private, on-device AI features on iPhone, iPad, and Mac with Foundation Models, Core ML, MLX Swift, or llama.cpp.

dpearson2699/swift-ios-skills1.2k—~3.4kAutomated safety check: PassUnknown2 mo ago
181

Builds with and operates Pi, the minimal terminal coding harness.

K-Dense-AI/scientific-agent-skills48k1 repo~2.1kAutomated safety check: PassMIT6 days ago
182

Select and verify the current region-specific serving container URI for a SageMaker model deployment.

waybarrios/opencode-power-pack534—~4.3kAutomated safety check: PassApache-2.05 days ago
183

Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts.

amd/Quark182—~4.8kAutomated safety check: PassMIT13 days ago
184

Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…

open-thoughts/OpenThoughts-Agent301—~1.2kAutomated safety check: PassApache-2.013 days ago
185

Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio, Reflex.

databricks/databricks-agent-skills345—~3.1kAutomated safety check: PassUnknownyesterday
186

Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs.

mathruffian-dot/claude-code-lazy-packs254—~118Automated safety check: PassMIT1 mo ago
187

Index of all PAI skills.

nirholas/PAI113—~333Automated safety check: PassUnknowntoday
188

Run OpenMed models fully on-device with the MLX (Apple Silicon), CoreML (iOS/macOS), or ONNX/WebGPU (cross-platform/browser) backends, including convert-quantize-run workflows.

maziyarpanahi/openmed5.5k—~2kAutomated safety check: PassApache-2.0today
189

A skill your agent uses when changing mesh-llm's llama.cpp patch queue, upstream pin, prepare/build scripts, or carried RPC, MoE, and mesh-hook llama.cpp patches.

Mesh-LLM/mesh-llm3.5k—~1.9kAutomated safety check: PassApache-2.0today
190

A skill your agent uses when changing mesh-llm's patched llama.cpp Skippy ABI, runtime hooks, model introspection, tensor filtering, activation-frame execution, GGUF writer surface, upstream pin, or…

Mesh-LLM/mesh-llm3.5k—~2.6kAutomated safety check: PassApache-2.0today
191

Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

sickn33/agentic-awesome-skills47k1 repo~1.9kAutomated safety check: PassApache-2.0yesterday
192

Inspect the availability of model serving on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed serving manifest.

affaan-m/ECC277k1 repo~1.5kAutomated safety check: PassMITtoday