Search
LLM inference and serving
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | 145.Facturas Úsalo cuando el usuario pida redactar una factura, una cuenta de cobro o una nota de cobro. | gustavoeenriquez/ | 212 | — | ~127 | Automated safety check: Pass | MIT | 2 days ago |
| 146 | Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups. | Orchestra-Research/ | 13k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 147 | Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA. | BBuf/ | 938 | — | ~7.5k | Automated safety check: Pass | No licence | 6 days ago |
| 148 | Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades. | MetaX-MACA/ | 180 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 149 | 149.Vss Deploy Deploys and manages VSS through setup.sh and its Docker Compose overlays. | open-edge-platform/ | 171 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 150 | Installation and configuration skill for Agent Brain document search system. | SpillwaveSolutions/ | 119 | — | ~7k | Automated safety check: Notes | MIT | 21 days ago |
| 151 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 182 | — | ~1.4k | Automated safety check: Pass | MIT | 13 days ago |
| 152 | A skill your agent uses when a new Ollama Cloud model is announced or available (e.g. | heypinchy/ | 182 | — | ~3.9k | Automated safety check: Notes | AGPL-3.0 | 20 days ago |
| 153 | This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. | vllm-project/ | 102 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 154 | 154.ML Engineer Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. | davila7/ | 33k | 9 repos | ~2.3k | Automated safety check: Pass | MIT | today |
| 155 | 155.Ollama Optimizer Optimize Ollama configuration for the current machine's hardware. | luongnv89/ | 131 | — | ~4.1k | Automated safety check: Notes | MIT | today |
| 156 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 157 | 157.Nemotron Ultra Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference. | NVIDIA-NeMo/ | 2.1k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 158 | A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence. | BBuf/ | 938 | — | ~1.5k | Automated safety check: Pass | No licence | 6 days ago |
| 159 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 160 | DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity… | open-thoughts/ | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 161 | Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 8 days ago |
| 162 | Create or debug an AReno multi-turn agentic dataset, runagent implementation, tool schemas, tool execution, reward, loss masks, interactive TUI game with OpenAI-compatible LLM inference, or agentic… | inclusionAI/ | 323 | — | ~435 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 163 | 163.Prime Agent A skill your agent uses when learning, configuring, or troubleshooting Prime Agent (PrimeIntellect-ai/prime-agent), including installation, providers, custom OpenAI-compatible models, local… | wcygan/ | 194 | — | ~1.8k | Automated safety check: Pass | No licence | yesterday |
| 164 | 164.Local LLM Free Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. | artokun/ | 803 | — | ~897 | Automated safety check: Notes | MIT | 6 days ago |
| 165 | A skill your agent uses when running quantization of a BF16/FP16 GGUF repo and Skippy layer-package creation as one local or Hugging Face Jobs workflow, publishing both artifacts to Hugging Face. | Mesh-LLM/ | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 166 | Adds an MCP server so the NanoClaw container agent can send prompts to local Ollama models, with optional tools to manage the model library. | nanocoai/ | 31k | — | ~3k | Automated safety check: Notes | MIT | yesterday |
| 167 | 167.Open Notebook Organizes research with the self-hosted Open Notebook alternative to NotebookLM. | K-Dense-AI/ | 48k | 1 repo | ~2.8k | Automated safety check: Pass | MIT | 6 days ago |
| 168 | 3DGS compression-to-deployment pipeline: quantization (scalar/VQ/mixed-precision), pruning (coreset/adaptive/variational/merge/Bayesian), progressive streaming & LoD, Web/WebGPU/mobile deployment… | jaccen/ | 161 | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 169 | Trim large MetaX model directories for dummy smoke tests or real-checkpoint loading on limited GPUs. | MetaX-MACA/ | 180 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 170 | 170.Llama Cpp llama.cpp local GGUF inference + HF Hub model discovery. An agent skill from Tommy-yw/RunbookHermes. | Tommy-yw/ | 546 | 4 repos | ~2.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 171 | 171.App Opinionated app components building on top of ./ui primitives | JakeATX/ | 166 | — | ~146 | Automated safety check: Pass | MIT | yesterday |
| 172 | Implement a production SageMaker endpoint with autoscaling, CloudWatch alarms, and tags. | waybarrios/ | 534 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 173 | L3 recipe that runs quark.onnx.AutoSearchPro end-to-end on a user .onnx model: intake → preset selection (or custom search space) → calibration / eval data reader → standalone autosearch script… | amd/ | 182 | — | ~3.4k | Automated safety check: Pass | MIT | 13 days ago |
| 174 | Diagnose and fix OCI AI Quick Actions (AQUA) issues including deployment failures, OOM errors, authorization problems, capacity issues, container errors, and policy misconfigurations. | oracle/ | 125 | — | ~1.8k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 175 | 175.Page Agent Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with… | Tommy-yw/ | 546 | 3 repos | ~2.3k | Automated safety check: Notes | MIT | 4 mo ago |
| 176 | Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic. | vllm-project/ | 7.1k | — | ~3k | Automated safety check: Pass | Apache-2.0 | today |
| 177 | Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and… | ai-dynamo/ | 8.3k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 178 | Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces. | darknecrocities/ | 239 | — | ~17k | Automated safety check: Pass | No licence | 5 days ago |
| 179 | Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. | vllm-project/ | 102 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 6 mo ago |
| 180 | Build private, on-device AI features on iPhone, iPad, and Mac with Foundation Models, Core ML, MLX Swift, or llama.cpp. | dpearson2699/ | 1.2k | — | ~3.4k | Automated safety check: Pass | Unknown | 2 mo ago |
| 181 | 181.Pi Agent Builds with and operates Pi, the minimal terminal coding harness. | K-Dense-AI/ | 48k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | 6 days ago |
| 182 | Select and verify the current region-specific serving container URI for a SageMaker model deployment. | waybarrios/ | 534 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 183 | 183.Quark Onnx Debug Diagnose failed Quark ONNX installation, calibration, quantization, custom-op compilation, or export attempts. | amd/ | 182 | — | ~4.8k | Automated safety check: Pass | MIT | 13 days ago |
| 184 | 184.Crud Archive Run Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL… | open-thoughts/ | 301 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 185 | Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio, Reflex. | databricks/ | 345 | — | ~3.1k | Automated safety check: Pass | Unknown | yesterday |
| 186 | 186.Cc Ollama Claude Code 安裝本地 AI Ollama。說「安裝 Ollama」「本地 AI」時載入. An agent skill from mathruffian-dot/claude-code-lazy-packs. | mathruffian-dot/ | 254 | — | ~118 | Automated safety check: Pass | MIT | 1 mo ago |
| 187 | 187.Pai Skills Root Index of all PAI skills. | nirholas/ | 113 | — | ~333 | Automated safety check: Pass | Unknown | today |
| 188 | Run OpenMed models fully on-device with the MLX (Apple Silicon), CoreML (iOS/macOS), or ONNX/WebGPU (cross-platform/browser) backends, including convert-quantize-run workflows. | maziyarpanahi/ | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 189 | A skill your agent uses when changing mesh-llm's llama.cpp patch queue, upstream pin, prepare/build scripts, or carried RPC, MoE, and mesh-hook llama.cpp patches. | Mesh-LLM/ | 3.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 190 | A skill your agent uses when changing mesh-llm's patched llama.cpp Skippy ABI, runtime hooks, model introspection, tensor filtering, activation-frame execution, GGUF writer surface, upstream pin, or… | Mesh-LLM/ | 3.5k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 191 | Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware. | sickn33/ | 47k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 192 | 192.Ito Inference Inspect the availability of model serving on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed serving manifest. | affaan-m/ | 277k | 1 repo | ~1.5k | Automated safety check: Pass | MIT | today |