Search
Qwen · LLM inference and serving
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation. | vllm-project/ | 7.1k | — | ~7.5k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release. | R6410418/ | 1.7k | — | ~1.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 3 | Run wally's LLM e2e on Apple Neural Engine (NeuRT) and Snapdragon Hexagon NPU (QHexRT) devices. | RunanywhereAI/ | 1.6k | — | ~845 | Automated safety check: Pass | MIT | today |
| 4 | Adapt and port new LLM model architectures to this xinfer project. | guoqingbao/ | 334 | — | ~4.2k | Automated safety check: Notes | MIT | 1 mo ago |
| 5 | 5.Resolve Always resolve Hugging Face models via model-shelf before any download. | alexziskind1/ | 130 | — | ~792 | Automated safety check: Pass | MIT | 1 mo ago |
| 6 | Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer. | guoqingbao/ | 334 | — | ~3.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 7 | Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares). | zhongkaifu/ | 568 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | today |
| 8 | Test LLM models served by xinfer for correctness, output quality, and performance. | guoqingbao/ | 334 | — | ~2.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 9 | 9.Research A skill your agent uses for web searches and current information lookups, finding sources, fact-checking, researching questions, comparing sources, or summarising web pages. | zhongkaifu/ | 568 | — | ~2.3k | Automated safety check: Warn | BSD-3-Clause | today |
| 10 | Guided workflow for adding a new model architecture to llama.cpp. | JakeATX/ | 166 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 11 | 11.Turbofit Hardware-aware adaptive Hermes runtime using portable Turbofiles, total usable memory, owned native llama.cpp residency, stable auto/active:main/active:aux routes, and evidence-backed promotion. | SouthpawIN/ | 107 | — | ~1.9k | Automated safety check: Pass | MIT | 3 days ago |
| 12 | Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. | amd/ | 408 | — | ~4k | Automated safety check: Notes | MIT | 2 days ago |
| 13 | Guide a tester through hipfire bring-up, serve smoke, claim-scoped harnesses, and benchmark reporting on AMD RDNA/CDNA GPUs. | warpfront/ | 658 | — | ~1.5k | Automated safety check: Pass | Unknown | yesterday |
| 14 | Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. | Orchestra-Research/ | 13k | 2 repos | ~2.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 15 | 15.Code Review Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. | JakeATX/ | 166 | — | ~5.6k | Automated safety check: Pass | MIT | yesterday |
| 16 | Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. | BBuf/ | 938 | — | ~3.9k | Automated safety check: Pass | No licence | 6 days ago |
| 17 | A skill your agent uses when a new Ollama Cloud model is announced or available (e.g. | heypinchy/ | 182 | — | ~3.9k | Automated safety check: Notes | AGPL-3.0 | 20 days ago |
| 18 | 18.App Opinionated app components building on top of ./ui primitives | JakeATX/ | 166 | — | ~146 | Automated safety check: Pass | MIT | yesterday |
| 19 | 19.Page Agent Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with… | Tommy-yw/ | 546 | 3 repos | ~2.3k | Automated safety check: Notes | MIT | 4 mo ago |
| 20 | 20.Ollama Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents. | Prism-Shadow/ | 2.5k | — | ~839 | Automated safety check: Notes | Apache-2.0 | yesterday |
| 21 | 21.Vllm Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads. | Prism-Shadow/ | 2.5k | — | ~1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 22 | Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP. | scragnog/ | 174 | — | ~6.6k | Automated safety check: Pass | MIT | 2 days ago |
| 23 | Low-memory file2file quantization for very large safetensors LLMs that cannot be loaded whole. | amd/ | 182 | — | ~2.3k | Automated safety check: Pass | MIT | 13 days ago |
| 24 | L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate. | amd/ | 182 | — | ~2.6k | Automated safety check: Pass | MIT | 13 days ago |
| 25 | Torch LLM PTQ workflow for AMD Quark — from model selection to quantized output. | amd/ | 182 | — | ~2.8k | Automated safety check: Pass | MIT | 13 days ago |
| 26 | Inspect a target model and prepare metadata for Quark PTQ planning. | amd/ | 182 | — | ~2k | Automated safety check: Pass | MIT | 13 days ago |
| 27 | Plan or review LoRA and edit-training work specifically for FLUX.2 Klein or Qwen-Image-Edit, including paired datasets, trainer-version contracts, and held-out fidelity checks. | AnastasiyaW/ | 154 | — | ~4.5k | Automated safety check: Pass | MIT | yesterday |
| 28 | Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request… | amd/ | 182 | — | ~2.3k | Automated safety check: Pass | MIT | 13 days ago |
| 29 | 29.Vllm Ascend vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 30 | Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode… | ascend-ai-coding/ | 174 | — | ~731 | Automated safety check: Pass | No licence | yesterday |
| 31 | 昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | yesterday |