Search

AI & LLM Engineering · llama.cpp

85 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

A skill your agent uses when inspecting GGUF models, planning layer ranges, generating or validating skippy package artifacts, fake packages for direct GGUFs, materialized stage cache behavior, or…

Mesh-LLM/mesh-llm3.5k—~588Automated safety check: PassApache-2.0today
50

Transcribes a single PCM16 WAV file to plain UTF-8 text locally through a standard-library Python wrapper around native SenseVoice Small F16 GGUF FunASR llama.cpp runtimes, preferring cross-vendor…

godot-fun/gai184—~707Automated safety check: PassMITtoday
51

Train or fine-tune language and vision models using TRL (Transformer Reinforcement Learning) or Unsloth with Hugging Face Jobs infrastructure.

sickn33/agentic-awesome-skills47k1 repo~1.1kAutomated safety check: PassApache-2.0yesterday
52

A skill your agent uses when testing or benchmarking target/draft GGUF pairs for speculative decoding compatibility, tokenizer agreement, draft acceptance rate, or staged verification behavior.

Mesh-LLM/mesh-llm3.5k—~260Automated safety check: PassApache-2.0today
53

A skill your agent uses when certifying a GGUF model family for skippy stage-split serving, reviewing capability data, promoting family evidence into topology policy, or updating staged split…

Mesh-LLM/mesh-llm3.5k—~568Automated safety check: PassApache-2.0yesterday
54

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio.

sickn33/agentic-awesome-skills47k2 repos~1.6kAutomated safety check: PassMIT2 days ago
55

Fine-tune and post-train LLMs with Unsloth Core on a single consumer GPU: VRAM sizing, LoRA/QLoRA, GRPO/DPO, chat-template correctness, and GGUF export.

sickn33/agentic-awesome-skills47k1 repo~4.1kAutomated safety check: PassApache-2.02 days ago
56

Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
57

Find and compare recommended Hugging Face models for a task using benchmarks, model size, and device constraints.

waybarrios/opencode-power-pack534—~1.4kAutomated safety check: PassApache-2.05 days ago
58

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

NVIDIA/skills3.6k1 repo~2.9kAutomated safety check: PassApache-2.0yesterday
59

Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

NVIDIA/skills3.6k1 repo~3.1kAutomated safety check: PassApache-2.0yesterday
60

Build WAN 2.2 First-Last-Frame video workflows. An agent skill from artokun/comfyui-mcp.

artokun/comfyui-mcp803—~5.1kAutomated safety check: PassMIT6 days ago
61

Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a…

ruvnet/ruflo74k—~336Automated safety check: NotesMITyesterday
62

Trace a commit to its published npm versions, including transitive SDK resolution with time-aware accuracy

tetherto/qvac685—~1.7kAutomated safety check: PassApache-2.0yesterday
63

Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP.

scragnog/HOT-Step-CPP174—~6.6kAutomated safety check: PassMITyesterday
64

[omh] Self-hosted LLM serving on GPUs: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the…

rlaope/oh-my-hermes3.2k—~2.2kAutomated safety check: PassMITyesterday
65

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago
66

Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy436—~6.1kAutomated safety check: PassMIT2 days ago
67

Converts one or more images into faithful text descriptions or OCR with the local 1.3B MiniCPM-V 4.6 GGUF model through llama.cpp, automatically preferring an available Vulkan GPU and falling back…

godot-fun/gai183—~813Automated safety check: PassMITyesterday
68

Redact, anonymize, sanitize, or remove PII locally with Distil-PII and llama.cpp; keep personal data and secret values out of model context, logs, and chat.

HybridAIOne/hybridclaw159—~1kAutomated safety check: PassMITyesterday
69

Serve an MLX or Hugging Face safetensors model already on this Mac with mlx-lm's server and join it to the person's fleet.

autonomous-ai/openharness1.2k—~1.9kAutomated safety check: PassMITyesterday
70

Reuse what Ollama already has on this computer: adopt a running Ollama server into the person's fleet, or serve an Ollama-downloaded model without Ollama.

autonomous-ai/openharness1.2k—~2.2kAutomated safety check: NotesMITyesterday
71

Run quick, offline, private LLM tasks on local models via llama.cpp, reusing models already downloaded by Ollama.

glebis/claude-skills391—~1.4kAutomated safety check: PassMIT3 days ago
72

Prepare export and downstream evaluation handoff for a planned or completed Quark PTQ run.

amd/Quark182—~1.5kAutomated safety check: PassMIT13 days ago
73

在本机 Mac 或 Apple Silicon 上部署 Gemma 4 12B。本地安装/升级 llama.cpp,下载 GGUF 量化模型,用 llama-server 暴露 OpenAI-compatible API,或用 Ollama 暴露本地模型服务;按用户需求在默认 Q4KM、64K/128K 长上下文、QAT Q40 @ 256K、左右对比演示之间选择,配置 tmux…

majiayu000/spellbook287—~875Automated safety check: NotesMIT2 days ago
74

Set up, install, and configure CONFIDE local de-identification — installs Python deps (natasha, scrubadub, phonenumbers, pymorphy2), ensures Ollama + pulls the default qwen2.5:3b model, detects…

glebis/claude-skills391—~1.1kAutomated safety check: PassMIT3 days ago
75

A skill your agent uses when running open-weight LLMs locally with Ollama — pulling and tagging models, calling the local API, picking a quantization or GGUF, writing Modelfiles, and sizing VRAM and…

ericrisco/rsc-harness180—~2.8kAutomated safety check: PassMITyesterday
76

Operate, configure, secure, and troubleshoot the LiteLLM AI gateway (proxy) and Python SDK: run the proxy (litellm --config), route to 100+ providers through one OpenAI-compatible API, configure…

magnus919/agent-skills115—~4.2kAutomated safety check: NotesMITyesterday
77
77.Vllm

Operate, configure, benchmark, and troubleshoot vLLM inference servers: Docker and Kubernetes deployment, quantization-aware model configuration (tensor parallelism, KV cache), OpenAI-compatible API…

magnus919/agent-skills115—~4.1kAutomated safety check: NotesMITyesterday
78

Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

AnastasiyaW/codex-claude-code-config154—~794Automated safety check: PassMITyesterday
79

Run quantized LLMs locally with llama.cpp — CPU+GPU inference, GGUF format, OpenAI-compatible server, and Python bindings.

AlexAI-MCP/hermes-CCC135—~2.3kAutomated safety check: PassMIT6 mo ago
80

Build Lightricks LTX-2 / LTX-2.3 video workflows covering text-to-video, image-to-video, GGUF and bundled checkpoints, distilled model, camera control LoRAs, synchronized audio, two-stage upscaling…

artokun/comfyui-mcp803—~6.8kAutomated safety check: WarnMIT6 days ago
81

A skill your agent uses when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then…

ericrisco/rsc-harness180—~3.8kAutomated safety check: PassMITyesterday
82

A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…

ericrisco/rsc-harness180—~3.6kAutomated safety check: PassMITyesterday
83

Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

magnus919/agent-skills115—~2.3kAutomated safety check: PassMITyesterday
84

Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…

magnus919/agent-skills115—~1.5kAutomated safety check: PassMITyesterday
85

使用本地 BiRefNet GGUF 模型完成图片或视频抠图、人物抠图、主体分割和背景移除,并输出透明 PNG、MOV 或 WebM。适用于用户提到图片抠图、照片去背景、人像透明图、视频抠图、透明视频、BiRefNet、JPG/PNG/BMP/WebP 图片,或 MP4/MOV/WebM 视频的场景;无需 Python、PyTorch 或 CUDA。

aiskillstore/marketplace433—~748Automated safety check: PassMITyesterday