Search
AI & LLM Engineering · For devops and sre engineers
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Step-by-step guide for adding a new observability/tracing provider to Agent Kernel. | yaalalabs/ | 192 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 98 | Diagnose and remediate per-node issues on a HyperPod cluster (EKS or Slurm) — a specific node is unhealthy, unresponsive, stuck, or needs replacing. | awslabs/ | 916 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 99 | Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. | NVIDIA/ | 3.6k | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 100 | A skill your agent uses for MolMIM, NVIDIA's BioNeMo NIM microservice for small-molecule latent-space generation and optimization. | NVIDIA/ | 3.6k | 1 repo | ~1.9k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 101 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | sickn33/ | 47k | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 102 | Deploy, monitor, and debug long GPU jobs on RENTED/remote instances (AutoDL, RunPod, vast.ai, Lambda, Slurm, K8s): teardown/billing safety, spot resilience, resumable checkpointing, OOM/NaN triage. | sickn33/ | 47k | 1 repo | ~5.8k | Automated safety check: Pass | MIT | 2 days ago |
| 103 | Deploys a baseline landing zone foundation for a Google Cloud Organization, establishing security guardrails using Organization Policies, resource hierarchy folders and projects, billing… | google/ | 21k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 104 | Sync Agent Kernel documentation from branch changes before a feature or bugfix is merged. | yaalalabs/ | 192 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 105 | Analyze the most expensive users in AI observability and explain why they cost so much. | PostHog/ | 40k | — | ~3.9k | Automated safety check: Pass | Unknown | yesterday |
| 106 | 106.Fastllm Routing Route requests across models in FastLLM — create, inspect or change frontend models, their default targets, weighted splits, and routing rules (by caller, prompt size, streaming, headers, budget… | azrtydxb/ | 108 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 107 | Adds Arize AX tracing to an LLM application for the first time. | github/ | 40k | — | ~6.2k | Automated safety check: Notes | MIT | 2 days ago |
| 108 | Port a published computer vision paper's official code and training recipe onto a customer's own dataset, or diagnose why such a transfer produced bad numbers. | NVIDIA/ | 3.6k | — | ~4.2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 109 | Take a working agent prototype to a scalable, observable production deployment on Microsoft Foundry. | microsoft/ | 77k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 110 | Take one working agent prototype go scalable, observable production deployment for Microsoft Foundry. | microsoft/ | 77k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 111 | 111.Qv Qip Triage Use during planning, implementation, PR review, or /qv-qip-triage when a change may affect public SDK API, native dependency, plugin contract, model registry contract, runtime, transport, storage… | tetherto/ | 685 | — | ~746 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 112 | Apply Karpathy-style minimalism and anti-dependency principles to code and system design. | LearnPrompt/ | 110 | — | ~1.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 113 | 113.Designing Change Resolve a non-trivial implementation or architecture decision before coding. | rapidaai/ | 745 | — | ~606 | Automated safety check: Pass | Unknown | 4 days ago |
| 114 | Add or modify telemetry/metrics instrumentation for assistant-api and integration-api flows while preserving metric schema compatibility and audit behavior. | rapidaai/ | 745 | — | ~736 | Automated safety check: Pass | Unknown | 4 days ago |
| 115 | Create or update repository documentation, runbooks, configuration references, and developer guidance. | rapidaai/ | 745 | — | ~584 | Automated safety check: Pass | Unknown | 4 days ago |
| 116 | Deploy AI and NLP-powered detection systems to identify business email compromise attacks by analyzing writing style, behavioral patterns, and contextual anomalies that evade traditional rule-based… | mukul975/ | 34k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 117 | Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment). | PostHog/ | 40k | — | ~5.7k | Automated safety check: Pass | Unknown | yesterday |
| 118 | Work with the Ambient Quality Agent (AQuA) added to this agents-cli project: augment the agents-cli agent with AQuA or attach the agent to an AQuA deployed elsewhere, read the quality insights it… | google/ | 10k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 119 | 119.Yao Meta Skill Create, refactor, evaluate, and package agent skills from workflows, prompts, transcripts, docs, or notes. | aiskillstore/ | 433 | 2 repos | ~806 | Automated safety check: Pass | MIT | yesterday |
| 120 | Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN… | grafana/ | 282 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 121 | 121.Flash Attention Speed up long-sequence transformer training and inference. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~2.7k | Automated safety check: Pass | MIT | 3 days ago |
| 122 | 122.Guidance Constrain LLM output with grammars; guarantee valid JSON. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~3.8k | Automated safety check: Pass | MIT | 3 days ago |
| 123 | 123.Saelens Train sparse autoencoders to interpret model features. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~3.7k | Automated safety check: Pass | MIT | 3 days ago |
| 124 | 124.Simpo Reference-free preference alignment, simpler than DPO. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~1.4k | Automated safety check: Pass | MIT | 3 days ago |
| 125 | 125.Slime RL post-training for LLMs with Megatron and SGLang. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~2.7k | Automated safety check: Pass | MIT | 3 days ago |
| 126 | 126.Cost Alert Review the configured cost alert rules and the alerts currently fired on the Agent Monitor dashboard, then explain exactly what tripped and why. | hoangsonww/ | 1.1k | — | ~772 | Automated safety check: Pass | MIT | yesterday |
| 127 | 127.Adapt Workflow A skill your agent uses when porting a workflow to a different AI provider, deployment environment, model tier, or organizational context. | sharpdeveye/ | 591 | — | ~597 | Automated safety check: Pass | MIT | 5 mo ago |
| 128 | Debug live Azure Kubernetes Service (AKS) incidents with a read-only, evidence-first investigation. | microsoft/ | 255 | — | ~3.2k | Automated safety check: Pass | MIT | yesterday |
| 129 | Plan and coordinate a model deployment to Amazon SageMaker, including serving stack and real-time versus async inference. | waybarrios/ | 534 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 130 | Check progress for a detached KERMT run (pretrain, finetune, or any kermtrundetached invocation). | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 131 | Build AI applications on Microsoft Foundry using the azure-ai-projects SDK. | aiskillstore/ | 433 | 4 repos | ~2.1k | Automated safety check: Pass | No licence | yesterday |
| 132 | Workflow specifically for creating a new resource, supporting both autogen and manual generation. | GoogleCloudPlatform/ | 974 | — | ~512 | Automated safety check: Pass | Unknown | yesterday |
| 133 | Operate and troubleshoot the UAV Mission Compute SDK for PX4 telemetry, camera streaming, missions, computer vision, and edge AI demonstrations. | open-edge-platform/ | 140 | — | ~1.1k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 134 | A skill your agent uses when adding logging, metrics, or instrumentation — structured Logger usage, telemetry handler attachment, Ecto/Phoenix telemetry events. | j-morgan6/ | 167 | — | ~2.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 135 | Inspect GitHub Actions self-hosted runner queues when a developer provides a run URL, job URL, runner label, or reports a blocked CI job, or invokes /qv-devops-runner-queue. | tetherto/ | 685 | — | ~413 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 136 | Self-healing China mirror source resolver. An agent skill from LeoYeAI/openclaw-master-skills. | LeoYeAI/ | 2.2k | — | ~4.6k | Automated safety check: Notes | MIT | 2 mo ago |
| 137 | A Principal ML Engineer interviewer that simulates a FAANG-style ML system design interview covering the full lifecycle from data to production. | PrepLabsAI/ | 112 | — | ~4.2k | Automated safety check: Pass | MIT | 4 days ago |
| 138 | Migrate workloads from Google Cloud Platform to AWS — plus AI and agentic workloads from any provider. | aws/ | 2.8k | — | ~15k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 139 | 139.Dt Obs Genai Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup. | Dynatrace/ | 163 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 140 | Guides tenant isolation architecture in Qdrant for multi-tenant or multi-user applications. | qdrant/ | 254 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 141 | 141.Zero Token Zero token cost. An agent skill from LeoYeAI/openclaw-master-skills. | LeoYeAI/ | 2.2k | — | ~889 | Automated safety check: Pass | MIT | 2 mo ago |
| 142 | Sync Agent Kernel skills and documentation from a specific commit hash. | yaalalabs/ | 192 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 143 | Sync Agent Kernel skills from branch changes before a feature or bugfix is merged. | yaalalabs/ | 192 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 144 | How to swap the DeepStream CV detection model in the VSS Alerts Blueprint verification (2dcv) mode - covers ONNX export, custom bbox parsers, compose mount gotchas, nvinfer config, runtime TRT… | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | yesterday |