Search
AI & LLM Engineering · For devops and sre engineers
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract. | microsoft/ | 19k | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 2 | Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up. | NVIDIA/ | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior. | JuliusBrussee/ | 111k | 1 repo | ~2.6k | Automated safety check: Warn | Apache-2.0 | yesterday |
| 4 | Decide, don't guess — trigger on ANY combinatorial or ground-state decision where a plausible guess is worse than silence: rosters and on-call schedules, packing and placement, RAG passage… | brayonpi/ | 1.4k | — | ~5k | Automated safety check: Pass | Proprietary | 1 mo ago |
| 5 | Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups. | huggingface/ | 11k | 1 repo | ~6.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 6 | Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training. | AI45Lab/ | 236 | — | ~1.8k | Automated safety check: Pass | No licence | 16 days ago |
| 7 | Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. | BBuf/ | 925 | — | ~2.5k | Automated safety check: Pass | No licence | 4 days ago |
| 8 | 8.Langfuse Interact with Langfuse and access its documentation: tracing, monitoring, creating datasets, running experiments, and evaluating AI applications. | langfuse/ | 300 | — | ~2.1k | Automated safety check: Notes | MIT | 8 days ago |
| 9 | Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno. | inclusionAI/ | 323 | — | ~486 | Automated safety check: Pass | Apache-2.0 | today |
| 10 | Starts, rebuilds, and troubleshoots the Opik local dev stack, including an optional Comet Platform integration mode for the Opik team. | comet-ml/ | 22k | — | ~734 | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks. | NVIDIA-AI-Blueprints/ | 1.9k | — | ~8.6k | Automated safety check: Notes | Apache-2.0 | today |
| 12 | Give the human an off switch and a cost meter for the coding agents on this machine, using ClawMetry. | vivekchand/ | 426 | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 13 | Interactive scaffold generator for Orloj multi-agent systems. | OrlojHQ/ | 123 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 21 days ago |
| 14 | Creates an azd environment, checks RBAC and model quota, provisions an AI agent app on Azure with azd up and health-checks the deployed app. | Azure-Samples/ | 374 | — | ~4.7k | Automated safety check: Notes | MIT | yesterday |
| 15 | 当任务需要创建、修改、扩展或重组阿里云 SLS 的 dashboard JSON 或可导入的大盘配置时使用;尤其适用于线上大盘、强对比的分析看板、已校验的查询包,或需要专业中文标签与指标定义的运维向大盘。 | alibaba/ | 200 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 16 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 17 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 456 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 18 | Diagnose and fix Cosmos3 environment, installation, and runtime errors. | NVIDIA/ | 559 | — | ~1.3k | Automated safety check: Notes | Unknown | today |
| 19 | This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a. | brevdev/ | 146 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 20 | Read your own agent telemetry from ClawMetry (waste, progress, cost) and act on it before finishing a task. | vivekchand/ | 426 | — | ~515 | Automated safety check: Pass | MIT | yesterday |
| 21 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | today |
| 22 | 22.Eval Evaluate and score agent behavior against a golden reference. | agentevals-dev/ | 162 | — | ~904 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 23 | Propose an improved version of a prompt registered in a self-hosted AgentX (AgentX-trace-eval) instance, using real low-rated evaluation results as evidence, then publish it as a new version once… | AgentX-ai/ | 106 | — | ~2k | Automated safety check: Pass | Unknown | 2 days ago |
| 24 | Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud. | google/ | 21k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 25 | Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues. | climate-analytics-lab/ | 108 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 26 | Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout. | Orchestra-Research/ | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 27 | Add new LLM model pricing entries to Litefuse's default-model-prices.json. | litefuse/ | 100 | — | ~3.6k | Automated safety check: Pass | Unknown | today |
| 28 | Runs OpenAI Codex CLI as a non-interactive worker for CI, Docker, Kubernetes or remote servers, with sandbox modes and JSONL-friendly output. | XiaomiMiMo/ | 14k | — | ~2.7k | Automated safety check: Pass | MIT | yesterday |
| 29 | Generate preventive Well-Architected guardrails — AWS Config rules, Service Control Policies, permission boundaries, CloudWatch alarms, and IaC policy checks (CDK Aspects, cfn-guard, OPA/Sentinel) —… | aws-samples/ | 273 | — | ~2.8k | Automated safety check: Pass | MIT-0 | 3 days ago |
| 30 | 30.Budget Set Define a spend budget for Claude Code and, optionally, create a cost alert rule that fires when usage crosses the limit, via POST /api/alerts/rules on the Agent Monitor dashboard. | hoangsonww/ | 1.1k | — | ~1k | Automated safety check: Pass | MIT | yesterday |
| 31 | Tune and review Langfuse autoscaling for web, web-iso, and web-ingestion. | langfuse/ | 36k | — | ~4.4k | Automated safety check: Pass | Unknown | today |
| 32 | Naming conventions for SGLang speculative decoding identifiers. | sgl-project/ | 37k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 33 | Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost. | Orchestra-Research/ | 13k | 4 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 34 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 35 | Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change. | google/ | 21k | — | ~5k | Automated safety check: Pass | Apache-2.0 | today |
| 36 | Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package. | adithya-s-k/ | 456 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 37 | Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech. | qdrant/ | 254 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 38 | Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK. | datadog-labs/ | 177 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 39 | 39.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | today |
| 40 | Create a new Agent Skill following project standards and templates. | oocx/ | 174 | — | ~1.5k | Automated safety check: Pass | MIT | 2 days ago |
| 41 | 41.Verify Verify a Txtify change end-to-end. An agent skill from lkmeta/txtify. | lkmeta/ | 135 | — | ~583 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 42 | 42.Inspect Inspect and debug live streaming agent sessions to understand what the agent did. | agentevals-dev/ | 162 | — | ~534 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 43 | OpenVINO — real-time object detection via Docker (NCS2, Intel GPU, CPU) | SharpAI/ | 3.1k | — | ~1.3k | Automated safety check: Pass | MIT | 23 days ago |
| 44 | LLM observability platform for tracing, evaluation, and monitoring. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 45 | 45.Release Used to release all toolbox, integrations, agents. An agent skill from memgraph/ai-toolkit. | memgraph/ | 114 | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 46 | 基于 LoongSuite Pilot / AI Coding Agent 日志生成事件洞察、组织洞察、数据质量、研发效能和 AI Native 使用类 SLS 报表时使用;包含 AI Coding 事件表语义,以及团队报表可选的部门维表、deptuser 组织关系、指标口径和公共 CTE,通常与 sls-dashboard-builder 一起使用。 | alibaba/ | 200 | — | ~944 | Automated safety check: Pass | Apache-2.0 | today |
| 47 | When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks. | lyonzin/ | 292 | — | ~1.8k | Automated safety check: Pass | MIT | 5 days ago |
| 48 | INVOKE THIS SKILL when adding Arize AX tracing or observability to an app for the first time, or when the user wants to instrument their LLM app or get started with LLM observability. | boshi-xixixi/ | 275 | — | ~5.1k | Automated safety check: Notes | MIT | 5 mo ago |