Search

AI & LLM Engineering · For devops and sre engineers

188 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1
1.Agent LightningOfficial

Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract.

microsoft/agent-lightning19k—~1.7kAutomated safety check: PassMITtoday
2

Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.

NVIDIA/Megatron-LM18k—~2.8kAutomated safety check: PassApache-2.0today
3

Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.

JuliusBrussee/caveman111k1 repo~2.6kAutomated safety check: WarnApache-2.0yesterday
4

Decide, don't guess — trigger on ANY combinatorial or ground-state decision where a plausible guess is worse than silence: rosters and on-call schedules, packing and placement, RAG passage…

brayonpi/hexstellar1.4k—~5kAutomated safety check: PassProprietary1 mo ago
5

Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

huggingface/skills11k1 repo~6.9kAutomated safety check: PassApache-2.0yesterday
6

Integrate a benchmark or custom environment into SAfactory using fixed adapter templates and local contract tests, optionally run Docker/RJob evaluation, or prepare GRPO/RL training.

AI45Lab/SAfactory236—~1.8kAutomated safety check: PassNo licence16 days ago
7

Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

BBuf/AI-Infra-Auto-Driven-SKILLS925—~2.5kAutomated safety check: PassNo licence4 days ago
8

Interact with Langfuse and access its documentation: tracing, monitoring, creating datasets, running experiments, and evaluating AI applications.

langfuse/skills300—~2.1kAutomated safety check: NotesMIT8 days ago
9

Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.

inclusionAI/AReno323—~486Automated safety check: PassApache-2.0today
10

Starts, rebuilds, and troubleshoots the Opik local dev stack, including an optional Comet Platform integration mode for the Opik team.

comet-ml/opik22k—~734Automated safety check: PassApache-2.0today
11

Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~8.6kAutomated safety check: NotesApache-2.0today
12

Give the human an off switch and a cost meter for the coding agents on this machine, using ClawMetry.

vivekchand/clawmetry426—~1.1kAutomated safety check: PassMITyesterday
13

Interactive scaffold generator for Orloj multi-agent systems.

OrlojHQ/orloj123—~2.6kAutomated safety check: PassApache-2.021 days ago
14

Creates an azd environment, checks RBAC and model quota, provisions an AI agent app on Azure with azd up and health-checks the deployed app.

Azure-Samples/get-started-with-ai-agents374—~4.7kAutomated safety check: NotesMITyesterday
15

当任务需要创建、修改、扩展或重组阿里云 SLS 的 dashboard JSON 或可导入的大盘配置时使用;尤其适用于线上大盘、强对比的分析看板、已校验的查询包,或需要专业中文标签与指标定义的运维向大盘。

alibaba/loongsuite-pilot200—~1.8kAutomated safety check: PassApache-2.0today
16

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
17

Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs.

adithya-s-k/FineEnvs456—~2.1kAutomated safety check: PassApache-2.0yesterday
18

Diagnose and fix Cosmos3 environment, installation, and runtime errors.

NVIDIA/cosmos-framework559—~1.3kAutomated safety check: NotesUnknowntoday
19

This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a.

brevdev/workshop-build-an-agent146—~2.3kAutomated safety check: NotesApache-2.0yesterday
20

Read your own agent telemetry from ClawMetry (waste, progress, cost) and act on it before finishing a task.

vivekchand/clawmetry426—~515Automated safety check: PassMITyesterday
21

Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

dstackai/dstack2.3k—~403Automated safety check: PassMPL-2.0today
22
22.Eval

Evaluate and score agent behavior against a golden reference.

agentevals-dev/agentevals162—~904Automated safety check: PassApache-2.0yesterday
23

Propose an improved version of a prompt registered in a self-hosted AgentX (AgentX-trace-eval) instance, using real low-rated evaluation results as evidence, then publish it as a new version once…

AgentX-ai/AgentX-Trace-Eval106—~2kAutomated safety check: PassUnknown2 days ago
24

Writes Terraform alerting policies for AI agents that emit OpenTelemetry metrics, covering reliability, cost, safety, security and quality signals on Google Cloud.

google/skills21k—~4.2kAutomated safety check: PassApache-2.0today
25

Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues.

climate-analytics-lab/jax-gcm108—~2.5kAutomated safety check: PassApache-2.0yesterday
26

Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT3 mo ago
27

Add new LLM model pricing entries to Litefuse's default-model-prices.json.

litefuse/litefuse100—~3.6kAutomated safety check: PassUnknowntoday
28

Runs OpenAI Codex CLI as a non-interactive worker for CI, Docker, Kubernetes or remote servers, with sandbox modes and JSONL-friendly output.

XiaomiMiMo/MiMo-Code14k—~2.7kAutomated safety check: PassMITyesterday
29
29.Wa GuardrailsOfficial

Generate preventive Well-Architected guardrails — AWS Config rules, Service Control Policies, permission boundaries, CloudWatch alarms, and IaC policy checks (CDK Aspects, cfn-guard, OPA/Sentinel) —…

aws-samples/sample-well-architected-skills-and-steering273—~2.8kAutomated safety check: PassMIT-03 days ago
30

Define a spend budget for Claude Code and, optionally, create a cost alert rule that fires when usage crosses the limit, via POST /api/alerts/rules on the Agent Monitor dashboard.

hoangsonww/Claude-Code-Agent-Monitor1.1k—~1kAutomated safety check: PassMITyesterday
31

Tune and review Langfuse autoscaling for web, web-iso, and web-ingestion.

langfuse/langfuse36k—~4.4kAutomated safety check: PassUnknowntoday
32

Naming conventions for SGLang speculative decoding identifiers.

sgl-project/sglang37k2 repos~1.6kAutomated safety check: PassApache-2.0today
33

Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.4kAutomated safety check: PassMIT3 mo ago
34

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT3 mo ago
35

Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

google/skills21k—~5kAutomated safety check: PassApache-2.0today
36

Builds an Open Reward Standard (ORS) variant of an RL environment using the official openreward Python package.

adithya-s-k/FineEnvs456—~2.3kAutomated safety check: NotesApache-2.0yesterday
37
37.Qdrant AdvisorOfficial

Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech.

qdrant/skills254—~1.7kAutomated safety check: PassApache-2.0yesterday
38

Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK.

datadog-labs/agent-skills177—~2.3kAutomated safety check: PassMITyesterday
39

dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

dstackai/dstack2.3k—~6.2kAutomated safety check: WarnMPL-2.0today
40

Create a new Agent Skill following project standards and templates.

oocx/tfplan2md174—~1.5kAutomated safety check: PassMIT2 days ago
41

Verify a Txtify change end-to-end. An agent skill from lkmeta/txtify.

lkmeta/txtify135—~583Automated safety check: PassApache-2.01 mo ago
42

Inspect and debug live streaming agent sessions to understand what the agent did.

agentevals-dev/agentevals162—~534Automated safety check: PassApache-2.0yesterday
43

OpenVINO — real-time object detection via Docker (NCS2, Intel GPU, CPU)

SharpAI/DeepCamera3.1k—~1.3kAutomated safety check: PassMIT23 days ago
44

LLM observability platform for tracing, evaluation, and monitoring.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT3 mo ago
45

Used to release all toolbox, integrations, agents. An agent skill from memgraph/ai-toolkit.

memgraph/ai-toolkit114—~1.2kAutomated safety check: PassMITtoday
46

基于 LoongSuite Pilot / AI Coding Agent 日志生成事件洞察、组织洞察、数据质量、研发效能和 AI Native 使用类 SLS 报表时使用;包含 AI Coding 事件表语义,以及团队报表可选的部门维表、deptuser 组织关系、指标口径和公共 CTE,通常与 sls-dashboard-builder 一起使用。

alibaba/loongsuite-pilot200—~944Automated safety check: PassApache-2.0today
47

When the user reports a bug, error message, stack trace, unexpected behavior, or "why is this broken" question, search the corpus first for prior occurrences, known fixes, or related runbooks.

lyonzin/knowledge-rag292—~1.8kAutomated safety check: PassMIT5 days ago
48

INVOKE THIS SKILL when adding Arize AX tracing or observability to an app for the first time, or when the user wants to instrument their LLM app or get started with LLM observability.

boshi-xixixi/TraeSkill275—~5.1kAutomated safety check: NotesMIT5 mo ago