Agent skill

LLM Aiops Guide

by wentorai in wentorai/research-plugins

Papers on LLMs for IT operations and AIOps research. An agent skill from wentorai/research-plugins.

MITAuto-check passedAI & LLM Engineering

Install LLM Aiops Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill llm-aiops-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins llm-aiops-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/domains/cs/llm-aiops-guide .claude/skills/llm-aiops-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-aiops-guide
GitHub stars
298
Used in
1 other repo
Token cost
~3.3k tokens
SKILL.md length
486 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Papers on LLMs for IT operations and AIOps research. An agent skill from wentorai/research-plugins.

  • Works in 5 steps: Reproducibility: Every model version… → Evaluation-first: Define evaluation… → Gradual rollout: Never switch 100%… → …
  • Tasks that involve MLOps
  • SKILL.md covers Overview, Research Areas, Key Practices for LLM Operations and Toolchain Overview, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

LLM Aiops Guide is an agent skill from wentorai/research-plugins. Papers on LLMs for IT operations and AIOps research

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering MLOps and LLM inference and serving. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve MLOps
  • Tasks that involve LLM inference and serving

Example prompts

  • “Use the llm-aiops-guide skill to paper on LLMs for IT operations and AIOps research. An agent skill from wentorai/research-plugins”
  • “/llm-aiops-guide”

Requirements

  • Docker

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Reproducibility: Every model version must be traceable to its training data, hyperparameters, and base model. Use deterministic seeds and…
  2. Evaluation-first: Define evaluation criteria before training. Include both automated metrics and human evaluation protocols.
  3. Gradual rollout: Never switch 100% traffic to a new model instantly. Use canary deployments (1% -> 10% -> 50% -> 100%) with automatic…
  4. Feedback loops: Collect user feedback (explicit thumbs up/down, implicit engagement metrics) and route it back to evaluation datasets.
  5. Safety gates: Automated checks for toxic output, PII leakage, and prompt injection before any model promotion.

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org
    • github.com
    • docs.smith.langchain.com
    • mlflow.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Aiops Guide loads about 3.3k tokens when it runs. Until then it costs about 17 tokens; SKILL.md has 486 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~17
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 486 words, ~3,320 tokens.

Download SKILL.mdSave it as .claude/skills/llm-aiops-guide/SKILL.md (or your agent's skills folder).
name
llm-aiops-guide
description
Papers on LLMs for IT operations and AIOps research

LLM for AIOps Guide

Overview

A curated collection of research on applying LLMs to IT Operations (AIOps) — log analysis, anomaly detection, incident management, root cause analysis, and automated remediation. Tracks how foundation models are transforming traditional rule-based operations tooling into intelligent, adaptive systems. Relevant for CS researchers at the intersection of systems, NLP, and operations.

Research Areas

LLM for AIOps
├── Log Analysis
│   ├── Log parsing (template extraction)
│   ├── Anomaly detection (from log sequences)
│   ├── Log summarization
│   └── Root cause from logs
├── Incident Management
│   ├── Incident triage and routing
│   ├── Severity classification
│   ├── Similar incident retrieval
│   └── Resolution recommendation
├── Root Cause Analysis
│   ├── Topology-aware diagnosis
│   ├── Multi-signal correlation
│   └── Causal inference
├── Monitoring & Alerting
│   ├── Metric anomaly detection
│   ├── Alert correlation
│   ├── Noise reduction
│   └── Capacity planning
└── Automated Remediation
    ├── Runbook generation
    ├── Script generation
    ├── Self-healing systems
    └── Change impact analysis

Key Practices for LLM Operations

Model Monitoring
Production LLM monitoring dimensions:

QUALITY MONITORING
- Output quality scores: automated evaluation (LLM-as-judge, BERTScore, ROUGE)
- Hallucination rate: factual grounding checks against retrieval context
- Refusal rate: track over-cautious or under-cautious safety filters
- Latency percentiles: p50, p95, p99 for time-to-first-token and total generation
- Token usage: input/output token distributions, context window utilization

DRIFT DETECTION
- Input drift: embedding-space distribution shift (cosine distance, MMD)
- Output drift: topic/style distribution changes over time windows
- Performance drift: sliding-window accuracy on held-out evaluation sets
- Concept drift: monitor for domain vocabulary shifts in user queries
- Baseline comparison: periodically re-evaluate against golden test suites

OPERATIONAL HEALTH
- GPU utilization and memory pressure (per-device, per-replica)
- Request queue depth and timeout rates
- Cache hit rates (KV cache, semantic cache, prompt cache)
- Error rates by error category (OOM, context overflow, timeout, malformed output)
- Throughput: tokens/second per deployment, requests/minute
A/B Testing for LLMs
Designing valid A/B tests for LLM systems:

CHALLENGES UNIQUE TO LLMs
- High output variance: same prompt can produce different outputs
- Evaluation subjectivity: many tasks lack clear ground truth
- Latency-quality tradeoff: larger models are better but slower
- Cost confound: better model may cost 10x more per query

RECOMMENDED APPROACH
1. Define metrics BEFORE experiment:
   - Primary: task-specific quality (accuracy, user satisfaction, resolution rate)
   - Secondary: latency, cost per query, token efficiency
   - Guardrail: safety violations, hallucination rate

2. Traffic splitting strategy:
   - User-level randomization (not request-level) to avoid confusion
   - Minimum 1-2 weeks for stable estimates
   - Stratify by user segment (power users vs. new users)

3. Evaluation methods:
   - Automated scoring with LLM-as-judge (calibrated against human raters)
   - Blind human evaluation on sampled outputs (inter-rater agreement > 0.7)
   - Downstream business metrics (ticket resolution time, user retention)

4. Statistical rigor:
   - Bootstrap confidence intervals for LLM quality scores
   - Account for multiple comparisons when testing many variants
   - Report effect sizes, not just p-values

Toolchain Overview

Experiment Tracking and Model Registry
ToolFocusKey Capabilities
MLflowEnd-to-end ML lifecycleExperiment tracking, model registry, deployment, LLM evaluation
Weights & BiasesExperiment tracking + LLM monitoringTraces, prompt versioning, evaluation tables, sweeps
LangSmithLLM application observabilityTrace visualization, prompt playground, dataset management, online evaluation
Comet MLExperiment managementModel comparison, artifact tracking, LLM prompt tracking
Serving and Inference
ToolFocusKey Capabilities
vLLMHigh-throughput servingPagedAttention, continuous batching, tensor parallelism, speculative decoding
TGI (Text Generation Inference)Production servingQuantization, streaming, multi-LoRA, watermarking
OllamaLocal model runningEasy setup, model library, OpenAI-compatible API
TensorRT-LLMNVIDIA-optimized inferenceFP8 quantization, in-flight batching, custom kernels
SGLangStructured generation servingRadixAttention, constrained decoding, multi-modal support
Orchestration and Pipelines
ToolFocusKey Capabilities
LangChain / LangGraphLLM application frameworkChains, agents, tool use, stateful multi-actor workflows
HaystackNLP pipeline frameworkRAG pipelines, document processing, evaluation
Prefect / AirflowWorkflow orchestrationDAG scheduling, retry logic, observability
Ray ServeDistributed servingAuto-scaling, multi-model composition, batch inference

Typical LLMOps Pipeline Architecture

End-to-end LLMOps pipeline:

┌─────────────────────────────────────────────────────────────────┐
│                     DATA PREPARATION                            │
│  Raw data → Cleaning → Annotation → Train/Eval split            │
│  Tools: Label Studio, Argilla, Lilac, DVC                       │
└──────────────────────────┬──────────────────────────────────────┘
                           ▼
┌─────────────────────────────────────────────────────────────────┐
│                   MODEL DEVELOPMENT                             │
│  Base model selection → Fine-tuning (LoRA/QLoRA) → Evaluation   │
│  Tools: Hugging Face Transformers, Axolotl, LLaMA-Factory       │
│  Eval: lm-evaluation-harness, HELM, custom domain benchmarks    │
└──────────────────────────┬──────────────────────────────────────┘
                           ▼
┌─────────────────────────────────────────────────────────────────┐
│                   MODEL REGISTRY & CI                           │
│  Version control → Automated testing → Approval gates           │
│  Tools: MLflow Registry, W&B Model Registry, HF Hub             │
│  Tests: regression suite, safety checks, latency benchmarks     │
└──────────────────────────┬──────────────────────────────────────┘
                           ▼
┌─────────────────────────────────────────────────────────────────┐
│                   DEPLOYMENT                                    │
│  Quantization → Containerization → Canary rollout → Full deploy │
│  Tools: vLLM, TGI, Docker, Kubernetes, Terraform                │
│  Strategy: blue-green or canary with automatic rollback          │
└──────────────────────────┬──────────────────────────────────────┘
                           ▼
┌─────────────────────────────────────────────────────────────────┐
│                   PRODUCTION MONITORING                         │
│  Quality monitoring → Drift detection → Alerting → Feedback     │
│  Tools: LangSmith, W&B Weave, Prometheus + Grafana, PagerDuty  │
│  Loop: degradation detected → trigger re-evaluation → retrain   │
└─────────────────────────────────────────────────────────────────┘
Pipeline Design Principles
  1. Reproducibility: Every model version must be traceable to its training data, hyperparameters, and base model. Use deterministic seeds and pin library versions.
  2. Evaluation-first: Define evaluation criteria before training. Include both automated metrics and human evaluation protocols.
  3. Gradual rollout: Never switch 100% traffic to a new model instantly. Use canary deployments (1% -> 10% -> 50% -> 100%) with automatic rollback on quality regression.
  4. Feedback loops: Collect user feedback (explicit thumbs up/down, implicit engagement metrics) and route it back to evaluation datasets.
  5. Safety gates: Automated checks for toxic output, PII leakage, and prompt injection before any model promotion.
Show full SKILL.md (154 more words)Show less

Cost Optimization Strategies

Quantization
Reducing model size and inference cost:

QUANTIZATION METHODS
- GPTQ: Post-training quantization, good quality at 4-bit, widely supported
- AWQ (Activation-aware Weight Quantization): Better quality than GPTQ at 4-bit
- GGUF: CPU-friendly format, variable bit-width (Q4_K_M, Q5_K_M, Q8_0)
- FP8: NVIDIA H100/B200 native, minimal quality loss, 2x throughput vs FP16
- AQLM: Additive quantization, state-of-the-art at 2-bit

PRACTICAL GUIDANCE
- 8-bit: negligible quality loss for most tasks (~0.1% accuracy drop)
- 4-bit: slight quality loss, acceptable for many production uses (~1-3% accuracy drop)
- 2-3 bit: noticeable degradation, use only when cost is critical
- Always evaluate on YOUR task after quantization (general benchmarks can be misleading)
- Combine quantization with speculative decoding for further speedup
Caching Strategies
Multi-layer caching for LLM systems:

EXACT MATCH CACHE
- Hash the full prompt, return cached response for identical queries
- Hit rate: typically 5-15% for general-purpose, 30-60% for structured queries
- Tools: Redis, DragonflyDB, in-memory LRU

SEMANTIC CACHE
- Embed the prompt, return cached response for semantically similar queries
- Similarity threshold: 0.95+ cosine similarity (tune per use case)
- Tools: GPTCache, Redis with vector search, Qdrant
- Risk: semantically similar prompts may require different answers

KV CACHE OPTIMIZATION
- PagedAttention (vLLM): eliminates memory waste from pre-allocated KV cache
- Prefix caching: reuse KV cache for shared system prompts across requests
- Quantized KV cache: FP8 or INT8 KV values (H100+, ~2x context capacity)

PROMPT CACHING (API providers)
- Anthropic prompt caching: cache static prefix, pay reduced rate for cached tokens
- OpenAI cached context: automatic for repeated prefixes
- Design prompts with static prefix (system prompt, examples) + dynamic suffix (user query)
Intelligent Routing
Cost-quality optimization through model routing:

TIERED MODEL ROUTING
- Simple queries → small/fast model (e.g., GPT-4o-mini, Claude Haiku, Llama-8B)
- Complex queries → large/capable model (e.g., GPT-4o, Claude Sonnet, Llama-70B)
- Critical queries → frontier model (e.g., o3, Claude Opus)

ROUTING STRATEGIES
1. Classifier-based: Train a small classifier on query complexity
   - Features: query length, vocabulary complexity, domain signals
   - Labels: which model tier produces acceptable quality
   - Cost: classifier inference is negligible (<1ms, <$0.001)

2. Cascade (try-small-first):
   - Route to cheapest model first
   - Check output quality with a verifier
   - Escalate to larger model if quality is insufficient
   - Effective when >50% of queries are simple

3. Task-based routing:
   - Summarization, translation → mid-tier model
   - Code generation, math reasoning → high-tier model
   - Classification, extraction → small model or fine-tuned specialist

EXPECTED SAVINGS
- Typical 40-70% cost reduction vs. routing everything to the best model
- Quality degradation: <5% when routing thresholds are properly calibrated

Key Papers

PaperYearFocus
LogPPT2023Few-shot log parsing with prompt tuning
OpsEval2024Benchmark for evaluating LLMs in AIOps
D-Bot2024LLM-based database diagnosis
RCAgent2024Agent for root cause analysis
LogAgent2024Autonomous log analysis agent
AIOpsLab2024Holistic benchmark suite for AIOps agents
MonitorAssistant2024LLM-based alert correlation and noise reduction
LLM4Ops Survey2024Comprehensive survey of LLMs for IT operations

Use Cases

  1. Literature tracking: Follow LLM-AIOps research evolution
  2. System design: Learn intelligent operations patterns
  3. Benchmark comparison: Evaluate AIOps approaches
  4. Research planning: Identify under-explored AIOps problems
  5. Industry applications: Bridge research to production AIOps
  6. Cost modeling: Design cost-efficient LLM serving architectures
  7. Pipeline design: Architect end-to-end LLMOps workflows

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/domains/cs/llm-aiops-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LLM Aiops Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Aiops Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Aiops Guide this skillwentorai/research-plugins2981 repos~3.3kAutomated safety check: PassMIT
Databricks ML Trainingdatabricks/databricks-agent-skills345—~4.6kAutomated safety check: PassCustom licence
ML System Design Interviewcuriositech/some_claude_skills243—~3.4kAutomated safety check: PassMIT
Tensorrt LLMLuciole-Studio/Misaka-Agent1581 repos~1.3kAutomated safety check: PassMIT
AWS AI MLaws/agent-toolkit-for-aws2.8k—~1.7kAutomated safety check: PassApache-2.0
SageMaker Production Defaultshuggingface/skills11k1 repos~6.9kAutomated safety check: PassApache-2.0

Similar skills

  • Databricks ML Training

    databricks/databricks-agent-skills

    Official

    Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills.

    345 GitHub stars~4.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • ML System Design Interview

    curiositech/some_claude_skills

    Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring.

    243 GitHub stars~3.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Tensorrt LLM

    Luciole-Studio/Misaka-Agent

    High-throughput LLM inference on NVIDIA GPUs. An agent skill from Luciole-Studio/Misaka-Agent.

    158 GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • AWS AI ML

    aws/agent-toolkit-for-aws

    Official

    Selects, deploys, and customizes AI models on Amazon SageMaker.

    2.8k GitHub stars~1.7k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about LLM Aiops Guide

What does LLM Aiops Guide do?

Papers on LLMs for IT operations and AIOps research. An agent skill from wentorai/research-plugins. LLM Aiops Guide is an agent skill from wentorai/research-plugins.

When should I use LLM Aiops Guide?

LLM Aiops Guide fits situations like: tasks that involve MLOps; tasks that involve LLM inference and serving.

How do I install LLM Aiops Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill llm-aiops-guide -a claude-code`. Or copy the skill folder (skills/domains/cs/llm-aiops-guide in wentorai/research-plugins) into .claude/skills/llm-aiops-guide in your project. Claude Code loads it when a task matches its description.

How do I install LLM Aiops Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill llm-aiops-guide -a codex`. Or copy the skill folder (skills/domains/cs/llm-aiops-guide in wentorai/research-plugins) into .agents/skills/llm-aiops-guide in your project. Codex loads it when a task matches its description.

Can I use LLM Aiops Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill llm-aiops-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-aiops-guide, .gemini/skills/llm-aiops-guide, .github/skills/llm-aiops-guide and .opencode/skills/llm-aiops-guide in your project.

What does LLM Aiops Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: LLM Aiops Guide is instructions for the agent only. Our summary lists: Docker.

Does LLM Aiops Guide access the network?

SKILL.md names 4 domains. As links in the text: arxiv.org, github.com, docs.smith.langchain.com and mlflow.org. This is read from the text; nothing was executed.

Is LLM Aiops Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Aiops Guide use?

LLM Aiops Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Aiops Guide use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Aiops Guide?

Skills that share tags, products or a category with LLM Aiops Guide: Databricks ML Training (databricks/databricks-agent-skills, 345 stars), ML System Design Interview (curiositech/some_claude_skills, 243 stars), Tensorrt LLM (Luciole-Studio/Misaka-Agent, 158 stars) and AWS AI ML (aws/agent-toolkit-for-aws, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Aiops Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.