Agent skill

LLM Inference Scaling

by sickn33 in sickn33/agentic-awesome-skills

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

MITAuto-check passedAI & LLM Engineering

Install LLM Inference Scaling

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill llm-inference-scaling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills llm-inference-scaling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-inference-scaling .claude/skills/llm-inference-scaling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-inference-scaling
GitHub stars
47k
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
313 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

  • Tasks that involve LLM inference and serving
  • SKILL.md covers When to Use This Skill, Prerequisites, GPU Node Setup and vLLM Deployment with GPU…, plus 9 more sections
  • Calls helm and kubectl; reaches prometheus-server.monitoring and helm.ngc.nvidia.com; needs HUGGING_FACE_HUB_TOKEN
  • Tasks that involve Container orchestration

What it does

LLM Inference Scaling is an agent skill from sickn33/agentic-awesome-skills. Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

It sits in AI & LLM Engineering, covering LLM inference and serving and Container orchestration. It works with Kubernetes, vLLM and Prometheus. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve Container orchestration

Example prompts

  • “/llm-inference-scaling”

Requirements

  • A credential in HUGGING_FACE_HUB_TOKEN
  • Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • helm
    • kubectl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • prometheus-server.monitoring
    • helm.ngc.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HUGGING_FACE_HUB_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.

    From compatibility in the SKILL.md frontmatter.

Context cost

LLM Inference Scaling loads about 2.1k tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 313 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 313 words, ~2,087 tokens.

Download SKILL.mdSave it as .claude/skills/llm-inference-scaling/SKILL.md (or your agent's skills folder).
name
llm-inference-scaling
description
Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.
compatibility
Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled.
category
devops
risk
critical
source
https://github.com/BagelHole/DevOps-Security-Agent-Skills
source_repo
BagelHole/DevOps-Security-Agent-Skills
source_type
community
date_added
2026-09-20
license
MIT
license_source
https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
metadata.author
devops-skills
metadata.version
1.0

LLM Inference Scaling

Scale LLM inference horizontally on Kubernetes with GPU-aware autoscaling, request queuing, and cost-efficient spot instance strategies.

When to Use This Skill

Use this skill when:

  • LLM API traffic is unpredictable and you need to scale up/down automatically
  • Managing a fleet of vLLM or TGI inference pods on Kubernetes
  • Reducing inference costs with spot/preemptible GPU instances
  • Implementing queue-based autoscaling for batch inference jobs
  • Building a multi-model serving platform that shares GPU resources

Prerequisites

  • Kubernetes cluster with GPU nodes (NVIDIA operator installed)
  • KEDA (Kubernetes Event-Driven Autoscaler) installed
  • Prometheus with GPU metrics (dcgm-exporter or gpu-operator)
  • Helm 3+ for chart deployments

GPU Node Setup

bash
# Install NVIDIA GPU Operator (handles drivers, container toolkit, DCGM)
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update

helm install gpu-operator nvidia/gpu-operator \
  --namespace gpu-operator \
  --create-namespace \
  --set driver.enabled=true \
  --set dcgm.enabled=true \
  --set devicePlugin.enabled=true

# Verify GPU nodes are recognized
kubectl get nodes -l nvidia.com/gpu.present=true
kubectl describe node <gpu-node> | grep nvidia

vLLM Deployment with GPU Resources

yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-llama-8b
  labels:
    app: vllm
    model: llama-3.1-8b
spec:
  replicas: 1
  selector:
    matchLabels:
      app: vllm
      model: llama-3.1-8b
  template:
    metadata:
      labels:
        app: vllm
        model: llama-3.1-8b
    spec:
      nodeSelector:
        nvidia.com/gpu.present: "true"
      tolerations:
      - key: nvidia.com/gpu
        operator: Exists
        effect: NoSchedule
      containers:
      - name: vllm
        image: vllm/vllm-openai:latest
        args:
        - "--model"
        - "meta-llama/Llama-3.1-8B-Instruct"
        - "--tensor-parallel-size"
        - "1"
        - "--gpu-memory-utilization"
        - "0.90"
        - "--max-num-seqs"
        - "128"
        resources:
          requests:
            nvidia.com/gpu: "1"
            memory: "20Gi"
            cpu: "4"
          limits:
            nvidia.com/gpu: "1"
            memory: "24Gi"
            cpu: "8"
        ports:
        - containerPort: 8000
        readinessProbe:
          httpGet:
            path: /health
            port: 8000
          initialDelaySeconds: 60
          periodSeconds: 10
        env:
        - name: HUGGING_FACE_HUB_TOKEN
          valueFrom:
            secretKeyRef:
              name: hf-token
              key: token

KEDA Autoscaling on Prometheus Metrics

yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: vllm-scaledobject
spec:
  scaleTargetRef:
    name: vllm-llama-8b
  minReplicaCount: 1
  maxReplicaCount: 8
  cooldownPeriod: 300          # 5 min before scale-down
  pollingInterval: 15
  triggers:
  - type: prometheus
    metadata:
      serverAddress: http://prometheus-server.monitoring:9090
      metricName: vllm_num_requests_waiting
      threshold: "10"           # scale up if >10 requests waiting
      query: |
        sum(vllm:num_requests_waiting{deployment="vllm-llama-8b"})
  - type: prometheus
    metadata:
      serverAddress: http://prometheus-server.monitoring:9090
      metricName: vllm_gpu_cache_usage
      threshold: "0.8"          # scale up if KV cache >80% full
      query: |
        avg(vllm:gpu_cache_usage_perc{deployment="vllm-llama-8b"})

Queue-Based Scaling (Redis + KEDA)

yaml
# ScaledJob for async batch inference
apiVersion: keda.sh/v1alpha1
kind: ScaledJob
metadata:
  name: llm-batch-inference
spec:
  jobTargetRef:
    template:
      spec:
        containers:
        - name: inference-worker
          image: myapp/inference-worker:latest
          env:
          - name: REDIS_URL
            value: redis://redis:6379
          - name: QUEUE_NAME
            value: inference-jobs
        restartPolicy: OnFailure
  minReplicaCount: 0
  maxReplicaCount: 20
  pollingInterval: 5
  successfulJobsHistoryLimit: 3
  triggers:
  - type: redis
    metadata:
      address: redis:6379
      listName: inference-jobs
      listLength: "5"           # 1 worker per 5 queued jobs

Spot Instance Strategy

yaml
# Mixed node pool: on-demand + spot GPUs
apiVersion: v1
kind: ConfigMap
metadata:
  name: cluster-autoscaler-priority-config
data:
  priorities: |
    10:  # low priority = prefer
    - .*spot.*
    50:
    - .*on-demand.*
---
# Node affinity for spot with on-demand fallback
spec:
  affinity:
    nodeAffinity:
      preferredDuringSchedulingIgnoredDuringExecution:
      - weight: 80
        preference:
          matchExpressions:
          - key: node.kubernetes.io/lifecycle
            operator: In
            values: [spot]
      - weight: 20
        preference:
          matchExpressions:
          - key: node.kubernetes.io/lifecycle
            operator: In
            values: [on-demand]

Cluster Autoscaler for GPU Nodes

bash
# AWS EKS — enable cluster autoscaler for GPU node group
helm install cluster-autoscaler autoscaler/cluster-autoscaler \
  --namespace kube-system \
  --set autoDiscovery.clusterName=my-cluster \
  --set awsRegion=us-east-1 \
  --set rbac.serviceAccount.annotations."eks\.amazonaws\.com/role-arn"=arn:aws:iam::ACCOUNT:role/ClusterAutoscalerRole \
  --set extraArgs.skip-nodes-with-local-storage=false \
  --set extraArgs.expander=least-waste

# Annotate GPU node group for autoscaler
kubectl annotate node <node> \
  cluster-autoscaler.kubernetes.io/safe-to-evict="false"

Scaling Metrics to Monitor

bash
# Prometheus queries for scaling decisions
# Requests waiting in vLLM queue
sum(vllm:num_requests_waiting) by (model)

# GPU KV cache utilization (>80% = bottleneck)
avg(vllm:gpu_cache_usage_perc) by (pod)

# Tokens per second throughput
sum(rate(vllm:generation_tokens_total[5m])) by (model)

# P99 time-to-first-token
histogram_quantile(0.99, rate(vllm:time_to_first_token_seconds_bucket[5m]))

Common Issues

IssueCauseFix
Pods stuck in PendingNo GPU nodes availableCheck cluster autoscaler logs; verify node group limits
Scale-up too slowCluster autoscaler delay + model load timePre-warm replicas; increase minReplicaCount
GPU fragmentationMultiple small models on large GPUsUse MIG partitioning or consolidate model sizes
Spot eviction causes errorsSpot instance reclamationAdd PodDisruptionBudget; use graceful shutdown
KEDA not scalingPrometheus query returns no dataTest query in Prometheus UI first

Best Practices

  • Set minReplicaCount: 1 to avoid cold starts; scale to 0 only for batch jobs.
  • Use PodDisruptionBudget with minAvailable: 1 to survive spot evictions.
  • Pre-pull model weights into a shared PVC to speed up pod startup by 5–10×.
  • Separate model families across node pools (A10G for 7B, A100 for 70B).
  • Use Kubernetes VPA for CPU/memory right-sizing alongside KEDA for replica count.
  • vllm-server (vllm-server) - vLLM configuration and tuning
  • gpu-server-management (gpu-server-management) - GPU node setup
  • model-serving-kubernetes (model-serving-kubernetes) - KServe
  • kubernetes-ops (kubernetes-ops) - Core Kubernetes
  • llm-cost-optimization (llm-cost-optimization) - Cost strategies

Limitations

  • Infrastructure commands can disrupt services: confirm target host/scope and have backups/snapshots before mutating state.
  • Docs-only import: upstream scripts and templates not bundled.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/llm-inference-scaling of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LLM Inference Scaling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Inference Scaling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Inference Scaling this skillsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
LLM Inference ScalingBagelHole/DevOps-Security-Agent-Skills1.2k—~2kAutomated safety check: PassMIT
Vllm Deploy K8svllm-project/vllm-skills102—~2kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Eks Best Practicesaws-samples/appmod-blueprints115—~5kAutomated safety check: PassMIT-0
Gke Manifest Generationgoogle/skills21k—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Inference Scaling

    BagelHole/DevOps-Security-Agent-Skills

    Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

    1.2k GitHub stars~2k tokensUpdated 4 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    102 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eks Best Practices

    aws-samples/appmod-blueprints

    Official

    Advisory guidance for Amazon EKS architecture and configuration decisions — compute strategy, networking, security, reliability, cost, autoscaling, observability, multi-tenancy, and upgrade planning.

    115 GitHub stars~5k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Official

    Generates and updates secure, production-ready Kubernetes YAML manifests optimized for GKE Autopilot and GKE Standard clusters.

    21k GitHub stars~3.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 5 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about LLM Inference Scaling

What does LLM Inference Scaling do?

Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. LLM Inference Scaling is an agent skill from sickn33/agentic-awesome-skills. Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling.

When should I use LLM Inference Scaling?

LLM Inference Scaling fits situations like: tasks that involve LLM inference and serving; tasks that involve Container orchestration.

How do I install LLM Inference Scaling in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill llm-inference-scaling -a claude-code`. Or copy the skill folder (skills/llm-inference-scaling in sickn33/agentic-awesome-skills) into .claude/skills/llm-inference-scaling in your project. Claude Code loads it when a task matches its description.

How do I install LLM Inference Scaling in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill llm-inference-scaling -a codex`. Or copy the skill folder (skills/llm-inference-scaling in sickn33/agentic-awesome-skills) into .agents/skills/llm-inference-scaling in your project. Codex loads it when a task matches its description.

Can I use LLM Inference Scaling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill llm-inference-scaling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-inference-scaling, .gemini/skills/llm-inference-scaling, .github/skills/llm-inference-scaling and .opencode/skills/llm-inference-scaling in your project.

What does LLM Inference Scaling need to run?

Going by SKILL.md and its folder, LLM Inference Scaling needs the command-line tools its instructions call (helm and kubectl) and credentials named HUGGING_FACE_HUB_TOKEN. Our summary lists: A credential in HUGGING_FACE_HUB_TOKEN. Compatibility (from SKILL.md): Requires the relevant OS/platform tooling and privileged access where noted. Docs-only; helper scripts and templates not bundled..

Does LLM Inference Scaling access the network?

SKILL.md names 2 domains. In commands or code: prometheus-server.monitoring and helm.ngc.nvidia.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is LLM Inference Scaling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Inference Scaling use?

LLM Inference Scaling is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LLM Inference Scaling use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Inference Scaling?

Skills that share tags, products or a category with LLM Inference Scaling: LLM Inference Scaling (BagelHole/DevOps-Security-Agent-Skills, 1.2k stars), Vllm Deploy K8s (vllm-project/vllm-skills, 102 stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and Eks Best Practices (aws-samples/appmod-blueprints, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Inference Scaling?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.