Agent skill

LLM Inference Batching Scheduler

by lazyFrogLOL in lazyFrogLOL/Harness_Engineering

Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators.

No licenceAuto-check passedAI & LLM Engineering

Install LLM Inference Batching Scheduler

skills CLI
$ npx skills add lazyFrogLOL/Harness_Engineering --skill llm-inference-batching-scheduler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lazyFrogLOL/Harness_Engineering llm-inference-batching-scheduler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lazyFrogLOL/Harness_Engineering.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-inference-batching-scheduler .claude/skills/llm-inference-batching-scheduler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-inference-batching-scheduler
GitHub stars
128
Token cost
~2.2k tokens
SKILL.md length
822 words
Files
1
Skills in repo
32
Repo updated
First seen
Licence
None found

At a glance

Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators.

  • Works in 3 steps: Mathematical Analysis (Before Any Code) → Systematic Parameter Search → Implementation with Invariant Checking
  • Tasks involving batch planning
  • SKILL.md covers When to Apply This Skill, Core Concepts, Systematic Approach and Common Pitfalls and Mitigations, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

LLM Inference Batching Scheduler is an agent skill from lazyFrogLOL/Harness_Engineering. Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators. This skill applies when optimizing request batching to minimize cost while meeting latency thresholds, particularly when dealing with shape compilation costs, padding overhead, and multi-bucket request distributions. Use this skill for tasks involving batch planning, shape selection, generation-length bucketing, and cost-model-driven optimization for neural network inference.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving and Deep learning.

When your agent uses it

  • Tasks involving batch planning
  • Shape selection
  • Generation-length bucketing
  • Cost-model-driven optimization for neural network inference

Example prompts

  • “/llm-inference-batching-scheduler”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Mathematical Analysis (Before Any Code)
  2. Systematic Parameter Search
  3. Implementation with Invariant Checking

What it can do on your machine

Read from SKILL.md and the folder at commit cae3b25. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Inference Batching Scheduler loads about 2.2k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 822 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 822 words (~2,202 tokens).

“This skill provides systematic approaches for designing batching schedulers that optimize LLM inference workloads on compilation-based accelerators (TPUs, custom ASICs). The core challenge involves balancing multiple competing objectives: minimizing compilation cost (fewer shapes), reducing padding waste (tighter batches), and meeting…”

— opening of SKILL.md by lazyFrogLOL
name
llm-inference-batching-scheduler

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/llm-inference-batching-scheduler of lazyFrogLOL/Harness_Engineering.

Open the folder on GitHubat commit cae3b25

Compare with similar skills

LLM Inference Batching Scheduler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Inference Batching Scheduler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Inference Batching Scheduler this skilllazyFrogLOL/Harness_Engineering128—~2.2kAutomated safety check: PassNone
Quark Torch Quant Perfamd/Quark182—~3kAutomated safety check: PassMIT
RWKV Architecture GuideOrchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT
ML Engineerdavila7/claude-code-templates33k9 repos~2.3kAutomated safety check: PassMIT
Databricks ML Trainingdatabricks/databricks-agent-skills345—~4.6kAutomated safety check: PassCustom licence
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0

Similar skills

  • Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models.

    182 GitHub stars~3k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • RWKV Architecture Guide

    Orchestra-Research/AI-Research-SKILLs

    Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting.

    13k GitHub starsUsed in 2 repos~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • ML Engineer

    davila7/claude-code-templates

    Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks.

    33k GitHub starsUsed in 9 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Databricks ML Training

    databricks/databricks-agent-skills

    Official

    Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills.

    345 GitHub stars~4.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Model Builder

    qualcomm/qai-appbuilder

    QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.

    247 GitHub stars~4.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from lazyFrogLOL/Harness_Engineering

All 32 skills in this repo
  • Chess Best Move

    lazyFrogLOL/Harness_Engineering

    Guide for analyzing chess positions from images and determining optimal moves.

    128 GitHub stars~1.6k tokensUpdated 4 mo ago
    Auto-check passed
  • Crack 7z Hash

    lazyFrogLOL/Harness_Engineering

    This skill provides guidance for cracking 7z archive password hashes.

    128 GitHub stars~1.3k tokensUpdated 4 mo ago
    Auto-check passed
  • Distribution Search

    lazyFrogLOL/Harness_Engineering

    Guidance for finding probability distributions that satisfy specific statistical constraints such as KL divergence targets, entropy requirements, or moment conditions.

    128 GitHub stars~2.6k tokensUpdated 4 mo ago
    Auto-check passed
  • Feal Linear Cryptanalysis

    lazyFrogLOL/Harness_Engineering

    This skill provides guidance for FEAL cipher linear cryptanalysis tasks.

    128 GitHub stars~1.7k tokensUpdated 4 mo ago
    Auto-check passed
  • Gcode To Text

    lazyFrogLOL/Harness_Engineering

    Decode and interpret text content from G-code files by analyzing toolpath geometry and coordinate patterns.

    128 GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Gpt2 Codegolf

    lazyFrogLOL/Harness_Engineering

    Guidance for implementing neural network inference (like GPT-2) under extreme code size constraints.

    128 GitHub stars~1.8k tokensUpdated 4 mo ago
    Auto-check passed

Questions about LLM Inference Batching Scheduler

What does LLM Inference Batching Scheduler do?

Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators. LLM Inference Batching Scheduler is an agent skill from lazyFrogLOL/Harness_Engineering. Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators.

When should I use LLM Inference Batching Scheduler?

LLM Inference Batching Scheduler fits situations like: tasks involving batch planning; shape selection; generation-length bucketing; cost-model-driven optimization for neural network inference.

How do I install LLM Inference Batching Scheduler in Claude Code?

Run `npx skills add lazyFrogLOL/Harness_Engineering --skill llm-inference-batching-scheduler -a claude-code`. Or copy the skill folder (skills/llm-inference-batching-scheduler in lazyFrogLOL/Harness_Engineering) into .claude/skills/llm-inference-batching-scheduler in your project. Claude Code loads it when a task matches its description.

How do I install LLM Inference Batching Scheduler in Codex?

Run `npx skills add lazyFrogLOL/Harness_Engineering --skill llm-inference-batching-scheduler -a codex`. Or copy the skill folder (skills/llm-inference-batching-scheduler in lazyFrogLOL/Harness_Engineering) into .agents/skills/llm-inference-batching-scheduler in your project. Codex loads it when a task matches its description.

Can I use LLM Inference Batching Scheduler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lazyFrogLOL/Harness_Engineering --skill llm-inference-batching-scheduler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-inference-batching-scheduler, .gemini/skills/llm-inference-batching-scheduler, .github/skills/llm-inference-batching-scheduler and .opencode/skills/llm-inference-batching-scheduler in your project.

What does LLM Inference Batching Scheduler need to run?

SKILL.md names no scripts, command-line tools or credentials: LLM Inference Batching Scheduler is instructions for the agent only.

Does LLM Inference Batching Scheduler access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is LLM Inference Batching Scheduler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LLM Inference Batching Scheduler use?

No licence was found for LLM Inference Batching Scheduler or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does LLM Inference Batching Scheduler use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Inference Batching Scheduler?

Skills that share tags, products or a category with LLM Inference Batching Scheduler: Quark Torch Quant Perf (amd/Quark, 182 stars), RWKV Architecture Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), ML Engineer (davila7/claude-code-templates, 33k stars) and Databricks ML Training (databricks/databricks-agent-skills, 345 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Inference Batching Scheduler?

lazyFrogLOL (a GitHub user) maintains it in lazyFrogLOL/Harness_Engineering, which has 128 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on May 18, 2026.

Source: lazyFrogLOL/Harness_Engineering on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.