Caveman Gateway Setup
JuliusBrussee/caveman
Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.
Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract.
$ npx skills add microsoft/agent-lightning --skill agent-lightning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/agent-lightning agent-lightning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/agent-lightning.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-lightning .claude/skills/agent-lightning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-lightning" agent skill from https://github.com/microsoft/agent-lightning/tree/main/skills/agent-lightning into .claude/skills/agent-lightning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-lightning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/agent-lightning/tree/main/skills/agent-lightningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/agent-lightning --skill agent-lightning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/agent-lightning agent-lightning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/agent-lightning.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agent-lightning .agents/skills/agent-lightning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-lightning" agent skill from https://github.com/microsoft/agent-lightning/tree/main/skills/agent-lightning into .agents/skills/agent-lightning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-lightning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/agent-lightning --skill agent-lightning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/agent-lightning agent-lightning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/agent-lightning.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agent-lightning .cursor/skills/agent-lightning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-lightning" agent skill from https://github.com/microsoft/agent-lightning/tree/main/skills/agent-lightning into .cursor/skills/agent-lightning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-lightning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/agent-lightning.git --path skills/agent-lightning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/agent-lightning --skill agent-lightning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/agent-lightning agent-lightning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/agent-lightning.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agent-lightning .gemini/skills/agent-lightning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-lightning" agent skill from https://github.com/microsoft/agent-lightning/tree/main/skills/agent-lightning into .gemini/skills/agent-lightning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-lightning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/agent-lightning agent-lightningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/agent-lightning --skill agent-lightning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/agent-lightning.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agent-lightning .github/skills/agent-lightning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-lightning" agent skill from https://github.com/microsoft/agent-lightning/tree/main/skills/agent-lightning into .github/skills/agent-lightning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-lightning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/agent-lightning --skill agent-lightning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/agent-lightning agent-lightning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/agent-lightning.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agent-lightning .opencode/skills/agent-lightning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-lightning" agent skill from https://github.com/microsoft/agent-lightning/tree/main/skills/agent-lightning into .opencode/skills/agent-lightning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-lightning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-lightningProvides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract.
Agent Lightning is an agent skill from microsoft/agent-lightning, published by the product's own GitHub organization. Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract. Use when optimizing agent accuracy, cost, latency, or reliability.
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `.claude-plugin/plugin.json`).
It sits in DevOps & Cloud. The repository describes itself as: The absolute trainer to light up AI agents. The licence is MIT.
Read from SKILL.md and the folder at commit d381995. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Lightning loads about 1.7k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 759 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from microsoft/agent-lightning at commit d381995, republished under its MIT licence (© microsoft). 759 words, ~1,660 tokens.
.claude/skills/agent-lightning/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Agent optimization is a search over interacting choices. The useful question is not which architecture is most sophisticated, but which change moves the requested accuracy, cost, latency, and reliability frontier for this agent.
The optimizer's development budget and the resulting agent's per-run cost are different quantities. More development budget creates room to learn; it does not imply that the deployed agent should spend more on every task.
Your remaining development budget is reported at /artifacts/cost_budget.json
({spent, budget, remaining}), refreshed as you work. Read that file to see how
much is left, and keep iterating — measure, edit, re-score — while meaningful
budget remains; do not stop at the first plausible result. A run is finished not
because one change worked, but because further measured changes no longer improve
the frontier within the budget you still have. If remaining is large, there is
more search to do: ground more cases, probe a lever you have not tested, or add
reps to resolve a noisy comparison. Check remaining again after each expensive
step so the decision to stop is evidence-based, not a default.
| Lever | What it changes | Useful signal | Main tradeoff |
|---|---|---|---|
| Input grounding | Information and state visible to the model | Relevant deployment-visible context is missing | Longer context can distract or cost more |
| Output contract | Representation, types, schema, files, and terminal state | Work looks reasonable but is rejected or unreadable | Can overfit evaluator quirks |
| Prompt | Interpretation, priorities, and constraints | Instructions are misunderstood or important details are ignored | Prompt gains can be brittle |
| Tools | Deterministic inspection, computation, and execution | The model is approximating work a tool can do reliably | More code and new failure modes |
| Model | Base capability and knowledge | The primary cannot solve grounded cases | Cost, latency, and availability |
| Reasoning effort | Computation used by the primary call | Grounded hard cases remain | Cost and latency can grow nonlinearly |
| Failure isolation | Whether one failure damages other work | Individual tasks crash, time out, or corrupt shared state | Isolation does not recover the failed task |
| Conditional repair | A second attempt informed by failure evidence | Deployment-visible checks expose a recoverable failure | Extra calls and possible regressions |
| Routing | Different handling for different task classes | Difficulty or failure risk varies predictably | Router mistakes and operational complexity |
| Planning and interaction | State, ordering, and tool use across multiple steps | Long tasks lose goals or ignore observations | More state and control-flow overhead |
| Retrieval | Facts supplied from an available corpus | Correct answers depend on external knowledge | Retrieval errors and added latency |
| Critique or selection | Additional views or candidates | Independent attempts expose different useful information | Multiplied calls, cost, and latency |
Different failures expose different amounts of information:
Levers interact. More reasoning cannot recover information the model never sees. Repair cannot fix a systematic contract error when the retry receives no new evidence. Failure isolation preserves the batch but does not repair an item. A global model or effort increase and conditional escalation occupy different points on the frontier.
Agent evaluations are often stochastic. A score can improve while many individual cases regress, and a single strong result can be a lucky draw. Fixed cases, validation splits, repeated runs, frozen primary outputs, and small probes are different ways to reduce uncertainty; their value depends on the decision and the available budget.
Comparisons are easiest to interpret when the checkpoint, cases, model, effort, concurrency, scorer, and execution path are held constant except for the variable being studied. Accuracy is only one result: completion rate, downside, cost, latency, and variance can change the decision.
Development may expose labels, metadata, or tools that do not exist at deployment. An improvement that depends on them is not a deployed improvement. Likewise, measurement plumbing can fail independently of the agent being tested.
© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/agent-lightning of microsoft/agent-lightning.
Open the folder on GitHubat commit d381995
Agent Lightning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Lightning this skillmicrosoft/agent-lightning | 19k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Caveman Gateway SetupJuliusBrussee/caveman | 110k | 1 repos | ~2.6k | Automated safety check: Warn | Apache-2.0 | |
| Megatron-LM Base Image BumpNVIDIA/Megatron-LM | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Production Defaultshuggingface/skills | 11k | 1 repos | ~6.9k | Automated safety check: Pass | Apache-2.0 | |
| Opik Local Dev Environmentcomet-ml/opik | 22k | — | ~734 | Automated safety check: Pass | Apache-2.0 | |
| Agent Kill Switchvivekchand/clawmetry | 424 | — | ~1.1k | Automated safety check: Pass | MIT |
JuliusBrussee/caveman
Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.
NVIDIA/Megatron-LM
Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.
huggingface/skills
Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.
comet-ml/opik
Starts, rebuilds, and troubleshoots the Opik local dev stack, including an optional Comet Platform integration mode for the Opik team.
vivekchand/clawmetry
Give the human an off switch and a cost meter for the coding agents on this machine, using ClawMetry.
OrlojHQ/orloj
Interactive scaffold generator for Orloj multi-agent systems.
microsoft/agent-lightning
Prepare and publish stable Agent Lightning releases through the repository's version bump, pull-request checks, merge, tag, PyPI trusted-publishing, and versioned-documentation workflows.
Categories
Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract. Agent Lightning is an agent skill from microsoft/agent-lightning, published by the product's own GitHub organization. Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract.
Agent Lightning fits situations like: optimizing agent accuracy.
Run `npx skills add microsoft/agent-lightning --skill agent-lightning -a claude-code`. Or copy the skill folder (skills/agent-lightning in microsoft/agent-lightning) into .claude/skills/agent-lightning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add microsoft/agent-lightning --skill agent-lightning -a codex`. Or copy the skill folder (skills/agent-lightning in microsoft/agent-lightning) into .agents/skills/agent-lightning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/agent-lightning --skill agent-lightning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-lightning, .gemini/skills/agent-lightning, .github/skills/agent-lightning and .opencode/skills/agent-lightning in your project.
SKILL.md names no scripts, command-line tools or credentials: Agent Lightning is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agent Lightning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agent Lightning: Caveman Gateway Setup (JuliusBrussee/caveman, 110k stars), Megatron-LM Base Image Bump (NVIDIA/Megatron-LM, 18k stars), SageMaker Production Defaults (huggingface/skills, 11k stars) and Opik Local Dev Environment (comet-ml/opik, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/agent-lightning, which has 18,572 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 29, 2026.
Source: microsoft/agent-lightning on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.