Official agent skill

Agent Lightning

by microsoft in microsoft/agent-lightning

Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract.

OfficialMITAuto-check passedDevOps & Cloud

Install Agent Lightning

skills CLI
$ npx skills add microsoft/agent-lightning --skill agent-lightning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/agent-lightning agent-lightning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/agent-lightning.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-lightning .claude/skills/agent-lightning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-lightning
GitHub stars
19k
Token cost
~1.7k tokens
SKILL.md length
759 words
Files
2
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract.

  • Optimizing agent accuracy
  • SKILL.md covers Action space, Reading the evidence, Evaluation context and Boundaries
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Lightning is an agent skill from microsoft/agent-lightning, published by the product's own GitHub organization. Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract. Use when optimizing agent accuracy, cost, latency, or reliability.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `.claude-plugin/plugin.json`).

It sits in DevOps & Cloud. The repository describes itself as: The absolute trainer to light up AI agents. The licence is MIT.

When your agent uses it

  • Optimizing agent accuracy

Example prompts

  • “Use the agent-lightning skill to provide the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark…”
  • “/agent-lightning”

What it can do on your machine

Read from SKILL.md and the folder at commit d381995. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Lightning loads about 1.7k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 759 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/agent-lightning at commit d381995, republished under its MIT licence (© microsoft). 759 words, ~1,660 tokens.

Download SKILL.mdSave it as .claude/skills/agent-lightning/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
agent-lightning
description
Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract. Use when optimizing agent accuracy, cost, latency, or reliability.

Agent Lightning

Agent optimization is a search over interacting choices. The useful question is not which architecture is most sophisticated, but which change moves the requested accuracy, cost, latency, and reliability frontier for this agent.

The optimizer's development budget and the resulting agent's per-run cost are different quantities. More development budget creates room to learn; it does not imply that the deployed agent should spend more on every task.

Your remaining development budget is reported at /artifacts/cost_budget.json ({spent, budget, remaining}), refreshed as you work. Read that file to see how much is left, and keep iterating — measure, edit, re-score — while meaningful budget remains; do not stop at the first plausible result. A run is finished not because one change worked, but because further measured changes no longer improve the frontier within the budget you still have. If remaining is large, there is more search to do: ground more cases, probe a lever you have not tested, or add reps to resolve a noisy comparison. Check remaining again after each expensive step so the decision to stop is evidence-based, not a default.

Action space

LeverWhat it changesUseful signalMain tradeoff
Input groundingInformation and state visible to the modelRelevant deployment-visible context is missingLonger context can distract or cost more
Output contractRepresentation, types, schema, files, and terminal stateWork looks reasonable but is rejected or unreadableCan overfit evaluator quirks
PromptInterpretation, priorities, and constraintsInstructions are misunderstood or important details are ignoredPrompt gains can be brittle
ToolsDeterministic inspection, computation, and executionThe model is approximating work a tool can do reliablyMore code and new failure modes
ModelBase capability and knowledgeThe primary cannot solve grounded casesCost, latency, and availability
Reasoning effortComputation used by the primary callGrounded hard cases remainCost and latency can grow nonlinearly
Failure isolationWhether one failure damages other workIndividual tasks crash, time out, or corrupt shared stateIsolation does not recover the failed task
Conditional repairA second attempt informed by failure evidenceDeployment-visible checks expose a recoverable failureExtra calls and possible regressions
RoutingDifferent handling for different task classesDifficulty or failure risk varies predictablyRouter mistakes and operational complexity
Planning and interactionState, ordering, and tool use across multiple stepsLong tasks lose goals or ignore observationsMore state and control-flow overhead
RetrievalFacts supplied from an available corpusCorrect answers depend on external knowledgeRetrieval errors and added latency
Critique or selectionAdditional views or candidatesIndependent attempts expose different useful informationMultiplied calls, cost, and latency
Show full SKILL.md (332 more words)Show less

Reading the evidence

Different failures expose different amounts of information:

  • Exceptions, missing artifacts, invalid schemas, and timeouts are objective signals. They can support deterministic checks or focused recovery.
  • A valid-looking but wrong answer may expose no label-free repair signal. More calls with the same information can repeat the same mistake.
  • Repeated failures across different attempts suggest shared blindness, a contract mismatch, or missing capability. Diverse failures make routing, critique, or selection more plausible.
  • A development gain that disappears under validation may come from randomness, memorized examples, training-only fields, or a different deployment path.
  • An unavailable measurement is unknown evidence, not proof that a candidate improved or regressed.

Levers interact. More reasoning cannot recover information the model never sees. Repair cannot fix a systematic contract error when the retry receives no new evidence. Failure isolation preserves the batch but does not repair an item. A global model or effort increase and conditional escalation occupy different points on the frontier.

Evaluation context

Agent evaluations are often stochastic. A score can improve while many individual cases regress, and a single strong result can be a lucky draw. Fixed cases, validation splits, repeated runs, frozen primary outputs, and small probes are different ways to reduce uncertainty; their value depends on the decision and the available budget.

Comparisons are easiest to interpret when the checkpoint, cases, model, effort, concurrency, scorer, and execution path are held constant except for the variable being studied. Accuracy is only one result: completion rate, downside, cost, latency, and variance can change the decision.

Development may expose labels, metadata, or tools that do not exist at deployment. An improvement that depends on them is not a deployed improvement. Likewise, measurement plumbing can fail independently of the agent being tested.

Boundaries

  • Preserve the target's external interface and deployment environment.
  • Do not expose held-out labels or training-only answer fields to the deployed path.
  • Treat the scorer and evaluation contract as immutable measurement surfaces.
  • Leave a coherent measured checkpoint, not an unfinished or partially tested edit.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/agent-lightning of microsoft/agent-lightning.

  • SKILL.md
  • .claude-plugin/plugin.json

Open the folder on GitHubat commit d381995

Compare with similar skills

Agent Lightning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Lightning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Lightning this skillmicrosoft/agent-lightning19k—~1.7kAutomated safety check: PassMIT
Caveman Gateway SetupJuliusBrussee/caveman110k1 repos~2.6kAutomated safety check: WarnApache-2.0
Megatron-LM Base Image BumpNVIDIA/Megatron-LM18k—~2.8kAutomated safety check: PassApache-2.0
SageMaker Production Defaultshuggingface/skills11k1 repos~6.9kAutomated safety check: PassApache-2.0
Opik Local Dev Environmentcomet-ml/opik22k—~734Automated safety check: PassApache-2.0
Agent Kill Switchvivekchand/clawmetry424—~1.1kAutomated safety check: PassMIT

Similar skills

  • Caveman Gateway Setup

    JuliusBrussee/caveman

    Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior.

    110k GitHub starsUsed in 1 repo~2.6k tokens
    DevOps & CloudAuto-check: warnings
  • Megatron-LM Base Image Bump

    NVIDIA/Megatron-LM

    Official

    Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.

    18k GitHub stars~2.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed
  • Starts, rebuilds, and troubleshoots the Opik local dev stack, including an optional Comet Platform integration mode for the Opik team.

    22k GitHub stars~734 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Agent Kill Switch

    vivekchand/clawmetry

    Give the human an off switch and a cost meter for the coding agents on this machine, using ClawMetry.

    424 GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Orloj Generator

    OrlojHQ/orloj

    Interactive scaffold generator for Orloj multi-agent systems.

    123 GitHub stars~2.6k tokensUpdated 19 days ago
    DevOps & CloudAuto-check passed

More from microsoft/agent-lightning

  • Release

    microsoft/agent-lightning

    Official

    Prepare and publish stable Agent Lightning releases through the repository's version bump, pull-request checks, merge, tag, PyPI trusted-publishing, and versioned-documentation workflows.

    19k GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check passed

Questions about Agent Lightning

What does Agent Lightning do?

Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract. Agent Lightning is an agent skill from microsoft/agent-lightning, published by the product's own GitHub organization. Provides the action space, tradeoffs, and evaluation context for improving an editable AI agent against a benchmark while preserving its deployment contract.

When should I use Agent Lightning?

Agent Lightning fits situations like: optimizing agent accuracy.

How do I install Agent Lightning in Claude Code?

Run `npx skills add microsoft/agent-lightning --skill agent-lightning -a claude-code`. Or copy the skill folder (skills/agent-lightning in microsoft/agent-lightning) into .claude/skills/agent-lightning in your project. Claude Code loads it when a task matches its description.

How do I install Agent Lightning in Codex?

Run `npx skills add microsoft/agent-lightning --skill agent-lightning -a codex`. Or copy the skill folder (skills/agent-lightning in microsoft/agent-lightning) into .agents/skills/agent-lightning in your project. Codex loads it when a task matches its description.

Can I use Agent Lightning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/agent-lightning --skill agent-lightning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-lightning, .gemini/skills/agent-lightning, .github/skills/agent-lightning and .opencode/skills/agent-lightning in your project.

What does Agent Lightning need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Lightning is instructions for the agent only.

Does Agent Lightning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Lightning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Lightning use?

Agent Lightning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Lightning use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Lightning?

Skills that share tags, products or a category with Agent Lightning: Caveman Gateway Setup (JuliusBrussee/caveman, 110k stars), Megatron-LM Base Image Bump (NVIDIA/Megatron-LM, 18k stars), SageMaker Production Defaults (huggingface/skills, 11k stars) and Opik Local Dev Environment (comet-ml/opik, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Lightning?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/agent-lightning, which has 18,572 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on September 29, 2026.

Source: microsoft/agent-lightning on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.