Official agent skill

Review Agent Primitives

by microsoft in microsoft/Huabu

Review the quality of an agent's tool descriptions, system/agent prompts, or SKILL.md files against current agent-engineering best practices.

OfficialMITAuto-check passedAI & LLM Engineering

Install Review Agent Primitives

skills CLI
$ npx skills add microsoft/Huabu --skill review-agent-primitives -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/Huabu review-agent-primitives --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/Huabu.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/review-agent-primitives .claude/skills/review-agent-primitives && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review-agent-primitives
GitHub stars
157
Token cost
~2.3k tokens
SKILL.md length
1,077 words
Files
4 (incl. references)
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Review the quality of an agent's tool descriptions, system/agent prompts, or SKILL.md files against current agent-engineering best practices.

  • Works in 6 steps: Minimal high-signal tokens. Find the… → Right altitude. Avoid the two failure… → Zero ambiguity → zero guessing. Every… → …
  • Asked to review
  • SKILL.md covers When to use, When NOT to use, Shared foundation (applies to… and Review workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Review Agent Primitives is an agent skill from microsoft/Huabu, published by the product's own GitHub organization. Review the quality of an agent's tool descriptions, system/agent prompts, or SKILL.md files against current agent-engineering best practices. USE WHEN asked to review, audit, critique, score, or improve a tool description, agent prompt, system prompt, or skill; when a tool is called with wrong parameters or not called when it should be; when an agent makes redundant or repeated tool calls; or before shipping a new tool/prompt/skill; or when deciding which primitive a behavior belongs in (tool vs system prompt vs…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/agent-prompt.md`, `references/skill-quality.md` and `references/tool-description.md`).

It sits in AI & LLM Engineering, covering Prompt engineering and Skill authoring. The repository describes itself as: Huabu, where you and your agents think together. The licence is MIT.

When your agent uses it

  • Asked to review
  • Improve a tool description
  • A tool is called with wrong parameters
  • Not called when it should be

Example prompts

  • “/review-agent-primitives”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Minimal high-signal tokens. Find the smallest set of tokens that fully specifies the behavior. minimal ≠ short: include everything the…
  2. Right altitude. Avoid the two failure modes: (a) brittle hardcoded if-else logic that tries to script every case, and (b) vague high-level…
  3. Zero ambiguity → zero guessing. Every implicit assumption (formats, niche terms, resource relationships, boundaries) must be explicit…
  4. Self-consistency. The artifact must not contradict itself, the system prompt, or a sibling tool/skill. If a human expert can't say which…
  5. Fail-safe by design (poka-yoke). Prefer constraints, examples, and actionable error text that make mistakes hard to make, over prose that…
  6. Discoverable & load-on-demand. The name/description (for a skill) and the name/description (for tool search) are the discovery surface…

What it can do on your machine

Read from SKILL.md and the folder at commit dc0cfa9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Review Agent Primitives loads about 2.3k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 161 tokens; SKILL.md has 1,077 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~161
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/Huabu at commit dc0cfa9, republished under its MIT licence (© microsoft). 1,077 words, ~2,340 tokens.

Download SKILL.mdSave it as .claude/skills/review-agent-primitives/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
review-agent-primitives
description
Review the quality of an agent's tool descriptions, system/agent prompts, or SKILL.md files against current agent-engineering best practices. USE WHEN asked to review, audit, critique, score, or improve a tool description, agent prompt, system prompt, or skill; when a tool is called with wrong parameters or not called when it should be; when an agent makes redundant or repeated tool calls; or before shipping a new tool/prompt/skill; or when deciding which primitive a behavior belongs in (tool vs system prompt vs skill). Reviews the three artifacts together so shared context, overlap, and contradictions are caught.

Review Agent Primitives

Evaluate three artifacts that steer an agent's behavior — tool descriptions, agent/system prompts, and skills (SKILL.md) — against a single, source-backed rubric. Reviews all three with one shared foundation so overlap, contradictions, and duplicated context are caught in one pass.

When to use

  • Reviewing / auditing / scoring a tool description, agent or system prompt, or a SKILL.md.
  • Diagnosing behavior: a tool is called with wrong params, not called when it should be, or the agent makes redundant tool calls / burns turns guessing.
  • Deciding placement: given a piece of behavior or context, which primitive it belongs in (system prompt vs tool vs skill), and flagging content that currently sits in the wrong one.
  • Pre-ship polish of any of the three artifacts.

When NOT to use

  • Writing the artifact from scratch (do that first, then review).
  • Reviewing application/business logic — that is a normal code review.

Shared foundation (applies to all three)

These five principles come from the "context engineering" line of thinking and cut across every artifact. Score them once, here, before diving into the type-specific rubric — do not repeat them per section.

  1. Minimal high-signal tokens. Find the smallest set of tokens that fully specifies the behavior. minimal ≠ short: include everything the reader needs, nothing it doesn't. Cut restated rules, filler, and dead caveats.
  2. Right altitude. Avoid the two failure modes: (a) brittle hardcoded if-else logic that tries to script every case, and (b) vague high-level guidance that assumes shared context the model doesn't have. Aim for specific-but-flexible heuristics.
  3. Zero ambiguity → zero guessing. Every implicit assumption (formats, niche terms, resource relationships, boundaries) must be explicit. Ambiguity is what causes the agent to guess wrong and spend extra tool-call turns. Name things unambiguously (user_id, not user).
  4. Self-consistency. The artifact must not contradict itself, the system prompt, or a sibling tool/skill. If a human expert can't say which tool/instruction applies in a situation, the agent can't either.
  5. Fail-safe by design (poka-yoke). Prefer constraints, examples, and actionable error text that make mistakes hard to make, over prose that merely warns against them.
  6. Discoverable & load-on-demand. The name/description (for a skill) and the name/description (for tool search) are the discovery surface — keyword-rich and distinctive. For large tool/skill libraries, defer loading and keep only the few highest-use items resident, so context isn't spent on definitions the task doesn't need.

Review workflow

  1. Identify the artifact type. Tool description → prompt → skill. If mixed (e.g. a skill that also defines tools, or a prompt with an embedded tool spec), review each part with its own rubric.
  2. Score the shared foundation (6 principles above) → Pass / Warn / Fail each.
  3. Load the matching reference rubric and score its items:
    ArtifactReference
    Tool description / specreferences/tool-description.md
    Agent / system promptreferences/agent-prompt.md
    Skill (SKILL.md)references/skill-quality.md
  4. When all three exist together, additionally check cross-artifact issues: duplicated context across prompt and tool descriptions, procedures restated in the system prompt that belong in a skill, tools the prompt references but doesn't scope, and name/description keywords that collide between skills or similar tools. Also check same-type fan-out: one fact restated across files of the same type — a tool description and its own parameter/field-schema description strings, a SKILL.md and its sibling reference files, or two peer skills. The cross-artifact-type lens misses these because both copies are the same type; grep the distinctive phrase to find every copy, then flag lockstep-sync risk (see the fan-out signal under Placement mode).
  5. Report using the output format below.
Show full SKILL.md (504 more words)Show less

Placement mode (which primitive should this live in?)

Trigger this when the ask is "where should this go?" rather than "is this good?", or whenever a review surfaces content sitting in the wrong artifact. Decide with the table, then recommend moving misplaced content to its correct owner.

Put it in…When the content is…Keep OUT
System prompt / always-on instructionsGlobal behavior, constraints, tone, refusal style; small, stable "always do X" policies that apply every turnLong multi-step procedures — they bloat every turn and go brittle
ToolAn action on the world: calls external services/DBs, creates side effects, fetches live state. Narrowly scoped, strongly-typed inputs, explicit side effectsStatic procedural knowledge; anything with no side effect
Skill (SKILL.md)A reusable, multi-step procedure; branching/conditional workflow; needs scripts/templates/assets; used sometimes, not every turn; wants independent versioningOne-off tasks; always-on policy; pure live-data fetch
Few-shot examplesCanonical, diverse demonstrations of expected behaviorExhaustive edge-case dumps. Workflow-specific examples belong in the skill, not the system prompt

Decision signals:

  • Needed every turn + short + stable → system prompt.
  • Has a side effect or needs live state → tool.
  • Multi-step / branching / needs code or assets / only sometimes → skill.
  • Same content appears in two places → converge to a single owner (prefer the most specific: skill > instructions).
  • Same fact fanned out across many same-type files (tool desc + its field-schema descriptions, a SKILL.md + its references/, two peer skills) → this is a maintainability/lockstep-sync smell, not a type-placement question. Converge to one canonical statement + pointers unless the copies genuinely diverge by audience/execution path — then keep them but call out that they must change together, and make sure the review surface actually includes every copy (e.g. schema files, not just the tool-definition file). When the fanned-out fact is a machine-derivable enumeration (e.g. a prose list of command/variant names mirroring a schema union), the strongest form of "converge" is to derive it from the schema (codegen or a compile-time guard) so the copies cannot drift at all.
  • A whole procedure sitting in the system prompt → move it to a skill (preserves reuse, versioning, and on-demand loading).

Scoring model

Each checklist item gets one verdict, with a concrete finding that quotes the exact offending line:

  • Pass — meets the bar.
  • Warn — works but costs tokens, invites guessing, or will age badly.
  • Fail — will cause wrong/failed/redundant calls or mis-triggering. Must fix.

Do not invent numeric scores per line; aggregate to one headline verdict per dimension.

Output format

Always produce, in this order:

  1. Verdict line — Pass / Warn / Fail overall + the single highest-impact issue (reference it by ID, e.g. F1).
  2. Findings table — ID | Dimension | Verdict | Evidence (quoted line) | Fix. Give every issue a stable ID: F1, F2, … for each Fail and W1, W2, … for each Warn, numbered in order of impact. Pass rows need no ID.
  3. Prioritized fixes — reference findings by ID (F1 → …), ordered by impact on wrong/redundant tool calls first (all F before any W).
  4. Rewrite — only if asked, provide the corrected artifact.

Keep evidence quotes short. One row per real issue; do not pad the table to look thorough.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .agents/skills/review-agent-primitives of microsoft/Huabu.

  • SKILL.md
  • references/agent-prompt.md
  • references/skill-quality.md
  • references/tool-description.md

Open the folder on GitHubat commit dc0cfa9

Compare with similar skills

Review Agent Primitives next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Review Agent Primitives compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Review Agent Primitives this skillmicrosoft/Huabu157—~2.3kAutomated safety check: PassMIT
Prompt LabMathews-Tom/armory328—~2.1kAutomated safety check: PassMIT
Skill With Prompt EngineeringLeoYeAI/openclaw-master-skills2.2k—~4.1kAutomated safety check: PassMIT
A-Evolve Agent Improvementaiming-lab/AutoResearchClaw15k—~1.8kAutomated safety check: PassMIT
Workflow Schema Tuningbreaking-brake/cc-wf-studio5.4k—~1.3kAutomated safety check: PassCustom licence
Writing Quotationmathbullet/skills173—~1.7kAutomated safety check: PassMIT

Similar skills

  • Prompt Lab

    Mathews-Tom/armory

    LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites.

    328 GitHub stars~2.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Skill With Prompt Engineering

    LeoYeAI/openclaw-master-skills

    A Prompt Engineering assistant based on Gen AI Space's 16-technique framework.

    2.2k GitHub stars~4.1k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • A-Evolve Agent Improvement

    aiming-lab/AutoResearchClaw

    Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.

    15k GitHub stars~1.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Workflow Schema Tuning

    breaking-brake/cc-wf-studio

    Guides edits to cc-wf-studio's workflow schema so AI agents generate better workflows, treating schema text as prompt engineering rather than validation.

    5.4k GitHub stars~1.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Writing Quotation

    mathbullet/skills

    Formatting rules for quoting external sources (papers, articles, web pages, prompt templates) inside a Markdown document.

    173 GitHub stars~1.7k tokensUpdated 29 days ago
    Agent WorkflowsAuto-check passed
  • Prompt Improver

    severity1/claude-code-prompt-improver

    This skill enriches vague prompts with targeted research and clarification before execution.

    1.9k GitHub starsUsed in 2 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed

More from microsoft/Huabu

All 10 skills in this repo
  • Code Review Expert

    microsoft/Huabu

    Official

    Expert code review of current git changes with a senior engineer lens.

    157 GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Release

    microsoft/Huabu

    Official

    Use ONLY when the user explicitly asks to cut/publish a Huabu desktop release or tag a version.

    157 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Space

    microsoft/Huabu

    Official

    Space mental model (the infinite work surface), tool boundaries, and command reference.

    157 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Review Huabu Agent

    microsoft/Huabu

    Official

    Review the Huabu operate agent's three steering artifacts at their fixed repo locations — system prompt, tool descriptions, and skills.

    157 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Issue Tracker

    microsoft/Huabu

    Official

    Coordinate one or more GitHub issues through isolated Git worktrees, durable Huabu Tasks, and dedicated Fixing Agent Threads.

    157 GitHub stars~279 tokensUpdated today
    Auto-check passed
  • Deepv Slides Maker

    microsoft/Huabu

    Official

    Create and revise editable slide decks, PowerPoint files, and slide images with DeepV.

    157 GitHub stars~276 tokensUpdated today
    Auto-check: notes

Questions about Review Agent Primitives

What does Review Agent Primitives do?

Review the quality of an agent's tool descriptions, system/agent prompts, or SKILL.md files against current agent-engineering best practices. Review Agent Primitives is an agent skill from microsoft/Huabu, published by the product's own GitHub organization.md files against current agent-engineering best practices.

When should I use Review Agent Primitives?

Review Agent Primitives fits situations like: asked to review; improve a tool description; A tool is called with wrong parameters; not called when it should be.

How do I install Review Agent Primitives in Claude Code?

Run `npx skills add microsoft/Huabu --skill review-agent-primitives -a claude-code`. Or copy the skill folder (.agents/skills/review-agent-primitives in microsoft/Huabu) into .claude/skills/review-agent-primitives in your project. Claude Code loads it when a task matches its description.

How do I install Review Agent Primitives in Codex?

Run `npx skills add microsoft/Huabu --skill review-agent-primitives -a codex`. Or copy the skill folder (.agents/skills/review-agent-primitives in microsoft/Huabu) into .agents/skills/review-agent-primitives in your project. Codex loads it when a task matches its description.

Can I use Review Agent Primitives in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/Huabu --skill review-agent-primitives -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-agent-primitives, .gemini/skills/review-agent-primitives, .github/skills/review-agent-primitives and .opencode/skills/review-agent-primitives in your project.

What does Review Agent Primitives need to run?

SKILL.md names no scripts, command-line tools or credentials: Review Agent Primitives is instructions for the agent only.

Does Review Agent Primitives access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Review Agent Primitives safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Review Agent Primitives use?

Review Agent Primitives is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Review Agent Primitives use?

About 2.3k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Review Agent Primitives?

Skills that share tags, products or a category with Review Agent Primitives: Prompt Lab (Mathews-Tom/armory, 328 stars), Skill With Prompt Engineering (LeoYeAI/openclaw-master-skills, 2.2k stars), A-Evolve Agent Improvement (aiming-lab/AutoResearchClaw, 15k stars) and Workflow Schema Tuning (breaking-brake/cc-wf-studio, 5.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Review Agent Primitives?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/Huabu, which has 157 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 8, 2026.

Source: microsoft/Huabu on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.