Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).

MITAuto-check passedAI & LLM Engineering

Install Tool Design

skills CLI
$ npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill tool-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentailor/fullstack-langgraph-nextjs-agent tool-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentailor/fullstack-langgraph-nextjs-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/tool-design .claude/skills/tool-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tool-design
GitHub stars
132
Token cost
~3.2k tokens
SKILL.md length
1,698 words
Files
4 (incl. references)
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).

  • Works in 5 steps: Strategic selection → Clear naming and namespacing → Meaningful context return → …
  • Writing a new tool for an agent
  • SKILL.md covers Overview, The Five Principles, Validation Checklist and Testing, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tool Design is an agent skill from agentailor/fullstack-langgraph-nextjs-agent. Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise). Use when writing a new tool for an agent, reviewing or fixing an existing tool definition, deciding how to split capabilities into tools, writing tests or evals for a tool, checking that a tool's output matches what its description promised, or debugging why an agent misuses, mis-selects, misreads the output of, or floods…

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/examples.md`, `references/principles.md` and `references/testing.md`).

It sits in AI & LLM Engineering, covering Building AI agents, Structured output and tool calling and MCP servers. It works with Model Context Protocol, LangChain, LangGraph and Python. The repository describes itself as: Production-ready Next.js template for building AI agents with LangGraph.js. Features MCP integration for dynamic tool loading, human-in-the-loop tool approval, persistent… The licence is MIT.

When your agent uses it

  • Writing a new tool for an agent
  • Fixing an existing tool definition
  • Deciding how to split capabilities into tools
  • Evals for a tool

Example prompts

  • “/tool-design”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Strategic selection
  2. Clear naming and namespacing
  3. Meaningful context return
  4. Token efficiency
  5. Descriptions are prompts

What it can do on your machine

Read from SKILL.md and the folder at commit 40414f8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tool Design loads about 3.2k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 158 tokens; SKILL.md has 1,698 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~158
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentailor/fullstack-langgraph-nextjs-agent at commit 40414f8, republished under its MIT licence (© agentailor). 1,698 words, ~3,162 tokens.

Download SKILL.mdSave it as .claude/skills/tool-design/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
tool-design
description
Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise). Use when writing a new tool for an agent, reviewing or fixing an existing tool definition, deciding how to split capabilities into tools, writing tests or evals for a tool, checking that a tool's output matches what its description promised, or debugging why an agent misuses, mis-selects, misreads the output of, or floods its context with a tool. Applies equally to standalone tools and MCP-server tools — a tool is a tool.

Tool Design

Overview

A tool is a contract between a deterministic system and a non-deterministic caller. A normal API assumes a rational developer who reads the docs, handles error codes, and knows which endpoint to call. An agent breaks all of those assumptions: it may pick the wrong tool because two names look alike, pass malformed parameters despite a clear schema, pull back a dataset that blows its own context window, or misread a cryptic error and retry the same failing call.

So tools for agents are designed defensively: clear enough that the agent can't easily misuse them, informative enough to steer the agent toward a better next move, and lean enough to spend the context window carefully.

The payoff: agent and human ergonomics align. A tool that an agent uses well is almost always a tool a human finds intuitive too. Designing for a non-deterministic caller just produces a better API.

None of this is framework- or language-specific. The same five principles apply whether the tool is an MCP server tool, a LangChain/LangGraph tool, an OpenAI/Anthropic function-calling definition, or a plain function exposed to a model — and whether it's written in TypeScript, Python, or anything else. What varies is the syntax of name / description / parameters / returns; the design thinking does not. See references/examples.md for the same tool proven across languages and surfaces.

The Five Principles

Apply these when writing or reviewing any tool. The deep dive with worked schema shapes is in references/principles.md.

1. Strategic selection

Build tools around user workflows, not database schemas or API endpoints. Don't wrap every endpoint as its own tool — the agent then struggles to choose among near-duplicates and you spend prompt budget documenting all of them. Consolidate related operations into one well-parameterized tool when it maps to how a user thinks about the task.

Prefer one get_expenses(start_date, end_date, category?, ...) over get_expense_by_id + list_all_expenses + filter_by_category + search_expenses.

Ask: does this map to how users think about the task? Would merging it with a sibling reduce the number of decisions the agent has to make?

But don't consolidate on data alone. Two tools touching the same table can still be two workflows. A bounded row listing ("show me my Dining transactions in June") and an aggregate query ("how much did I spend on Dining last quarter") look like prime merge candidates — same data, adjacent phrasing. Merging them into one tool with a mode flag satisfies the principle's letter and makes selection harder: the agent now reasons about which mode on every call, and the two have genuinely different return shapes and safety properties. Consolidate operations that share a workflow, not operations that merely share data.

2. Clear naming and namespacing

Names and descriptions decide whether the agent picks the right tool. Use descriptive, action-oriented names — never bare search, fetch, process. When an agent has tools from several sources, collisions cause mis-selection, so namespace with a consistent prefix — by service (slack_search, notion_search) or resource (expenses_get, expenses_summarize).

Check whether your surface prefixes for you: many MCP clients auto-prepend the server name, so hardcoding the prefix too yields agentailor_agentailor_search. Let the surface prefix, or prefix yourself — not both. See references/principles.md.

A description should answer three questions:

  • What does it do? (1–2 sentences)
  • When should the agent reach for it? (a few example queries)
  • How is it used? (key parameters, constraints, what it returns)

Every parameter gets a description with its format, an example, and constraints — "Start date in ISO format (YYYY-MM-DD). Example: '2025-01-01'", not start_date: string.

3. Meaningful context return

Return information the agent can reason about directly, not identifiers it must resolve with another call. Include the expense's description and category, not just a UUID. Attach lightweight metadata (totals, the date range, which filters were applied) so the agent knows what it's looking at.

Make verbosity configurable: a response_format of concise (essential fields) vs detailed (full metadata) lets the agent trade detail against tokens per task. When a result feeds the next tool call, return the exact value that next call needs (e.g. the canonical URL/id to pass along), so the agent can chain without a lookup.

4. Token efficiency

Every token in a tool response is a token unavailable for reasoning. Give collection-returning tools sensible default limits (e.g. 50) plus pagination and filter parameters. When you truncate, say so and say how to continue — "Found 847, returning 50; narrow the date range or add a category filter" beats silently dropping rows. Validate inputs that would produce huge responses (e.g. reject a >1-year range) before running the query.

5. Descriptions are prompts

The tool's name, description, and parameter docs are prompt engineering — every word shapes how the agent uses the tool. Be explicit about when to use it, required vs optional parameters, and the shape of the output. Refine this text iteratively against real agent behavior: unclear wording is the most common reason an agent mis-selects or mis-calls a tool.

Validation Checklist

When writing or reviewing a tool, confirm:

  • Strategic — consolidates a real workflow; not one-tool-per-endpoint, not overlapping near-duplicates.
  • Named — descriptive, action-oriented, namespaced to avoid collisions; not a bare verb.
  • Described — the description says what / when / how, with example queries.
  • Parameterized — every parameter documents format + example + constraints; required vs optional is intentional; optionals have sensible defaults.
  • Contextual — returns human-readable info and metadata, not just IDs; verbosity is configurable when responses can be large.
  • Efficient — default limits, pagination/filters, truncation messages that guide the next query; guards against responses that would blow the context window.
  • Helpful on failure — errors explain what went wrong and what to do next, with an example; no cryptic codes.
  • Tested — the checklist items above that are mechanically checkable have deterministic tests (unit at minimum, integration where real I/O decides correctness); the behaviors only a model can exercise have evals (or are knowingly deferred).
Show full SKILL.md (741 more words)Show less

Testing

The checklist above says what a good tool does. Nothing so far says how you know it still does — and a tool description is a contract with a non-deterministic caller, so an unverified contract is the default failure. The implementation drifts from what the description promised, and the agent, which has no way to check, believes the description.

Two layers, split on whether a model is in the loop, and neither subsumes the other:

Deterministic tests — cheap, repeatable, always. Whatever runs without a model and gives a stable verdict: unit tests as the floor, plus integration tests when correctness depends on real I/O (a multi-step write, a transaction boundary, a third-party API). Most of the checklist is mechanically testable this way: that truncation is signalled rather than merely applied, that errors return structured actionable objects, that the payload really contains the fields the description promised, that defaults behave as documented. Assert on the tool's returned payload — the surface the agent sees, and the same surface an eval grades later, so the assertions survive.

A tool that silently caps rows at 50 while its description promises a bounded list will report 262 matches as count: 50. A test that stubs a full page against a larger total and asserts truncated === true catches that in milliseconds. See references/testing.md.

Evals — needed, harness-agnostic. No deterministic test can catch mis-selection (nothing in a unit test decides whether the agent reached for query_transactions or run_sql — the model chooses, from your descriptions) or mis-interpretation (a unit test proves truncated: true is present; only an eval proves the agent noticed it and didn't sum a partial page into a confident wrong total). Run the real agent against a seeded fixture and grade the transcript on tools called and final answer. Deferring this layer is a resource decision, not evidence the risk is absent.

Layout: keep a tool's deterministic tests beside the tool, so tool + tests lift into any project as one unit. Evals are the opposite — cross-cutting, spanning several tools, and centralized. The two layers don't live in the same place.

Workflows

Writing a new tool — Start from the workflow the user cares about (Principle 1), not the data model. Draft the name + description + parameters as if they were prompt text (Principles 2 & 5). Decide the return shape and metadata (Principle 3), then add limits, filters, and guard rails (Principle 4). Run the checklist. Then test both layers (see Testing): unit-test the contract — truncation signalled, errors structured, promised fields present — and give the agent 3–5 realistic requests to watch whether it selects, calls, and interprets the tool correctly; refine the description where it stumbles.

Reviewing an existing tool — Walk the checklist top to bottom. For each miss, name the severity, quote the exact line, explain why it trips an agent, and show the fixed version. The most common real failures: bare/colliding names, param: type with no description, returning IDs the agent then has to resolve, unbounded responses, and cryptic error codes.

Failure Modes to Avoid

  • One tool per endpoint — fragments the decision and burns prompt budget. Consolidate.
  • Bare or colliding names (search, get, fetch) — the agent can't tell yours apart from another server's. Namespace and be specific.
  • Undocumented parameters — start_date: string tells the agent nothing about format or bounds. Add format + example + constraints.
  • Returning IDs instead of context — forces an extra lookup call and wastes tokens. Return what the agent can reason about.
  • Unbounded responses — no limit, no pagination; one call floods the context window. Cap and paginate by default.
  • Cryptic errors (ERR_INVALID_DATE, TOO_MANY_RESULTS) — the agent can't self-correct. Say what happened and what to try next.
  • Framework tunnel vision — assuming these ideas only apply to MCP, or only to your language. They apply to every tool surface; see references/examples.md.
  • An unverified contract — the description promises one thing, the implementation returns another, and nothing notices. Unit-test the checklist items that are mechanically checkable; eval the ones that aren't.

References

  • references/principles.md — the five principles in depth, each illustrated with a neutral name/description/parameters/returns schema shape (no framework assumed).
  • references/examples.md — the same tool proven across surfaces and languages: one tool in TypeScript (Zod) and Python (Pydantic) side by side, one standalone (non-MCP) tool, and one real MCP-server tool annotated principle by principle.
  • references/testing.md — verifying the contract: which checklist items are mechanically testable and what each assertion looks like, a worked truncation bug (found in real shipped code) with the test that catches it, and what evals catch that no deterministic test can.

© agentailor, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .agents/skills/tool-design of agentailor/fullstack-langgraph-nextjs-agent.

  • SKILL.md
  • references/examples.md
  • references/principles.md
  • references/testing.md

Open the folder on GitHubat commit 40414f8

Compare with similar skills

Tool Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tool Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tool Design this skillagentailor/fullstack-langgraph-nextjs-agent132—~3.2kAutomated safety check: PassMIT
Add Example AgentGetBindu/Bindu10k—~1.1kAutomated safety check: NotesCustom licence
Agent Inspectrajudandigam/agent-inspect165—~424Automated safety check: PassMIT
Strandsstrands-agents/harness-sdk8.7k—~1kAutomated safety check: PassApache-2.0
Ydc Openai Agent SDK IntegrationLeoYeAI/openclaw-master-skills2.2k—~4.4kAutomated safety check: NotesMIT
Magic ResumeMagic-Resume/Magic-Resume101—~663Automated safety check: PassMIT

Similar skills

  • Add Example Agent

    GetBindu/Bindu

    Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.

    10k GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Inspect

    rajudandigam/agent-inspect

    Local evidence debugger and trajectory-test toolkit for TypeScript AI agents.

    165 GitHub stars~424 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Strands

    strands-agents/harness-sdk

    Build, extend, evaluate, or migrate applications with Strands Agents in Python or TypeScript.

    8.7k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ydc Openai Agent SDK Integration

    LeoYeAI/openclaw-master-skills

    Integrate OpenAI Agents SDK with You.com MCP server - Hosted and Streamable HTTP support for Python and TypeScript.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Magic Resume

    Magic-Resume/Magic-Resume

    How AI agents integrate with Magic Resume — read and safely edit a user's resumes through the native MCP server (@magic-resume/mcp).

    101 GitHub stars~663 tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from agentailor/fullstack-langgraph-nextjs-agent

  • Agent Prompt Engineering

    agentailor/fullstack-langgraph-nextjs-agent

    Comprehensive guide for designing, refining, and auditing system prompts for autonomous AI agents based on Anthropic's production practices.

    132 GitHub stars~3.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Eval Cases

    agentailor/fullstack-langgraph-nextjs-agent

    Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

    132 GitHub stars~5.3k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Tool Design

What does Tool Design do?

Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise). Tool Design is an agent skill from agentailor/fullstack-langgraph-nextjs-agent. Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).

When should I use Tool Design?

Tool Design fits situations like: writing a new tool for an agent; fixing an existing tool definition; deciding how to split capabilities into tools; evals for a tool.

How do I install Tool Design in Claude Code?

Run `npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill tool-design -a claude-code`. Or copy the skill folder (.agents/skills/tool-design in agentailor/fullstack-langgraph-nextjs-agent) into .claude/skills/tool-design in your project. Claude Code loads it when a task matches its description.

How do I install Tool Design in Codex?

Run `npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill tool-design -a codex`. Or copy the skill folder (.agents/skills/tool-design in agentailor/fullstack-langgraph-nextjs-agent) into .agents/skills/tool-design in your project. Codex loads it when a task matches its description.

Can I use Tool Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentailor/fullstack-langgraph-nextjs-agent --skill tool-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tool-design, .gemini/skills/tool-design, .github/skills/tool-design and .opencode/skills/tool-design in your project.

What does Tool Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Tool Design is instructions for the agent only.

Does Tool Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tool Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tool Design use?

Tool Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tool Design use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.3k tokens, read only when the agent opens those files.

What are the alternatives to Tool Design?

Skills that share tags, products or a category with Tool Design: Add Example Agent (GetBindu/Bindu, 10k stars), Agent Inspect (rajudandigam/agent-inspect, 165 stars), Strands (strands-agents/harness-sdk, 8.7k stars) and Ydc Openai Agent SDK Integration (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tool Design?

agentailor (a GitHub organization) maintains it in agentailor/fullstack-langgraph-nextjs-agent, which has 132 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 4, 2026.

Source: agentailor/fullstack-langgraph-nextjs-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.