Agent skill

Ag2 Testing

by ag2ai in ag2ai/build-with-ag2

Test AG2 beta agents and tools without hitting a real LLM provider.

Apache-2.0Auto-check passedTesting & QA

Install Ag2 Testing

skills CLI
$ npx skills add ag2ai/build-with-ag2 --skill ag2-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ag2ai/build-with-ag2 ag2-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ag2ai/build-with-ag2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ag2-testing .claude/skills/ag2-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ag2-testing
GitHub stars
252
Token cost
~1.3k tokens
SKILL.md length
327 words
Files
1
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Test AG2 beta agents and tools without hitting a real LLM provider.

  • The user is writing pytest tests for an Agent
  • SKILL.md covers When to use, 60-second recipe — mock an LLM…, Simulate a successful tool call and Test tool error paths, plus 4 more sections
  • Needs OPENAI_API_KEY
  • Tasks that involve Unit testing

What it does

Ag2 Testing is an agent skill from ag2ai/build-with-ag2. Test AG2 beta agents and tools without hitting a real LLM provider. Pass TestConfig(...) from autogen.beta.testing as the agent's config (or per-ask) to mock LLM responses, inject ToolCallEvents to simulate tool execution, and assert success / error paths. Use when the user is writing pytest tests for an Agent or Tool.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Unit testing and Building AI agents. It works with pytest. The repository describes itself as: Sample code and application showcases to get you going with AG2 (formally AutoGen). The licence is Apache-2.0.

When your agent uses it

  • The user is writing pytest tests for an Agent
  • Tasks that involve Unit testing
  • Tasks that involve Building AI agents

Example prompts

  • “/ag2-testing”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 29eeac3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ag2 Testing loads about 1.3k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 327 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ag2ai/build-with-ag2 at commit 29eeac3, republished under its Apache-2.0 licence (© ag2ai). 327 words, ~1,273 tokens.

Download SKILL.mdSave it as .claude/skills/ag2-testing/SKILL.md (or your agent's skills folder).
name
ag2-testing
description
Test AG2 beta agents and tools without hitting a real LLM provider. Pass `TestConfig(...)` from `autogen.beta.testing` as the agent's config (or per-`ask`) to mock LLM responses, inject `ToolCallEvent`s to simulate tool execution, and assert success / error paths. Use when the user is writing pytest tests for an Agent or Tool.
license
Apache-2.0

Testing agents and tools

When to use

Writing tests for code that builds AG2 beta Agents, custom @tool functions, middleware, or response schemas — anywhere you don't want to make real LLM API calls.

60-second recipe — mock an LLM response

python
import pytest
from autogen.beta import Agent
from autogen.beta.testing import TestConfig

@pytest.mark.asyncio
async def test_mocked_response():
    agent = Agent("test_agent")
    reply = await agent.ask("Hi!", config=TestConfig("This is a mocked response."))
    assert reply.body == "This is a mocked response."

TestConfig(*responses) replaces the model client. Each positional arg is the mocked response for the next sequential turn — strings for text replies, ToolCallEvent for tool dispatches.

Simulate a successful tool call

Pass a ToolCallEvent first (the model "decides" to call the tool), then the final answer:

python
import pytest
from autogen.beta import Agent
from autogen.beta.events import ToolCallEvent
from autogen.beta.testing import TestConfig

@pytest.mark.asyncio
async def test_tool_success():
    def my_tool() -> str:
        return "tool execution result"

    agent = Agent("test_agent", tools=[my_tool])
    config = TestConfig(
        ToolCallEvent(name="my_tool"),
        "final result",
    )
    reply = await agent.ask("Please use my_tool", config=config)
    assert reply.body == "final result"

Test tool error paths

If a tool raises, the exception propagates to ask():

python
@pytest.mark.asyncio
async def test_tool_raises():
    def failing_tool() -> str:
        raise ValueError("Something went wrong")

    config = TestConfig(
        ToolCallEvent(name="failing_tool"),
        "result",
    )
    agent = Agent("test_agent", config=config, tools=[failing_tool])

    with pytest.raises(ValueError, match="Something went wrong"):
        await agent.ask("Hi!")

Tool not found

If the LLM calls a tool the agent doesn't have, the framework raises ToolNotFoundError:

python
from autogen.beta.exceptions import ToolNotFoundError

@pytest.mark.asyncio
async def test_tool_not_found():
    config = TestConfig(ToolCallEvent(name="unregistered_tool"))
    agent = Agent("test_agent", config=config)
    with pytest.raises(ToolNotFoundError, match="Tool `unregistered_tool` not found"):
        await agent.ask("Hi!")

Useful test patterns

Override Depends dependencies
python
def get_production_db():
    raise Exception("Do not call in tests!")

@tool
def read_data(db: Annotated[object, Depends(get_production_db)]) -> str:
    return "Data"

agent = Agent("test", tools=[read_data])
agent.dependency_provider.override(get_production_db, lambda: "mock_db")
Override Inject dependencies

Just pass dependencies={...} to agent.ask(...):

python
await agent.ask("Read", dependencies={"database_pool": fake_pool})
Capture stream events
python
from autogen.beta import MemoryStream
from autogen.beta.events import ToolCallEvent

stream = MemoryStream()
collected: list[ToolCallEvent] = []
stream.where(ToolCallEvent).subscribe(lambda e: collected.append(e))

await agent.ask("Test", stream=stream)
assert collected[0].name == "expected_tool"
Multi-turn mock

Each positional arg in TestConfig(...) corresponds to one model response. For a multi-turn test, supply enough responses for each turn the test exercises.

Going deeper

  • Source doc: website/docs/beta/testing.mdx.
  • Test markers / async config — repo pyproject.toml. Use @pytest.mark.asyncio (the project uses pytest-asyncio).
  • Streams (for asserting events): website/docs/beta/advanced/stream.mdx.

Common pitfalls

  • Forgetting @pytest.mark.asyncio — the test will skip or fail oddly.
  • Mismatched response count — TestConfig runs out of responses if the agent makes more LLM calls than you expect (e.g. tool error → another LLM call). Add more positional args or assert that the call sequence is what you intended.
  • Mocking the LLM but not the tool — your tool function still runs (and may hit real APIs / disk). Mock the tool if you're isolating LLM behaviour, or override its Depends to inject test doubles.
  • Asserting on reply.body when you set a response_schema — body is the raw text. Use await reply.content() for the validated value.
  • Sharing Agent instances across async tests — agents carry mutable state (variables, dependencies). Construct fresh agents per test for isolation.
  • Using real provider clients in CI — wrap the provider config with TestConfig per-test or via a fixture; never rely on OPENAI_API_KEY etc. being available in test environments.

© ag2ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/ag2-testing of ag2ai/build-with-ag2.

Open the folder on GitHubat commit 29eeac3

Compare with similar skills

Ag2 Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ag2 Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ag2 Testing this skillag2ai/build-with-ag2252—~1.3kAutomated safety check: PassApache-2.0
Adk Agent Buildergoogle/adk-python22k—~879Automated safety check: PassApache-2.0
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence
Test GuardamElnagdy/guard-skills1.3k2 repos~2.1kAutomated safety check: PassMIT
Fla Optimization Loopfla-org/flash-linear-attention5.8k—~2.6kAutomated safety check: PassMIT

Similar skills

  • Adk Agent Builder

    google/adk-python

    Official

    Builds ADK (Agent Development Kit) Python agents: LLM agents with tools, graph workflows of function and agent nodes, conditional routing, fan-out and join, schema-validated delegation between…

    22k GitHub stars~879 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Test Guard

    amElnagdy/guard-skills

    Reviews newly written or edited tests against nine rules that cut test bloat, such as mock-heavy checks and near-duplicate cases, before they are committed.

    1.3k GitHub starsUsed in 2 repos~2.1k tokens
    Testing & QAAuto-check passed
  • Fla Optimization Loop

    fla-org/flash-linear-attention

    Disciplined, reproducible loop for making an FLA kernel faster (Triton, Gluon, TileLang, CuTe) without ever breaking or gaming correctness.

    5.8k GitHub stars~2.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Pytest Runner

    saleor/saleor

    Run pytest tests with automatic virtual environment activation. Use this skill whenever running tests, executing pytest, or when asked to "run tests", "test…

    23k GitHub stars~251 tokensUpdated yesterday
    Testing & QAAuto-check passed

More from ag2ai/build-with-ag2

All 16 skills in this repo
  • Ag2 Add Custom Tool

    ag2ai/build-with-ag2

    Add a custom Python tool to an AG2 beta Agent using the @tool decorator.

    252 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Middleware

    ag2ai/build-with-ag2

    Intercept the AG2 beta agent loop with BaseMiddleware — wrap full turns (onturn), each LLM call (onllmcall), each tool execution (ontoolexecution), or each human-input request (onhumaninput).

    252 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Use Builtin Tools

    ag2ai/build-with-ag2

    Wire AG2 beta's shipped tools into an Agent — both provider-native server-side tools (web search, web fetch, code execution, MCP, image generation, memory) and locally-executed common toolkits…

    252 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Knowledge And Memory

    ag2ai/build-with-ag2

    Persist agent state across runs, shape what the LLM sees per turn, and cap history to fit a context window.

    252 GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Observers And Alerts

    ag2ai/build-with-ag2

    Monitor an AG2 beta agent's stream — log events, detect repeated tool calls, track token spend, build trigger-driven observers, route observer alerts to the model, and halt on FATAL conditions.

    252 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Quickstart

    ag2ai/build-with-ag2

    Build a minimal AG2 beta Agent end to end — pick a model provider, set a prompt, call agent.ask(), then continue the conversation with reply.ask() (multi-turn).

    252 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check: notes

Works with

Categories

Questions about Ag2 Testing

What does Ag2 Testing do?

Test AG2 beta agents and tools without hitting a real LLM provider. Ag2 Testing is an agent skill from ag2ai/build-with-ag2. Test AG2 beta agents and tools without hitting a real LLM provider.

When should I use Ag2 Testing?

Ag2 Testing fits situations like: the user is writing pytest tests for an Agent; tasks that involve Unit testing; tasks that involve Building AI agents.

How do I install Ag2 Testing in Claude Code?

Run `npx skills add ag2ai/build-with-ag2 --skill ag2-testing -a claude-code`. Or copy the skill folder (.agents/skills/ag2-testing in ag2ai/build-with-ag2) into .claude/skills/ag2-testing in your project. Claude Code loads it when a task matches its description.

How do I install Ag2 Testing in Codex?

Run `npx skills add ag2ai/build-with-ag2 --skill ag2-testing -a codex`. Or copy the skill folder (.agents/skills/ag2-testing in ag2ai/build-with-ag2) into .agents/skills/ag2-testing in your project. Codex loads it when a task matches its description.

Can I use Ag2 Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ag2ai/build-with-ag2 --skill ag2-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ag2-testing, .gemini/skills/ag2-testing, .github/skills/ag2-testing and .opencode/skills/ag2-testing in your project.

What does Ag2 Testing need to run?

Going by SKILL.md and its folder, Ag2 Testing needs credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Ag2 Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ag2 Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ag2 Testing use?

Ag2 Testing is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ag2 Testing use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ag2 Testing?

Skills that share tags, products or a category with Ag2 Testing: Adk Agent Builder (google/adk-python, 22k stars), Adk Verify Snippets (google/adk-python, 22k stars), Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars) and Test Guard (amElnagdy/guard-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ag2 Testing?

ag2ai (a GitHub organization) maintains it in ag2ai/build-with-ag2, which has 252 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 6, 2026.

Source: ag2ai/build-with-ag2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.