Agent skill

Autonomous Testing

by alinaqi in alinaqi/maggy

AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type

MITAuto-check passed

Install Autonomous Testing

skills CLI
$ npx skills add alinaqi/maggy --skill autonomous-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alinaqi/maggy autonomous-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alinaqi/maggy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/autonomous-testing .claude/skills/autonomous-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autonomous-testing
GitHub stars
707
Token cost
~1.1k tokens
SKILL.md length
70 words
Files
1
Skills in repo
71
Repo updated
First seen
Licence
MIT

At a glance

AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type

  • Works in 6 steps: Discover — What Needs Testing? → Generate — AI-Written Tests → Execute — Run Everything → …
  • SKILL.md covers Overview, Pipeline, Phase 1: Discover — What Needs… and Phase 2: Generate — AI-Written…, plus 7 more sections
  • Calls python, npx and pytest

What it does

Autonomous Testing is an agent skill from alinaqi/maggy. AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: What started as an opinionated Claude Code setup kit is now an autonomous AI engineering command center. The licence is MIT.

Example prompts

  • “/autonomous-testing”

Requirements

  • Python 3
  • Node.js

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Discover — What Needs Testing?
  2. Generate — AI-Written Tests
  3. Execute — Run Everything
  4. Evaluate — AI-Powered Assessment
  5. Fix — Autonomous Repair
  6. Report — Structured Output

What it can do on your machine

Read from SKILL.md and the folder at commit 72a456e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • npx
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autonomous Testing loads about 1.1k tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 70 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from alinaqi/maggy at commit 72a456e, republished under its MIT licence (© alinaqi). 70 words, ~1,145 tokens.

Download SKILL.mdSave it as .claude/skills/autonomous-testing/SKILL.md (or your agent's skills folder).
name
autonomous-testing
description
AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type
when-to-use
When setting up automated test generation or running an autonomous test-fix loop across Python, TypeScript, API, or web projects
user-invocable
false
effort
high

Autonomous Testing Agent

Overview

An AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type. Inspired by the edubites autonomous test runner pattern, generalized for Claude Bootstrap + Maggy.

Pipeline

Source Scan → Discover Gaps → Generate Tests → Execute → Evaluate → Report → Fix Loop

Phase 1: Discover — What Needs Testing?

Auto-detect project type:
  Python    → scan for *.py files, extract public functions/classes
  TypeScript → scan for *.ts/*.tsx files, extract exports
  API       → scan FastAPI/Express routes, extract endpoints + methods
  Web       → scan React/Vue components, extract user flows

Map existing tests:
  Python    → pytest --collect-only
  TypeScript → vitest --list
  API       → scan tests/ for endpoint coverage

Compute coverage gaps:
  - Functions with 0 tests
  - API endpoints with 0 tests
  - Components with 0 tests
  - Branches with <80% coverage

Phase 2: Generate — AI-Written Tests

For each uncovered function/endpoint/component:
  1. Read source code → understand inputs, outputs, edge cases
  2. Generate test scaffold using ~/bin/deepseek --pro
  3. Include: happy path, error cases, edge cases, auth checks
  4. Write to appropriate test directory

Model routing for generation:
  - Simple functions    → ~/bin/deepseek --flash (cheap, fast)
  - Complex logic       → ~/bin/deepseek --pro (thorough)
  - Auth/security tests → ~/bin/deepseek --pro (quality-critical)

Phase 3: Execute — Run Everything

bash
# Python
pytest -x --cov --cov-report=json

# TypeScript
npx vitest run --coverage

# E2E (if Playwright detected)
npx playwright test

# Parse results → structured TestRun { pass/fail, coverage, duration, failures[] }

Phase 4: Evaluate — AI-Powered Assessment

For each test failure:
  1. Capture: test name, error message, stack trace, source code diff
  2. Classify failure:
     - TEST_BUG: test is wrong (outdated expectation, bad mock)
     - CODE_BUG: code is wrong (regression, edge case)
     - ENV_BUG: environment issue (missing dep, config)
  3. AI evaluation: ~/bin/deepseek --pro analyzes failure and classifies

For E2E/web tests:
  - Capture screenshots at failure points
  - ~/bin/gemini --flash evaluates visual state (multimodal)

Phase 5: Fix — Autonomous Repair

TEST_BUG → regenerate test with corrected expectation
CODE_BUG → propose fix with ~/bin/deepseek --pro, apply, re-run
ENV_BUG → report to user with fix instructions

Auto-fix loop:
  while test_failures > 0 and attempts < 3:
    for each failure:
      classify → fix → re-run
    if fixed: record as "auto-fixed"
    if not: escalate to CLAUDE tier

Phase 6: Report — Structured Output

json
{
  "project": "my-app",
  "timestamp": "2026-05-16T12:00:00Z",
  "summary": {
    "tests_run": 247,
    "passed": 231,
    "failed": 12,
    "auto_fixed": 8,
    "needs_manual": 4,
    "coverage": 0.83
  },
  "gaps_found": 15,
  "tests_generated": 15,
  "next_actions": [
    "4 manual fixes needed in auth module",
    "Coverage gap: src/payment.py has 0 tests",
    "3 E2E flows untested: signup, checkout, profile-edit"
  ]
}

Integration with Maggy

Maggy Dashboard → Testing tab shows:
  - Coverage trend over time
  - Auto-generated test count
  - Failure classification (TEST_BUG vs CODE_BUG)
  - "Generate tests for gaps" one-click button

Heartbeat job: auto-generate tests weekly for new untested code
Auto-review hook: triggers test generation after significant PR merges

Usage

bash
# Discover test gaps
maggy test discover

# Generate tests for all gaps
maggy test generate --all

# Generate tests for specific module
maggy test generate --module auth

# Run full test cycle (discover → generate → execute → fix → report)
maggy test autonomous

# Watch mode — auto-test on file changes
maggy test watch

Configuration

json
// ~/.claude/testing-config.json
{
  "auto_generate": true,
  "auto_fix": true,
  "max_fix_attempts": 3,
  "min_coverage": 0.8,
  "generate_model": "deepseek-pro",
  "evaluate_model": "gemini-flash",
  "fix_model": "deepseek-pro",
  "exclude_patterns": ["*/migrations/*", "*/node_modules/*"]
}

© alinaqi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/autonomous-testing of alinaqi/maggy.

Open the folder on GitHubat commit 72a456e

Compare with similar skills

Autonomous Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autonomous Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autonomous Testing this skillalinaqi/maggy707—~1.1kAutomated safety check: PassMIT
Executealirezarezvani/claude-skills28k—~831Automated safety check: PassMIT
Generatealirezarezvani/claude-skills28k1 repos~1.1kAutomated safety check: PassMIT
Fal Generatenexu-io/open-design100k—~306Automated safety check: PassApache-2.0
Generate Then Executemohitagw15856/pm-claude-skills1.4k—~938Automated safety check: PassMIT
Video Generationbytedance/deer-flow83k4 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Execute

    alirezarezvani/claude-skills

    /cs:execute <decision — Generate a 90-day execution plan with weekly milestones, DRIs, and check-in cadence from an approved decision.

    28k GitHub stars~831 tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • Generate

    alirezarezvani/claude-skills

    Generate Playwright tests. An agent skill from alirezarezvani/claude-skills.

    28k GitHub starsUsed in 1 repo~1.1k tokens
    Testing & QAAuto-check passed
  • Fal Generate

    nexu-io/open-design

    Generate images and videos using fal.ai AI models. An agent skill from nexu-io/open-design.

    100k GitHub stars~306 tokensUpdated today
    Media & CreativeAuto-check passed
  • Generate Then Execute

    mohitagw15856/pm-claude-skills

    Separate raw idea-generation from judgment so creativity isn't strangled by your inner critic — diverge with zero evaluation, then switch to hard critique.

    1.4k GitHub stars~938 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    83k GitHub starsUsed in 4 repos~1.4k tokens
    Media & CreativeAuto-check passed
  • Official

    Debug failed or wrong-output workflow executions using executions tools.

    207k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed

More from alinaqi/maggy

All 71 skills in this repo
  • AI Models

    alinaqi/maggy

    Latest AI models reference - Claude, OpenAI, Gemini, Eleven Labs, Replicate

    707 GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Azure Cosmosdb

    alinaqi/maggy

    Azure Cosmos DB partition keys, consistency levels, change feed, SDK patterns

    707 GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • LLM Patterns

    alinaqi/maggy

    AI-first application patterns, LLM testing, prompt management

    707 GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Woocommerce

    alinaqi/maggy

    WooCommerce REST API - products, orders, customers, webhooks

    707 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check: notes
  • Aeo Optimization

    alinaqi/maggy

    AI Engine Optimization - semantic triples, page templates, content clusters for AI citations

    707 GitHub stars~3.7k tokensUpdated 14 days ago
    Auto-check passed
  • Agent Teams

    alinaqi/maggy

    Claude Code Agent Teams - default team-based development with strict TDD pipeline enforcement

    707 GitHub stars~5k tokensUpdated 14 days ago
    Auto-check: notes

Questions about Autonomous Testing

What does Autonomous Testing do?

AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type. Autonomous Testing is an agent skill from alinaqi/maggy.

How do I install Autonomous Testing in Claude Code?

Run `npx skills add alinaqi/maggy --skill autonomous-testing -a claude-code`. Or copy the skill folder (skills/autonomous-testing in alinaqi/maggy) into .claude/skills/autonomous-testing in your project. Claude Code loads it when a task matches its description.

How do I install Autonomous Testing in Codex?

Run `npx skills add alinaqi/maggy --skill autonomous-testing -a codex`. Or copy the skill folder (skills/autonomous-testing in alinaqi/maggy) into .agents/skills/autonomous-testing in your project. Codex loads it when a task matches its description.

Can I use Autonomous Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alinaqi/maggy --skill autonomous-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autonomous-testing, .gemini/skills/autonomous-testing, .github/skills/autonomous-testing and .opencode/skills/autonomous-testing in your project.

What does Autonomous Testing need to run?

Going by SKILL.md and its folder, Autonomous Testing needs the command-line tools its instructions call (python, npx and pytest). Our summary lists: Python 3; Node.js.

Does Autonomous Testing access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Autonomous Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Autonomous Testing use?

Autonomous Testing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Autonomous Testing use?

About 1.1k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Autonomous Testing?

Skills that share tags, products or a category with Autonomous Testing: Execute (alirezarezvani/claude-skills, 28k stars), Generate (alirezarezvani/claude-skills, 28k stars), Fal Generate (nexu-io/open-design, 100k stars) and Generate Then Execute (mohitagw15856/pm-claude-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autonomous Testing?

alinaqi (a GitHub user) maintains it in alinaqi/maggy, which has 707 GitHub stars. The repository holds 71 skills in this directory. The repository was last updated on September 24, 2026.

Source: alinaqi/maggy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.