Agent skill

Gaia Debugging

by ruvnet in ruvnet/ruflo

Diagnose why a GAIA question failed — extract trace, classify failure mode, and propose a fix.

MITAuto-check: notesDevelopment

Install Gaia Debugging

skills CLI
$ npx skills add ruvnet/ruflo --skill gaia-debugging -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo gaia-debugging --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ruflo-workflows/skills/gaia-debugging .claude/skills/gaia-debugging && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gaia-debugging
GitHub stars
74k
Token cost
~1.1k tokens
SKILL.md length
339 words
Files
1
Skills in repo
264
Repo updated
First seen
Licence
MIT

At a glance

Diagnose why a GAIA question failed — extract trace, classify failure mode, and propose a fix.

  • Works in 5 steps: Load the question trace → Classify the failure → Re-run with extended logging → …
  • A GAIA benchmark run reports a failed/incorrect taskid and you need to root-cause it before resubmitting
  • SKILL.md covers When to use, Failure mode taxonomy, Diagnostic workflow and Quick reference: tool…, plus 1 more section
  • Calls node and npx; needs GOOGLE_AI_API_KEY

What it does

Gaia Debugging is an agent skill from ruvnet/ruflo. Diagnose why a GAIA question failed — extract trace, classify failure mode, and propose a fix. Use when a GAIA benchmark run reports a failed/incorrect taskid and you need to root-cause it before resubmitting.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Debugging and Root cause analysis. The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

When your agent uses it

  • A GAIA benchmark run reports a failed/incorrect taskid and you need to root-cause it before resubmitting
  • Tasks that involve Debugging
  • Tasks that involve Root cause analysis

Example prompts

  • “/gaia-debugging”

Requirements

  • Node.js
  • A credential in GOOGLE_AI_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read, mcp__plugin_ruflo-core_ruflo__memory_search, mcp__plugin_ruflo-core_ruflo__memory_store, mcp__plugin_ruflo-core_ruflo__agentdb_pattern_search, mcp__plugin_ruflo-core_ruflo__agentdb_pattern_store

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Load the question trace
  2. Classify the failure
  3. Re-run with extended logging
  4. Apply targeted fix
  5. Verify fix and store pattern

What it can do on your machine

Read from SKILL.md and the folder at commit 6051f67. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • mcp__plugin_ruflo-core_ruflo__memory_search
    • mcp__plugin_ruflo-core_ruflo__memory_store
    • mcp__plugin_ruflo-core_ruflo__agentdb_pattern_search
    • mcp__plugin_ruflo-core_ruflo__agentdb_pattern_store

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_AI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gaia Debugging loads about 1.1k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 339 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, mcp__plugin_ruflo-core_ruflo__memory_search, mcp__plugin_ruflo-core_ruflo__memory_store,

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit 6051f67, republished under its MIT licence (© ruvnet). 339 words, ~1,102 tokens.

Download SKILL.mdSave it as .claude/skills/gaia-debugging/SKILL.md (or your agent's skills folder).
name
gaia-debugging
description
Diagnose why a GAIA question failed — extract trace, classify failure mode, and propose a fix. Use when a GAIA benchmark run reports a failed/incorrect task_id and you need to root-cause it before resubmitting.
allowed-tools
Bash, Read, mcp__plugin_ruflo-core_ruflo__memory_search, mcp__plugin_ruflo-core_ruflo__memory_store, mcp__plugin_ruflo-core_ruflo__agentdb_pattern_search, mcp__plugin_ruflo-core_ruflo__agentdb_pattern_store
argument-hint
<task_id> [--results=<path>]

GAIA Debugging Skill

When a GAIA question fails, systematically diagnose the root cause and propose a targeted fix.

When to use

  • A specific task_id returns the wrong answer or times out
  • Pass-rate dropped between two runs and you need to find the regression
  • You want to understand why a particular question class is consistently failing

Failure mode taxonomy

CodeModeSymptomFix direction
TGTool GapAgent lacks a required tool (no image OCR, no PDF reader)Add tool to catalogue
RMReasoning MissAgent has the right data but draws wrong conclusionImprove system prompt, add CoT instruction
EBExtraction BugAnswer is in the trace but FINAL_ANSWER: regex failsFix answer extraction pattern
LILoop IssueAgent loops (re-asks same tool call) and hits turn limitIncrease max-turns or add loop-detection
DSDataset ShiftGround truth differs from what web currently showsFlag for HAL dataset audit
ATAPI TimeoutTool call times out; agent never gets the resultIncrease per-turn timeout

Diagnostic workflow

Step 1 — Load the question trace
bash
# Find the result for the task_id in the latest run
RESULTS=~/.cache/ruflo/gaia/results-latest.json
node -e "
  const r = JSON.parse(require('fs').readFileSync('$RESULTS'));
  const q = r.results.find(x => x.task_id === '$TASK_ID');
  console.log(JSON.stringify(q, null, 2));
"
Step 2 — Classify the failure

Look at the trace output:

  1. No tools called at all → RM or configuration issue
  2. Tool called but returned error → TG or AT
  3. Tool returned data, wrong answer → RM or EB
  4. Correct answer in trace but marked wrong → EB
  5. max-turns hit → LI or question too hard for current model
Step 3 — Re-run with extended logging
bash
node v3/@claude-flow/cli/bin/cli.js gaia-bench run \
  --level 1 --limit 1 \
  --task-id $TASK_ID \
  --models claude-sonnet-4-6 \
  --max-turns 20 \
  --output json
Step 4 — Apply targeted fix
FailureAction
TG — missing web_browseVerify gaia-tools/index.ts exports web_browse; check tool registration
TG — missing image OCRAdd image_describe tool call; verify GOOGLE_AI_API_KEY
RM — reasoningAdd a system prompt instruction: "Before answering, list all facts you have gathered"
EB — extractionTest the FINAL_ANSWER_RE regex against the trace manually
LI — loopAdd a tool-call deduplication guard in gaia-agent.ts
AT — timeoutSet DEFAULT_PER_TURN_TIMEOUT_MS higher or use --max-turns flag
Step 5 — Verify fix and store pattern
bash
# Re-run the single question
node … gaia-bench run --task-id $TASK_ID --models $MODEL --output json

# If now passing, store the pattern
npx @claude-flow/cli@latest memory store \
  --namespace gaia-debug-patterns \
  --key "fix-$FAILURE_CODE-$(date +%Y%m%d)" \
  --value "task_id=$TASK_ID, mode=$FAILURE_CODE, fix=$FIX_DESCRIPTION"

Quick reference: tool catalogue check

bash
node -e "
  const { createDefaultToolCatalogue } = require('./v3/@claude-flow/cli/src/benchmarks/gaia-tools/index.js');
  const cat = createDefaultToolCatalogue({});
  console.log('Tools registered:', cat.definitions.map(t => t.name));
"

Expected: web_search, file_read, web_browse, image_describe, python_exec

Pattern storage

After resolving a debugging session, store the finding:

bash
npx @claude-flow/cli@latest memory store \
  --namespace gaia-debug-patterns \
  --key "session-$(date +%Y%m%d-%H%M)" \
  --value '{"task_id":"$TASK_ID","failure_mode":"$CODE","fix":"$FIX","verified":true}'

Search for similar past failures:

bash
npx @claude-flow/cli@latest memory search \
  --namespace gaia-debug-patterns \
  --query "extraction bug final answer regex"

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ruflo-workflows/skills/gaia-debugging of ruvnet/ruflo.

Open the folder on GitHubat commit 6051f67

Compare with similar skills

Gaia Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gaia Debugging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gaia Debugging this skillruvnet/ruflo74k—~1.1kAutomated safety check: NotesMIT
OpenLogi macOS Permissions TriageAprilNEA/OpenLogi23k—~2.5kAutomated safety check: NotesApache-2.0
Bug Finder for daisyUIsaadeghi/daisyui43k—~2.3kAutomated safety check: PassMIT
Root Cause Debugginggarrytan/gstack136k—~1.4kAutomated safety check: PassMIT
Graph-Based Bug Tracingtirth8205/code-review-graph32k1 repos~287Automated safety check: PassMIT
Systematic DebuggingChrisWiles/claude-code-showcase6.1k3 repos~1.2kAutomated safety check: PassNone

Similar skills

  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes
  • Bug Finder for daisyUI

    saadeghi/daisyui

    Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.

    43k GitHub stars~2.3k tokensUpdated 8 days ago
    DevelopmentAuto-check passed
  • Root Cause Debugging

    garrytan/gstack

    Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.

    136k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Graph-Based Bug Tracing

    tirth8205/code-review-graph

    Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.

    32k GitHub starsUsed in 1 repo~287 tokens
    DevelopmentAuto-check passed
  • Systematic Debugging

    ChrisWiles/claude-code-showcase

    Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written.

    6.1k GitHub starsUsed in 3 repos~1.2k tokens
    DevelopmentAuto-check passed
  • Debugging and Error Recovery

    addyosmani/agent-skills

    Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

    103k GitHub starsUsed in 1 repo~2.6k tokens
    DevelopmentAuto-check passed

More from ruvnet/ruflo

All 264 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 3 repos~830 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 2 repos~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 2 repos~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 2 repos~779 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Agent Coordination

    ruvnet/ruflo

    Reference for spawning, listing, monitoring and stopping agents with claude-flow commands, with agent type families, routing codes and coordination tips.

    74k GitHub starsUsed in 2 repos~519 tokens
    Auto-check passed

Categories

Questions about Gaia Debugging

What does Gaia Debugging do?

Diagnose why a GAIA question failed — extract trace, classify failure mode, and propose a fix. Gaia Debugging is an agent skill from ruvnet/ruflo. Diagnose why a GAIA question failed — extract trace, classify failure mode, and propose a fix.

When should I use Gaia Debugging?

Gaia Debugging fits situations like: A GAIA benchmark run reports a failed/incorrect taskid and you need to root-cause it before resubmitting; tasks that involve Debugging; tasks that involve Root cause analysis.

How do I install Gaia Debugging in Claude Code?

Run `npx skills add ruvnet/ruflo --skill gaia-debugging -a claude-code`. Or copy the skill folder (plugins/ruflo-workflows/skills/gaia-debugging in ruvnet/ruflo) into .claude/skills/gaia-debugging in your project. Claude Code loads it when a task matches its description.

How do I install Gaia Debugging in Codex?

Run `npx skills add ruvnet/ruflo --skill gaia-debugging -a codex`. Or copy the skill folder (plugins/ruflo-workflows/skills/gaia-debugging in ruvnet/ruflo) into .agents/skills/gaia-debugging in your project. Codex loads it when a task matches its description.

Can I use Gaia Debugging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill gaia-debugging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gaia-debugging, .gemini/skills/gaia-debugging, .github/skills/gaia-debugging and .opencode/skills/gaia-debugging in your project.

What does Gaia Debugging need to run?

Going by SKILL.md and its folder, Gaia Debugging needs the command-line tools its instructions call (node and npx) and credentials named GOOGLE_AI_API_KEY. Our summary lists: Node.js; A credential in GOOGLE_AI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, mcp__plugin_ruflo-core_ruflo__memory_search, mcp__plugin_ruflo-core_ruflo__memory_store, mcp__plugin_ruflo-core_ruflo__agentdb_pattern_search, mcp__plugin_ruflo-core_ruflo__agentdb_pattern_store.

Does Gaia Debugging access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Gaia Debugging safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Gaia Debugging use?

Gaia Debugging is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gaia Debugging use?

About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gaia Debugging?

Skills that share tags, products or a category with Gaia Debugging: OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars), Bug Finder for daisyUI (saadeghi/daisyui, 43k stars), Root Cause Debugging (garrytan/gstack, 136k stars) and Graph-Based Bug Tracing (tirth8205/code-review-graph, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gaia Debugging?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,089 GitHub stars. The repository holds 264 skills in this directory. The repository was last updated on October 8, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.