Agent skill

Gaia Submission

by ruvnet in ruvnet/ruflo

Walk through a complete GAIA benchmark→submit flow — from key resolution through HAL-compatible package generation

MITAuto-check: notes

Install Gaia Submission

skills CLI
$ npx skills add ruvnet/ruflo --skill gaia-submission -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ruvnet/ruflo gaia-submission --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ruvnet/ruflo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ruflo-workflows/skills/gaia-submission .claude/skills/gaia-submission && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gaia-submission
GitHub stars
74k
Token cost
~1.2k tokens
SKILL.md length
424 words
Files
1
Skills in repo
265
Repo updated
First seen
Licence
MIT

At a glance

Walk through a complete GAIA benchmark→submit flow — from key resolution through HAL-compatible package generation

  • Works in 6 steps: Validate environment → Estimate cost and confirm → Run the benchmark → …
  • SKILL.md covers When to use, Prerequisites, Phase 1 — Validate environment and Phase 2 — Estimate cost and…, plus 5 more sections
  • Calls npx and node; needs ANTHROPIC_API_KEY and HF_TOKEN

What it does

Gaia Submission is an agent skill from ruvnet/ruflo. Walk through a complete GAIA benchmark→submit flow — from key resolution through HAL-compatible package generation

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory…. The licence is MIT.

Example prompts

  • “/gaia-submission”

Requirements

  • Node.js
  • A credential in ANTHROPIC_API_KEY
  • Pre-approved tools (allowed-tools): Bash, mcp__plugin_ruflo-core_ruflo__memory_store, mcp__plugin_ruflo-core_ruflo__memory_search, mcp__plugin_ruflo-core_ruflo__memory_list, mcp__plugin_ruflo-core_ruflo__hooks_post_task, mcp__plugin_ruflo-core_ruflo__hooks_pre_task

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Validate environment
  2. Estimate cost and confirm
  3. Run the benchmark
  4. Package for submission
  5. Compare and report
  6. Persist learnings

What it can do on your machine

Read from SKILL.md and the folder at commit 6c04654. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • mcp__plugin_ruflo-core_ruflo__memory_store
    • mcp__plugin_ruflo-core_ruflo__memory_search
    • mcp__plugin_ruflo-core_ruflo__memory_list
    • mcp__plugin_ruflo-core_ruflo__hooks_post_task
    • mcp__plugin_ruflo-core_ruflo__hooks_pre_task

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY
    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gaia Submission loads about 1.2k tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 424 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, mcp__plugin_ruflo-core_ruflo__memory_store, mcp__plugin_ruflo-core_ruflo__memory_search, mcp__

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ruvnet/ruflo at commit 6c04654, republished under its MIT licence (© ruvnet). 424 words, ~1,239 tokens.

Download SKILL.mdSave it as .claude/skills/gaia-submission/SKILL.md (or your agent's skills folder).
name
gaia-submission
description
Walk through a complete GAIA benchmark→submit flow — from key resolution through HAL-compatible package generation
allowed-tools
Bash, mcp__plugin_ruflo-core_ruflo__memory_store, mcp__plugin_ruflo-core_ruflo__memory_search, mcp__plugin_ruflo-core_ruflo__memory_list, mcp__plugin_ruflo-core_ruflo__hooks_post_task, mcp__plugin_ruflo-core_ruflo__hooks_pre_task
argument-hint
[level] [limit] [models]

GAIA Submission Skill

Walk Claude Code through every step needed to go from a clean environment to a signed, HAL-compatible submission package ready to upload to the Princeton GAIA leaderboard.

When to use

When the user wants to:

  • Run a benchmark and submit results to the HAL leaderboard
  • Package an existing results file into a submission archive
  • Confirm their environment is ready for a benchmark run

Prerequisites

Before starting, confirm these are available:

RequirementCheck
ANTHROPIC_API_KEYecho ${ANTHROPIC_API_KEY:0:8}… (should show sk-ant-…)
HF_TOKENecho ${HF_TOKEN:0:5}… (should show hf_…)
Node.js 20+node --version
CLI builtnode v3/@claude-flow/cli/bin/cli.js --version

Phase 1 — Validate environment

bash
# Run all pre-flight checks
/gaia validate

If any check fails, resolve it before continuing.

Phase 2 — Estimate cost and confirm

Ask the user for their configuration:

  • Level (default: 1)
  • Question limit (default: 53 for a quick run, 165 for the full L1 set)
  • Models (default: claude-sonnet-4-6)
  • Self-consistency voting (default: 1; use 3 for L2/L3)
bash
/gaia cost --level=$LEVEL --limit=$LIMIT --models=$MODELS --voting=$VOTING

If projected cost > $5, show the estimate and ask: "This run will cost approximately $X. Proceed? (y/N)"

Phase 3 — Run the benchmark

bash
/gaia run --level=$LEVEL --limit=$LIMIT --models=$MODELS --voting=$VOTING

While running, progress is reported every 5 questions:

[12/53] 22.7% (5 passed of 22 scored) — est. remaining: $0.18

Store the run summary in memory for history tracking:

bash
npx @claude-flow/cli@latest memory store \
  --namespace gaia-runs \
  --key "run-$(date +%Y%m%d-%H%M)" \
  --value '{"level":$LEVEL,"model":"$MODEL","total":$TOTAL,"passed":$PASSED,"pass_rate":$RATE,"est_cost_usd":$COST}'

Phase 4 — Package for submission

bash
/gaia submit --results=~/.cache/ruflo/gaia/results-latest.json

This produces:

submission-<date>-<sha>/
├── results.jsonl        ← HAL-compatible, one JSON per line
├── trajectories.jsonl   ← full agent traces
├── metadata.json        ← harness info, model, tool catalogue
├── audit-report.json    ← ADR-167 pre-submission exploit-audit report
├── manifest.md.json     ← Ed25519-signed witness (signs audit-report.json's hash)
└── README.md            ← human summary + leaderboard comparison
Show full SKILL.md (230 more words)Show less
Integrity gate — the audit runs before signing (ADR-167)

Post-RDI (UC Berkeley broke 8 agent benchmarks — GAIA to ~98% — without solving a task), a signature alone is not enough: it proves the bytes are untampered, not that the score was earned. /gaia submit therefore runs a deterministic, $0 exploit audit before signing and refuses to build the leaderboard package on a CRITICAL failure unless --allow-dirty is passed. The audit report is signed into the witness manifest as an ADR-103 fix marker, so a ruflo GAIA submission attests both transport-integrity and earning-integrity.

If the gate blocks, treat it as a real finding — inspect audit-report.json (answer-leakage, no-work pass, oracle leakage, grader monkey-patching, an answer-key read outside the dataset dir, or dynamic eval/exec of task content in the runner) rather than reaching for --allow-dirty. The static source-scan family (answer-key-reads, dynamic-eval, judge-injection) enforces today with no trajectory instrumentation; the trajectory-fed checks the current schema cannot feed are reported as harness_gaps (ADR-167 §7), not passes.

Phase 5 — Compare and report

bash
/gaia leaderboard --level=$LEVEL
/gaia history

Interpret the gap between ruflo's score and the leaderboard top-10. Identify the primary failure mode (tool gap, reasoning miss, extraction bug) using the /gaia-debugging skill if needed.

Phase 6 — Persist learnings

bash
npx @claude-flow/cli@latest hooks post-task \
  --task-id "gaia-submission-$(date +%Y%m%d)" \
  --success true \
  --train-neural true

Store any discovered patterns:

bash
npx @claude-flow/cli@latest memory store \
  --namespace gaia-patterns \
  --key "submission-notes-$(date +%Y%m%d)" \
  --value "Level $LEVEL, $MODEL: $NOTES"

Extensibility note

This skill is intentionally structured to be benchmark-agnostic. The phase headers (validate → estimate → run → package → compare → learn) apply to SWE-bench, WebArena, and HumanEval with only phase 3-4 details changing.

© ruvnet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/ruflo-workflows/skills/gaia-submission of ruvnet/ruflo.

Open the folder on GitHubat commit 6c04654

Compare with similar skills

Gaia Submission next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gaia Submission compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gaia Submission this skillruvnet/ruflo74k—~1.2kAutomated safety check: NotesMIT
Benchmarkaffaan-m/ECC277k3 repos~654Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k—~412Automated safety check: PassMIT
Benchmarkaffaan-m/ECC276k—~330Automated safety check: PassMIT
Benchmarkandroidx/androidx6.1k—~1.1kAutomated safety check: PassApache-2.0
Benchmarksamchon/typia5.9k—~1.2kAutomated safety check: PassMIT

Similar skills

  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    277k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed
  • Benchmark

    affaan-m/ECC

    このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

    276k GitHub stars~412 tokensUpdated yesterday
    Auto-check passed
  • Benchmark

    affaan-m/ECC

    使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。

    276k GitHub stars~330 tokensUpdated yesterday
    Auto-check passed
  • Benchmark

    androidx/androidx

    Benchmarking and improving the performance of Jetpack Compose.

    6.1k GitHub stars~1.1k tokensUpdated yesterday
    MobileAuto-check passed
  • Benchmark

    samchon/typia

    Defines typia benchmark fixture integrity, result reporting, and publication safeguards.

    5.9k GitHub stars~1.2k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from ruvnet/ruflo

All 265 skills in this repo
  • Stores, searches, and retrieves successful patterns with HNSW-indexed semantic search so agents can reuse past solutions instead of relearning them.

    74k GitHub starsUsed in 1 repo~830 tokens
    Auto-check passed
  • Sets up and drives Ruflo, an npm-installed orchestration layer for multi-agent swarms, persistent memory, routing, hooks and its MCP tool catalog.

    74k GitHub starsUsed in 1 repo~975 tokens
    Auto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 1 repo~823 tokens
    Auto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 1 repo~829 tokens
    Auto-check passed
  • Coordinates a hierarchical swarm of specialized agents through the claude-flow CLI for work that spans several files or modules at once.

    74k GitHub starsUsed in 1 repo~779 tokens
    Auto-check passed
  • Finds models on the Hugging Face router that lack descriptions in chat-ui's prod.yaml and dev.yaml, researches each one and adds short descriptions.

    74k GitHub starsUsed in 1 repo~600 tokens
    Auto-check passed

Questions about Gaia Submission

What does Gaia Submission do?

Walk through a complete GAIA benchmark→submit flow — from key resolution through HAL-compatible package generation. Gaia Submission is an agent skill from ruvnet/ruflo.

How do I install Gaia Submission in Claude Code?

Run `npx skills add ruvnet/ruflo --skill gaia-submission -a claude-code`. Or copy the skill folder (plugins/ruflo-workflows/skills/gaia-submission in ruvnet/ruflo) into .claude/skills/gaia-submission in your project. Claude Code loads it when a task matches its description.

How do I install Gaia Submission in Codex?

Run `npx skills add ruvnet/ruflo --skill gaia-submission -a codex`. Or copy the skill folder (plugins/ruflo-workflows/skills/gaia-submission in ruvnet/ruflo) into .agents/skills/gaia-submission in your project. Codex loads it when a task matches its description.

Can I use Gaia Submission in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ruvnet/ruflo --skill gaia-submission -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gaia-submission, .gemini/skills/gaia-submission, .github/skills/gaia-submission and .opencode/skills/gaia-submission in your project.

What does Gaia Submission need to run?

Going by SKILL.md and its folder, Gaia Submission needs the command-line tools its instructions call (npx and node) and credentials named ANTHROPIC_API_KEY and HF_TOKEN. Our summary lists: Node.js; A credential in ANTHROPIC_API_KEY. Its frontmatter pre-approves these tools: Bash, mcp__plugin_ruflo-core_ruflo__memory_store, mcp__plugin_ruflo-core_ruflo__memory_search, mcp__plugin_ruflo-core_ruflo__memory_list, mcp__plugin_ruflo-core_ruflo__hooks_post_task, mcp__plugin_ruflo-core_ruflo__hooks_pre_task.

Does Gaia Submission access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Gaia Submission safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Gaia Submission use?

Gaia Submission is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gaia Submission use?

About 1.2k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gaia Submission?

Skills that share tags, products or a category with Gaia Submission: Benchmark (affaan-m/ECC, 277k stars), Benchmark (affaan-m/ECC, 276k stars), Benchmark (affaan-m/ECC, 276k stars) and Benchmark (androidx/androidx, 6.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gaia Submission?

ruvnet (a GitHub user) maintains it in ruvnet/ruflo, which has 74,222 GitHub stars. The repository holds 265 skills in this directory. The repository was last updated on October 10, 2026.

Source: ruvnet/ruflo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.