Agent skill

Research Fuzzer

by ARA-Labs in ARA-Labs/Agent-Native-Research-Artifact

Treat an open-ended investigation the way a fuzzer treats a program.

MITAuto-check passedSecurity

Install Research Fuzzer

skills CLI
$ npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-fuzzer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ARA-Labs/Agent-Native-Research-Artifact research-fuzzer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ARA-Labs/Agent-Native-Research-Artifact.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/research-fuzzer .claude/skills/research-fuzzer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
research-fuzzer
GitHub stars
691
Token cost
~2.4k tokens
SKILL.md length
1,133 words
Files
4 (incl. scripts, references)
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Treat an open-ended investigation the way a fuzzer treats a program.

  • ANY investigation without a known map: research experiments
  • SKILL.md covers Works anywhere — by construction, The notebook, The rules (hard — these ARE… and The panel, plus 1 more section
  • Runs Python scripts from its folder; calls python3
  • Market research

What it does

Research Fuzzer is an agent skill from ARA-Labs/Agent-Native-Research-Artifact. Treat an open-ended investigation the way a fuzzer treats a program. After every action, reflect on two things a fuzzer always knows and an agent never does: did anything NEW happen, and where have I NOT been yet. Keeps an append-only notebook of predictions and outcomes; reports what you explored, the leads you saw but never followed, unexplained results, and a going-in-circles alarm. Use for ANY investigation without a known map: research experiments, debugging, data analysis, literature or market research…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `scripts/log.py` and `scripts/tally.py`).

It sits in Security, covering Fuzzing. The repository describes itself as: Research Artifact Protocol for Rigorous and Trustworthy AI Scientists. The licence is MIT.

When your agent uses it

  • ANY investigation without a known map: research experiments
  • Market research

Example prompts

  • “/research-fuzzer”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit e52a925. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Research Fuzzer loads about 2.4k tokens when it runs, and up to ~2.8k if it reads all its reference files. Until then it costs about 182 tokens; SKILL.md has 1,133 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~182
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ARA-Labs/Agent-Native-Research-Artifact at commit e52a925, republished under its MIT licence (© ARA-Labs). 1,133 words, ~2,441 tokens.

Download SKILL.mdSave it as .claude/skills/research-fuzzer/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
research-fuzzer
description
Treat an open-ended investigation the way a fuzzer treats a program. After every action, reflect on two things a fuzzer always knows and an agent never does: did anything NEW happen, and where have I NOT been yet. Keeps an append-only notebook of predictions and outcomes; reports what you explored, the leads you saw but never followed, unexplained results, and a going-in-circles alarm. Use for ANY investigation without a known map: research experiments, debugging, data analysis, literature or market research, evaluations. Fire it (1) when starting an investigation, (2) after every action or batch of actions that returned results, (3) before stating any conclusion. Skip it for trivial single-step tasks.

research-fuzzer — the fuzzer's feedback loop, for any investigation

A greybox fuzzer almost never finds the bug on a given run. It still wins, because every single run answers one cheap question: did that reach somewhere new? The answer steers the next input. An agent investigating an open-ended question has no such loop by default: it reacts to its last result, grinds the same corner, and stops when the budget runs out — sampling, not searching.

This skill gives you the fuzzer's loop. The mapping is exact:

the fuzzer hasyou keep (the panel's word)plain words
coverage mapexplored — settled bets, per sub-questionwhere you have already walked
seed queueleads — noticed but never trieddoors you passed, never opened
crash reportsunexplained — outcomes you cannot explainthe confusion ledger
"no new coverage"novelty — recent results that taught you nothingthe going-in-circles alarm
triage before reportinggate — before any conclusionno claim until the books are clean

Fuzzing vocabulary stops here: the panel and the rules below use the neutral words, so the skill does not nudge you toward software-shaped experiments when your world is biology, markets, or people.

Nobody — you included — can know what fraction of the world you have covered; that number needs a god's-eye view that does not exist. Every reading above is computed from your own footprints instead. That is the whole trick, and it is the same trick fuzzers use: coverage is always measured against what you have seen, never against all possible behaviors.

Works anywhere — by construction

  • Any domain. An "action" is anything that returns information: an experiment, a query, a benchmark run, an interview, a paper read, a grep. Predictions may be quantitative ("loss < 0.5") or qualitative ("most users will cite price"); they only need to be falsifiable.
  • Any agent, any harness. This file is the skill. Everything below is executable by hand with no tooling at all; scripts/tally.py (stdlib-only Python) is an optional convenience that computes the same panel faster. No network, no packages, no framework.
  • Any timescale. The notebook is one append-only file at the investigation root — it survives context loss, session restarts, and handoffs to other agents. A new session starts by reading the notebook and printing the panel.
  • When not to use it. Single-step lookups and trivial fixes. The loop earns its overhead only when the answer is genuinely unknown and multiple actions will be needed.
  • Investigation, not synthesis. This loop is built for probing a world that already exists — why is X happening, what law governs Y, where is the bug. For creative work (designing a system, constructing a proof, writing), apply it only to the investigative episodes inside the work — "will this design choice survive load?" is a bet; the act of creation itself is not.

The notebook

One append-only file at the investigation root: fuzz-notebook.jsonl. Never edit or delete past lines — append corrections as new lines.

Prefer the writer script when Python is available — it owns the format (escaping, ids, timestamps) and rejects rule violations at write time, so you never hand-craft JSON:

bash
python3 scripts/log.py bet    --action "interview 5 churned users" \
                              --prediction "most cite price" --confidence 0.6 --target "price?"
python3 scripts/log.py settle --bet b1 --outcome "4/5: need disappeared" \
                              --verdict surprised --update "churn may not be dissatisfaction" \
                              --door "interview RETAINED users"

Hand-append JSON lines only when no tooling exists — one record per line, newlines inside strings escaped as \n. Five record types (entries may be written in any language):

jsonl
{"type":"frame","v":1,"question":"why are users churning?","subq":["price?","UX?","competitors?"]}
{"type":"bet","id":"b1","action":"interview 5 churned users","prediction":"most cite price","confidence":0.6,"target":"price?"}
{"type":"settle","bet":"b1","outcome":"4/5 said their need disappeared; price never mentioned","verdict":"surprised","update":"churn may not be dissatisfaction — add 'need evaporated?' branch","doors":["interview RETAINED users","segment churn by tenure"]}
{"type":"reframe","from":1,"to":2,"trigger":["b1","b4","b7"],"change":"added subq: need-evaporation; demoted price"}
{"type":"claim","statement":"churn is driven by X","scope":"within frame v2 and the doors entered","kill_shot":"b14","limitations":["door 'interview RETAINED users' never entered"]}
Show full SKILL.md (599 more words)Show less

The rules (hard — these ARE the skill)

1. No action without a bet. Before every action, append one bet line: what you will do, what you predict will happen, how confident, and which sub-question it serves. The prediction is written BEFORE you see the result — that ordering is the entire defense against fooling yourself. A result you would have "expected either way" teaches nothing; make the prediction sharp enough to be wrong.

2. Every settle picks exactly one verdict.

  • confirmed — the prediction held. You learned little. A streak of confirmed is a leave signal, not an achievement: no new coverage here.
  • surprised — the prediction broke. You MUST write the update: which belief changed, in one sentence. A surprise without an update line is not settled.
  • anomaly — you cannot explain it. It goes on the unexplained-results ledger, may never be silently dropped, and blocks the gate (rule 5) until resolved or explicitly carried as a limitation.

3. Every settle asks: what did this open? Each result exposes leads you now know you could chase (a segment you noticed, a condition you held fixed, a control you lack). Append them to doors — even when you do not take them. This queue is your only honest measure of "how much is left": never "40% of the world" (unknowable), always "9 leads seen and never tried" (exact).

4. Three unresolved anomalies force a reframe. When unexplained results cluster, stop asking "which belief is wrong" and ask: "what axis is missing from my frame for these to make sense?" Publish frame v(N+1) with the new sub-question. A reframe is not an embarrassment; it is usually the most valuable single discovery of the run — record it as one.

5. The gate: no claim without passing it. Before stating any conclusion:

  • it must account for every settled bet, or the exceptions are listed;
  • the unexplained-results ledger is empty, or each open item is named in limitations;
  • you have bet and settled a kill shot — the one action you believe most likely to break the claim, aimed somewhere you have not yet tested;
  • the claim carries its scope: "within frame vN and the doors entered" — never an unscoped "the answer is X".

The panel

After each batch, produce the panel — run python scripts/tally.py fuzz-notebook.jsonl if Python is available, otherwise compute it by hand from the notebook (every reading is a count):

PANEL — frame v2 · 23 bets settled
  explored    : price? 6 · UX? 4 · need-evaporated? 8
  leads       : 9 untried · 5 followed
                oldest untried: "interview RETAINED users" (seen 12 settles ago)
  novelty     : last 10 settles → 8 confirmed, 1 surprised, 1 anomaly
                nothing new here lately — consider a far lead
  unexplained : 1 open (b17: retention spike in oldest cohort)
  spread      : 8 of your last 10 bets probed the same area (need-evaporated?)
  gate        : BLOCKED (1 unexplained open; no kill-shot bet on record)

The statistics window adapts to your pace: with only a handful of settles (slow, expensive experiments), the panel reports plain counts and withholds trend advice — a "last 10" reading means nothing when you have run 3.

Read it, then decide. The panel informs; it never commands. It only makes your own footprints visible: where you walked, what you saw and skipped, whether the last stretch taught you anything. State in one sentence what the readings imply before placing the next bet — a reading that is never read is not feedback.

What this is not

  • The panel steers; it never certifies. Passing the gate means "not yet proven wrong within what you explored" — never "correct".
  • Confirmed streaks feel like progress. They are the fuzzer's strongest signal to move: this region yields no new coverage.
  • The queue only contains doors you noticed. Doors you never saw are not on it — which is why conclusions carry scope, and why a fat queue is a sign of good peripheral vision, not of failure.
  • Do not tune the instrument mid-run: no reclassifying anomalies as confirmed, no editing old lines, no writing bets after seeing results. Editing the instrument to please the reading is the exact failure this skill exists to prevent.

© ARA-Labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in skills/research-fuzzer of ARA-Labs/Agent-Native-Research-Artifact.

  • SKILL.md
  • references/example-notebook.jsonl
  • scripts/log.py
  • scripts/tally.py

Open the folder on GitHubat commit e52a925

Compare with similar skills

Research Fuzzer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Research Fuzzer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Research Fuzzer this skillARA-Labs/Agent-Native-Research-Artifact691—~2.4kAutomated safety check: PassMIT
Fizzpashov/skills1.2k2 repos~11kAutomated safety check: PassMIT
Fizz Syncpashov/skills1.2k2 repos~3.9kAutomated safety check: PassMIT
Vuln Researchtanweai/xianzhi-research185—~847Automated safety check: PassNone
Binary Reverse Engineering Audittihanyin/REx-skill108—~5.1kAutomated safety check: PassMIT
Harness Design Fuzzingprovos/ironcurtain612—~5.7kAutomated safety check: PassApache-2.0

Similar skills

  • Fizz

    pashov/skills

    Generate Echidna/Medusa-compatible Solidity fuzz suites from Foundry or Hardhat projects.

    1.2k GitHub starsUsed in 2 repos~11k tokens
    SecurityAuto-check passed
  • Fizz Sync

    pashov/skills

    Reconcile an existing Fizz harness with a changed source tree.

    1.2k GitHub starsUsed in 2 repos~3.9k tokens
    SecurityAuto-check passed
  • Vuln Research

    tanweai/xianzhi-research

    安全研究元思考方法论 - 从先知社区5600+篇安全文档中提炼的漏洞挖掘方法论框架. An agent skill from tanweai/xianzhi-research.

    185 GitHub stars~847 tokensUpdated 8 mo ago
    SecurityAuto-check passed
  • Guides evidence-first reverse engineering of compiled programs to find and prove defects, from triage and decompilation to fuzzing, patch diffing and firmware.

    108 GitHub stars~5.1k tokensUpdated 17 days ago
    SecurityAuto-check passed
  • Harness Design Fuzzing

    provos/ironcurtain

    Reference vocabulary for designing instrumented harnesses that drive vulnerability discovery — design classes (trigger-driven vs coverage-driven), tiered scope (T1 isolated function / T2…

    612 GitHub stars~5.7k tokensUpdated 4 days ago
    SecurityAuto-check passed
  • Run Fuzzing Campaign

    opensage-agent/opensage-adk

    Run a fuzzing campaign using AFL++ with optional seeds; supports --custommutatorpath (you can write your own custom mutator and use this to execute).

    127 GitHub stars~542 tokensUpdated 2 mo ago
    SecurityAuto-check passed

More from ARA-Labs/Agent-Native-Research-Artifact

  • Submit Ara

    ARA-Labs/Agent-Native-Research-Artifact

    ARA Submitter. An agent skill from ARA-Labs/Agent-Native-Research-Artifact.

    691 GitHub stars~2.4k tokensUpdated 4 days ago
    Auto-check passed
  • Rigor Reviewer

    ARA-Labs/Agent-Native-Research-Artifact

    ARA Seal Level 2: Semantic Epistemic Review. An agent skill from ARA-Labs/Agent-Native-Research-Artifact.

    691 GitHub stars~4.7k tokensUpdated 4 days ago
    Auto-check passed
  • Compiler

    ARA-Labs/Agent-Native-Research-Artifact

    Universal ARA Compiler. An agent skill from ARA-Labs/Agent-Native-Research-Artifact.

    691 GitHub stars~6.4k tokensUpdated 4 days ago
    Auto-check passed
  • Research Foresight

    ARA-Labs/Agent-Native-Research-Artifact

    ARA World Model — read-only reasoning engine over ONE Agent-Native Research Artifact (ARA), run LOCALLY with the coding agent itself as the LLM (no SDK, no API key).

    691 GitHub stars~626 tokensUpdated 4 days ago
    Auto-check passed
  • Research Manager

    ARA-Labs/Agent-Native-Research-Artifact

    End-of-turn research process recorder with progressive crystallization.

    691 GitHub stars~9k tokensUpdated 4 days ago
    Auto-check passed

Questions about Research Fuzzer

What does Research Fuzzer do?

Treat an open-ended investigation the way a fuzzer treats a program. Research Fuzzer is an agent skill from ARA-Labs/Agent-Native-Research-Artifact. Treat an open-ended investigation the way a fuzzer treats a program.

When should I use Research Fuzzer?

Research Fuzzer fits situations like: ANY investigation without a known map: research experiments; market research.

How do I install Research Fuzzer in Claude Code?

Run `npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-fuzzer -a claude-code`. Or copy the skill folder (skills/research-fuzzer in ARA-Labs/Agent-Native-Research-Artifact) into .claude/skills/research-fuzzer in your project. Claude Code loads it when a task matches its description.

How do I install Research Fuzzer in Codex?

Run `npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-fuzzer -a codex`. Or copy the skill folder (skills/research-fuzzer in ARA-Labs/Agent-Native-Research-Artifact) into .agents/skills/research-fuzzer in your project. Codex loads it when a task matches its description.

Can I use Research Fuzzer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-fuzzer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/research-fuzzer, .gemini/skills/research-fuzzer, .github/skills/research-fuzzer and .opencode/skills/research-fuzzer in your project.

What does Research Fuzzer need to run?

Going by SKILL.md and its folder, Research Fuzzer needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Research Fuzzer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Research Fuzzer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Research Fuzzer use?

Research Fuzzer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Research Fuzzer use?

About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 312 tokens, read only when the agent opens those files.

What are the alternatives to Research Fuzzer?

Skills that share tags, products or a category with Research Fuzzer: Fizz (pashov/skills, 1.2k stars), Fizz Sync (pashov/skills, 1.2k stars), Vuln Research (tanweai/xianzhi-research, 185 stars) and Binary Reverse Engineering Audit (tihanyin/REx-skill, 108 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Research Fuzzer?

ARA-Labs (a GitHub organization) maintains it in ARA-Labs/Agent-Native-Research-Artifact, which has 691 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 7, 2026.

Source: ARA-Labs/Agent-Native-Research-Artifact on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.