Hypothesis Generation
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
Source-backed longform research workflow for AI agents: scope, sources, claims, and independent review.
$ npx skills add rrrrrredy/research-toolkit --skill research-toolkit -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install rrrrrredy/research-toolkit research-toolkit --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "research-toolkit" agent skill from https://github.com/rrrrrredy/research-toolkit/tree/main into .claude/skills/research-toolkit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-toolkit", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rrrrrredy/research-toolkit --skill research-toolkit -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install rrrrrredy/research-toolkit research-toolkit --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "research-toolkit" agent skill from https://github.com/rrrrrredy/research-toolkit/tree/main into .agents/skills/research-toolkit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-toolkit", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rrrrrredy/research-toolkit --skill research-toolkit -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install rrrrrredy/research-toolkit research-toolkit --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "research-toolkit" agent skill from https://github.com/rrrrrredy/research-toolkit/tree/main into .cursor/skills/research-toolkit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-toolkit", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rrrrrredy/research-toolkit --skill research-toolkit -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install rrrrrredy/research-toolkit research-toolkit --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "research-toolkit" agent skill from https://github.com/rrrrrredy/research-toolkit/tree/main into .gemini/skills/research-toolkit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-toolkit", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install rrrrrredy/research-toolkit research-toolkitInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add rrrrrredy/research-toolkit --skill research-toolkit -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "research-toolkit" agent skill from https://github.com/rrrrrredy/research-toolkit/tree/main into .github/skills/research-toolkit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-toolkit", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rrrrrredy/research-toolkit --skill research-toolkit -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install rrrrrredy/research-toolkit research-toolkit --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "research-toolkit" agent skill from https://github.com/rrrrrredy/research-toolkit/tree/main into .opencode/skills/research-toolkit/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-toolkit", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
research-toolkitSource-backed longform research workflow for AI agents: scope, sources, claims, and independent review.
Research Toolkit is an agent skill from rrrrrredy/research-toolkit. Source-backed longform research workflow for AI agents: scope, sources, claims, and independent review.
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 621 other files, including scripts and reference files (for example `.agents/plugins/marketplace.json`, `.github/ISSUE_TEMPLATE/failure-case.yml` and `.github/workflows/framework-checks.yml`).
It sits in Research & Science. The repository describes itself as: Research methods, workflow tools, and independent review for AI agents, packaged as a Skill, plugin, and local MCP server. The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 10b51c2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
RESEARCH_TOOLKIT_REVIEW_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Research Toolkit loads about 3.5k tokens when it runs, and up to ~40k if it reads all its reference files. Until then it costs about 30 tokens; SKILL.md has 1,681 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from rrrrrredy/research-toolkit at commit 10b51c2, republished under its MIT licence (© rrrrrredy). 1,681 words, ~3,464 tokens.
.claude/skills/research-toolkit/SKILL.md (or your agent's skills folder). This skill also uses 616 other files; get the full folder from GitHub.Produce the research deliverable the user requested. Use this entry point throughout the task; load the relevant methods at the stages below. The research standard retains the complete rules, state contract, examples, and failure guidance. Paths in commands and code spans are relative to the installed toolkit or the research task, as indicated.
Use for substantial industry, market, company, product, technology, policy and ecosystem research reports. Not for quick facts or short summaries.
research_start(profile="full") is the default. Select profile="lite" (CLI: start --profile lite) for a bounded task; the saved profile applies to later actions and cannot be silently changed when resuming.
| Task size, use and evidence | Recommended profile |
|---|---|
| Up to about 2,000 words/Chinese characters; an internal briefing or exploratory answer; a few directly readable sources | lite |
| Longer or multi-section comparison; many or conflicting sources; a report intended for publication or decisions | full |
| Consequential decisions, explicitly required independent review, or frozen evaluation, regardless of length | full with explicitly configured independent review |
Lite retains brief clarification, source-backed claim registration, section-by-section drafting and the final checklist. It skips research_review and the full delivery hard gates with explicit SKIP logs. Call research_finish for the checklist, then submit every item with passed: true and a concrete evidence locator. The result is saved in state/final_checklist.json; it is not a full-gate receipt. Use qualified conclusions such as “the available sources suggest”; disclose that there was no independent review. These profile rules govern any full-only review/delivery requirements in the stage references.
Before collecting sources, read research workflow and check the existing conversation and materials for:
Use the context to propose concrete coverage, questions to answer, and what the user will receive. The agent translates the request into a research plan; the user need not define research terms or choose abstract scope or depth labels. For consequential choices that cannot be inferred, explain the alternatives and how they change the result, then ask one compact batch. When the context is sufficient, proceed without a confirmation round. Keep proposed defaults distinct from explicit user requirements; do not repeat answered questions.
Record the agreed brief and outline in state/task_spec.md. If an unanswered question would materially change the object, scope, evidence standard, or deliverable, keep dependent work pending and continue only unaffected work. Use and record reasonable defaults for non-critical details. If the user delegates a choice, record that choice and proceed.
The agent owns reading the applicable methods, maintaining records, executing reviews, recovering failures, and checking completion. Ask the user for research decisions or necessary access; do not make them supervise these execution duties.
state/requirements.jsonl; waivers and accepted unfinished obligations need the specific user decision.Read the relevant file or linked section before its stage. Keep already-read instructions in use; do not reread every reference on every turn.
| When | Read | Required result |
|---|---|---|
| Starting or resuming | Workflow and recovery; state transitions | Brief and current state are clear; resume unfinished work without restarting completed stages |
| Collecting and analyzing | Sources and claims | Required reading is tracked; claims, uncertainty, and counter-evidence support the actual questions |
| Selecting an analysis method | Optional lenses; horizontal/vertical analysis only if selected | A useful method for this question, without forcing a universal report structure |
| Drafting and editing | Writing style; operating loop | Bounded sections meet the agreed depth; reader editing follows stable evidence, coverage, and argument |
| Delegating or reviewing | Roles and review; review records | Bounded assignments, effective content reviews and evidenced handling of findings; audit methods when expressly required |
| Closing a stage or delivering | Quality gates; delivery verification | Current content and records meet the applicable completion conditions |
| Diagnosing repeated drift | Gotchas; research lessons | Correct the specific failure without expanding the task |
For additional rules, consult the matching section of the research standard. Keep records proportionate; use existing task files rather than adding locks, transactions, or extra control systems.
Keep state/findings.jsonl, state/directions_tried.json, state/iteration_log.jsonl and logs/work.jsonl only when they help the task. They are optional history, not delivery prerequisites. Uncertainty can stay with the claims; a separate data/uncertainty_registry.csv is optional.
| Reviewer tier | Setup cost | Review strength | Suitable use |
|---|---|---|---|
self | None; explicitly switch roles in the author context | Degraded; retains author blind spots, marked reviewer: self | Small reports and a zero-configuration start |
external | Configure one command or endpoint | An external model response; quality depends on the backend and evidence | Routine reports with another model's feedback |
independent | Configure fresh reviewer contexts and any task-required audits | Existing independent review and version-binding requirements | Consequential reports and frozen evaluations |
For self, research_review returns a built-in critical-review prompt, the frozen report/evidence and input_version. Perform that review in the current context, then call the tool again with the structured result and input_version in self_review. The tool never fabricates a review or an external execution. Every self-review record and delivery receipt retains review_strength: degraded; describe conclusions as provisional and disclose the missing independent opinion.
Cover requirements, evidence, adversarial reasoning, structure/depth, reader usefulness, process language and natural expression. Retain the original response, located coverage, findings and version bindings for every tier. A negative review is complete but unresolved necessary corrections block Full delivery. Preserve valid negative results; use evidenced no-change decisions for mistaken findings rather than rerunning for a better verdict.
Check configuration with research_check_reviewer or python scripts/research_workflow.py check-reviewer. An explicitly requested external/independent reviewer that cannot run stays incomplete: restore its configured access and resume the same assignment. Do not substitute self for a declared independent requirement. Extra audits and sampling remain required for frozen evaluations and explicitly audited plans.
Reviewer tier describes the review relationship; backend describes the command or endpoint that receives the report and evidence and returns a structured review. With no backend configured, a new Full report defaults to self; a configured backend defaults to independent. Explicit tiers and saved independent/audited plans never silently fall back after an error.
| Backend | Configuration cost | Review strength | Material destination and usage |
|---|---|---|---|
self | None | Degraded author-context review | Current author context; its existing model usage; no additional reviewer service |
codex — Codex CLI | Set RESEARCH_TOOLKIT_REVIEW_BACKEND=codex; install and sign in to the CLI; optional RESEARCH_TOOLKIT_REVIEW_MODEL | Fresh external context; supports independent review | Configured CLI model service and account quota; forward the intended CODEX_HOME when customized |
claude — Claude Code CLI | Set RESEARCH_TOOLKIT_REVIEW_BACKEND=claude; install and authenticate the CLI; optional model | Fresh external context, without tools; supports independent review | Service/account configured by the CLI; consumes that account's subscription or API usage |
generic — compatible HTTP endpoint | Set backend to generic, RESEARCH_TOOLKIT_REVIEW_BASE_URL, RESEARCH_TOOLKIT_REVIEW_MODEL; optional RESEARCH_TOOLKIT_REVIEW_API_KEY | Fresh request with the full assignment; supports independent review, not guaranteed statistical independence | Complete materials go to your chosen endpoint; its account is billed; local endpoints follow your own hosting policy |
| Trusted custom command | RESEARCH_TOOLKIT_REVIEW_CONFIG with an argv list and format: "json" | Depends on the adapter's actual context and evidence | Destination and usage are controlled by that command; document them before sending materials |
The configuration check is local and does not establish service availability or quota. Credentials belong in the runtime environment, never in reports or versioned configuration. The generic adapter uses /chat/completions; a base URL may include /v1 or the complete route. See the backend interface for requests, responses and recovery.
progress.json.stage uses brief, collect, analyze, draft, review, revise, and final.
progress.json.status accepts in_progress, paused, blocked, and complete.
Record the actual stage, open issues, and next action. On recovery, read the task specification, progress, requirement ledger if present, and any existing research notes useful for recovery. A checkpoint remains a checkpoint.
For Full reader-ready final delivery (Lite uses its checklist above):
state/final_delivery.json. The coherent terminal state is stage: final and status: complete.python scripts/check_delivery.py <task-directory>A missing required input, failed gate, stale review, or unavailable required check cannot become final completion. Resolve it or give an accurate checkpoint with the remaining work. Hashes and offline PASS results establish record consistency, not authentic model execution or sound research judgments. This Skill supplies instructions and checks; it does not automatically execute or enforce the workflow.
© rrrrrredy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 616 other files (scripts, references) in the repository root of rrrrrredy/research-toolkit.
Open the folder on GitHubat commit 10b51c2
Research Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Research Toolkit this skillrrrrrredy/research-toolkit | 182 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Hypothesis Generationspacering-net/codeg | 3.9k | 14 repos | ~3.6k | Automated safety check: Notes | MIT | |
| GitHub Deep Researchbytedance/deer-flow | 84k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Nature Paper CardYuan1z0825/nature-skills | 47k | 2 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Content Research Writerweapp-tailwindcss/weapp-tailwindcss | 1.9k | 25 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Last30daysmvanhorn/last30days-skill | 64k | — | ~7.9k | Automated safety check: Notes | MIT |
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
bytedance/deer-flow
Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.
Yuan1z0825/nature-skills
Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.
weapp-tailwindcss/weapp-tailwindcss
Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.
mvanhorn/last30days-skill
Research what people actually say about any topic in the last 30 days.
spacering-net/codeg
Structured manuscript/grant review with checklist-based evaluation.
Categories
Source-backed longform research workflow for AI agents: scope, sources, claims, and independent review. Research Toolkit is an agent skill from rrrrrredy/research-toolkit. Source-backed longform research workflow for AI agents: scope, sources, claims, and independent review.
Research Toolkit fits situations like: research & Science work in your project.
Run `npx skills add rrrrrredy/research-toolkit --skill research-toolkit -a claude-code`. Or copy the skill folder (the rrrrrredy/research-toolkit repository) into .claude/skills/research-toolkit in your project. Claude Code loads it when a task matches its description.
Run `npx skills add rrrrrredy/research-toolkit --skill research-toolkit -a codex`. Or copy the skill folder (the rrrrrredy/research-toolkit repository) into .agents/skills/research-toolkit in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rrrrrredy/research-toolkit --skill research-toolkit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/research-toolkit, .gemini/skills/research-toolkit, .github/skills/research-toolkit and .opencode/skills/research-toolkit in your project.
Going by SKILL.md and its folder, Research Toolkit needs the command-line tools its instructions call (python) and credentials named RESEARCH_TOOLKIT_REVIEW_API_KEY. Our summary lists: Python 3; A credential in RESEARCH_TOOLKIT_REVIEW_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Research Toolkit is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 36k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Research Toolkit: Hypothesis Generation (spacering-net/codeg, 3.9k stars), GitHub Deep Research (bytedance/deer-flow, 84k stars), Nature Paper Card (Yuan1z0825/nature-skills, 47k stars) and Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
rrrrrredy (a GitHub user) maintains it in rrrrrredy/research-toolkit, which has 182 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 10, 2026.
Source: rrrrrredy/research-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.