Scholar Open
joshzyj/open-scholar-skill
Implement open science practices for a social science study.
Turns papers, repositories, logs or notes into an Agent-Native Research Artifact with claims, concepts, configs, an exploration graph and grounded evidence.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ara-compiler -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ara-compiler --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/22-agent-native-research-artifact/compiler .claude/skills/ara-compiler && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ara-compiler" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/22-agent-native-research-artifact/compiler into .claude/skills/ara-compiler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ara-compiler", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/22-agent-native-research-artifact/compilerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ara-compiler -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ara-compiler --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/22-agent-native-research-artifact/compiler .agents/skills/ara-compiler && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ara-compiler" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/22-agent-native-research-artifact/compiler into .agents/skills/ara-compiler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ara-compiler", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ara-compiler -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ara-compiler --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/22-agent-native-research-artifact/compiler .cursor/skills/ara-compiler && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ara-compiler" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/22-agent-native-research-artifact/compiler into .cursor/skills/ara-compiler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ara-compiler", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Orchestra-Research/AI-Research-SKILLs.git --path 22-agent-native-research-artifact/compiler--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ara-compiler -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ara-compiler --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/22-agent-native-research-artifact/compiler .gemini/skills/ara-compiler && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ara-compiler" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/22-agent-native-research-artifact/compiler into .gemini/skills/ara-compiler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ara-compiler", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ara-compilerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ara-compiler -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .github/skills && cp -r skills-src/22-agent-native-research-artifact/compiler .github/skills/ara-compiler && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ara-compiler" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/22-agent-native-research-artifact/compiler into .github/skills/ara-compiler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ara-compiler", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill ara-compiler -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Orchestra-Research/AI-Research-SKILLs ara-compiler --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/22-agent-native-research-artifact/compiler .opencode/skills/ara-compiler && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ara-compiler" agent skill from https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/22-agent-native-research-artifact/compiler into .opencode/skills/ara-compiler/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ara-compiler", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ara-compilerTurns papers, repositories, logs or notes into an Agent-Native Research Artifact with claims, concepts, configs, an exploration graph and grounded evidence.
The compiler takes almost anything that holds research knowledge: PDF papers and arXiv links, GitHub repositories, code and notebooks, experiment logs and configs, raw notes and chat threads, or just a conversation with you. It identifies what it was given, cross-references every source, asks you when the input is ambiguous, and marks fields it cannot fill as not available from the provided input rather than guessing.
The output is a validated ARA with a cognitive layer of claims, concepts and heuristics, a physical layer of configs and code stubs, an exploration graph and grounded evidence. `--output` sets the target directory, which defaults to `./ara-output/`, and `--rubric` maps coverage to a PaperBench rubric. Reference files supply the ARA schema, an exploration-tree specification and a validation checklist.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
ARA Research Compiler loads about 3.7k tokens when it runs, and up to ~9.3k if it reads all its reference files. Until then it costs about 128 tokens; SKILL.md has 1,773 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 1,773 words, ~3,702 tokens.
.claude/skills/ara-compiler/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.You are the ARA Universal Compiler. Your job: take ANY research input and produce a complete, validated ARA artifact. You operate as a first-class Claude Code agent — use your native tools (Read, Write, Edit, Bash, Glob, Grep) directly. No API wrapper needed.
The compiler is open-ended. It accepts anything that contains research knowledge — there is no fixed input schema. Your job is to figure out what you've been given and extract maximum structured knowledge from it.
Possible inputs include (but are NOT limited to):
.py, .ipynb, .rs, .cpp, etc.)When arguments are provided ($ARGUMENTS), interpret them flexibly:
--output <dir> → where to write the ARA (default: ./ara-output/)--rubric <path> → PaperBench rubric for coverage mappingAdapt to whatever you receive:
1. READ all inputs
2. REASON through the 4-stage epistemic protocol (see below)
3. GENERATE all ARA files using Write tool
4. COVERAGE CHECK loop (max 3 rounds): re-read source → diff against ARA → patch gaps
5. VALIDATE by running Seal Level 1
6. FIX any failures, re-validate
7. REPORT summary to userRead ALL provided inputs thoroughly before generating anything. For PDFs, read every page, including appendices — appendices often carry reproduction-critical content and should be treated with the same priority as main-text pages.
For repos, prioritize: README → core algorithm files → configs → environment files.
Before writing any files, reason through these 4 stages. Think carefully about each stage.
Stage 1 — Semantic Deconstruction Strip narrative framing. Extract the raw knowledge atoms:
Before moving on, perform an evidence capture pass:
Table 2, Figure 4, etc.)Table N unless it reproduces the original source table faithfullyStage 2 — Cognitive Mapping
Map extracted atoms to /logic/:
Appendix content (worked examples, prompt templates, enumerated taxonomies, annotation schemas, extended analyses, prescriptive content) should be routed into the ARA layers where it fits best, preserving the granularity the source uses. Never silently drop an appendix section.
When writing claims:
Statement at the strongest level directly supported by the cited evidenceEvidence basisInterpretationrelated_work.md should reflect the paper's full citation footprint, not only the
closest predecessors. Works with a specific technical delta get full RW blocks; remaining
citations from the paper's References list should still be captured (more briefly) so the
intellectual neighborhood is preserved.
Stage 3 — Physical Stubbing
Generate /src/:
rubric/requirements.md mapping every leaf nodeStage 4 — Exploration Graph Extraction
Reconstruct the research DAG for /trace/exploration_tree.yaml:
also_depends_on for DAG convergence pointsexplicit from source material or inferred from reconstructionWrite ALL mandatory files. See references/ara-schema.md for the complete directory structure and field-level requirements for every file.
Mandatory files (all must exist and be non-trivial):
PAPER.md — YAML frontmatter (title, authors, year, venue, doi, ara_version, domain, keywords, claims_summary, abstract) + Layer Indexlogic/problem.md — Observations (O1, O2...), Gaps (G1, G2...), Key Insight, Assumptionslogic/claims.md — Claims (C01, C02...) each with Statement, Status, Falsification criteria, Proof, Evidence basis, Interpretation, Dependencies, Tagslogic/concepts.md — ≥5 concepts each with Notation, Definition, Boundary conditions, Related conceptslogic/experiments.md — ≥3 experiments (E01, E02...) each with Verifies, Setup, Procedure, Metrics, Expected outcome (directional only!), Baselines, Dependencieslogic/solution/architecture.md — Component graph with inputs/outputslogic/solution/algorithm.md — Math formulation + pseudocode + complexitylogic/solution/constraints.md — Boundary conditions and limitationslogic/solution/heuristics.md — Heuristics (H01, H02...) each with Rationale, Sensitivity, Bounds, Code ref, Sourcelogic/related_work.md — Related work (RW01, RW02...) each with DOI, Type, Delta, Claims affectedsrc/configs/training.md — Hyperparameters with Value, Rationale, Search range, Sensitivity, Sourcesrc/configs/model.md — Model/architecture configssrc/execution/{module}.py — ≥1 code stub with typed signaturessrc/environment.md — Python version, framework, hardware, dependencies, seedstrace/exploration_tree.yaml — Research DAG (≥8 nodes, nested YAML)evidence/README.md — Index table mapping every evidence file to claimsevidence/tables/*.md — ALL result tables (exact cell values, never rounded)evidence/figures/*.md — ALL quantitative figures (extracted data points)Evidence-generation rules:
table3_...) must match that source object's caption and contentsderived_, subset_, or equivalent, and the file must state what it was derived fromBefore running Seal validation, verify that the ARA faithfully covers the source material. Repeat up to 3 rounds; stop early if a round produces no patches.
Each round: re-read the source, identify anything not yet captured or only shallowly captured in the ARA, patch those gaps, then note how many fixes were made. If zero, exit early. Pay particular attention to appendix content and to citations from the paper's References list, which are easy to miss on the first pass.
The coverage loop does not replace validation — it ensures the ARA is semantically complete before structural checks run.
Run ARA Seal Level 1 validation. Perform these checks:
logic/, logic/solution/, src/, src/configs/, trace/, evidence/support_level: explicit|inferredsupport_level: explicit include source referencesFor each validation failure:
Typically converges in 2-3 rounds.
Print a summary:
experiments.md contains only directional/relative expected outcomes. Exact numbers go in evidence/Table N or Figure N unless it faithfully reproduces the original source objectFor detailed schema specifications, load these on demand:
© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in 22-agent-native-research-artifact/compiler of Orchestra-Research/AI-Research-SKILLs.
Open the folder on GitHubat commit 773a529
ARA Research Compiler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| ARA Research Compiler this skillOrchestra-Research/AI-Research-SKILLs | 13k | — | ~3.7k | Automated safety check: Pass | MIT | |
| Scholar Openjoshzyj/open-scholar-skill | 168 | — | ~14k | Automated safety check: Pass | Custom licence | |
| Read GitHubAgentTeam-TaichuAI/ScienceClaw | 671 | 2 repos | ~638 | Automated safety check: Pass | None | |
| Light Research OrchestratorLight0305/Light-skills | 640 | — | ~3.8k | Automated safety check: Pass | MIT | |
| News to Research Idea BriefingOpenLAIR/dr-claw | 1.2k | — | ~1.3k | Automated safety check: Notes | Custom licence | |
| Paper Reproductionjerry609/PaperBot | 108 | — | ~1k | Automated safety check: Pass | MIT |
joshzyj/open-scholar-skill
Implement open science practices for a social science study.
AgentTeam-TaichuAI/ScienceClaw
Read and search GitHub repository documentation via gitmcp.io MCP service.
Light0305/Light-skills
Coordinates and recovers multi-stage Light research projects from a single passport file, with checkpoints, stale-work tracking and rerouting only when you approve.
OpenLAIR/dr-claw
Clusters the latest news-feed results by topic and writes a briefing of research idea seeds with citations, plus a structured seeds file, without crawling new sources.
jerry609/PaperBot
This skill should be used when the user asks to "reproduce a paper", "implement paper code", "paper2code", "replicate research results", "run experiment from paper", "implement the algorithm from…
mizchi/skills
Evaluate the reproducibility of technical articles. An agent skill from mizchi/skills.
Orchestra-Research/AI-Research-SKILLs
Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Orchestra-Research/AI-Research-SKILLs
Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.
Works with
Categories
Turns papers, repositories, logs or notes into an Agent-Native Research Artifact with claims, concepts, configs, an exploration graph and grounded evidence. The compiler takes almost anything that holds research knowledge: PDF papers and arXiv links, GitHub repositories, code and notebooks, experiment logs and configs, raw notes and chat threads, or just a conversation with you. It identifies what it was given, cross-references every source, asks you when the input is ambiguous, and marks fields it cannot fill as not available from the provided input rather than guessing.
ARA Research Compiler fits situations like: converting a paper PDF into a structured, machine-readable knowledge package; building an ARA from a codebase and its experiment logs; capturing the decisions and dead ends from research notes in a falsifiable form.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill ara-compiler -a claude-code`. Or copy the skill folder (22-agent-native-research-artifact/compiler in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/ara-compiler in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill ara-compiler -a codex`. Or copy the skill folder (22-agent-native-research-artifact/compiler in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/ara-compiler in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill ara-compiler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ara-compiler, .gemini/skills/ara-compiler, .github/skills/ara-compiler and .opencode/skills/ara-compiler in your project.
SKILL.md names no scripts, command-line tools or credentials: ARA Research Compiler is instructions for the agent only. Our summary lists: Read and write access to the input files and an output directory; Network access if you want URLs fetched or repositories cloned.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
ARA Research Compiler is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with ARA Research Compiler: Scholar Open (joshzyj/open-scholar-skill, 168 stars), Read GitHub (AgentTeam-TaichuAI/ScienceClaw, 671 stars), Light Research Orchestrator (Light0305/Light-skills, 640 stars) and News to Research Idea Briefing (OpenLAIR/dr-claw, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,374 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.
Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.