Recording
codewhale-hq/Codewhale
Capture screenshots on registered computers, record on macOS or HarmonyOS, and manage saved captures.
End-of-turn research process recorder with progressive crystallization.
$ npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-manager -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ARA-Labs/Agent-Native-Research-Artifact research-manager --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ARA-Labs/Agent-Native-Research-Artifact.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/research-manager .claude/skills/research-manager && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "research-manager" agent skill from https://github.com/ARA-Labs/Agent-Native-Research-Artifact/tree/main/skills/research-manager into .claude/skills/research-manager/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-manager", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ARA-Labs/Agent-Native-Research-Artifact/tree/main/skills/research-managerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-manager -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ARA-Labs/Agent-Native-Research-Artifact research-manager --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ARA-Labs/Agent-Native-Research-Artifact.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/research-manager .agents/skills/research-manager && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "research-manager" agent skill from https://github.com/ARA-Labs/Agent-Native-Research-Artifact/tree/main/skills/research-manager into .agents/skills/research-manager/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-manager", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-manager -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ARA-Labs/Agent-Native-Research-Artifact research-manager --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ARA-Labs/Agent-Native-Research-Artifact.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/research-manager .cursor/skills/research-manager && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "research-manager" agent skill from https://github.com/ARA-Labs/Agent-Native-Research-Artifact/tree/main/skills/research-manager into .cursor/skills/research-manager/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-manager", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ARA-Labs/Agent-Native-Research-Artifact.git --path skills/research-manager--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-manager -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ARA-Labs/Agent-Native-Research-Artifact research-manager --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ARA-Labs/Agent-Native-Research-Artifact.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/research-manager .gemini/skills/research-manager && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "research-manager" agent skill from https://github.com/ARA-Labs/Agent-Native-Research-Artifact/tree/main/skills/research-manager into .gemini/skills/research-manager/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-manager", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ARA-Labs/Agent-Native-Research-Artifact research-managerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-manager -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ARA-Labs/Agent-Native-Research-Artifact.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/research-manager .github/skills/research-manager && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "research-manager" agent skill from https://github.com/ARA-Labs/Agent-Native-Research-Artifact/tree/main/skills/research-manager into .github/skills/research-manager/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-manager", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-manager -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ARA-Labs/Agent-Native-Research-Artifact research-manager --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ARA-Labs/Agent-Native-Research-Artifact.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/research-manager .opencode/skills/research-manager && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "research-manager" agent skill from https://github.com/ARA-Labs/Agent-Native-Research-Artifact/tree/main/skills/research-manager into .opencode/skills/research-manager/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "research-manager", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
research-managerEnd-of-turn research process recorder with progressive crystallization.
Research Manager is an agent skill from ARA-Labs/Agent-Native-Research-Artifact. End-of-turn research process recorder with progressive crystallization. Invoked at the END of EVERY turn, after the user's current request has been fully addressed and before yielding control back to the user. Reviews what happened in the turn, extracts research-significant events, and writes them into the ara/ artifact through a three-stage pipeline: Context Harvester → Event Router → Maturity Tracker. Trace events (decisions, experiments, dead ends, pivots) are recorded immediately as journey facts. Knowledge…
Its SKILL.md is about 9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/event-taxonomy.md`, `references/taste-comments.md` and `templates/reader-report.md`).
The repository describes itself as: Research Artifact Protocol for Rigorous and Trustworthy AI Scientists. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e52a925. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditGlobGrepFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Research Manager loads about 9k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 257 tokens; SKILL.md has 3,391 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ARA-Labs/Agent-Native-Research-Artifact at commit e52a925, republished under its MIT licence (© ARA-Labs). 3,391 words, ~9,019 tokens.
.claude/skills/research-manager/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.You are the Live PM. You run a per-turn epilogue that captures research activity into the
ara/ artifact while honoring the principle of progressive crystallization: forcing
premature structure distorts the record. Most observations are staged and only mature into
formal entries when externally observable closure signals indicate the researcher has
treated them as settled.
The artifact has two mutability regimes. Honor them strictly.
ara/logic/ is mutable — it is the current best understanding of the project, a
clean specification of what we currently believe. Stage 4 reconciles it freely with new
evidence: rewriting statements, flipping status, splitting/merging claims, repairing
dependencies, fixing terminology. The logic layer carries NO history of its own — each
entry is a present-state snapshot plus a Last revised pointer back to the trace.ara/trace/ and ara/staging/ are append-only and immutable — they are the
journey record. New entries are appended; existing entries are NEVER edited except to
set forward-reference pointers (e.g. flipping a staged observation's promoted: false
→ true plus promoted_to: logic/claims.md:C07, or appending to a session record's
events for the current turn). Prior entries' content is never rewritten. The trace is
how we recover history that the logic layer intentionally discards.This split lets claims.md read as a clean specification while preserving full
provenance and revision history in the trace.
ara/ while still working on the user's request.┌──────────────────┐ ┌──────────────┐ ┌──────────────────┐ ┌──────────────────────┐
│Context Harvester │->│ Event Router │->│ Maturity Tracker │->│ Logic Layer │
│ (extract what │ │ (classify + │ │ (crystallize on │ │ Reconciliation │
│ happened) │ │ route) │ │ closure signal) │ │ (reconcile current │
│ │ │ │ │ │ │ state w/ this turn)│
└──────────────────┘ └──────────────┘ └──────────────────┘ └──────────────────────┘Scan THIS TURN only (the user's most recent message + your tool calls and results since the previous epilogue). Identify research-significant activity in two categories:
contradiction_reports produced against
this ARA by a reader engine (e.g. research-foresight PREDICT §7) and supplied as input this
turn — open reader-report issues on the ARA's repository, or report files handed to this run
(shape: templates/reader-report.md). Each is a candidate event, NEVER an edit to apply.Output a flat list of candidate events with raw context.
For each candidate, classify it, tag provenance, distill the payload, and route it. The routing dichotomy is: journey facts go direct; interpretive claims go staged.
→ Use references/event-taxonomy.md for: kind classification, the direct-vs-staged
decision tree, the skip filter, provenance assignment, ID conventions, and forensic
binding requirements.
Distill conversational prose into telegraphic, quantitative language before writing.
Walk staging/observations.yaml and decide which staged observations are mature. Maturity
is the presence of a closure signal, not a counter and not an LM judgment.
A staged observation crystallizes when at least one of these signals is present:
Topic abandonment — observation's topic has no events in the last k=5 turns AND
open_threads does not reference it. Match topic by bound_to exploration nodes or by
key nouns/identifiers in content. Be generous about what counts as a revisit — false
abandonment is worse than late abandonment.
Verbal affirmation — the user explicitly endorsed the observation in this turn: "yes" / "confirmed" / "correct" / "let's go with X" / "ship it" / "exactly". The adoption must be FIRST-PERSON. Silence is not affirmation. "Maybe" / "probably" is not affirmation.
Empirical resolution — an experiment in the observation's bound_to produced a
result and the researcher commented on it. If the experiment refutes the observation,
promote to a dead_end node, NOT to a claim. The observation is closed either way.
Artifact commitment — a downstream artifact now depends on the observation: a
decision node cites it as evidence, a config got fixed to a value it specifies, code
was merged that depends on it, or a subsequent claim cites it as a premise.
Default to non-promotion. If no signal is clearly present, leave it staged. Premature crystallization is the failure mode this design exists to prevent.
When a signal fires for O{XX}:
content, context, potential_type, provenance, bound_to.Statement/Rationale, ground it per "Number grounding" below — open the source, copy the
matched line verbatim into Sources, then write the number as a copy of that quote. Carry
forward provenance. Verbal-affirmation upgrades ai-suggested → user-revised (or user if
reproduced verbatim). The other three signals do not upgrade provenance.Crystallized via: <signal>, From staging: O{XX}.[pending] + TODO if a binding cannot be made now.promoted: true, promoted_to: <layer>:<id>, crystallized_via: <signal>.
Do not delete the observation — the trail from raw to typed is part of the record.Every load-bearing number in a Statement (or a heuristic's Rationale/Sensitivity/Bounds)
is grounded the way code is — transcribed from an open source, never written from memory:
Sources (<value> ← <source ref> «matched line» [input|result]).
The number you then write in the prose is a copy of the value inside that quote — not a value
recalled and back-cited. An entry with a bare path and no «quote» is invalid.[input] (a value you set — cite the source that defines it)
or [result] (a value the run produced — cite the log/output that reports it). Don't cite a
measured outcome to the config meant to produce it, or vice versa.[pending] beats a guess. Can't open or locate a source this turn? Write
<value> ← [pending: what's missing]. An unverified-but-plausible path is fabrication and is
worse than [pending].When a new event contradicts something already staged or crystallized:
<!-- CONFLICT: see {other-id} --> (or # CONFLICT: in YAML).unresolved decision node to the exploration tree referencing both, with
provenance reflecting who introduced the contradiction.Reader reports are adjudicated by this manager in the turn they arrive — every report leaves
the turn with a verdict. This is the one sanctioned exception to the contradiction trigger's defer
rule, scoped to reader reports only (the manager's own mid-research contradictions still defer as
above). A report targeting nothing in logic/ is simply staged as an ordinary observation
(provenance: ai-suggested, report ref recorded in context). For a report targeting a
crystallized entry:
basis refs and re-read the targeted entry. The report is
upheld only when evidence resolvable inside the ARA (trace nodes, evidence files, session
records) corroborates the observation and genuinely contradicts the cited clause. Reader-side
pointers the manager cannot resolve are recorded but do not count toward upholding.ai-suggested, report ref recorded), record full before/after under logic_revisions:, and
append a decision node (status: resolved) referencing both the entry and the report. Status
changes follow the ordinary transition rules — an upheld report counts as empirical resolution.decision
node (status: resolved) recording the verdict and its specific reason.<!-- CONFLICT: see reader-report <ref> -->) and append an unresolved decision node referencing both, carrying the report's
repro for a future run to execute. A possibly-true dispute stays visible on the entry rather
than dying in the session record.[PM] summary line (e.g.
reader-report on C02 upheld → Conditions revised; reader-report on C04 rejected: evidence does not contradict clause; reader-report on C07 unverifiable → CONFLICT flagged, repro preserved). The manager reaches a verdict every time.The single-writer rule is unchanged: readers never write the ARA — this manager is the only writer, and a report is INPUT to it, not an edit.
A staged observation that has neither been promoted nor referenced for 3+ session-days
gets stale: true. Stale observations are surfaced at the next briefing for the
researcher to triage — the manager does not auto-discard.
Reconcile logic/ (the current best understanding) with this turn's events so it stays
internally consistent and faithful to present evidence. Operates only on already-crystallized
entries — staged observations belong to Stage 3. (History lives in the trace; see Layer Mutability.)
Status field when evidence warrants.Statement, Rationale, or definition when new
evidence narrows scope, terminology changed, or wording no longer matches what's
actually supported. Keep Statement a generalized mechanism/relationship and sharpen
Conditions as the regime becomes clearer; new run numbers update Proof/evidence,
never the Statement. A rewrite re-grounds every number it now contains (Number grounding);
any changed value gets its own fresh Sources «quote», never a carried-over one.Dependencies are those narrower claims and whose
Proof spans their evidence — keep the narrower claims in place; the new claim sits
above them, not instead of them (only when a signal this turn makes the relationship
evident — never a routine sweep).concepts.md, dependency loops.hypothesis ──► testing ──► supported
│ │ ▲
│ └──► weakened┘
├────────────────► refuted (terminal, empirical)
├────────────────► withdrawn (terminal, non-empirical)
└─ any ─────────► revised (Statement rewritten; reset to testing/hypothesis)hypothesis: just crystallized; no evidence gathered yet (default for new claims)untested: deliberately deferred — work not started, not currently plannedtesting: an experiment that bears on the claim is in progresssupported: empirical evidence confirms the claimweakened: evidence is mixed, partial, or weaker than requiredrefuted: empirical evidence disproves — terminalwithdrawn: researcher dropped the claim for non-empirical reasons (pivot, scope cut) — terminalrevised: a transition marker, not a resting state — after recording the revision in
the trace, the claim's Status settles to testing if prior evidence still applies,
else hypothesisrefuted and withdrawn are terminal unless the user explicitly revives the claim (in
which case route through revised).
For each crystallized entry in logic/, check this turn for:
Proof refs or bound_to
nodes produced a result this turn AND the researcher commented on it.supported (or one step toward it)weakened, and consider rewriting the
Statement to match the actual scope supportedrefuted AND append a dead_end node referencing the claimhypothesis → testing (the commitment IS the test); does NOT reach
supported alone.concepts.md this turn refines or
renames a term the entry uses. Update the wording for consistency.unresolved decision node, defer.When a signal fires for entry E (claim, heuristic, or concept):
- **Last revised**: YYYY-MM-DD (turn-id) on the entry.- **Status**: to the new value.refuted, ensure a dead_end node exists in
exploration_tree.yaml referencing the entry (create one if not).withdrawn with
Merged into: C{XX}, redirect cross-references.Dependencies
to the narrower claims, and leave those claims in place (they remain its grounding).logic_revisions:
(see schema below). This is the ONLY place the prior wording is preserved — the
logic file does not keep it.pm_reasoning_log.yaml explaining which signal fired AND any
signal you considered but rejected (near-misses are the most useful continuity record).provenance: userprovenance: user-revisedprovenance: ai-suggested. The researcher can revert at any future turn by saying so.hypothesis → supported in a single
turn requires BOTH empirical resolution AND verbal affirmation in the same turn.refuted or withdrawn by
inference from silence or staleness.supported → weakened on a single new event — flag as
contradiction instead and let the researcher adjudicate.Statement must remain a
falsifiable assertion with intact Falsification criteria. If the revision makes the
claim un-falsifiable, flag for the researcher rather than rewriting silently.pm_reasoning_log.yaml.1. Read existing ara/ files (current state, next IDs).
2. Stage 1 — harvest this turn's candidate events.
3. Stage 2 — classify/route each (per event-taxonomy.md): journey facts direct to trace/; interpretive events staged to staging/observations.yaml.
4. Stage 3 — crystallize staged observations whose closure signal fired; flag contradictions; mark 3+-day-idle observations stale.
5. Stage 4 — for each crystallized logic/ entry, apply status/content/structural edits when a signal fires; run the cross-ref consistency pass; record before/after in the session record; log near-misses.
6. Append turn events to today's session record; update session_index.yaml; append a line to pm_reasoning_log.yaml.
7. Print one-line summary, e.g.:
[PM] Turn captured: 1 decision (direct), 2 observations staged, 1 claim crystallized via affirmation, C03 testing→supported, C07 revised (scope narrowed).
Or, for empty turns:
[PM] Turn skipped: no research events.ara/
PAPER.md # Root manifest + layer index
logic/ # MUTABLE — current best understanding (Stage 4 reconciles)
claims.md problem.md concepts.md experiments.md related_work.md
solution/ # constraints.md + method files per the compiler's domain profile
src/ # How (artifacts) — configs/code/data per domain profile; always environment.md
trace/ # APPEND-ONLY — the journey, never rewritten
exploration_tree.yaml # Research DAG: decisions, experiments, dead_ends, pivots, questions
pm_reasoning_log.yaml # Manager's own organizational decisions per turn
taste_log.yaml # OPTIONAL — researcher's taste comments on trace nodes (pointer-only, never edits the node)
sessions/
session_index.yaml # Master session index (one entry per calendar day)
YYYY-MM-DD_NNN.yaml # Per-day session record, incl. logic_revisions
evidence/ # APPEND-ONLY — raw proof
README.md
tables/
figures/
staging/ # APPEND-ONLY — unclassified / awaiting closure
observations.yaml # The crystallization buffertrace/exploration_tree.yaml)Nested DAG. Each node may have children:. Use also_depends_on: [N{XX}] for cross-edges.
The tree's shape stays recoverable from a flat append log through two fields you already write: mark
each level/phase boundary as a pivot (or question) node (it opens a new branch), and list what
a node builds on in also_depends_on. Only when a node resumes an earlier branch — rather than
continuing the step right before it — add an explicit parent: N{XX} to point back; in the common
case its place is already implied and no extra field is needed.
tree:
- id: N01
type: question | decision | experiment | dead_end | pivot
title: "{short title}"
provenance: user | ai-suggested | ai-executed | user-revised
timestamp: "YYYY-MM-DDTHH:MM"
# type-specific fields:
description: > # question
choice: > # decision
alternatives: [] # decision
evidence: [] # decision, experiment
result: > # experiment
hypothesis: > # dead_end
failure_mode: > # dead_end
lesson: > # dead_end
from: "" # pivot
to: "" # pivot
trigger: "" # pivot
status: open | resolved | unresolved # unresolved used for contradiction-decision nodes
also_depends_on: [] # cross-edges (ids) — what this node builds on
parent: N{XX} # OPTIONAL — only to point back to an earlier branch; omit when implied
children:
- { ... }logic/claims.md) — crystallized only## C{XX}: {generalized title — the takeaway, not a recipe name}
- **Statement**: {the generalized, mechanistic conclusion; subject = a mechanism/relationship, never a named recipe; carries NO run numbers}
- **Conditions**: {under what conditions it holds; the regime; the known untested boundary}
- **Sources**: [{one entry per load-bearing number in the claim (now in `Conditions`/`Proof`): `<value> ← <file:line | trace-node:field> «verbatim line copied from source» [input|result]`, or `<value> ← [pending: reason]`}] # see "Number grounding"; a bare path with no «quote» is invalid
- **Status**: hypothesis | untested | testing | supported | weakened | refuted | withdrawn
- **Provenance**: user | ai-suggested | user-revised
- **Falsification**: {a concrete observation that would disprove it — for a mechanism claim, about the system/world; for a methodological/regime claim, about the benchmark's behavior. NOT a tautology or a re-run of the same gate ("if the recipe fails the gate")}
- **Proof**: [{evidence refs (→ evidence/) or "pending"; run numbers/IDs/scores live HERE, not in Statement}]
- **Dependencies**: [C{YY}, ...]
- **Tags**: {comma-separated}
- **Last revised**: YYYY-MM-DD (turn-id) # pointer back to the trace; absent until first revision
- **Taste** (optional): # researcher's own reactions; see references/taste-comments.md — absent until the first one
- [YYYY-MM-DD] `endorse | uncertain | reject` on `claim | evidence | framing | priority` — {free-text comment}The Statement is the generalized conclusion the evidence supports — a mechanism or relationship,
not a restatement of run numbers. What keeps it falsifiable and honest is Conditions (the regime
it holds in + the untested boundary) plus a Falsification, not a narrowed sentence. Numbers (run
IDs, n, scores, step counts) belong in Proof → evidence/ (grounded per Number grounding), never
in Statement. Conditions is mandatory: a generalized Statement with no Conditions is an unbounded
slogan.
Calibrate the Statement to what the evidence actually separates. Do not assert a distinction the
design cannot disentangle (confounded factors — e.g. matrix "shape" vs "role" when they co-vary), or
a law from a single instance. When that's the case, hedge in the Statement itself — name the
unseparated factors together, or say "shown once here" — rather than only burying it in Conditions.
Conditions bounds where the claim applies; it is not a license for the Statement's verb to
over-reach. The Statement/Conditions may be sharpened on a later turn (Stage 4 content revision) as
the mechanism becomes clearer — no new closure signal is needed.
Current-state snapshot only — no prior statements, no From staging/Crystallized via
notes. Crystallization and every edit are recorded in the trace (trace/sessions/… under
logic_revisions: with before/after; source observation stays in staging/; reasoning in
pm_reasoning_log.yaml). refuted/withdrawn are terminal and revised is a transition
marker, not a resting state — see Stage 4.
logic/solution/heuristics.md) — crystallized only## H{XX}: {title}
- **Rationale**: {current best explanation of why this works}
- **Sources**: [{one entry per load-bearing number in `Rationale`/`Sensitivity`/`Bounds`, same format as claims — see "Number grounding"}]
- **Status**: active | weakened | retired
- **Provenance**: user | ai-suggested | user-revised
- **Sensitivity**: low | medium | high | unknown # "unknown" until the turn establishes it — never guess
- **Code ref**: [{file paths, or "pending"}]
- **Last revised**: YYYY-MM-DD (turn-id) # absent until first revision
- **Taste** (optional): # researcher's own reactions; see references/taste-comments.md — absent until the first one
- [YYYY-MM-DD] `endorse | uncertain | reject` on `claim | evidence | framing | priority` — {free-text comment}Current-state snapshot only (same as claims); history lives in the trace.
staging/observations.yaml) — stagedobservations:
- id: O{XX}
timestamp: "YYYY-MM-DDTHH:MM"
provenance: user | ai-suggested | ai-executed | user-revised
content: "{raw observation, factually distilled}"
context: "{what was happening this turn}"
potential_type: claim | heuristic | concept | constraint | architecture | unknown
bound_to: [N{XX}, ...] # exploration nodes this depends on
promoted: false
promoted_to: null # e.g., "logic/claims.md:C07" once crystallized
crystallized_via: null # which closure signal fired
stale: falsetrace/sessions/YYYY-MM-DD_NNN.yaml) — turns append within the daysession:
id: "YYYY-MM-DD_NNN"
date: "YYYY-MM-DD"
started: "YYYY-MM-DDTHH:MM"
last_turn: "YYYY-MM-DDTHH:MM"
turn_count: 0
summary: "{rolling one-line summary}"
events_logged:
- turn: 1
type: decision | experiment | dead_end | pivot | observation | ...
id: "{N/O}{XX}"
routing: direct | staged | crystallized
provenance: user | ai-suggested | ai-executed | user-revised
summary: "{telegraphic what}"
ai_actions:
- turn: 1
action: "{what AI did}"
provenance: ai-executed
files_changed: ["{paths}"]
claims_touched:
- id: C{XX}
action: created | crystallized | advanced | weakened | confirmed | refuted | withdrawn | revised | split | merged
turn: 1
logic_revisions: # full before/after for every edit Stage 4 makes
- turn: 1
entry: C{XX} # or H{XX}, concept id, etc.
field: Statement | Status | Rationale | Dependencies | id | ...
before: "{prior value, verbatim}"
after: "{new value, verbatim}"
signal: empirical-resolution | verbal-declaration | dependency-change | artifact-commitment | terminology-drift | user-directive
provenance: user | ai-suggested | user-revised
note: "{one-line why, optional}"
# structural changes record both endpoints, e.g. for a split:
- turn: 1
entry: C07
field: split
before: "C07 covered both training and inference"
after: "C07 = training-time claim; C12 = inference-time claim"
signal: verbal-declaration
provenance: user-revised
key_context:
- turn: 1
excerpt: "{quote or paraphrase capturing decisive exchange}"
open_threads:
- "{what needs follow-up}"
ai_suggestions_pending:
- "{unconfirmed AI suggestions still awaiting closure}"trace/sessions/session_index.yaml)sessions:
- id: "YYYY-MM-DD_NNN"
date: "YYYY-MM-DD"
summary: "{main outcome}"
turn_count: {N}
events_count: {N}
claims_touched: [C{XX}, ...]
open_threads: {N}trace/pm_reasoning_log.yaml) — self-continuityA few lines per turn explaining the manager's own organizational decisions. Cheap on tokens, prevents organizational drift.
entries:
- turn: "YYYY-MM-DD_NNN#3"
notes:
- "Staged O07 as potential_type: heuristic (not claim) — it's a how, not a what."
- "Did NOT crystallize O05 despite affirmation-like language: user said 'maybe' not 'yes'."
- "Routed N12 as dead_end rather than experiment — code was abandoned mid-run."trace/taste_log.yaml) — optional, append-onlyResearcher's taste comments on trace nodes. Never edits exploration_tree.yaml — points at
it instead, the same way a promoted observation points at its logic-layer destination
without rewriting itself. See references/taste-comments.md for trigger detection, target
resolution, and the confirm-before-write procedure. File does not exist until the first entry.
entries:
- id: T{XX}
timestamp: "YYYY-MM-DDTHH:MM"
target: N{XX} # trace node this comments on; never edited
tag: endorse | uncertain | reject
object: claim | evidence | framing | priority
comment: "{free-text comment}"ara/ does not exist)Create the structure on the first turn that contains research-significant activity. Do not ask unprompted on a purely conversational opener.
mkdir -p ara/{logic/solution,src,trace/sessions,evidence/{tables,figures},staging}Seed:
ara/PAPER.md — root manifest (infer title, authors, venue from project context)ara/trace/sessions/session_index.yaml — sessions: []ara/trace/exploration_tree.yaml — tree: []ara/trace/pm_reasoning_log.yaml — entries: []ara/staging/observations.yaml — observations: []ara/logic/claims.md — # Claimsara/logic/problem.md — # Problemara/logic/solution/heuristics.md — # Heuristicsara/evidence/README.md — # Evidence IndexThen run the per-turn procedure normally.
On the first turn of a new conversation (not every turn), silently read:
summary, open_threads, ai_suggestions_pending, key_contextclaims.md status countsstaging/observations.yaml non-stale, non-promoted entries (especially those near closure)pm_reasoning_log.yaml last few entries (organizational continuity)Surface relevant pieces only when they bear on the user's first task — never lead with a formal briefing the researcher did not ask for. If the user asks "where did we leave off", deliver the full briefing.
Separate from the four-stage pipeline above. When the user reacts evaluatively to a specific
claim, heuristic, or trace node this turn, record it as a taste comment: always
provenance: user, never staged, never affects Status or crystallization. Every taste
comment carries both an attitude (endorse | uncertain | reject) and an object of judgment
(claim | evidence | framing | priority) — the two are independent axes, not one label. →
Use references/taste-comments.md for trigger detection, target resolution, the
confirm-before-write procedure, and the tag rules; schemas above.
Taste is additive, never a substitute for the normal pipeline: if the same utterance also
introduces new research content, that content is routed through Stage 1–4 as its own event
regardless of the taste comment (see references/taste-comments.md).
Runs inline within the normal epilogue when triggered — not a separate interactive prompt, and not asked about on turns where it doesn't come up.
ai-suggested holds until explicit user affirmation.refuted/withdrawn) need explicit triggers, never silence/staleness. Log near-misses.logic/ overwrites in place; trace/ and staging/ are append-only except forward-reference pointers. Every logic edit gets a logic_revisions: before/after in the session record — the only place pre-edit content is kept.unresolved decision node, defer.[pending]+TODO if not yet bindable. Keep YAML valid; summary line terse.taste_log.yaml and never edits the node.© ARA-Labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/research-manager of ARA-Labs/Agent-Native-Research-Artifact.
Open the folder on GitHubat commit e52a925
Research Manager next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Research Manager this skillARA-Labs/Agent-Native-Research-Artifact | 692 | — | ~9k | Automated safety check: Pass | MIT | |
| Recordingcodewhale-hq/Codewhale | 41k | — | ~540 | Automated safety check: Pass | MIT | |
| Architecture Decision Recordsaffaan-m/ECC | 276k | 4 repos | ~1.8k | Automated safety check: Pass | MIT | |
| Architecture Decision Recordsaffaan-m/ECC | 276k | 1 repos | ~863 | Automated safety check: Pass | MIT | |
| Architecture Decision Recordsaffaan-m/ECC | 276k | — | ~1.1k | Automated safety check: Pass | MIT | |
| Browser Recordruvnet/ruflo | 74k | — | ~735 | Automated safety check: Notes | MIT |
codewhale-hq/Codewhale
Capture screenshots on registered computers, record on macOS or HarmonyOS, and manage saved captures.
affaan-m/ECC
Capture architectural decisions as numbered ADR markdown files in docs/adr/ with context, alternatives considered, consequences, and an index README.
affaan-m/ECC
在Claude Code会话期间,将做出的架构决策捕获为结构化的架构决策记录(ADR)。自动检测决策时刻,记录上下文、考虑的替代方案和理由。维护一个ADR日志,以便未来的开发人员理解代码库为何以当前方式构建。
affaan-m/ECC
コーディングセッション中にアーキテクチャ決定を構造化ADRとして記録し、自動的に決定の瞬間を検出し、コンテキスト、検討された代替案、根拠を記録します。今後の開発者がコードベースの形成理由を理解するためのADRログを維持します。
ruvnet/ruflo
Open a named, traced browser session into an RVF cognitive container with a ruvector trajectory recording every action
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing image assets, markup, and CDN or build transforms related to Use progressive JPEG encoding.
ARA-Labs/Agent-Native-Research-Artifact
Treat an open-ended investigation the way a fuzzer treats a program.
ARA-Labs/Agent-Native-Research-Artifact
ARA Submitter. An agent skill from ARA-Labs/Agent-Native-Research-Artifact.
ARA-Labs/Agent-Native-Research-Artifact
ARA Seal Level 2: Semantic Epistemic Review. An agent skill from ARA-Labs/Agent-Native-Research-Artifact.
ARA-Labs/Agent-Native-Research-Artifact
Universal ARA Compiler. An agent skill from ARA-Labs/Agent-Native-Research-Artifact.
ARA-Labs/Agent-Native-Research-Artifact
ARA World Model — read-only reasoning engine over ONE Agent-Native Research Artifact (ARA), run LOCALLY with the coding agent itself as the LLM (no SDK, no API key).
End-of-turn research process recorder with progressive crystallization. Research Manager is an agent skill from ARA-Labs/Agent-Native-Research-Artifact. End-of-turn research process recorder with progressive crystallization.
Run `npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-manager -a claude-code`. Or copy the skill folder (skills/research-manager in ARA-Labs/Agent-Native-Research-Artifact) into .claude/skills/research-manager in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-manager -a codex`. Or copy the skill folder (skills/research-manager in ARA-Labs/Agent-Native-Research-Artifact) into .agents/skills/research-manager in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ARA-Labs/Agent-Native-Research-Artifact --skill research-manager -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/research-manager, .gemini/skills/research-manager, .github/skills/research-manager and .opencode/skills/research-manager in your project.
SKILL.md names no scripts, command-line tools or credentials: Research Manager is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Write, Edit, Glob, Grep.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Research Manager is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 9k tokens (SKILL.md is roughly 36k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Research Manager: Recording (codewhale-hq/Codewhale, 41k stars), Architecture Decision Records (affaan-m/ECC, 276k stars), Architecture Decision Records (affaan-m/ECC, 276k stars) and Architecture Decision Records (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ARA-Labs (a GitHub organization) maintains it in ARA-Labs/Agent-Native-Research-Artifact, which has 692 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 7, 2026.
Source: ARA-Labs/Agent-Native-Research-Artifact on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.