Agent skill

Deli Autoresearch

by rongxinzy in rongxinzy/RongxinAI

A protocol framework for long-horizon autonomous research tasks.

AGPL-3.0Auto-check passedResearch & Science

Install Deli Autoresearch

skills CLI
$ npx skills add rongxinzy/RongxinAI --skill deli-autoresearch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rongxinzy/RongxinAI deli-autoresearch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rongxinzy/RongxinAI.git skills-src && mkdir -p .claude/skills && cp -r skills-src/SKILLs/deli-autoresearch .claude/skills/deli-autoresearch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deli-autoresearch
GitHub stars
154
Token cost
~3.2k tokens
SKILL.md length
1,378 words
Files
3
Skills in repo
94
Repo updated
First seen
Licence
AGPL-3.0

At a glance

A protocol framework for long-horizon autonomous research tasks.

  • Works in 11 steps: Runtime Mapping (this environment) → Motivation → Behavioral Constraints → …
  • The user asks for academic research
  • SKILL.md covers 0. Runtime Mapping (this…, Controlled Shortcut Completion, 1. Motivation and 2. Behavioral Constraints, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Deli Autoresearch is an agent skill from rongxinzy/RongxinAI. A protocol framework for long-horizon autonomous research tasks. Targets three empirically-observed failure modes — cognitive loops, stalling, runtime fragility — by prescribing state management, stall detection, and watchdog mechanisms. Use when the user asks for academic research, literature surveys, paper writing, or any unattended multi-day research task. Triggers: academic research, 学术研究, literature review, 文献综述, paper writing, 论文写作, ICLR survey, autonomous research.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `zhiyuan/metadata.yaml`).

It sits in Research & Science, covering Autonomous loops, Scientific writing and State management. The repository describes itself as: An all-in-one local AI Agent workspace with a fully self-developed stack. The licence is AGPL-3.0.

When your agent uses it

  • The user asks for academic research
  • Literature surveys
  • Any unattended multi-day research task

Example prompts

  • “/deli-autoresearch”

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. Runtime Mapping (this environment)
  2. Motivation
  3. Behavioral Constraints
  4. Architecture
  5. State Files
  6. Usage
  7. Stall Detection & Pivoting
  8. Heartbeat Watchdog
  9. Subagent Scheduling Patterns
  10. Engineering Constraints
  11. Validation & Limits

What it can do on your machine

Read from SKILL.md and the folder at commit 9c64865. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deli Autoresearch loads about 3.2k tokens when it runs. Until then it costs about 124 tokens; SKILL.md has 1,378 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rongxinzy/RongxinAI at commit 9c64865, republished under its AGPL-3.0 licence (© rongxinzy). 1,378 words, ~3,174 tokens.

Download SKILL.mdSave it as .claude/skills/deli-autoresearch/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
deli-autoresearch
description
A protocol framework for long-horizon autonomous research tasks. Targets three empirically-observed failure modes — cognitive loops, stalling, runtime fragility — by prescribing state management, stall detection, and watchdog mechanisms. Use when the user asks for academic research, literature surveys, paper writing, or any unattended multi-day research task. Triggers: academic research, 学术研究, literature review, 文献综述, paper writing, 论文写作, ICLR survey, autonomous research.
metadata.version
1.1
metadata.category
research
<!-- zhiyuan_AutoResearch: protocol framework for long-horizon autonomous tasks. Ships no executable code; prescribes conventions for state persistence, stall detection, and layered guardians. -->

zhiyuan_AutoResearch

This skill is a protocol framework for long-horizon autonomous tasks (days to weeks). It ships no executable code; instead it prescribes a set of battle-tested conventions: how state is persisted, how stalls are detected, how guardians are layered, and what constraints bind agent behavior. Implementation details are left to the adopter's environment.

0. Runtime Mapping (this environment)

The protocol's abstract mechanisms map onto concrete tools here:

Protocol mechanismConcrete tool
Orchestrator loop (/loop)agent_loop tool — start a goal-mode loop, one iteration per research cycle
Work agent (Agent tool)subagent tool — single / parallel / chain delegation to researcher-scout-planner-reviewer profiles
Fresh session per iterationeach subagent call is an isolated session with no shared conversation memory
Durable heartbeat watchdogSQLite-backed scheduled tasks; in Work sessions, agent_loop carries the iteration cadence instead
Nudge / direction injectionthe parent session steers between loop iterations by editing state files

State files and logs below are tool-agnostic: they are what makes every iteration resumable.

Controlled Shortcut Completion

When this skill is selected through the Academic Research sidebar shortcut, the runtime applies the Deep Research evidence gate in addition to this protocol. Before requesting completion, save the final cited research report as a .md deliverable and a readable .md, .txt, or .json validation report in the selected workspace. Record both with workflow_state; plans, delegates, and source URLs are necessary evidence but never substitute for the final report.

1. Motivation

Long-running code agents exhibit three recurring failure modes:

  1. Cognitive loop — successive iterations try similar directions with diminishing returns, unable to escape a local optimum on their own.
  2. Stalling — the agent finishes a chunk of work, outputs a summary, and waits for user feedback. Externally the session looks alive and polling runs, but work has effectively stopped. Run logs show this is more common than crashes.
  3. Runtime fragility — context compaction silently breaks the loop; closing a session takes down the timers parasitic on it. Failures go unnoticed by default.

The common cause of all three is missing engineering scaffolding, not insufficient model capability. Every mechanism in this framework targets the failure modes above.

2. Behavioral Constraints

  1. Zero interaction — no prompting the user during a run: no Plan Mode, no question tool, no ending on a question. Continue working until the user stops you. Resolve ambiguity yourself and write the reasoning to the log (level=decision).
  2. Ready means execute — the most common hidden violation: finishing all preparation and then asking "should I submit?". The purpose of preparation is execution; submitting, resubmitting, fixing, and starting monitors are all routine operations needing no confirmation.
  3. Callback means report-alive — after context compaction the loop dies silently. The first action of every callback is to update its own last_seen, then check liveness; on detecting failure it restarts immediately and logs it.
  4. Persist state to files — all progress is written to state/ files, not conversation memory. Each iteration starts a fresh session, injecting only curated state; never use resume.
  5. Guardian / worker separation — a heartbeat patrol may take only three actions on tasks that are not its own: liveness-check, restart, nudge. It does not read their data, modify their state files, or report to the user on their behalf.

3. Architecture

┌── Orchestrator (current session / durable cron) ──┐
│ monitor state files → detect stalls → inject direction │
└────┬─────────────┬─────────────┬────────────┘
  [Task A]      [Task B]      [Task C]   ← each its own fresh session

Core design decisions:

  • Separate execution from evaluation — the agent doing the work does not judge its own progress; stall determination is made by the orchestration layer based on quantitative metrics.
  • Fresh session over resume — context accumulation is the primary cause of cognitive loops. Each iteration starts with fresh context; state is injected via files.
  • Enforced direction diversity — before each iteration, read the list of tried directions; a new direction must differ from all history.

4. State Files

{task}/state/
├── task_spec.md           # goal / milestones / success criteria
├── progress.json          # {iteration, total_findings, status, stale_count}
├── findings.jsonl         # accumulated findings (append-only)
├── directions_tried.json  # directions already tried
└── iteration_log.jsonl    # per-iteration summary

{task}/logs/
├── work.jsonl             # written by work agent; decisions tagged level=decision
├── orchestrator.jsonl     # written by orchestrator
└── heartbeat.jsonl        # written by heartbeat watchdog

Log line format: {"ts":"...", "source":"...", "level":"info|warn|error|decision", "event":"...", "detail":"..."}

5. Usage

# 1. Initialize the task directory, write state/task_spec.md and an initial progress.json

# 2. Start the orchestrator loop with the agent_loop tool (goal mode):
#    goal = "every task's progress.json advances each iteration".
#    Each iteration: (1) read progress.json for all tasks;
#    (2) if stale_count>=3 generate a fresh direction; (3) launch a work agent
#    via the subagent tool (with explicit goal and completion criteria);
#    (4) write results back to state files; (5) call agent_loop next.
#    Zero interaction.

# 3. Register a durable SQLite-backed scheduled-task heartbeat watchdog
# (survives across sessions):
# hourly patrol: write a timestamp; check each loop's last_seen against interval×3,
# restart if exceeded; check each task's progress for stalls over 2h, nudge if stalled.
# Zero interaction. In Work sessions, agent_loop itself is the cadence — keep
# iterations short enough that a stalled one is visible in the session.

6. Stall Detection & Pivoting

MechanismRule
Stall detectionan iteration with 0 new findings or a metric drop → stale_count + 1
Forced pivotstale_count >= 2 → change a structural constraint, not tactical parameters; >= 4 → flag for human attention
Direction diversitya new direction must differ from every tried one; after a stall, inject a perturbation strategy
Round capa single work session caps at 15 rounds or 30 minutes

"Pivot structure, not tactics" comes from practice: when a task stalls repeatedly within a frame, the decisive gain usually comes from correcting the environment/structural constraint itself, not from tuning strategy parameters harder inside the existing frame.

Show full SKILL.md (464 more words)Show less

7. Heartbeat Watchdog

The business loop is itself unreliable and needs an independent guardian layer. Three mutually-checking layers (V3):

LayerFormDepends onRole
L0resident shell guardno sessionheartbeat stale > 2h → spin up an emergency patrol via a headless agent
L1durable cron, hourlya living interactive sessioncheck each loop's last_seen, restart timed-out loops, detect stalling and nudge
L2business loopeach its own sessionfirst line of each callback updates its own last_seen

Any one layer dying can be detected and recovered by another.

Stall detection: if progress has no update for over 2 hours and the last output is a question → judged stalled, launch a nudge subagent. Three consecutive nudges with no progress → judged structurally stuck; stop nudging and reopen with a new direction. The 2h threshold is deliberately shorter than the 4h stuck-task threshold.

8. Subagent Scheduling Patterns

All patterns below are expressed with the subagent tool: one call per work agent, parallel mode for fan-out, chain mode for staged pipelines (later steps receive earlier output via {previous}).

PatternUseKey idea
A Goal-drivenresearch iterationinject tried directions, require verifiable findings, write back to findings.jsonl
B Parallel explorationcomplex sub-problemsone subagent parallel call: investigation, refutation, cross-domain analogy
C Experiment runlong compute jobschain mode: submit → minute-level polling → auto-diagnose errors, fix, resubmit
D Verificationpost-iteration QAan independent reviewer subagent audits the evidence chain of findings

A subagent task prompt should include: background, a verifiable deliverable, working directory, file/line caps, and completion criteria.

9. Engineering Constraints

  1. At most 5 large files per iteration; no single file over 300 lines.
  2. State is injected via files, not conversation history.
  3. Validation (test / compile / check) must run between iterations.
  4. Citation-like content is verified every 20 entries, never batched up.
  5. With multiple candidate directions, prefer adding diversity over digging one deeper.
  6. Unresolvable external-dependency failures escalate (full report + notify the owner + poll for a reply); never abandon silently.

10. Validation & Limits

The framework has carried several heterogeneous tasks: academic paper writing, long-horizon research, etc. Paper-track output:

PaperPagesCitationsSelf-rated
Autonomous Research Agents592288.0/10
Continual Learning653268.0/10
Long-Horizon Decision-Making553848.0/10
Self-Play (285B RL experiment + theory hardening)752178.6/10

Limits:

  1. Scores come from in-framework multi-persona simulated review; comparable only longitudinally within the same protocol, not an external quality claim.
  2. The longest continuous run on record was 72 hours, with 6 directional human inputs during it — zero operational intervention, directional intervention retained.
  3. Fabricated citations and data artifacts originate from the LLM itself; the framework makes external checking a mechanical step in the process, it does not remove the error source.
  4. Separation of duties relies on protocol constraints, not model self-discipline; removing the constraints brings overstepping behavior back.

© rongxinzy, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in SKILLs/deli-autoresearch of rongxinzy/RongxinAI.

  • SKILL.md
  • zhiyuan/icon.svg
  • zhiyuan/metadata.yaml

Open the folder on GitHubat commit 9c64865

Compare with similar skills

Deli Autoresearch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deli Autoresearch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deli Autoresearch this skillrongxinzy/RongxinAI154—~3.2kAutomated safety check: PassAGPL-3.0
Integrity Forensicswanshuiyin/Auto-claude-code-research-in-sleep17k—~4.2kAutomated safety check: NotesMIT
Academic Paper Writing PipelineImbad0202/academic-research-skills51k—~16kAutomated safety check: PassCustom licence
Academic Research PipelineImbad0202/academic-research-skills51k—~15kAutomated safety check: PassCustom licence
Academic Research Suite for CodexImbad0202/academic-research-skills-codex12k—~12kAutomated safety check: PassCustom licence
Econ Writehanlulong/econ-writing-skill6512 repos~14kAutomated safety check: PassMIT

Similar skills

  • Integrity Forensics

    wanshuiyin/Auto-claude-code-research-in-sleep

    Run the Anti-Autoresearch integrity-forensics sweep (span-anchored evidence ledger → GPT auditors propose findings → a rules-only reporter that lists every proposal with what the auditor said about…

    17k GitHub stars~4.2k tokensUpdated 4 days ago
    Research & ScienceAuto-check: notes
  • Academic Paper Writing Pipeline

    Imbad0202/academic-research-skills

    Runs a 12-agent pipeline that plans, drafts, cites, reviews and formats academic papers, with modes for revision, rebuttals, abstracts and citation checks.

    51k GitHub stars~16k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Academic Research Pipeline

    Imbad0202/academic-research-skills

    Orchestrates a ten-stage academic workflow from research to finished manuscript, including integrity checks, two rounds of peer review and revision.

    51k GitHub stars~15k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Academic Research Suite for Codex

    Imbad0202/academic-research-skills-codex

    A router skill that sends academic work such as literature reviews, drafting, citation checks, peer review and revision to the right workflow in the ARS suite.

    12k GitHub stars~12k tokensUpdated 8 days ago
    Research & ScienceAuto-check passed
  • Econ Write

    hanlulong/econ-writing-skill

    Expert economics paper writing assistant synthesizing advice from 50+ top guides by Cochrane, McCloskey, Shapiro, Head, Bellemare, Goldin, Glaeser, Kremer, and other leading economists.

    651 GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed
  • Autonomous Research

    federicodeponte/opendraft

    An 18-agent pipeline that turns one topic line into a drafted research paper, literature review, or thesis chapter.

    507 GitHub stars~8.2k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed

More from rongxinzy/RongxinAI

All 94 skills in this repo
  • SaaS Metrics Coach

    rongxinzy/RongxinAI

    SaaS financial health advisor. An agent skill from rongxinzy/RongxinAI.

    154 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • Churn Prevention

    rongxinzy/RongxinAI

    Reduce voluntary and involuntary churn through cancel flow design, save offers, exit surveys, and dunning sequences.

    154 GitHub starsUsed in 3 repos~2.6k tokens
    Auto-check passed
  • Presentation Studio

    rongxinzy/RongxinAI

    The only skill for creating a new PowerPoint deck. An agent skill from rongxinzy/RongxinAI.

    154 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Zhiyuan Expert Manager

    rongxinzy/RongxinAI

    ZhiYuan Agent expert package lifecycle manager for the pi engine.

    154 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Ziwei Doushu

    rongxinzy/RongxinAI

    Professional Ziwei Doushu consultation skill with an offline calculation engine.

    154 GitHub stars~746 tokensUpdated yesterday
    Auto-check passed
  • Lark Mail

    rongxinzy/RongxinAI

    飞书邮箱:Use when user mentions 起草邮件、写邮件、草稿、发送/回复/转发邮件、查阅邮件、看邮件、搜索邮件、邮件文件夹、邮件标签、邮件联系人、监听新邮件、邮件收信规则等;use for mail/email intent only.

    154 GitHub starsUsed in 3 repos~4.1k tokens
    Auto-check: warnings

Questions about Deli Autoresearch

What does Deli Autoresearch do?

A protocol framework for long-horizon autonomous research tasks. Deli Autoresearch is an agent skill from rongxinzy/RongxinAI. A protocol framework for long-horizon autonomous research tasks.

When should I use Deli Autoresearch?

Deli Autoresearch fits situations like: the user asks for academic research; literature surveys; any unattended multi-day research task.

How do I install Deli Autoresearch in Claude Code?

Run `npx skills add rongxinzy/RongxinAI --skill deli-autoresearch -a claude-code`. Or copy the skill folder (SKILLs/deli-autoresearch in rongxinzy/RongxinAI) into .claude/skills/deli-autoresearch in your project. Claude Code loads it when a task matches its description.

How do I install Deli Autoresearch in Codex?

Run `npx skills add rongxinzy/RongxinAI --skill deli-autoresearch -a codex`. Or copy the skill folder (SKILLs/deli-autoresearch in rongxinzy/RongxinAI) into .agents/skills/deli-autoresearch in your project. Codex loads it when a task matches its description.

Can I use Deli Autoresearch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rongxinzy/RongxinAI --skill deli-autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deli-autoresearch, .gemini/skills/deli-autoresearch, .github/skills/deli-autoresearch and .opencode/skills/deli-autoresearch in your project.

What does Deli Autoresearch need to run?

SKILL.md names no scripts, command-line tools or credentials: Deli Autoresearch is instructions for the agent only.

Does Deli Autoresearch access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Deli Autoresearch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deli Autoresearch use?

Deli Autoresearch is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deli Autoresearch use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deli Autoresearch?

Skills that share tags, products or a category with Deli Autoresearch: Integrity Forensics (wanshuiyin/Auto-claude-code-research-in-sleep, 17k stars), Academic Paper Writing Pipeline (Imbad0202/academic-research-skills, 51k stars), Academic Research Pipeline (Imbad0202/academic-research-skills, 51k stars) and Academic Research Suite for Codex (Imbad0202/academic-research-skills-codex, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deli Autoresearch?

rongxinzy (a GitHub organization) maintains it in rongxinzy/RongxinAI, which has 154 GitHub stars. The repository holds 94 skills in this directory. The repository was last updated on October 10, 2026.

Source: rongxinzy/RongxinAI on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.