Agent skill

Autoresearch

by grandamenium in grandamenium/cortextos

The analyst has assigned you a research cycle, or you have identified a metric you want to improve through systematic experimentation.

MITAuto-check passedAgent Workflows

Install Autoresearch

skills CLI
$ npx skills add grandamenium/cortextos --skill autoresearch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install grandamenium/cortextos autoresearch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/grandamenium/cortextos.git skills-src && mkdir -p .claude/skills && cp -r skills-src/community/skills/autoresearch .claude/skills/autoresearch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autoresearch
GitHub stars
101
Token cost
~1.9k tokens
SKILL.md length
665 words
Files
1
Skills in repo
55
Repo updated
First seen
Licence
MIT

At a glance

The analyst has assigned you a research cycle, or you have identified a metric you want to improve through systematic experimentation.

  • Works in 6 steps: Gather Context → Evaluate Previous Experiment → Hypothesize → …
  • Tasks that involve Autonomous loops
  • SKILL.md covers What It Is, The Experiment Loop, Measurement Methods and Setting Up a Cycle, plus 1 more section
  • Calls jq and bash

What it does

Autoresearch is an agent skill from grandamenium/cortextos. The analyst has assigned you a research cycle, or you have identified a metric you want to improve through systematic experimentation. You will form a hypothesis, make a targeted change, measure the outcome against a baseline, and decide whether to keep or discard the change. You repeat this loop until the metric improves or you exhaust viable hypotheses. This is not ad-hoc research — it is structured scientific iteration with a defined metric, a hypothesis, and a measurable result.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Autonomous loops and A/B testing. The licence is MIT.

When your agent uses it

  • Tasks that involve Autonomous loops
  • Tasks that involve A/B testing

Example prompts

  • “/autoresearch”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Gather Context
  2. Evaluate Previous Experiment
  3. Hypothesize
  4. Create Experiment
  5. Make Changes and Run
  6. Wait

What it can do on your machine

Read from SKILL.md and the folder at commit 6f93838. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autoresearch loads about 1.9k tokens when it runs. Until then it costs about 125 tokens; SKILL.md has 665 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~125
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from grandamenium/cortextos at commit 6f93838, republished under its MIT licence (© grandamenium). 665 words, ~1,927 tokens.

Download SKILL.mdSave it as .claude/skills/autoresearch/SKILL.md (or your agent's skills folder).
name
autoresearch
description
The analyst has assigned you a research cycle, or you have identified a metric you want to improve through systematic experimentation. You will form a hypothesis, make a targeted change, measure the outcome against a baseline, and decide whether to keep or discard the change. You repeat this loop until the metric improves or you exhaust viable hypotheses. This is not ad-hoc research — it is structured scientific iteration with a defined metric, a hypothesis, and a measurable result.
triggers
experiment, autoresearch, hypothesis, research cycle, optimize, improve metric, run experiment, test hypothesis, measure improvement, scientific loop…

Autoresearch

You are a scientist. Autoresearch is how you systematically improve specific aspects of your work by running experiments, measuring results, and learning from outcomes.

What It Is

You have research cycles assigned to you (check experiments/config.json). Each cycle has:

  • A metric you are optimizing (the dependent variable)
  • A surface you are experimenting on (the independent variable - what you change)
  • A direction (higher or lower = better)
  • A measurement window (how long to wait before measuring)
  • A measurement method (how to get the metric value)

You cannot autonomously modify your own cycle configuration. If the user asks you to modify a cycle, you can. Otherwise, the analyst (via theta wave) is the one who creates, modifies, or removes cycles. You CAN and SHOULD run experiments within your assigned cycles.

The Experiment Loop

When your experiment cron fires, execute these steps:

Step 1: Gather Context
bash
cortextos bus gather-context --agent $CTX_AGENT_NAME --format markdown

Read the output carefully. Pay attention to:

  • What experiments have been tried before
  • What was kept (these patterns work - build on them)
  • What was discarded (these approaches failed - avoid repeating)
  • Your current keep rate and trajectory
Step 2: Evaluate Previous Experiment

If there is an active experiment (check experiments/active.json):

  • Compare ALL relevant aspects: the surface changes you made, the context around those changes, and the output metric
  • Measure the metric using the configured measurement method
  • Run evaluate-experiment:
bash
cortextos bus evaluate-experiment <experiment_id> <measured_value> --justification "Why this result makes sense"

For qualitative metrics, use --score <1-10> with a written justification.

Step 3: Hypothesize

Based on accumulated learnings:

  • Review what worked (keeps) and what failed (discards)
  • Identify patterns - what themes appear in successful experiments?
  • Consider untested approaches
  • Form a specific, testable hypothesis
  • Your hypothesis must be evidence-backed (cite past results or research)

Exploit vs Explore: If something has been kept 3+ times in a row, exploit that pattern further. If you have been discarding 3+ times, try something more radically different.

Step 4: Create Experiment
bash
cortextos bus create-experiment "<metric_name>" "<your hypothesis>" --surface <path> --direction <higher|lower> --window <duration>

If approval_required is true in experiments/config.json, you must manually create an approval before proceeding:

bash
APPR_ID=$(cortextos bus create-approval "Run experiment: <hypothesis>" experiments "Cycle: <cycle_name>, Metric: <metric_name>, Surface: <surface>")
cortextos bus send-telegram $CTX_TELEGRAM_CHAT_ID "Approval needed to run experiment for <metric_name> — check dashboard"
# Block until approved, then continue to Step 5
Step 5: Make Changes and Run

Apply your hypothesized changes to the surface file. Then:

bash
cortextos bus run-experiment <experiment_id> "Description of what you changed"

This creates a git commit with your changes (the experiment commit) so they can be cleanly reverted if the experiment fails.

Step 6: Wait

The cycle ends. Your next cron trigger picks up at Step 1, where you will evaluate this experiment.

Measurement Methods

Quantitative (scripted)

A script returns a number. Example: API scrape for engagement rate.

bash
bash connectors/measure-instagram.sh
# Output: metric_value: 3.2
Quantitative (computed)

You calculate from existing data. Example: task completion rate.

bash
COMPLETED=$(cortextos bus list-tasks --agent $CTX_AGENT_NAME --status completed | jq length)
TOTAL=$(cortextos bus list-tasks --agent $CTX_AGENT_NAME | jq length)
RATE=$(echo "scale=2; $COMPLETED / $TOTAL * 100" | bc)
Show full SKILL.md (264 more words)Show less
Qualitative (subjective)

You evaluate output quality on a 1-10 scale. You MUST write a justification.

bash
cortextos bus evaluate-experiment <id> 0 --score 7 --justification "Output is more concise and actionable than baseline, but loses some nuance"
Qualitative (comparative)

You compare baseline vs experiment output side by side and score 1-10.

Setting Up a Cycle

If the user asks you to set up autoresearch, collect these 8 things:

  1. Metric — what to optimize (e.g., "engagement_rate", "task_completion_rate", "briefing_quality")
  2. Metric type — quantitative (a number you can script/compute) or qualitative (a 1-10 score you evaluate)
  3. Surface — the file to experiment on (e.g., experiments/surfaces/engagement/current.md for a prompt, or SOUL.md for behavior)
  4. Direction — higher or lower is better
  5. Measurement — how to get the metric value (a script, computed from tasks, or self-evaluation)
  6. Window — how long to wait before measuring the result (e.g., 24h, 48h)
  7. Loop interval — how often to run the experiment loop (the cron frequency — often same as window)
  8. Approval — should you need approval before running each experiment?

Then create the cycle and surface directory:

bash
# Create surface directory and baseline file
mkdir -p "experiments/surfaces/<metric>"
cat > "experiments/surfaces/<metric>/current.md" << 'EOF'
# <metric> — Baseline

[Describe the current approach being tested]
EOF

# Register the cycle
cortextos bus manage-cycle create $CTX_AGENT_NAME \
  --cycle "<metric_name>" \
  --metric "<metric_name>" \
  --metric-type "<quantitative|qualitative>" \
  --surface "experiments/surfaces/<metric>/current.md" \
  --direction "<higher|lower>" \
  --window "<e.g. 24h>" \
  --measurement "<how to measure>" \
  --loop-interval "<e.g. 48h>"

# Update approval setting in config if needed (default is true)
# Only set to false if user explicitly says no approval needed

Then add the experiment cron via the bus (persistent across restarts):

bash
cortextos bus add-cron $CTX_AGENT_NAME experiment-<metric> <loop_interval> "Read .claude/skills/autoresearch/SKILL.md and execute the experiment loop."

To modify a cycle when the user asks:

bash
cortextos bus manage-cycle modify $CTX_AGENT_NAME --cycle "<name>" \
  --window "<new>" \
  --loop-interval "<new>" \
  --enabled <true|false>

Use --enabled false to pause a cycle without deleting it.

Important Rules

  1. Never autonomously modify your own cycle config. If the user asks you to, you can.
  2. You MUST log learnings for EVERY experiment, including failures. Negative learnings are equally valuable.
  3. You MUST respect the measurement window - do not evaluate early.
  4. If approval_required is true, WAIT for approval before running.
  5. Never repeat a hypothesis that was already discarded. Find a new angle.
  6. Keep experiments focused - change one thing at a time when possible.

© grandamenium, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in community/skills/autoresearch of grandamenium/cortextos.

Open the folder on GitHubat commit 6f93838

Compare with similar skills

Autoresearch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autoresearch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autoresearch this skillgrandamenium/cortextos101—~1.9kAutomated safety check: PassMIT
KapsoLeeroo-AI/kapso121—~642Automated safety check: PassMIT
Autoresearchericosiu/ai-marketing-skills3.6k2 repos~2.2kAutomated safety check: PassMIT
Autoresearchgithub/awesome-copilot40k1 repos~2.8kAutomated safety check: PassMIT
Edt MCP Tool DescriptionsDitriXNew/EDT-MCP296—~3kAutomated safety check: PassAGPL-3.0
Evo MemoryEvoScientist/EvoSkills4783 repos~4.8kAutomated safety check: PassApache-2.0

Similar skills

  • Kapso

    Leeroo-AI/kapso

    Optimize code using KAPSO (Knowledge-Grounded Optimization).

    121 GitHub stars~642 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Autoresearch

    ericosiu/ai-marketing-skills

    Run Karpathy-style autoresearch optimization on any content.

    3.6k GitHub starsUsed in 2 repos~2.2k tokens
    Marketing & SEOAuto-check passed
  • Autoresearch

    github/awesome-copilot

    Official

    Autonomous iterative experimentation loop for any programming task.

    40k GitHub starsUsed in 1 repo~2.8k tokens
    Agent WorkflowsAuto-check passed
  • Edt MCP Tool Descriptions

    DitriXNew/EDT-MCP

    How to size, write and A/B-test the text of a tool — its description and its inputSchema parameter prose — so that cutting it does not cost call quality, and so that a tool that IS getting called…

    296 GitHub stars~3k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Evo Memory

    EvoScientist/EvoSkills

    Manages persistent research memory across ideation and experimentation cycles.

    478 GitHub starsUsed in 3 repos~4.8k tokens
    Agent WorkflowsAuto-check passed
  • Spec Optimize

    leo-kuang-ai/spec-first

    Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.

    107 GitHub stars~13k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from grandamenium/cortextos

All 55 skills in this repo
  • Cortext Self Diagnosis

    grandamenium/cortextos

    Diagnose cortextOS itself when the framework misbehaves — an agent has gone silent or wedged, agents are crash-looping, Telegram or agent-to-agent messages are not arriving, crons did not fire, an…

    101 GitHub stars~3.7k tokensUpdated 18 days ago
    Auto-check passed
  • Activity Channel

    grandamenium/cortextos

    You have completed something significant and want the whole org — all agents and the user — to know about it.

    101 GitHub stars~624 tokensUpdated 18 days ago
    Auto-check passed
  • Agentcard Purchase

    grandamenium/cortextos

    You need to make a purchase on behalf of the user — buy a SaaS subscription, pay for an API, purchase a domain, or any transaction requiring a credit card.

    101 GitHub stars~1.1k tokensUpdated 18 days ago
    Auto-check passed
  • Claude To Codex Migration

    grandamenium/cortextos

    Migrate ANY cortextOS agent from the claude-code runtime to the live codex-app-server runtime.

    101 GitHub stars~12k tokensUpdated 18 days ago
    Auto-check: warnings
  • Bus Reference

    grandamenium/cortextos

    Complete cortextos bus CLI reference - all available commands with examples.

    101 GitHub stars~3.8k tokensUpdated 18 days ago
    Auto-check passed
  • Business News Monitor

    grandamenium/cortextos

    Daily cron-driven scan of news/forums/social in a domain to surface market shifts, new competitors, regulatory changes, and net-new opportunities.

    101 GitHub stars~1.3k tokensUpdated 18 days ago
    Auto-check passed

Questions about Autoresearch

What does Autoresearch do?

The analyst has assigned you a research cycle, or you have identified a metric you want to improve through systematic experimentation. Autoresearch is an agent skill from grandamenium/cortextos. The analyst has assigned you a research cycle, or you have identified a metric you want to improve through systematic experimentation.

When should I use Autoresearch?

Autoresearch fits situations like: tasks that involve Autonomous loops; tasks that involve A/B testing.

How do I install Autoresearch in Claude Code?

Run `npx skills add grandamenium/cortextos --skill autoresearch -a claude-code`. Or copy the skill folder (community/skills/autoresearch in grandamenium/cortextos) into .claude/skills/autoresearch in your project. Claude Code loads it when a task matches its description.

How do I install Autoresearch in Codex?

Run `npx skills add grandamenium/cortextos --skill autoresearch -a codex`. Or copy the skill folder (community/skills/autoresearch in grandamenium/cortextos) into .agents/skills/autoresearch in your project. Codex loads it when a task matches its description.

Can I use Autoresearch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add grandamenium/cortextos --skill autoresearch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autoresearch, .gemini/skills/autoresearch, .github/skills/autoresearch and .opencode/skills/autoresearch in your project.

What does Autoresearch need to run?

Going by SKILL.md and its folder, Autoresearch needs the command-line tools its instructions call (jq and bash).

Does Autoresearch access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Autoresearch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Autoresearch use?

Autoresearch is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Autoresearch use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Autoresearch?

Skills that share tags, products or a category with Autoresearch: Kapso (Leeroo-AI/kapso, 121 stars), Autoresearch (ericosiu/ai-marketing-skills, 3.6k stars), Autoresearch (github/awesome-copilot, 40k stars) and Edt MCP Tool Descriptions (DitriXNew/EDT-MCP, 296 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autoresearch?

grandamenium (a GitHub user) maintains it in grandamenium/cortextos, which has 101 GitHub stars. The repository holds 55 skills in this directory. The repository was last updated on September 23, 2026.

Source: grandamenium/cortextos on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.