Agent skill

Coral New Task

by Human-Agent-Society in Human-Agent-Society/CORAL

End-to-end recipe for adding a new task under examples/ — the three pieces that have to line up (task.yaml, seed/, and grader/), what to put in each, the TaskGrader API surface, the coral validate →…

Apache-2.0Auto-check passedTesting & QA

Install Coral New Task

skills CLI
$ npx skills add Human-Agent-Society/CORAL --skill coral-new-task -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Human-Agent-Society/CORAL coral-new-task --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Human-Agent-Society/CORAL.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/coral-new-task .claude/skills/coral-new-task && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coral-new-task
GitHub stars
1.1k
Token cost
~3.4k tokens
SKILL.md length
1,093 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

End-to-end recipe for adding a new task under examples/ — the three pieces that have to line up (task.yaml, seed/, and grader/), what to put in each, the TaskGrader API surface, the coral validate →…

  • Works in 5 steps: The seed → The grader → The task.yaml → …
  • The user wants to add a new CORAL task
  • SKILL.md covers Reference implementations, 1. The seed, 2. The grader and 3. The task.yaml, plus 3 more sections
  • Calls uv

What it does

Coral New Task is an agent skill from Human-Agent-Society/CORAL. End-to-end recipe for adding a new task under examples/ — the three pieces that have to line up (task.yaml, seed/, and grader/), what to put in each, the TaskGrader API surface, the coral validate → smoke-test loop, and the common mistakes (repopath pointing at the wrong dir, score direction backwards, hidden answer keys leaking into seed/, grader writing to codebasepath which the daemon force-removes, private-vs-public confusion, missing run() signature). Use whenever the user wants to add a new CORAL task or…

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering QA and bug reports. The repository describes itself as: Open-source autoresearch powered by autonomous coding agents. Run Claude Code, OpenCode, and Codex with grading, shared knowledge, and multi-agent evolution. Accepted at COLM 2026. The licence is Apache-2.0.

When your agent uses it

  • The user wants to add a new CORAL task
  • Port an existing benchmark into CORAL

Example prompts

  • “/coral-new-task”

Requirements

  • Python 3
  • Docker

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. The seed
  2. The grader
  3. The task.yaml
  4. Validate before running agents
  5. Wire it into the index

What it can do on your machine

Read from SKILL.md and the folder at commit 0123dfb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Coral New Task loads about 3.4k tokens when it runs. Until then it costs about 146 tokens; SKILL.md has 1,093 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~146
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Human-Agent-Society/CORAL at commit 0123dfb, republished under its Apache-2.0 licence (© Human-Agent-Society). 1,093 words, ~3,361 tokens.

Download SKILL.mdSave it as .claude/skills/coral-new-task/SKILL.md (or your agent's skills folder).
name
coral-new-task
description
End-to-end recipe for adding a new task under `examples/` — the three pieces that have to line up (`task.yaml`, `seed/`, and `grader/`), what to put in each, the `TaskGrader` API surface, the `coral validate` → smoke-test loop, and the common mistakes (repo_path pointing at the wrong dir, score direction backwards, hidden answer keys leaking into seed/, grader writing to codebase_path which the daemon force-removes, private-vs-public confusion, missing `run()` signature). Use whenever the user wants to add a new CORAL task or port an existing benchmark into CORAL.

Creating a new CORAL task

A CORAL task is three things that must line up:

examples/<task>/
├── task.yaml      # config: name, description, grader entrypoint, agent count
├── seed/          # starter code agents see when they begin (the repo_path)
│   └── solution.py
└── grader/        # standalone Python package
    ├── pyproject.toml
    └── src/<task>_grader/
        ├── __init__.py
        └── grader.py     # class Grader(TaskGrader): ...

The packaged form is the only supported form. The package gives the grader its own venv and ships everything the eval needs — grader code, helper modules, and hidden data (see "Hidden data" below).

Reference implementations

Look at these before writing anything new — copy the closest one and edit:

ReferenceWhen to copy it
examples/erdos/Minimal packaged grader, single grader file, numpy-only deps
examples/dna_design/Packaged grader with bundled data files (importlib.resources) and [ml] optional-deps for heavy libs
examples/swebench-verified/Tiered eval (different instance counts per tier), private answer keys, harbor integration
examples/circle_packing/Smallest packaged task end-to-end — single solution file, single grader file
examples/mnist/Packaged grader with a hidden answer key (note: secret data belongs under grader.private in a taskdata/ sibling of grader/, never inside the grader package)

1. The seed

Whatever lives in seed/ is what the agent sees on first checkout — it's the working directory the grader will later score. The contract between seed/ and the grader is the program file: a Python file with a function the grader imports and calls.

The convention across examples is:

  • solution.py (or initial_program.py) defining a top-level run() function.
  • The grader passes program_file: "solution.py" via grader.args.
  • run()'s signature is whatever the grader expects — usually () -> result or (input_path) -> result.

Put a real, runnable baseline here. Agents should be able to coral eval immediately and get a non-zero score, so they have a starting point to improve. A no-op skeleton that crashes is not a good baseline.

If the task needs data files at runtime (training data, fixtures), put them under seed/data/ and reference them by relative path from solution.py. The grader will see them at <codebase_path>/data/....

2. The grader

grader/
├── pyproject.toml
└── src/<task>_grader/
    ├── __init__.py
    └── grader.py

pyproject.toml is a thin Hatchling package. Crib from examples/erdos/grader/pyproject.toml:

toml
[project]
name = "<task>-grader"
version = "0.1.0"
description = "CORAL grader for the <task> task."
requires-python = ">=3.11"
dependencies = ["coral", "numpy"]   # Whatever the grader actually imports.

[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"

[tool.hatch.build.targets.wheel]
packages = ["src/<task>_grader"]

Subclass TaskGrader and implement evaluate():

python
# grader/src/<task>_grader/grader.py
from coral.grader import TaskGrader
from coral.types import ScoreBundle


class Grader(TaskGrader):
    def evaluate(self) -> float | ScoreBundle:
        program_file = self.args.get("program_file", "solution.py")
        # self.codebase_path  — the agent's commit checked out detached
        # self.private_dir    — .coral/private/ (your hidden answer keys live here)
        # self.args           — dict from task.yaml grader.args
        # self.timeout        — grader.timeout in seconds (or None)
        # self.eval_logs_dir  — write subprocess logs / artifacts the agent should see post-grade

        try:
            result = run_program_and_score(...)
        except TimeoutError:
            return self.fail(f"Evaluation timed out after {self.timeout}s")
        except Exception as e:
            return self.fail(f"Evaluation failed: {e}")

        return self.score(result, explanation=f"score={result:.4f}")

What you have available on self:

Attribute / methodUse it for
self.codebase_pathPath to the commit being graded (detached worktree). Read-only — anything written here is discarded after the eval.
self.private_dir.coral/private/. Your answer keys, hidden test data, anything from grader.private lives here.
self.argsdict from task.yaml::grader.args. Use self.args.get("program_file", "solution.py") etc.
self.timeoutEval timeout in seconds (or None if grader.timeout: 0).
self.eval_logs_dirPer-attempt directory for logs/artifacts that should outlive the grader. Symlinked into each agent worktree as <shared_dir>/eval_logs/<hash>/.
self.score(value, explanation=...)Build a single-task ScoreBundle from a numeric score.
self.fail(reason)Return a fail ScoreBundle with reason as feedback.
self.get_python_command()List for the python binary inside the codebase's env (uses uv run if a pyproject.toml is present). Always use this instead of sys.executable so task-specific deps are visible.
self.run_program(filename, *args)Convenience: runs <codebase_path>/<filename> as a subprocess via get_python_command().
Bundling data files with the grader

If the grader needs reference files (model weights, ground-truth answers, scoring fixtures), ship them inside the package and load via importlib.resources:

python
import importlib.resources

scorer_dir = str(importlib.resources.files("<task>_grader.scorers"))

examples/dna_design/grader/src/dna_design_grader/grader.py is the canonical pattern — note the scorers/ subpackage. Add the directory to [tool.hatch.build.targets.wheel] if it has non-Python files.

Heavy / optional dependencies

If the grader wants torch, grelu, etc., put them in optional-dependencies and have the grader fall back gracefully when missing — see examples/dna_design/grader/pyproject.toml. Then grader.setup becomes ["uv pip install -e ./grader[ml]"] for the full version.

Hidden data: put secrets under grader.private

Answer keys, hidden test fixtures, and anything the agent must not see go under grader.private in task.yaml — CORAL copies those paths into .coral/private/<name> (which every agent runtime is denied read access to) and the grader reads them via self.private_dir:

yaml
grader:
  private:
    - "taskdata"   # a sibling of grader/ → .coral/private/taskdata, read via Path(self.private_dir) / "taskdata"

The single rule: everything inside the grader/ package is visible to agents, and secrets live in grader.private outside grader/. The whole grader source is surfaced read-only to agents at <shared_dir>/grader/ (a symlink to the real package) so they can read how they're scored — so anything under grader/ is readable, including a grader.private path that points inside it (it gets copied to .coral/private/ and leaked through the surfaced source). Keep taskdata/ as a sibling of grader/ (declare it taskdata, resolving to <task_dir>/taskdata); coral validate errors if a private path is inside the package.

Non-secret bundled data (lookup tables, reference configs, helper modules) may live inside grader/ and be read via Path(__file__).parent / ... — just remember it's visible, so never put a secret there. E.g. examples/ADRS/txn_scheduling/ puts helper modules on sys.path that way.

Show full SKILL.md (394 more words)Show less

3. The task.yaml

Fields that must be set; everything else has a sensible default.

yaml
task:
  name: "My Task"                       # shows up in results/<slug>/
  description: |                        # rendered into CORAL.md, agents read this
    What the agent should do.
    Reference the program file by name (e.g. solution.py and its run() signature).
  tips: |                               # optional, also rendered into CORAL.md
    - Eval timeout is N seconds.
    - Constraints / scoring details / known baselines.

grader:
  entrypoint: "<task>_grader.grader:Grader"   # required
  setup:
    - "uv pip install -e ./grader"            # runs once in .coral/private/grader_venv/
  timeout: 600                          # seconds; 0 disables, default 300
  direction: maximize                   # or minimize — controls leaderboard ordering
  args:                                 # arbitrary dict, read inside grader as self.args
    program_file: "solution.py"
  private: []                           # extra files copied into .coral/private/ (hidden from agents)
  parallel:
    max_workers: 1                      # bump only when the grader is concurrency-safe
  max_pending_per_agent: 1              # cap on in-flight submissions per agent

agents:
  count: 1                              # raise once the task is known to be stable
  runtime: claude_code                  # claude_code | codex | cursor_agent | kiro | opencode
  model: sonnet                         # default depends on runtime; see coral/agent/registry.py

workspace:
  results_dir: "./results"              # where each run lands
  repo_path: "./examples/<task>/seed"   # MUST point at the seed/ dir
  setup:                                # runs in each agent worktree before agents start
    - "uv pip install numpy"            # task-runtime deps go here, NOT in grader/setup

run:
  verbose: false
  ui: false
  session: tmux                         # local | tmux | docker

The examples/README.md documents the full schema with every default. When in doubt, look there before adding fields.

4. Validate before running agents

bash
coral validate examples/<task>

This:

  1. Parses task.yaml and reports schema errors.
  2. Bootstraps .coral/private/grader_venv/ and runs grader.setup.
  3. Copies seed/ into a tempdir and runs the grader against it once.
  4. Prints the resulting score and explanation.

If coral validate succeeds, the grader can score the seed. That's the single most important checkpoint — most "agent stuck" issues trace back to a grader that crashes on the seed.

After validation, smoke-test with one agent:

bash
coral start -c examples/<task>/task.yaml agents.count=1 run.session=local
# Wait for one eval, then:
coral stop

5. Wire it into the index

Add a one-line entry to the table in examples/README.md and a short ### <task> section under "Details" with the bullet points the others use (Agents / Timeout / Session). Skip this only for throwaway local tasks.

Common mistakes

MistakeSymptomFix
repo_path points at examples/<task>/ instead of examples/<task>/seed/Grader sees task.yaml and grader/ in codebase_pathAlways point repo_path at the seed dir.
direction: maximize for a loss / minimize for a benchmark ratioLeaderboard ordered backwardsScore = "ratio against benchmark, >1 is better" → maximize. Score = "raw error" → minimize.
Hidden answer key under seed/ or anywhere inside the grader/ packageAgents read it and game the score — seed/ is their repo, and the whole grader/ source is surfaced at <shared_dir>/grader/Put it under grader.private outside grader/ (e.g. a sibling taskdata/), read via self.private_dir. coral validate errors on a private path inside grader/.
Grader writes results under self.codebase_path and reads them laterFiles vanish — daemon force-removes the worktree after each evalWrite under self.eval_logs_dir.
Grader uses sys.executable to run the agent's programMisses task-specific deps installed via workspace.setupUse self.get_python_command() (it switches to uv run --project when the codebase has a pyproject.toml).
Heavy deps in main grader dependenciescoral validate is slow / fails on machines without the GPU stackMove to optional-dependencies and fall back gracefully. See dna_design's [ml] extra.
grader.setup tries to install task-runtime depsThe grader venv has them but the agent's worktree doesn'tTask-runtime deps go in workspace.setup. Grader-only deps go in grader.setup.
parallel.max_workers > 1 with a non-concurrency-safe graderSporadic failures when two evals collide on Docker ports / GPU / scratch dirsLeave at 1 unless the grader is provably safe.
Forgetting coral validateAgents start, fail every eval with the same errorAlways validate first.

© Human-Agent-Society, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/coral-new-task of Human-Agent-Society/CORAL.

Open the folder on GitHubat commit 0123dfb

Compare with similar skills

Coral New Task next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Coral New Task compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Coral New Task this skillHuman-Agent-Society/CORAL1.1k—~3.4kAutomated safety check: PassApache-2.0
Acceptance Evidence for Deliverieslobehub/lobehub83k—~9.7kAutomated safety check: PassApache-2.0
Senpi Agent QA Harnesscode-yeongyu/senpi474—~2.7kAutomated safety check: NotesMIT
tmux Real User TestingQwenLM/qwen-code28k—~2.3kAutomated safety check: PassApache-2.0
StandardsItamarZand88/CLI-Anything-WEB231—~4.2kAutomated safety check: PassMIT
Dev IssueFHIR/fhir-codegen155—~4kAutomated safety check: PassMIT

Similar skills

  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Senpi Agent QA Harness

    code-yeongyu/senpi

    Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.

    474 GitHub stars~2.7k tokensUpdated today
    Testing & QAAuto-check: notes
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Standards

    ItamarZand88/CLI-Anything-WEB

    Runs Phase 4 review/publish/verify for a cli-web- CLI: implementation review by 3 parallel agents, the tiered quality checklist (Tier 1 critical fail-fast, then comprehensive), pip install + smoke…

    231 GitHub stars~4.2k tokensUpdated 10 days ago
    Testing & QAAuto-check passed
  • Dev Issue

    FHIR/fhir-codegen

    Publishes a slot's feature request or bug report to GitHub as an issue, and keeps that issue in sync, in the role of a release-minded engineer.

    155 GitHub stars~4k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Testing Scope MCP

    daydreamlive/scope

    Test Daydream Scope through its stdio MCP server, including disconnected startup, connecttoscope, direct Python stdio client fallback, and lightweight smoke tests.

    452 GitHub stars~1.2k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed

More from Human-Agent-Society/CORAL

  • Creating A Coral Task

    Human-Agent-Society/CORAL

    Author a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout…

    1.1k GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Running Coral Experiments

    Human-Agent-Society/CORAL

    Run and manage CORAL experiments from the operator side — launch agents with coral start (dotlist overrides, model/count, tmux vs local), monitor with coral status / coral log / coral show / the web…

    1.1k GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Coral Debug

    Human-Agent-Society/CORAL

    Verify and debug changes to CORAL itself — smallest reproduce loop per area (grader / daemon / CLI / hooks / manager / workspace / hub / template / config / web), where to look when something breaks…

    1.1k GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Promoting Dev To Main

    Human-Agent-Society/CORAL

    A skill your agent uses when preparing, reviewing, resolving conflicts for, or merging a CORAL release pull request from the long-lived dev branch into main.

    1.1k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • Setting Up Coral

    Human-Agent-Society/CORAL

    One-time machine setup after installing the coral CLI — register local agent runtimes as named bindings with coral setup / coral setup agent, validate them with coral agents doctor (incl.

    1.1k GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Coral Extend

    Human-Agent-Society/CORAL

    Add a new component to the CORAL framework itself — a new agent runtime under coral/agent/builtin/ (claudecode/codex/cursoragent style), a new CLI command in coral/cli/, a new bundled skill or…

    1.1k GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Coral New Task

What does Coral New Task do?

End-to-end recipe for adding a new task under examples/ — the three pieces that have to line up (task.yaml, seed/, and grader/), what to put in each, the TaskGrader API surface, the coral validate →…. Coral New Task is an agent skill from Human-Agent-Society/CORAL.yaml, seed/, and grader/), what to put in each, the TaskGrader API surface, the coral validate → smoke-test loop, and the common mistakes (repopath pointing at the wrong dir, score direction backwards, hidden answer keys leaking into seed/, grader writing to codebasepath which the daemon force-removes, private-vs-public confusion, missing run() signature).

When should I use Coral New Task?

Coral New Task fits situations like: the user wants to add a new CORAL task; port an existing benchmark into CORAL.

How do I install Coral New Task in Claude Code?

Run `npx skills add Human-Agent-Society/CORAL --skill coral-new-task -a claude-code`. Or copy the skill folder (.claude/skills/coral-new-task in Human-Agent-Society/CORAL) into .claude/skills/coral-new-task in your project. Claude Code loads it when a task matches its description.

How do I install Coral New Task in Codex?

Run `npx skills add Human-Agent-Society/CORAL --skill coral-new-task -a codex`. Or copy the skill folder (.claude/skills/coral-new-task in Human-Agent-Society/CORAL) into .agents/skills/coral-new-task in your project. Codex loads it when a task matches its description.

Can I use Coral New Task in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Human-Agent-Society/CORAL --skill coral-new-task -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coral-new-task, .gemini/skills/coral-new-task, .github/skills/coral-new-task and .opencode/skills/coral-new-task in your project.

What does Coral New Task need to run?

Going by SKILL.md and its folder, Coral New Task needs the command-line tools its instructions call (uv). Our summary lists: Python 3; Docker.

Does Coral New Task access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Coral New Task safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Coral New Task use?

Coral New Task is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Coral New Task use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Coral New Task?

Skills that share tags, products or a category with Coral New Task: Acceptance Evidence for Deliveries (lobehub/lobehub, 83k stars), Senpi Agent QA Harness (code-yeongyu/senpi, 474 stars), tmux Real User Testing (QwenLM/qwen-code, 28k stars) and Standards (ItamarZand88/CLI-Anything-WEB, 231 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Coral New Task?

Human-Agent-Society (a GitHub organization) maintains it in Human-Agent-Society/CORAL, which has 1,060 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 8, 2026.

Source: Human-Agent-Society/CORAL on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.