Evaluation Framework
athola/claude-night-market
Provides weighted scoring, rubrics, and decision-threshold patterns.
Author a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout…
$ npx skills add Human-Agent-Society/CORAL --skill creating-a-coral-task -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Human-Agent-Society/CORAL creating-a-coral-task --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Human-Agent-Society/CORAL.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/creating-a-coral-task .claude/skills/creating-a-coral-task && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "creating-a-coral-task" agent skill from https://github.com/Human-Agent-Society/CORAL/tree/main/plugin/skills/creating-a-coral-task into .claude/skills/creating-a-coral-task/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-a-coral-task", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Human-Agent-Society/CORAL/tree/main/plugin/skills/creating-a-coral-taskType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Human-Agent-Society/CORAL --skill creating-a-coral-task -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Human-Agent-Society/CORAL creating-a-coral-task --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Human-Agent-Society/CORAL.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugin/skills/creating-a-coral-task .agents/skills/creating-a-coral-task && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "creating-a-coral-task" agent skill from https://github.com/Human-Agent-Society/CORAL/tree/main/plugin/skills/creating-a-coral-task into .agents/skills/creating-a-coral-task/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-a-coral-task", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Human-Agent-Society/CORAL --skill creating-a-coral-task -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Human-Agent-Society/CORAL creating-a-coral-task --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Human-Agent-Society/CORAL.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugin/skills/creating-a-coral-task .cursor/skills/creating-a-coral-task && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "creating-a-coral-task" agent skill from https://github.com/Human-Agent-Society/CORAL/tree/main/plugin/skills/creating-a-coral-task into .cursor/skills/creating-a-coral-task/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-a-coral-task", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Human-Agent-Society/CORAL.git --path plugin/skills/creating-a-coral-task--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Human-Agent-Society/CORAL --skill creating-a-coral-task -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Human-Agent-Society/CORAL creating-a-coral-task --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Human-Agent-Society/CORAL.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugin/skills/creating-a-coral-task .gemini/skills/creating-a-coral-task && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "creating-a-coral-task" agent skill from https://github.com/Human-Agent-Society/CORAL/tree/main/plugin/skills/creating-a-coral-task into .gemini/skills/creating-a-coral-task/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-a-coral-task", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Human-Agent-Society/CORAL creating-a-coral-taskInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Human-Agent-Society/CORAL --skill creating-a-coral-task -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Human-Agent-Society/CORAL.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugin/skills/creating-a-coral-task .github/skills/creating-a-coral-task && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "creating-a-coral-task" agent skill from https://github.com/Human-Agent-Society/CORAL/tree/main/plugin/skills/creating-a-coral-task into .github/skills/creating-a-coral-task/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-a-coral-task", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Human-Agent-Society/CORAL --skill creating-a-coral-task -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Human-Agent-Society/CORAL creating-a-coral-task --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Human-Agent-Society/CORAL.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugin/skills/creating-a-coral-task .opencode/skills/creating-a-coral-task && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "creating-a-coral-task" agent skill from https://github.com/Human-Agent-Society/CORAL/tree/main/plugin/skills/creating-a-coral-task into .opencode/skills/creating-a-coral-task/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-a-coral-task", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
creating-a-coral-taskAuthor a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout…
Creating A Coral Task is an agent skill from Human-Agent-Society/CORAL. Author a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout float, test pass-rate, ratio-vs-baseline, multi-metric, or an LLM rubric judge). Use whenever the user wants to create a CORAL task, write or wire a grader, port a benchmark into CORAL, score open-ended outputs (reports/memos) with a judge, or debug a grader that crashes on the seed / ranks the leaderboard backwards / leaks…
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/cookbook.md`, `references/grader-api.md` and `references/rubric-judges.md`).
It sits in Testing & QA, covering QA and bug reports and Quizzes and assessments. The repository describes itself as: Open-source autoresearch powered by autonomous coding agents. Run Claude Code, OpenCode, and Codex with grading, shared knowledge, and multi-agent evolution. Accepted at COLM 2026. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 0123dfb. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
coral.compounding-intelligence.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Creating A Coral Task loads about 2.2k tokens when it runs, and up to ~8.1k if it reads all its reference files. Until then it costs about 174 tokens; SKILL.md has 895 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Human-Agent-Society/CORAL at commit 0123dfb, republished under its Apache-2.0 licence (© Human-Agent-Society). 895 words, ~2,204 tokens.
.claude/skills/creating-a-coral-task/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.A CORAL task is three things that must line up. Scaffold them with coral init, then iterate edit → coral validate until the grader scores the seed.
my-task/
├── task.yaml # config: name, description, grader entrypoint, agent count
├── seed/ # starter code agents see at t=0 (this is workspace.repo_path)
│ └── solution.py
└── grader/ # standalone Python package — gets its own isolated venv
├── pyproject.toml
└── src/my_task_grader/
├── __init__.py
└── grader.py # class Grader(TaskGrader): ...The packaged grader is the only supported form — it gives the grader an isolated venv and bundles everything the eval needs (grader code, helpers, hidden answer keys). There is no eval/grader.py auto-discovery anymore.
Optimizing code the user already has? Scaffold inside a
.coral_workspace/at the root of their project (gitignored), and copy the code to optimize intoseed/— keeps CORAL's task/results out of their source tree. Thecoral-quickstartskill has the end-to-end.coral_workspace/flow; this skill covers the grader you'll write once the code is inseed/."Optimize this" is a build instruction, not a question — never answer it with a process menu. A 1/2/3 like "point me to a task / create one / optimize outside coral" is the failure mode; do not produce it. The absence of a
task.yamlis not ambiguity — it just means you build one from the current repo. Concretely: (1) dig for what's already measurable — a research/framework repo almost always ships an eval/benchmark script, a test suite, or a metric in its README/paper; that's your target and metric. (2) If no single number is obvious, construct one by wrapping the repo's existing evaluation — don't conclude "no measurable objective" just because there's no CORAL scaffold. (3) Scaffold the most plausible target and start building (a.coral_workspace/+ draft grader is cheap and reversible); state your assumption in one line and proceed. (4) Only as a last resort, if you've actually read the repo and it exposes nothing scorable, propose 2-3 concrete optimization targets you found (each with its metric), pick the most likely, and scaffold that — still not a process menu.
coral init my-task # scaffold all three pieces (a runnable end-to-end example)
cd my-task
# ... edit the three pieces for your problem ...
coral validate . # bootstraps the grader venv, runs the grader on seed/, prints a score
# repeat edit → validate until the seed scores as you expectcoral validate succeeding is the one checkpoint that matters — it proves the grader can score the seed. Most "agents are stuck, every eval fails" reports trace to a grader that crashes on the seed, which validate would have caught. Always start from coral init rather than hand-writing the layout; the generated files are the canonical minimal example.
1. The seed (seed/) — what the agent checks out at t=0 and what the grader later scores. The contract between seed and grader is the program file: a file (e.g. solution.py) with a function or stdout convention the grader invokes, named in grader.args.program_file. Put a real, runnable baseline here — agents should coral eval immediately and get a non-zero score to beat. A skeleton that crashes is a bad baseline. Runtime data goes under seed/data/ and is read by relative path.
2. The grader (grader/) — subclass TaskGrader, implement evaluate(), return a number (or ScoreBundle). The minimum:
from coral.grader import TaskGrader
class Grader(TaskGrader):
def evaluate(self) -> float:
result = self.run_program(self.args.get("program_file", "solution.py"))
if result.returncode != 0:
return self.fail(f"crashed: {result.stderr[:200]}")
try:
return float(result.stdout.strip())
except ValueError:
return self.fail(f"expected a float, got {result.stdout[:80]!r}")This stdout-float shape is one of several. Pick the pattern that matches how your task scores → references/cookbook.md:
| Score by... | Pattern |
|---|---|
| A number the program prints | stdout float |
| Fraction of hidden tests passing | test pass-rate |
| Improvement over a baseline | ratio vs baseline |
| Several weighted criteria | multi-metric ScoreBundle |
| An LLM judging a report/memo/doc | rubric judge → references/rubric-judges.md |
Full TaskGrader surface — every attribute (self.codebase_path, self.private_dir, self.args, self.eval_logs_dir, self.tune) and method (run_program, run_script, run_script_json, score, fail, bundle) — is in references/grader-api.md.
3. The task.yaml — wiring. The fields that must be right are grader.entrypoint, grader.direction, and workspace.repo_path: ./seed. Full annotated schema (agents, islands, sharing, gateway, all defaults) → references/task-yaml.md.
The single rule: everything inside the grader/ package is visible to agents — the whole grader source is surfaced read-only at <shared_dir>/grader/ so they can read how they're scored — so secrets go in grader.private, in a dir outside grader/. CORAL copies those paths into .coral/private/ (which every runtime is denied read access to) and the grader reads them via self.private_dir. Declare a sibling, conventionally taskdata (resolving to <task_dir>/taskdata); a grader.private path inside grader/ would be copied to .coral/private/ and leaked via the surfaced source, so coral validate errors on it. Non-secret bundled data (lookup tables, helper modules) may sit inside grader/ and be read via Path(__file__).parent / ... — it's visible, so never put a secret there. Never put an answer key under seed/ either — agents read seed/ and will game the score.
coral start -c task.yaml agents.count=1 run.session=local # one agent, foreground
# watch for one real eval, confirm the score moves, then:
coral stopOnce one agent evals cleanly, raise agents.count. Driving the run from here is the running-coral-experiments skill.
| Mistake | Symptom | Fix |
|---|---|---|
repo_path points at the task root, not ./seed | Grader sees task.yaml/grader/ in codebase_path | Point repo_path at ./seed. |
direction backwards | Leaderboard ordered upside down | "ratio, higher better" → maximize; "raw error/latency" → minimize. |
Answer key under seed/ or anywhere inside grader/ | Agents read it and game the score — seed/ is their repo and the whole grader/ source is surfaced at <shared_dir>/grader/ | Put it under grader.private outside grader/ (sibling taskdata/), read via self.private_dir; coral validate errors on a private path inside grader/. |
Grader writes under self.codebase_path and re-reads it | Files vanish — daemon force-removes the worktree after each eval | Write under self.eval_logs_dir. |
Grader uses sys.executable | Misses task deps from workspace.setup | Use self.get_python_command() / self.run_program / self.run_script. |
Runtime deps in grader.setup | Validate passes, the run fails every eval | Runtime deps → workspace.setup; grader-only deps → grader.setup. |
| Scoring speed without a correctness gate | Agents "optimize" by returning garbage fast | Gate on correctness first, then score the metric. |
parallel.max_workers > 1 with an unsafe grader | Sporadic port/GPU/scratch collisions | Leave at 1 unless provably concurrency-safe. |
Skipping coral validate | Agents start, fail every eval identically | Always validate first. |
When in doubt, run coral init throwaway and read the generated files. Full config schema: https://coral.compounding-intelligence.ai/docs/api/config
© Human-Agent-Society, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in plugin/skills/creating-a-coral-task of Human-Agent-Society/CORAL.
Open the folder on GitHubat commit 0123dfb
Creating A Coral Task next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Creating A Coral Task this skillHuman-Agent-Society/CORAL | 1.1k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Evaluation Frameworkathola/claude-night-market | 341 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Reproduce Chat Statesdifferent-ai/openwork | 24k | — | ~673 | Automated safety check: Pass | Custom licence | |
| Dynamo Jira TicketDynamoDS/Dynamo | 2k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Moav E2EMotherofallVPNs/MoaV | 449 | — | ~1.9k | Automated safety check: Notes | MIT | |
| Launch Rlmarin-community/marin | 3.9k | — | ~894 | Automated safety check: Pass | Apache-2.0 |
athola/claude-night-market
Provides weighted scoring, rubrics, and decision-threshold patterns.
different-ai/openwork
Fires known chat states in the running OpenWork desktop app, such as provider errors, retries and tool steps, so you can check how each renders.
DynamoDS/Dynamo
Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.
MotherofallVPNs/MoaV
Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.
marin-community/marin
Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.
hardisgroupcom/sfdx-hardis
Walks the sfdx-hardis training course end to end as a learner would, against a real Developer Edition org and fork, fixing broken steps and screenshots that no longer match.
Human-Agent-Society/CORAL
Run and manage CORAL experiments from the operator side — launch agents with coral start (dotlist overrides, model/count, tmux vs local), monitor with coral status / coral log / coral show / the web…
Human-Agent-Society/CORAL
Verify and debug changes to CORAL itself — smallest reproduce loop per area (grader / daemon / CLI / hooks / manager / workspace / hub / template / config / web), where to look when something breaks…
Human-Agent-Society/CORAL
End-to-end recipe for adding a new task under examples/ — the three pieces that have to line up (task.yaml, seed/, and grader/), what to put in each, the TaskGrader API surface, the coral validate →…
Human-Agent-Society/CORAL
A skill your agent uses when preparing, reviewing, resolving conflicts for, or merging a CORAL release pull request from the long-lived dev branch into main.
Human-Agent-Society/CORAL
One-time machine setup after installing the coral CLI — register local agent runtimes as named bindings with coral setup / coral setup agent, validate them with coral agents doctor (incl.
Human-Agent-Society/CORAL
Add a new component to the CORAL framework itself — a new agent runtime under coral/agent/builtin/ (claudecode/codex/cursoragent style), a new CLI command in coral/cli/, a new bundled skill or…
Categories
Author a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout…. Creating A Coral Task is an agent skill from Human-Agent-Society/CORAL.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout float, test pass-rate, ratio-vs-baseline, multi-metric, or an LLM rubric judge).
Creating A Coral Task fits situations like: the user wants to create a CORAL task; port a benchmark into CORAL; score open-ended outputs (reports/memos) with a judge; debug a grader that crashes on the seed / ranks the leaderboard backwards / leaks the answer key.
Run `npx skills add Human-Agent-Society/CORAL --skill creating-a-coral-task -a claude-code`. Or copy the skill folder (plugin/skills/creating-a-coral-task in Human-Agent-Society/CORAL) into .claude/skills/creating-a-coral-task in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Human-Agent-Society/CORAL --skill creating-a-coral-task -a codex`. Or copy the skill folder (plugin/skills/creating-a-coral-task in Human-Agent-Society/CORAL) into .agents/skills/creating-a-coral-task in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Human-Agent-Society/CORAL --skill creating-a-coral-task -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/creating-a-coral-task, .gemini/skills/creating-a-coral-task, .github/skills/creating-a-coral-task and .opencode/skills/creating-a-coral-task in your project.
SKILL.md names no scripts, command-line tools or credentials: Creating A Coral Task is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: coral.compounding-intelligence.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Creating A Coral Task is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Creating A Coral Task: Evaluation Framework (athola/claude-night-market, 341 stars), Reproduce Chat States (different-ai/openwork, 24k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars) and Moav E2E (MotherofallVPNs/MoaV, 449 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Human-Agent-Society (a GitHub organization) maintains it in Human-Agent-Society/CORAL, which has 1,060 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 8, 2026.
Source: Human-Agent-Society/CORAL on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.