Orca CLI
stablyai/orca
Operate Orca-managed worktrees, folder contexts, terminals, repos, automations, artifacts, skill sharing, worktree comments, and Orca's embedded browser…
Generates several independent candidate solutions in parallel worktrees, judges them once against one explicit rubric, and applies the winning candidate only after it passes verification.
$ npx skills add codewhale-hq/Codewhale --skill best-of-n -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install codewhale-hq/Codewhale best-of-n --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/codewhale-hq/Codewhale.git skills-src && mkdir -p .claude/skills && cp -r skills-src/crates/tui/assets/skills/best-of-n .claude/skills/best-of-n && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "best-of-n" agent skill from https://github.com/codewhale-hq/Codewhale/tree/main/crates/tui/assets/skills/best-of-n into .claude/skills/best-of-n/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "best-of-n", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/codewhale-hq/Codewhale/tree/main/crates/tui/assets/skills/best-of-nType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add codewhale-hq/Codewhale --skill best-of-n -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install codewhale-hq/Codewhale best-of-n --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/codewhale-hq/Codewhale.git skills-src && mkdir -p .agents/skills && cp -r skills-src/crates/tui/assets/skills/best-of-n .agents/skills/best-of-n && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "best-of-n" agent skill from https://github.com/codewhale-hq/Codewhale/tree/main/crates/tui/assets/skills/best-of-n into .agents/skills/best-of-n/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "best-of-n", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add codewhale-hq/Codewhale --skill best-of-n -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install codewhale-hq/Codewhale best-of-n --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/codewhale-hq/Codewhale.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/crates/tui/assets/skills/best-of-n .cursor/skills/best-of-n && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "best-of-n" agent skill from https://github.com/codewhale-hq/Codewhale/tree/main/crates/tui/assets/skills/best-of-n into .cursor/skills/best-of-n/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "best-of-n", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/codewhale-hq/Codewhale.git --path crates/tui/assets/skills/best-of-n--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add codewhale-hq/Codewhale --skill best-of-n -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install codewhale-hq/Codewhale best-of-n --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/codewhale-hq/Codewhale.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/crates/tui/assets/skills/best-of-n .gemini/skills/best-of-n && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "best-of-n" agent skill from https://github.com/codewhale-hq/Codewhale/tree/main/crates/tui/assets/skills/best-of-n into .gemini/skills/best-of-n/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "best-of-n", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install codewhale-hq/Codewhale best-of-nInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add codewhale-hq/Codewhale --skill best-of-n -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/codewhale-hq/Codewhale.git skills-src && mkdir -p .github/skills && cp -r skills-src/crates/tui/assets/skills/best-of-n .github/skills/best-of-n && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "best-of-n" agent skill from https://github.com/codewhale-hq/Codewhale/tree/main/crates/tui/assets/skills/best-of-n into .github/skills/best-of-n/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "best-of-n", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add codewhale-hq/Codewhale --skill best-of-n -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install codewhale-hq/Codewhale best-of-n --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/codewhale-hq/Codewhale.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/crates/tui/assets/skills/best-of-n .opencode/skills/best-of-n && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "best-of-n" agent skill from https://github.com/codewhale-hq/Codewhale/tree/main/crates/tui/assets/skills/best-of-n into .opencode/skills/best-of-n/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "best-of-n", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
best-of-nGenerates several independent candidate solutions in parallel worktrees, judges them once against one explicit rubric, and applies the winning candidate only after it passes verification.
This skill is the preferred ensemble pattern for a high-stakes or ambiguous task that has several plausible approaches, such as a consequential design, implementation, explanation or debugging problem, and it explicitly avoids tiny changes or cases where the user already picked an approach. Before any candidate starts, it fixes one task description, one evidence packet and one explicit scoring rubric covering correctness, fit to the request, simplicity, risk and verification, and a candidate count between 2 and 4, defaulting to 3, with an explicit search mode scaling up to 16 live candidates under a worker concurrency gate.
Every candidate gets the identical task and rubric, distinguished only by a candidate number, and is started as a parallel background agent worker so the parent session stays free; candidates that implement code each get their own git worktree and a write authority bounded to the same file paths, and parallel writers are never run in the parent checkout. No candidate sees another's answer before generation finishes, and each builder returns a structured contract: candidate id, hypothesis, paths touched, commands run, a self-verdict, risks and artifact references, with the self-verdict treated as evidence to check rather than a final answer.
A single read-only reviewer, or the parent session for a small result, then judges every candidate against the original rubric by citing evidence from each one rather than voting on style, and only the winning candidate's change is applied after it passes verification.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 0ea319a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are json).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Best-of-N Candidate Tournament loads about 1.2k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 601 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from codewhale-hq/Codewhale at commit 0ea319a, republished under its MIT licence (© codewhale-hq). 601 words, ~1,177 tokens.
.claude/skills/best-of-n/SKILL.md (or your agent's skills folder).Use this skill when a consequential design, implementation, explanation, or debugging task has several plausible solutions and comparison is worth the extra model work. In Operate mode this is the preferred ensemble pattern for high-stakes or ambiguous approaches. Do not use it for a tiny change or when the user has already chosen the approach.
N from 2 to 4 for a quick comparison (default 3). For an explicit
experimental search, use the Workflow search option: 2–16 live candidates,
with larger validated populations queued at the Workflow host's 16-worker
concurrency gate rather than launched at once.create_goal or active /goal) when the tournament
spans more than one parent turn.Start the candidates as parallel background agent workers and return agent_ids
immediately so the parent stays free. For proposals, reviews, or research, keep
them read-only:
{
"action": "start",
"name": "candidate_1",
"prompt": "Produce candidate 1 for the task below. Return the proposal, evidence, risks, and rubric self-score. Do not edit files.\n\n<TASK AND RUBRIC>",
"type": "worker",
"model_strength": "same",
"write_authority": "read_only"
}Launch the remaining candidates with the same contract, then use agent wait
or completion events to collect every result. Do not show one candidate another
candidate's answer before generation finishes.
When candidates must implement code, give each one:
type: "builder"worktree: truewrite_authority: "worktree_write"write_roots or exact_filesNever run parallel writers in the parent checkout. Each builder must return the structured candidate contract (candidate id, hypothesis, paths, commands, self-verdict, risks, and artifact references). A self-verdict is evidence to inspect, not a hard-gate result.
Optional diversity: pin different model / Fleet fleet_profile values when
the project has multiple capable routes; otherwise keep model strength same.
Use one read-only reviewer worker, or the parent when the result is small, to score all candidates against the original rubric. The judge must:
Do not ask candidates to vote for themselves. Do not silently merge incompatible approaches into a new unreviewed solution.
For proposal-only work, return the winning answer with a compact score summary. For code work:
NONE is valid when every candidate fails.The checked-in operate_best_of_n.workflow.js recipe supports
strategy: "search" for structured 2–16 candidate generation and review. It
does not yet turn prompt-listed commands into hidden runtime gates. Do not
advertise those gates until a runtime evaluator host consumes a frozen search
spec.
Stop early when one candidate reveals a hard constraint that invalidates the tournament. Report the negative result rather than spending the remaining budget to manufacture variety.
© codewhale-hq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in crates/tui/assets/skills/best-of-n of codewhale-hq/Codewhale.
Open the folder on GitHubat commit 0ea319a
Best-of-N Candidate Tournament next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Best-of-N Candidate Tournament this skillcodewhale-hq/Codewhale | 41k | — | ~1.2k | Automated safety check: Pass | MIT | |
| Orca CLIstablyai/orca | 89k | 2 repos | ~593 | Automated safety check: Pass | MIT | |
| Agent of Empires Session Manageragent-of-empires/agent-of-empires | 3.3k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Agent Manager Fleet TUIYoanWai/agent-manager | 582 | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Agtx Task Sweepfynnfluegge/agtx | 1.7k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| PRP Workstream OrchestratorWirasm/prp | 2.3k | — | ~3.5k | Automated safety check: Pass | MIT |
stablyai/orca
Operate Orca-managed worktrees, folder contexts, terminals, repos, automations, artifacts, skill sharing, worktree comments, and Orca's embedded browser…
agent-of-empires/agent-of-empires
Starts, monitors and organizes coding agent sessions that run in tmux through the aoe command, including groups, profiles and worktree-based parallel work.
YoanWai/agent-manager
Runs several coding-agent CLIs as real tmux sessions in one terminal UI, color-coded by whether each is working, waiting, idle or blocked.
fynnfluegge/agtx
Breaks a conversation's results into feature-level tasks and pushes them to the agtx kanban board, where each task gets its own worktree and agent session.
Wirasm/prp
Coordinates several PRP workstreams in isolated Git worktrees from one session, verifying proof, holding merge gates and sequencing the merges.
microsoft/waza
Shared collaboration rules for a team of squad agents covering worktree awareness, writing decisions to an inbox, cross-agent requests and reviewer lockout.
codewhale-hq/Codewhale
Proves a Codewhale change in the real product: a stamped release build, an atomic local install, fresh-shell verification and manual QA that automated gates cannot cover.
codewhale-hq/Codewhale
Writes a paste-ready handoff for the next agent session, opening with a state-check command block and separating done, suspected and blocked work.
codewhale-hq/Codewhale
Decides how verified work should reach main, directly, in a worktree or on an integration branch, while keeping contributor credit and respecting merge gates.
codewhale-hq/Codewhale
Guides when and how to split multi-step coding, research or verification work into focused sub-agent runs while the parent keeps integration and final checks.
codewhale-hq/Codewhale
Triages and manages Codewhale fleet runs and workers with typed commands, classifying failures and choosing a safe restart, resume or escalation.
codewhale-hq/Codewhale
Moves a list of GitHub issues into a milestone or assigns them to owners with the gh CLI, checking each one before and after the change.
Categories
Generates several independent candidate solutions in parallel worktrees, judges them once against one explicit rubric, and applies the winning candidate only after it passes verification. This skill is the preferred ensemble pattern for a high-stakes or ambiguous task that has several plausible approaches, such as a consequential design, implementation, explanation or debugging problem, and it explicitly avoids tiny changes or cases where the user already picked an approach. Before any candidate starts, it fixes one task description, one evidence packet and one explicit scoring rubric covering correctness, fit to the request, simplicity, risk and verification, and a candidate count between 2 and 4, defaulting to 3, with an explicit search mode scaling up to 16 live candidates under a worker concurrency gate.
Best-of-N Candidate Tournament fits situations like: comparing several plausible implementations of a consequential feature; resolving an ambiguous design or debugging approach by trying multiple solutions; running a wider experimental search across many candidate approaches; judging competing proposals against one fixed, explicit rubric.
Run `npx skills add codewhale-hq/Codewhale --skill best-of-n -a claude-code`. Or copy the skill folder (crates/tui/assets/skills/best-of-n in codewhale-hq/Codewhale) into .claude/skills/best-of-n in your project. Claude Code loads it when a task matches its description.
Run `npx skills add codewhale-hq/Codewhale --skill best-of-n -a codex`. Or copy the skill folder (crates/tui/assets/skills/best-of-n in codewhale-hq/Codewhale) into .agents/skills/best-of-n in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add codewhale-hq/Codewhale --skill best-of-n -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/best-of-n, .gemini/skills/best-of-n, .github/skills/best-of-n and .opencode/skills/best-of-n in your project.
SKILL.md names no scripts, command-line tools or credentials: Best-of-N Candidate Tournament is instructions for the agent only. Our summary lists: Git worktree support for isolating each candidate's changes; A background multi-agent runner able to start, wait on and collect parallel workers.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Best-of-N Candidate Tournament is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Best-of-N Candidate Tournament: Orca CLI (stablyai/orca, 89k stars), Agent of Empires Session Manager (agent-of-empires/agent-of-empires, 3.3k stars), Agent Manager Fleet TUI (YoanWai/agent-manager, 582 stars) and Agtx Task Sweep (fynnfluegge/agtx, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
codewhale-hq (a GitHub organization) maintains it in codewhale-hq/Codewhale, which has 41,082 GitHub stars. The repository holds 63 skills in this directory. The repository was last updated on October 11, 2026.
Source: codewhale-hq/Codewhale on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.