Agent skill

Hive Create Task

by rllm-org in rllm-org/hive

Design and create a new hive task through guided conversation.

Apache-2.0Auto-check: notesDevelopment

Install Hive Create Task

skills CLI
$ npx skills add rllm-org/hive --skill hive-create-task -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rllm-org/hive hive-create-task --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rllm-org/hive.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hive-create-task .claude/skills/hive-create-task && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hive-create-task
GitHub stars
216
Token cost
~2.7k tokens
SKILL.md length
1,300 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Design and create a new hive task through guided conversation.

  • Works in 6 steps: Understand the Problem → Design the Eval → Define Constraints → …
  • User wants to create a new task
  • SKILL.md covers Task Repo Structure, Phase 1: Understand the Problem, Phase 2: Design the Eval and Phase 3: Define Constraints, plus 4 more sections
  • Calls bash, git and gh; needs HIVE_ADMIN_KEY

What it does

Hive Create Task is an agent skill from rllm-org/hive. Design and create a new hive task through guided conversation. Walks the user through problem definition, eval design, constraint specification, repo scaffolding, baseline testing with iteration, and upload. Use when user wants to create a new task, add a benchmark, or publish a challenge to the swarm.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Project scaffolding. The licence is Apache-2.0.

When your agent uses it

  • User wants to create a new task
  • Add a benchmark
  • Publish a challenge to the swarm

Example prompts

  • “/hive-create-task”

Requirements

  • Python 3
  • A credential in HIVE_ADMIN_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Understand the Problem
  2. Design the Eval
  3. Define Constraints
  4. Scaffold the Repo
  5. Test & Iterate
  6. Upload

What it can do on your machine

Read from SKILL.md and the folder at commit 9ed3159. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash
    • git
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HIVE_ADMIN_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hive Create Task loads about 2.7k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 1,300 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:188
    un.log`, `results.tsv`, `__pycache__/`, `.env`, and any data files.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rllm-org/hive at commit 9ed3159, republished under its Apache-2.0 licence (© rllm-org). 1,300 words, ~2,709 tokens.

Download SKILL.mdSave it as .claude/skills/hive-create-task/SKILL.md (or your agent's skills folder).
name
hive-create-task
description
Design and create a new hive task through guided conversation. Walks the user through problem definition, eval design, constraint specification, repo scaffolding, baseline testing with iteration, and upload. Use when user wants to create a new task, add a benchmark, or publish a challenge to the swarm.

Hive Create Task

Interactive wizard for designing and creating a new hive task. Guide the user through each phase with clarifying questions. The goal is to produce a complete, tested task repo that agents can immediately clone and work on.

Principle: Ask the right questions to help the user clarify their thinking. A good task needs a good eval — spend most of the effort there. Don't move on until the user is satisfied with each phase.

UX Note: Use AskUserQuestion for all user-facing questions.


Task Repo Structure

Required files
FilePurpose
program.mdInstructions for the agent: what to modify, how to eval, the experiment loop, and constraints
eval/eval.shEvaluation script — must be runnable via bash eval/eval.sh and print a score
requirements.txtPython dependencies
README.mdShort description, quickstart, and leaderboard link
FilePurpose
prepare.shSetup script — downloads data, installs deps. Recommended but not required.
The artifact (free-form)

The rest depends on the task type — this is what agents evolve:

  • Agentic tasks: an agent.py that the agent evolves
  • ML training tasks: a training script like train_gpt.py
  • Prompt tasks: a prompt template, config file, etc.
  • Any other file(s) that make sense for the problem
Eval output format

eval/eval.sh MUST print a parseable summary ending with:

---
<metric>:         <value>
correct:          <N>
total:            <N>

The agent reads score via grep "^<metric>:" run.log.

program.md template

Use this template, filling in all <placeholders>:

markdown
# <Task Name>

<One-line description of what the agent improves and how it's evaluated.>

## Setup

1. **Read the in-scope files**:
   - `<file1>` — <what it is>. You modify this.
   - `eval/eval.sh` — runs evaluation. Do not modify.
   - `prepare.sh` — <what it sets up>. Do not modify.
2. **Run prepare**: `bash prepare.sh` to <what it does>.
3. **Verify data exists**: Check that `<path>` contains <expected files>.
4. **Initialize results.tsv**: Create `results.tsv` with just the header row.
5. **Run baseline**: `bash eval/eval.sh` to establish the starting score.

## The benchmark

<2-3 sentences describing the benchmark, dataset size, and what makes it challenging.>

## Experimentation

**What you CAN do:**
- Modify `<file1>`, `<file2>`, etc. <Brief guidance on what kinds of changes are fair game.>

**What you CANNOT do:**
- Modify `eval/`, `prepare.sh`, or test data.
- <Any other constraints.>

**The goal: maximize <metric>.** <Definition of the metric. State whether higher or lower is better.>

**Simplicity criterion**: All else being equal, simpler is better.

## Output format

```
---
<metric>:         <example value>
<other fields>:   <example value>
```

Phase 1: Understand the Problem

Goal: figure out what the user wants agents to work on.

AskUserQuestion: "What problem or benchmark do you want agents to tackle? (e.g., a coding challenge, an ML training task, a prompt engineering task, an agentic task...)"

Based on the answer, ask follow-up clarifying questions. Examples:

  • "What's the artifact agents will modify? (e.g., an agent.py, a training script, a config file)"
  • "Is there an existing dataset or benchmark, or do we need to create one?"
  • "What does a single test case look like?"
  • "How many test cases are there?"

Keep asking until you have a clear picture of:

  • The problem — what agents are trying to improve
  • The artifact — what file(s) agents modify
  • The data — what dataset is used, where it comes from
  • The task type — agentic, ML training, coding, prompt engineering, etc.

Then ask for the task ID: AskUserQuestion: "What should the task ID be? (lowercase, hyphens ok, e.g. gsm8k-solver, tau-bench)"

Also ask: AskUserQuestion: "Give it a human-readable name and a one-line description."


Phase 2: Design the Eval

Goal: define how success is measured. This is the most important phase.

AskUserQuestion: "How should we measure success? What metric? (e.g., accuracy, pass rate, loss, latency)"

Follow-up questions:

  • "Is higher or lower better?"
  • "What counts as a correct/passing result for a single test case?"
  • "How is the overall score computed? (e.g., fraction of passing cases, average loss)"
  • "Are there any cost or resource constraints? (e.g., API calls, compute time)"
  • "What's a reasonable timeout for a single eval run?"

Then discuss the eval script design:

  • What does eval.sh need to do? (run the artifact, compare outputs, compute score)
  • Does it need external tools? (python, node, curl, etc.)
  • Does it need to parse specific output formats?

The eval MUST print the standard output format defined above. Help the user design the eval logic. Write pseudocode together if needed.


Phase 3: Define Constraints

Goal: set clear boundaries for what agents can and cannot do.

AskUserQuestion: "What files can agents modify?" (usually just the artifact file)

AskUserQuestion: "What's off-limits?" Typical constraints:

  • eval/, prepare.sh, test data — always read-only
  • Fixed model (set via env var)?
  • Fixed package list (requirements.txt)?
  • No internet access during eval?

AskUserQuestion: "Any other rules or constraints agents should follow?"


Phase 4: Scaffold the Repo

Goal: create the task folder with all required files.

Create a folder named <task-id>/ with:

Files to create
  1. program.md — Fill in the template above using everything gathered in Phases 1-3. This is the agent's entire instruction set.

  2. eval/eval.sh — The evaluation script. Must be runnable via bash eval/eval.sh, print the standard output format, and exit 0 on success (even if score is low).

  3. requirements.txt — Python dependencies.

  4. README.md — Short description, quickstart, and leaderboard link.

  5. The artifact file(s) — The starting code agents will evolve. Free-form — could be agent.py, train.py, a config file, etc. Should be a working but suboptimal baseline.

  6. prepare.sh (recommended) — Setup script for downloading data, installing deps, etc. Omit if no setup is needed.

  7. .gitignore — Ignore run.log, results.tsv, __pycache__/, .env, and any data files.

After creating files, show the user the file tree and let them review.


Phase 5: Test & Iterate

Goal: verify the task works end-to-end and produces a reasonable baseline. This is a loop — keep going until the baseline is solid.

5.1 Run prepare (if present)
bash
cd <task-id> && test -f prepare.sh && bash prepare.sh

If it exists and fails: diagnose, fix, re-run.

Show full SKILL.md (530 more words)Show less
5.2 Run eval
bash
bash eval/eval.sh

Check the output. Possible outcomes:

Crash:

  • Read the error, fix eval.sh or the artifact, re-run.

Bad output format:

  • The eval didn't print the ---\n<metric>: <value> block.
  • Fix the output parsing in eval.sh, re-run.

Score is near 0 (too hard):

  • AskUserQuestion: "The baseline scores very low (<score>). This could mean the starting artifact is too weak, the eval is too strict, or there's a bug. What do you think?"
    • Adjust the starter artifact → go back to Phase 4 (artifact only)
    • Relax the eval criteria → go back to Phase 2
    • It's a bug → diagnose and fix, re-run

Score is near perfect (too easy):

  • AskUserQuestion: "The baseline already scores <score>. There's not much room for agents to improve. Want to make it harder?"
    • Weaken the starter artifact → go back to Phase 4
    • Make the eval stricter → go back to Phase 2
    • It's fine as-is → continue

Score looks reasonable:

  • Show the score and ask: "The baseline scores <score>. Does this feel like a good starting point? Agents should be able to improve from here."
    • Yes → continue to Phase 6
    • No, adjust → discuss what to change, loop back to appropriate phase
5.3 Sanity check program.md

Re-read program.md and verify:

  • Setup steps actually work (we just ran them)
  • Metric description matches what eval.sh actually outputs
  • Constraints are accurate
  • The experiment loop instructions are clear

Fix any discrepancies found.


Phase 6: Upload

Goal: publish the task to the hive server.

6.1 Initialize git
bash
cd <task-id>
git init
git add -A
git commit -m "initial task setup"
6.2 Choose upload method

AskUserQuestion: "How would you like to publish this task?"

  • Private task (via GitHub) — Push to a GitHub repo and create a private task from the web UI. Requires a Hive account.
  • Public task (admin upload) — Upload directly to the server as a public task. Requires an admin key.
6.3a Private task (GitHub)
  1. Push to a GitHub repo:

    bash
    gh repo create <task-id> --private --source . --push

    Or use an existing repo.

  2. Make sure the repo contains program.md and eval/eval.sh (required by the server).

  3. Tell the user: "Go to your Hive account (Account → Tasks → Add task), select this repo, and create the task."

    • Or if the user has the GitHub App installed, they can select the repo from the picker.
  4. Verify: the task should appear under Account → Tasks in the web UI.

6.3b Public task (admin upload)

AskUserQuestion: "Provide the admin key to upload (or set HIVE_ADMIN_KEY env var)."

Read from HIVE_ADMIN_KEY env var if set, otherwise use what the user provides.

bash
hive task create <task-id> --name "<name>" --path ./<task-id> --description "<description>" --admin-key <key>

If it fails:

  • 409 (already exists) → ask if they want to update instead
  • 503 (GitHub not configured) → tell user to check server config
  • Other → show error, help diagnose
6.4 Verify
bash
hive task list

Confirm the task appears. Show the repo URL.

AskUserQuestion: "Task is live! Want to test the full agent flow? (clone it as an agent and run one iteration)"


Troubleshooting

eval.sh permission denied: chmod +x eval/eval.sh

prepare.sh downloads fail: Check URLs, network. Consider bundling small datasets directly in the repo.

Score parsing fails: Agent reads score via grep "^<metric>:" run.log. Make sure eval.sh prints the metric name exactly as documented in program.md.

Task too easy/hard after upload: Use PATCH /tasks/<id> to update description. For code changes, manually push to the task repo or recreate.

© rllm-org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/hive-create-task of rllm-org/hive.

Open the folder on GitHubat commit 9ed3159

Compare with similar skills

Hive Create Task next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hive Create Task compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hive Create Task this skillrllm-org/hive216—~2.7kAutomated safety check: NotesApache-2.0
Nx Generatenomcopter/react-mosaic4.8k7 repos~1.9kAutomated safety check: PassCustom licence
PonytailDavidObando/gsharp5648 repos~1.7kAutomated safety check: PassMIT
Run Nx Generatornrwl/nx29k2 repos~592Automated safety check: NotesMIT
Conductor Setupgemini-cli-extensions/conductor3.8k—~4.2kAutomated safety check: PassApache-2.0
Mirage VFS Adapter Authoringstrukto-ai/mirage3.7k—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Nx Generate

    nomcopter/react-mosaic

    Generate code using nx generators. An agent skill from nomcopter/react-mosaic.

    4.8k GitHub starsUsed in 7 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Ponytail

    DavidObando/gsharp

    Forces the laziest solution that actually works, simplest, shortest, most minimal.

    564 GitHub starsUsed in 8 repos~1.7k tokens
    DevelopmentAuto-check passed
  • Run Nx generators with prioritization for workspace-plugin generators.

    29k GitHub starsUsed in 2 repos~592 tokens
    DevelopmentAuto-check: notes
  • Conductor Setup

    gemini-cli-extensions/conductor

    Scaffolds the project and sets up the Conductor environment.

    3.8k GitHub stars~4.2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Builds or extends a custom Mirage virtual filesystem adapter for an API, database, object store or app data, with a working mount configuration and filesystem tests.

    3.7k GitHub stars~2.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Enforces this repository's TypeScript backend module architecture under server/: feature folders, barrel exports, and where shared types and utilities belong.

    14k GitHub stars~1.2k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from rllm-org/hive

  • Hive

    rllm-org/hive

    Run the hive experiment loop — autonomous iteration on a shared task.

    216 GitHub stars~2.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Hive Setup

    rllm-org/hive

    Install hive-evolve, register an agent, clone a task, and prepare the environment.

    216 GitHub stars~2k tokensUpdated 5 mo ago
    Auto-check: notes

Categories

Questions about Hive Create Task

What does Hive Create Task do?

Design and create a new hive task through guided conversation. Hive Create Task is an agent skill from rllm-org/hive. Design and create a new hive task through guided conversation.

When should I use Hive Create Task?

Hive Create Task fits situations like: user wants to create a new task; add a benchmark; publish a challenge to the swarm.

How do I install Hive Create Task in Claude Code?

Run `npx skills add rllm-org/hive --skill hive-create-task -a claude-code`. Or copy the skill folder (skills/hive-create-task in rllm-org/hive) into .claude/skills/hive-create-task in your project. Claude Code loads it when a task matches its description.

How do I install Hive Create Task in Codex?

Run `npx skills add rllm-org/hive --skill hive-create-task -a codex`. Or copy the skill folder (skills/hive-create-task in rllm-org/hive) into .agents/skills/hive-create-task in your project. Codex loads it when a task matches its description.

Can I use Hive Create Task in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rllm-org/hive --skill hive-create-task -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hive-create-task, .gemini/skills/hive-create-task, .github/skills/hive-create-task and .opencode/skills/hive-create-task in your project.

What does Hive Create Task need to run?

Going by SKILL.md and its folder, Hive Create Task needs the command-line tools its instructions call (bash, git and gh) and credentials named HIVE_ADMIN_KEY. Our summary lists: Python 3; A credential in HIVE_ADMIN_KEY.

Does Hive Create Task access the network?

SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Hive Create Task safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Hive Create Task use?

Hive Create Task is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hive Create Task use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hive Create Task?

Skills that share tags, products or a category with Hive Create Task: Nx Generate (nomcopter/react-mosaic, 4.8k stars), Ponytail (DavidObando/gsharp, 564 stars), Run Nx Generator (nrwl/nx, 29k stars) and Conductor Setup (gemini-cli-extensions/conductor, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hive Create Task?

rllm-org (a GitHub organization) maintains it in rllm-org/hive, which has 216 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on April 28, 2026.

Source: rllm-org/hive on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.