Prompt Improver
severity1/claude-code-prompt-improver
This skill enriches vague prompts with targeted research and clarification before execution.
Fully automated self-improving loop — takes a project prompt, designs a team, runs the benchmark, analyzes results, optimizes framework code, rebuilds, and repeats until target grade is reached.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add 0x0funky/vibehq-hub --skill benchmark-loop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install 0x0funky/vibehq-hub benchmark-loop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/0x0funky/vibehq-hub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/benchmark-loop .claude/skills/benchmark-loop && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmark-loop" agent skill from https://github.com/0x0funky/vibehq-hub/tree/main/.claude/skills/benchmark-loop into .claude/skills/benchmark-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-loop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/0x0funky/vibehq-hub/tree/main/.claude/skills/benchmark-loopType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add 0x0funky/vibehq-hub --skill benchmark-loop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install 0x0funky/vibehq-hub benchmark-loop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/0x0funky/vibehq-hub.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/benchmark-loop .agents/skills/benchmark-loop && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmark-loop" agent skill from https://github.com/0x0funky/vibehq-hub/tree/main/.claude/skills/benchmark-loop into .agents/skills/benchmark-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-loop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add 0x0funky/vibehq-hub --skill benchmark-loop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install 0x0funky/vibehq-hub benchmark-loop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/0x0funky/vibehq-hub.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/benchmark-loop .cursor/skills/benchmark-loop && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmark-loop" agent skill from https://github.com/0x0funky/vibehq-hub/tree/main/.claude/skills/benchmark-loop into .cursor/skills/benchmark-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-loop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/0x0funky/vibehq-hub.git --path .claude/skills/benchmark-loop--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add 0x0funky/vibehq-hub --skill benchmark-loop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install 0x0funky/vibehq-hub benchmark-loop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/0x0funky/vibehq-hub.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/benchmark-loop .gemini/skills/benchmark-loop && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmark-loop" agent skill from https://github.com/0x0funky/vibehq-hub/tree/main/.claude/skills/benchmark-loop into .gemini/skills/benchmark-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-loop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install 0x0funky/vibehq-hub benchmark-loopInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add 0x0funky/vibehq-hub --skill benchmark-loop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/0x0funky/vibehq-hub.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/benchmark-loop .github/skills/benchmark-loop && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-loop" agent skill from https://github.com/0x0funky/vibehq-hub/tree/main/.claude/skills/benchmark-loop into .github/skills/benchmark-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-loop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add 0x0funky/vibehq-hub --skill benchmark-loop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install 0x0funky/vibehq-hub benchmark-loop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/0x0funky/vibehq-hub.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/benchmark-loop .opencode/skills/benchmark-loop && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-loop" agent skill from https://github.com/0x0funky/vibehq-hub/tree/main/.claude/skills/benchmark-loop into .opencode/skills/benchmark-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-loop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmark-loopFully automated self-improving loop — takes a project prompt, designs a team, runs the benchmark, analyzes results, optimizes framework code, rebuilds, and repeats until target grade is reached.
Benchmark Loop is an agent skill from 0x0funky/vibehq-hub. Fully automated self-improving loop — takes a project prompt, designs a team, runs the benchmark, analyzes results, optimizes framework code, rebuilds, and repeats until target grade is reached.
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Prompt engineering. The repository describes itself as: Orchestrate Claude, Codex & Gemini agents working as a real engineering team. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 5f2964b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
nodenpxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmark Loop loads about 5.3k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 1,641 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
Then immediately proceed to Step 2 (do NOT wait for user confirmation — this is full auto).Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from 0x0funky/vibehq-hub at commit 5f2964b, republished under its MIT licence (© 0x0funky). 1,641 words, ~5,299 tokens.
.claude/skills/benchmark-loop/SKILL.md (or your agent's skills folder).You are an autonomous benchmark runner for VibeHQ. Given a single project prompt, you design the team, run the benchmark, analyze results, optimize the framework, and repeat — fully unattended.
┌─────────────────────────────────────────────────────┐
│ 0. Parse prompt → design team → generate configs │
│ for each iteration (v1, v2, v3, ...): │
│ 1. Start hub + spawn agents │
│ 2. Wait for benchmark completion │
│ 3. Analyze results (vibehq-analyze) │
│ 4. Check stop conditions │
│ 5. Run /optimize-protocol to fix issues │
│ 6. Rebuild (npx tsup) │
│ 7. Update loop state → next iteration │
│ ─────────────────────────────────────────────────── │
└─────────────────────────────────────────────────────┘ALWAYS start here. Read ~/.vibehq/analytics/optimizations/loop-state.json if it exists.
phase is NOT "completed", resume from the saved phase. Skip to the appropriate step. The team config, dirs, and everything are already saved in loop-state."completed", this is a fresh run. Continue to Step 1.Also read ~/.vibehq/analytics/optimizations/history.jsonl for previous optimization context.
You are a professional technical recruiter and team architect. Your job is to analyze the project requirements, determine the minimum effective team composition, and assign the right specialists to the right domains. You don't blindly hire — you evaluate what the project actually needs, avoid redundant roles, and ensure every team member has a clearly independent workstream. Overstaffing wastes budget and creates coordination overhead; understaffing creates bottlenecks. Find the right balance.
Key decision framework:
Input: The user's $ARGUMENTS contains the project description and optional flags.
Parse the arguments:
-- flags is the project prompt--target <grade> — target grade (default: B)--port <number> — hub port (default: 3013)--max-iterations <number> — max iterations (default: 8)Read the project prompt and determine:
Project name (short, kebab-case, e.g., chat-app, ecommerce, blog-platform)
Team size — determine by counting distinct, independent work domains:
Core principle: 1 agent = 1 independent work domain = 1 directory. Never put 2 agents in the same directory — they will overwrite each other's files and cause conflicts.
How to count domains:
Sizing guidelines:
| Domains | Team size | When |
|---|---|---|
| 1 | 2 (PM + 1) | Single-stack project (API only, CLI tool, script) |
| 2 | 3 (PM + 2) | Typical full-stack (backend + frontend), or backend + data |
| 3 | 4 (PM + 3) | Full-stack + separate infra/data/design domain |
| 4+ | 5 max (PM + 4) | Large multi-stack project. Cap at 5 to control cost |
Anti-patterns to avoid:
backend/ — they'll conflict on shared files (types, index.ts, package.json)backend/) — shared models/types cause conflictsCost awareness: Each Opus agent costs $8-12 per benchmark run. A 3-person team ($25) vs 5-person team (~$50) — prefer smaller teams unless domains are truly independent.
Examples:
Agent names — assign human names (Emma, Sam, Alex, Jordan, Taylor, Riley, etc.)
Directory structure — each non-PM agent gets a unique subdirectory matching their domain (e.g., backend/, frontend/, data/, infra/). No two agents share a directory.
For the PM/Orchestrator, generate a system prompt that includes:
You are <Name>, the Project Manager for: <project prompt>
Project scope:
<break down the user's prompt into concrete deliverables>
Your workflow has TWO phases:
## Phase 1: Research
Before any implementation, create RESEARCH tasks for each domain that needs investigation.
Research tasks should ask team members to investigate and produce spec documents.
Examples of research tasks:
- "Research available free APIs for <domain>. Investigate endpoints, rate limits, auth requirements, response formats. Produce a spec document as a shared file: <domain>-research.md"
- "Research UX patterns and component libraries for <use case>. Produce ui-research.md"
- "Research best practices for <technical challenge>. Produce architecture-research.md"
Each research task MUST:
- Be assigned to the domain expert on the team
- Require a shared file as output (the spec/research document)
- Complete BEFORE any implementation tasks in that domain
## Phase 2: Implementation
After research tasks are done, READ the research output documents, then create implementation tasks.
Implementation tasks MUST:
- Reference the research output using `consumes` field
- Have specific acceptance criteria based on the research findings
- Require REAL integrations (real APIs, real libraries) — not mock/placeholder data
- Specify: "Mock data is only acceptable as a fallback when real API is unavailable"
## General rules:
1. Create a project brief first (publish_artifact)
2. Use depends_on to enforce: research tasks → implementation tasks
3. Use consumes to link implementation tasks to research output files
4. Track progress via list_tasks, unblock agents, ensure quality
5. When reviewing completed tasks: reject if using only mock data when real API was available
6. When all tasks are done, publish a final status report
Team:
<list each teammate with their role>
You are a COORDINATOR. Never write code. Only use MCP coordination tools.For worker agents, do NOT generate custom system prompts — the spawner's built-in role presets are sufficient. Workers automatically know how to use MCP tools and work on assigned tasks.
Write the PM's system prompt to a temp file:
/tmp/vibehq-loop-pm-prompt.mdGenerate and write the spawn config to /tmp/vibehq-loop-config.json:
{
"team": "<project-name>-benchmark",
"hubPort": <port>,
"agents": [
{
"name": "Emma",
"role": "Project Manager",
"subdir": "",
"systemPromptFile": "/tmp/vibehq-loop-pm-prompt.md"
},
{
"name": "Sam",
"role": "Product Designer",
"subdir": "design",
"systemPromptFile": null
},
{
"name": "Alex",
"role": "Backend Engineer",
"subdir": "backend",
"systemPromptFile": null
},
{
"name": "Jordan",
"role": "Frontend Engineer",
"subdir": "frontend",
"systemPromptFile": null
}
]
}CRITICAL: The team field MUST include the iteration number (e.g., <project-name>-benchmark-v1). Each iteration uses a completely fresh team name so that hub state, shared files, and MCP server names don't carry over from previous iterations. The baseTeam field stores the base name for reference.
{
"team": "<project-name>-benchmark-v1",
"baseTeam": "<project-name>-benchmark",
"projectPrompt": "<the full user prompt>",
"currentIteration": 1,
"phase": "benchmarking",
"targetGrade": "<target>",
"maxIterations": <max>,
"hubPort": <port>,
"baseDir": "D:\\<project-name>-benchmark",
"agents": [<copy from spawn config>],
"iterationDir": "D:\\<project-name>-benchmark-v1",
"history": []
}Save to ~/.vibehq/analytics/optimizations/loop-state.json.
========================================
Team designed for: <project prompt>
========================================
Team: <project-name>-benchmark
Port: <port>
Target: <grade>
Agents:
- Emma (Project Manager) → D:\<project>-benchmark-v1\
- Sam (Product Designer) → D:\<project>-benchmark-v1\design\
- Alex (Backend Engineer) → D:\<project>-benchmark-v1\backend\
- Jordan (Frontend Engineer) → D:\<project>-benchmark-v1\frontend\
Starting iteration 1...
========================================Then immediately proceed to Step 2 (do NOT wait for user confirmation — this is full auto).
For iteration N, create a brand new directory tree:
ITER_DIR="D:\<project-name>-benchmark-v<N>"
mkdir -p "$ITER_DIR"
# Create subdirectories for each agent that has a subdir
mkdir -p "$ITER_DIR/design"
mkdir -p "$ITER_DIR/backend"
mkdir -p "$ITER_DIR/frontend"Also delete the hub-state.json for the team if it exists:
~/.vibehq/teams/<team-name>/hub-state.json
node dist/bin/hub.js --port <hubPort> --team <team-name> &Run with run_in_background: true. Wait 3 seconds for startup.
Write a Node.js spawn script to /tmp/vibehq-loop-spawn.js that reads the config from loop-state and spawns all agents. The script should:
Read loop-state.json to get team config, iteration dir, hub port
For each agent, build the spawn command with these flags:
--name, --role, --team, --hub ws://localhost:<port>--skip-permissions — benchmark mode, no human approval--auto-kickstart — CRITICAL: auto-injects initial prompt after 8s so agents start working immediately--system-prompt-file (if applicable)Platform-specific terminal management:
Windows: Write a .cmd launcher file per agent:
@echo off
chcp 65001 >nul
set CLAUDECODE=
cd /d "<agent-cwd>"
vibehq-spawn --name "<name>" --role "<role>" --team "<team>" --hub "ws://localhost:<port>" --skip-permissions --auto-kickstart [--system-prompt-file "<path>"] -- claude
pauseLaunch with: wt -w new --title "<name>" cmd /k "<launcher-path>"
macOS/Linux: Use tmux to manage all agents in one session:
const sessionName = `vibehq-${team}`;
// Kill existing session if any
try { execSync(`tmux kill-session -t "${sessionName}" 2>/dev/null`); } catch {}
// First agent: create new session
execSync(`tmux new-session -d -s "${sessionName}" -n "${agent.name}" "${spawnCmd}"`);
// Subsequent agents: new window in same session
execSync(`tmux new-window -t "${sessionName}" -n "${agent.name}" "${spawnCmd}"`);
// After all agents: join windows into tiled panes
for (let w = agents.length - 1; w >= 1; w--) {
execSync(`tmux join-pane -s "${sessionName}:${w}" -t "${sessionName}:0" -h`);
}
execSync(`tmux select-layout -t "${sessionName}:0" tiled`);On macOS/Linux, set CLAUDECODE= in the spawn command (env var prefix).
Wait 3 seconds between each agent spawn
CRITICAL: The .cmd files must use Windows syntax (>nul not >/dev/null, \r\n line endings). Use Node.js fs.writeFileSync() and child_process.exec() — do NOT use bash heredocs to write .cmd files.
CRITICAL: Include set CLAUDECODE= in every launcher (Windows .cmd) or as env prefix (macOS/Linux) to clear the env var that prevents nested Claude Code sessions.
CRITICAL: Always include --auto-kickstart — without it, agents spawn but sit idle waiting for manual input.
Run the spawn script:
node /tmp/vibehq-loop-spawn.jsAfter spawning, print the tmux attach command (macOS/Linux):
tmux attach -t <sessionName> # to view agents
tmux kill-session -t <sessionName> # to stop allSet phase: "benchmarking", save loop-state.json.
Poll ~/.vibehq/teams/<team-name>/hub-state.json every 30 seconds.
Completion check logic:
"done" or "rejected" → COMPLETEWrite a poll script to /tmp/vibehq-loop-poll.js:
const fs = require('fs');
const path = require('path');
const home = process.env.USERPROFILE || process.env.HOME;
const team = process.argv[2] || 'default';
const statePath = path.join(home, '.vibehq', 'teams', team, 'hub-state.json');
if (!fs.existsSync(statePath)) {
console.log('NO_STATE');
process.exit(0);
}
const state = JSON.parse(fs.readFileSync(statePath, 'utf-8'));
const tasks = Object.values(state.tasks || {});
const total = tasks.length;
const done = tasks.filter(t => t.status === 'done' || t.status === 'rejected').length;
const agents = Object.values(state.agents || {});
console.log('Agents: ' + agents.map(a => a.name + '(' + a.status + ')').join(', '));
console.log('Tasks: ' + done + '/' + total);
for (const t of tasks) {
const icon = t.status === 'done' ? 'v' : t.status === 'in_progress' ? '>' : t.status === 'rejected' ? 'x' : '.';
console.log(' [' + icon + '] ' + t.title + ' -> ' + t.status + ' (' + (t.assignee || 'unassigned') + ')');
}
if (total > 0 && done === total) console.log('\nCOMPLETE');
else if (total === 0) console.log('\nNO_TASKS');
else console.log('\nWAITING');Use: node /tmp/vibehq-loop-poll.js <team-name>
Polling pattern: Use sleep 30 && node /tmp/vibehq-loop-poll.js <team> with a 60s timeout. Repeat until COMPLETE or 20 minutes elapsed.
Timeout: If waiting > 20 minutes, stop and proceed to analysis. Benchmark is likely stuck.
First, find the agent JSONL log files. They are in ~/.claude/projects/ under directories matching the agent working directories (path separators replaced with -). Read ~/.vibehq/teams/<team-name>/agent-logs.json to find recorded log paths.
Run the analyzer in static mode (no --with-llm):
node dist/bin/analyze.js <log1.jsonl> <log2.jsonl> ... --team <team-name> --save --run-id <project-name>-v<N>Do NOT call external LLM APIs. Instead, read the analysis outputs and hub-state directly, then produce the report card yourself:
~/.vibehq/analytics/runs/<project-name>-v<N>/run_metrics.json — durations, tokens, per-agent stats, utilization~/.vibehq/analytics/runs/<project-name>-v<N>/detected_flags.json — flag counts and details~/.vibehq/teams/<team-name>/hub-state.json — task details, team updates, artifactsfind <iterationDir> -name "*.ts" -o -name "*.tsx" | grep -v node_modules and wc -lEvaluate on 4 dimensions (each 0-100):
Grade scale: A (90+), A- (85-89), B+ (80-84), B (75-79), B- (70-74), C+ (65-69), C (60-64), D (50-59), F (<50)
Save to ~/.vibehq/analytics/runs/<project-name>-v<N>/report_card.json with this structure:
{
"overall_grade": "<grade>",
"score": <0-100>,
"analyzedBy": "claude-code-direct",
"grade_reasoning": "<summary>",
"coordination_assessment": { ... },
"output_assessment": { "total_loc": N, "total_files": N, "frontend_builds": bool, ... },
"token_assessment": { ... },
"per_agent_scores": [ { "agent_id": "...", "score": N, "strengths": [...], "issues": [...] } ],
"improvement_suggestions": [ { "priority": "P1|P2|P3", "target": "framework|orchestrator_prompt|analyzer_bug", "suggestion": "...", "expected_impact": "..." } ],
"fix_actions": [ { "priority": "P1|P2|P3", "target_file": "...", "action": "modify|fix|add", "description": "...", "detection_rule": "..." } ]
}Add this iteration to the history array and set phase: "analyzed".
========================================
Iteration <N> complete
Grade: <grade> | Score: <score>/100
Duration: <time> | Tasks: <done>/<total> | Cost: $<cost>
Parallel Efficiency: <value>% | LOC: <loc> | Files: <files>
Flags: C:<n> H:<n> M:<n> L:<n>
Top issues:
- <issue 1>
- <issue 2>
History:
v1: B+ (9m, $35, 57% eff)
→ v<N>: <grade> (<time>, $<cost>, <eff>% eff)
========================================| Condition | Action |
|---|---|
| Grade >= targetGrade | SUCCESS — target reached |
| 2 consecutive iterations with no grade improvement | PLATEAU — incremental fixes aren't working |
| Grade dropped for 2 consecutive iterations | REGRESSION — stop and alert user |
| currentIteration >= maxIterations | LIMIT — safety cap reached |
| Previous optimize produced 0 code changes | EXHAUSTED — nothing left to fix |
If stopping:
phase: "completed" in loop-stateIf continuing, proceed to Step 6.
Windows:
wmic process where "commandline like '%vibehq-spawn%'" call terminate 2>/dev/null
wmic process where "commandline like '%hub.js%--port <hubPort>%'" call terminate 2>/dev/nullmacOS/Linux:
tmux kill-session -t vibehq-<team-name> 2>/dev/null
pkill -f 'vibehq-spawn' 2>/dev/null
pkill -f 'hub.js.*<hubPort>' 2>/dev/nullSet phase: "optimizing" in loop-state.
Option A (preferred): Inline optimization
Read and follow .claude/skills/optimize-protocol/SKILL.md with run-id <project-name>-v<N>.
Option B (fallback): If context is getting large (>50% window)
Save state, tell user to run /optimize-protocol <project-name>-v<N> then /benchmark-loop to resume.
npx tsupMust succeed. Fix any build errors before continuing.
Increment currentIteration, update team to include new iteration number (e.g., <baseTeam>-v<N+1>), update iterationDir, set phase: "benchmarking", save loop-state.
Go back to Step 2.
set CLAUDECODE= in launcher .cmd files.project-benchmark-v1, project-benchmark-v2). This ensures fresh hub state, shared files, and MCP server names. Never reuse a team name across iterations — agents will see stale tasks/artifacts from previous runs.© 0x0funky, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/benchmark-loop of 0x0funky/vibehq-hub.
Open the folder on GitHubat commit 5f2964b
Benchmark Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmark Loop this skill0x0funky/vibehq-hub | 195 | — | ~5.3k | Automated safety check: Warn | MIT | |
| Prompt Improverseverity1/claude-code-prompt-improver | 1.9k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Prompt Engineering Patternsynulihao/AgentSkillOS | 617 | 14 repos | ~1.7k | Automated safety check: Pass | None | |
| Patch CreationPiebald-AI/tweakcc | 2.5k | — | ~1.6k | Automated safety check: Pass | MIT | |
| LLM Application DevMoizIbnYousaf/ai-agent-skills | 1.1k | 2 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Senior Prompt Engineermaslennikov-ig/claude-code-orchestrator-kit | 259 | 4 repos | ~1.4k | Automated safety check: Pass | Custom licence |
severity1/claude-code-prompt-improver
This skill enriches vague prompts with targeted research and clarification before execution.
ynulihao/AgentSkillOS
Master advanced prompt engineering techniques to maximize LLM performance, reliability, and controllability in production.
Piebald-AI/tweakcc
Create and register new patches for tweakcc. An agent skill from Piebald-AI/tweakcc.
MoizIbnYousaf/ai-agent-skills
Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration.
maslennikov-ig/claude-code-orchestrator-kit
Provides reference guides and Python scripts for prompt optimization, RAG evaluation, and agent orchestration when building or tuning LLM systems.
baskduf/FableCodex
Apply a Claude Fable 5 inspired operating style inside Codex.
0x0funky/vibehq-hub
Autonomous framework engineer — reads VibeHQ post-run analysis, understands root causes of multi-agent coordination failures, then designs and implements real code changes (new features, refactors…
0x0funky/vibehq-hub
Run a single team session to build a project from a prompt. An agent skill from 0x0funky/vibehq-hub.
Categories
Fully automated self-improving loop — takes a project prompt, designs a team, runs the benchmark, analyzes results, optimizes framework code, rebuilds, and repeats until target grade is reached. Benchmark Loop is an agent skill from 0x0funky/vibehq-hub. Fully automated self-improving loop — takes a project prompt, designs a team, runs the benchmark, analyzes results, optimizes framework code, rebuilds, and repeats until target grade is reached.
Benchmark Loop fits situations like: tasks that involve Prompt engineering.
Run `npx skills add 0x0funky/vibehq-hub --skill benchmark-loop -a claude-code`. Or copy the skill folder (.claude/skills/benchmark-loop in 0x0funky/vibehq-hub) into .claude/skills/benchmark-loop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add 0x0funky/vibehq-hub --skill benchmark-loop -a codex`. Or copy the skill folder (.claude/skills/benchmark-loop in 0x0funky/vibehq-hub) into .agents/skills/benchmark-loop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add 0x0funky/vibehq-hub --skill benchmark-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-loop, .gemini/skills/benchmark-loop, .github/skills/benchmark-loop and .opencode/skills/benchmark-loop in your project.
Going by SKILL.md and its folder, Benchmark Loop needs the command-line tools its instructions call (node and npx). Our summary lists: Node.js.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.
Benchmark Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Benchmark Loop: Prompt Improver (severity1/claude-code-prompt-improver, 1.9k stars), Prompt Engineering Patterns (ynulihao/AgentSkillOS, 617 stars), Patch Creation (Piebald-AI/tweakcc, 2.5k stars) and LLM Application Dev (MoizIbnYousaf/ai-agent-skills, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
0x0funky (a GitHub user) maintains it in 0x0funky/vibehq-hub, which has 195 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on March 26, 2026.
Source: 0x0funky/vibehq-hub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.