Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch.
$ npx skills add agentscope-ai/OpenJudge --skill bootstrap -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentscope-ai/OpenJudge bootstrap --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval_pipeline/08-bootstrap .claude/skills/bootstrap && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bootstrap" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/08-bootstrap into .claude/skills/bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bootstrap", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/08-bootstrapType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentscope-ai/OpenJudge --skill bootstrap -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentscope-ai/OpenJudge bootstrap --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval_pipeline/08-bootstrap .agents/skills/bootstrap && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bootstrap" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/08-bootstrap into .agents/skills/bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bootstrap", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill bootstrap -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentscope-ai/OpenJudge bootstrap --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval_pipeline/08-bootstrap .cursor/skills/bootstrap && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bootstrap" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/08-bootstrap into .cursor/skills/bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bootstrap", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentscope-ai/OpenJudge.git --path skills/eval_pipeline/08-bootstrap--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentscope-ai/OpenJudge --skill bootstrap -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentscope-ai/OpenJudge bootstrap --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval_pipeline/08-bootstrap .gemini/skills/bootstrap && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bootstrap" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/08-bootstrap into .gemini/skills/bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bootstrap", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentscope-ai/OpenJudge bootstrapInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentscope-ai/OpenJudge --skill bootstrap -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval_pipeline/08-bootstrap .github/skills/bootstrap && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bootstrap" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/08-bootstrap into .github/skills/bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bootstrap", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill bootstrap -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentscope-ai/OpenJudge bootstrap --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval_pipeline/08-bootstrap .opencode/skills/bootstrap && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bootstrap" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/08-bootstrap into .opencode/skills/bootstrap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bootstrap", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bootstrapA skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch.
Bootstrap is an agent skill from agentscope-ai/OpenJudge. Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. Also use when the user says "I need to start evaluating my app but don't know where to begin," "I want to set up eval for a new product," or has just identified failure modes and needs to turn them into principles. Outputs a v0 grader in 30 minutes using OpenJudge SimpleRubricsGenerator, plus a roadmap to reach calibrated evaluation.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
dashscope.aliyuncs.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bootstrap loads about 2k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 566 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 566 words, ~2,013 tokens.
.claude/skills/bootstrap/SKILL.md (or your agent's skills folder).<HARD-GATE>
NO v0 grader deployed WITHOUT explicitly marking it as uncalibrated.
NO synthetic labels — LLM can generate eval inputs, but labels MUST come from real system output + human judgment.
NO principle without a source label documenting where it came from.
</HARD-GATE>
Cold-start an evaluation system when you have nothing. In 30 minutes you get a working v0 grader and a clear path to a calibrated, trustworthy evaluation.
Requires OpenJudge (
pip install py-openjudge) for the grader generators (SimpleRubricsGenerator/IterativeRubricsGenerator). The interview, stratification, and calibration-roadmap methodology is SDK-independent.
You MUST create a task for each item and complete them in order:
Ask the user to describe their system in one go:
To bootstrap your evaluation, I need to understand what you're building.
Please describe (all at once):
- What does your system do? Who uses it?
- What are 3 examples of perfect outputs?
- What are 3 things the system must never do?
- What failures worry you most?Don't drip-feed these questions. One prompt, one answer. If the user provides a spec doc or design document instead, read that directly.
Use OpenJudge's SimpleRubricsGenerator to create a zero-shot grader from the
product description:
import asyncio
from openjudge.models.openai_chat_model import OpenAIChatModel
from openjudge.generator.simple_rubric.generator import (
SimpleRubricsGenerator,
SimpleRubricsGeneratorConfig,
)
from openjudge.runner.grading_runner import GradingRunner
# OpenAIChatModel reads OPENAI_API_KEY / OPENAI_BASE_URL from the environment.
# For Aliyun DashScope (Bailian): set OPENAI_BASE_URL to
# https://dashscope.aliyuncs.com/compatible-mode/v1 and OPENAI_API_KEY to your key.
model = OpenAIChatModel(model="qwen-plus") # or "gpt-4o", etc.
config = SimpleRubricsGeneratorConfig(
grader_name="Initial Quality Grader",
model=model,
task_description="<summarize from the interview>",
scenario="<usage context from interview>",
min_score=0,
max_score=1,
)
generator = SimpleRubricsGenerator(config)
grader = await generator.generate(
dataset=[],
sample_queries=[
"<example query 1 from interview>",
"<example query 2 from interview>",
"<example query 3 from interview>",
],
)Why zero-shot instead of asking the user to write criteria? At this stage, the user doesn't know what "good" means operationally. The generator produces a reasonable starting point. The user refines it after seeing v0 results.
Generate 30 test inputs with stratification. Use 3 different prompt templates for diversity:
Template 1: "Generate a typical {domain} query for a {user_type}"
Template 2: "Create an ambiguous {domain} query where intent is unclear"
Template 3: "Generate an edge-case {domain} query that's unusual but realistic"Target distribution:
Critical: Generate inputs ONLY. Never generate labels. The labels come from running the actual system and getting human judgments.
# The dataset format for GradingRunner
dataset = [
{
"query": "What's the status of my order #12345?",
"response": "<will be filled by running the system>",
},
# ... 30 inputs
]Plug the generated grader into GradingRunner:
from openjudge.runner.grading_runner import GradingRunner
from openjudge.graders.schema import GraderScore, GraderError
runner = GradingRunner(
grader_configs={"v0_quality": grader},
max_concurrency=8,
)
results = await runner.arun(dataset)
scores = [r.score for r in results["v0_quality"] if isinstance(r, GraderScore)]
errors = [r for r in results["v0_quality"] if isinstance(r, GraderError)]
print(f"V0 Results: avg={sum(scores)/len(scores):.2f}, errors={len(errors)}")The v0 grader is uncalibrated — you don't know its TPR/TNR yet. Give the user an exact path to trustworthiness:
Your v0 evaluation is ready. Here's the path to a calibrated system:
Phase 1 (now): Run the v0 grader on 30 inputs to get a baseline.
→ The grader is UNCALIBRATED. Treat scores as directional, not definitive.
Phase 2 (1-2 weeks): Collect 50 human-labeled examples (25 pass + 25 fail).
→ For each system output, have a human mark pass/fail against the criterion.
→ Store labels in labels/<grader_name>.jsonl
Phase 3: When you have 50 labels, run 03-align-human to:
→ Measure TPR/TNR of the v0 grader
→ Detect biases (position, verbosity, self-enhancement)
→ Get a human-reduction roadmap
Phase 4: When TPR >= 0.8 and TNR >= 0.8:
→ The grader is calibrated and can be used as a production gateIterativeRubricsGenerator
instead of SimpleRubricsGenerator for data-driven grader creation:from openjudge.generator.iterative_rubric.generator import (
IterativeRubricsGenerator,
IterativePointwiseRubricsGeneratorConfig,
)
config = IterativePointwiseRubricsGeneratorConfig(
grader_name="Data-Driven Grader",
model=model,
task_description="<from interview>",
min_score=0, max_score=1,
max_epochs=3,
batch_size=10,
)
generator = IterativeRubricsGenerator(config)
grader = await generator.generate(dataset=labeled_data) # 20+ labeled examplesAfter 08-bootstrap:
03-align-human: Once 50 human labels are collected, calibrate the grader.01-eval-design: If you want a properly stratified dataset beyond the v0 30 inputs.02-metric-design: If you need multiple graders for different dimensions.© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/eval_pipeline/08-bootstrap of agentscope-ai/OpenJudge.
Open the folder on GitHubat commit d1e0642
Bootstrap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bootstrap this skillagentscope-ai/OpenJudge | 868 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 6 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| Peft Fine TuningOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.1k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.8k | 15 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
agentscope-ai/OpenJudge
A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…
agentscope-ai/OpenJudge
A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.
agentscope-ai/OpenJudge
A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.
agentscope-ai/OpenJudge
Automatically evaluate and compare multiple AI models or agents without pre-existing test data.
agentscope-ai/OpenJudge
A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…
agentscope-ai/OpenJudge
Build custom LLM evaluation pipelines using the OpenJudge framework.
Categories
A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. Bootstrap is an agent skill from agentscope-ai/OpenJudge. Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch.
Bootstrap fits situations like: the user has nothing — no traces; no eval set — and needs to build a v0 evaluation from scratch; the user says I need to start evaluating my app but dont know where to begin; I want to set up eval for a new product.
Run `npx skills add agentscope-ai/OpenJudge --skill bootstrap -a claude-code`. Or copy the skill folder (skills/eval_pipeline/08-bootstrap in agentscope-ai/OpenJudge) into .claude/skills/bootstrap in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentscope-ai/OpenJudge --skill bootstrap -a codex`. Or copy the skill folder (skills/eval_pipeline/08-bootstrap in agentscope-ai/OpenJudge) into .agents/skills/bootstrap in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill bootstrap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bootstrap, .gemini/skills/bootstrap, .github/skills/bootstrap and .opencode/skills/bootstrap in your project.
Going by SKILL.md and its folder, Bootstrap needs the command-line tools its instructions call (pip) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.
SKILL.md names 1 domain. In commands or code: dashscope.aliyuncs.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bootstrap is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bootstrap: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 868 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.
Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.