Agent skill

Bootstrap

by agentscope-ai in agentscope-ai/OpenJudge

A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Bootstrap

skills CLI
$ npx skills add agentscope-ai/OpenJudge --skill bootstrap -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/OpenJudge bootstrap --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval_pipeline/08-bootstrap .claude/skills/bootstrap && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bootstrap
GitHub stars
868
Token cost
~2k tokens
SKILL.md length
566 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch.

  • Works in 5 steps: Product Interview (One Shot) → Generate v0 Grader → Synthesize Eval Inputs → …
  • The user has nothing — no traces
  • SKILL.md covers Checklist, Step 1: Product Interview (One…, Step 2: Generate v0 Grader and Step 3: Synthesize Eval Inputs, plus 6 more sections
  • Calls pip; reaches dashscope.aliyuncs.com; needs OPENAI_API_KEY

What it does

Bootstrap is an agent skill from agentscope-ai/OpenJudge. Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. Also use when the user says "I need to start evaluating my app but don't know where to begin," "I want to set up eval for a new product," or has just identified failure modes and needs to turn them into principles. Outputs a v0 grader in 30 minutes using OpenJudge SimpleRubricsGenerator, plus a roadmap to reach calibrated evaluation.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.

When your agent uses it

  • The user has nothing — no traces
  • No eval set — and needs to build a v0 evaluation from scratch
  • The user says I need to start evaluating my app but dont know where to begin
  • I want to set up eval for a new product

Example prompts

  • “I need to start evaluating my app but don”
  • “I want to set up eval for a new product,”
  • “/bootstrap”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Product Interview (One Shot)
  2. Generate v0 Grader
  3. Synthesize Eval Inputs
  4. Run v0 Evaluation
  5. Output Roadmap

What it can do on your machine

Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • dashscope.aliyuncs.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bootstrap loads about 2k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 566 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 566 words, ~2,013 tokens.

Download SKILL.mdSave it as .claude/skills/bootstrap/SKILL.md (or your agent's skills folder).
name
bootstrap
description
Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. Also use when the user says "I need to start evaluating my app but don't know where to begin," "I want to set up eval for a new product," or has just identified failure modes and needs to turn them into principles. Outputs a v0 grader in 30 minutes using OpenJudge SimpleRubricsGenerator, plus a roadmap to reach calibrated evaluation.
<HARD-GATE>
NO v0 grader deployed WITHOUT explicitly marking it as uncalibrated.
NO synthetic labels — LLM can generate eval inputs, but labels MUST come from real system output + human judgment.
NO principle without a source label documenting where it came from.
</HARD-GATE>

Bootstrap

Cold-start an evaluation system when you have nothing. In 30 minutes you get a working v0 grader and a clear path to a calibrated, trustworthy evaluation.

Requires OpenJudge (pip install py-openjudge) for the grader generators (SimpleRubricsGenerator / IterativeRubricsGenerator). The interview, stratification, and calibration-roadmap methodology is SDK-independent.

Checklist

You MUST create a task for each item and complete them in order:

  1. Understand the product — one-shot interview, not question-by-question
  2. Generate v0 grader — use OpenJudge SimpleRubricsGenerator
  3. Synthesize eval inputs — 30 inputs with 60/30/10 stratification
  4. Run v0 evaluation — GradingRunner with the generated grader
  5. Output roadmap — exactly how to reach 50 labels → calibrate

Step 1: Product Interview (One Shot)

Ask the user to describe their system in one go:

To bootstrap your evaluation, I need to understand what you're building.
Please describe (all at once):

- What does your system do? Who uses it?
- What are 3 examples of perfect outputs?
- What are 3 things the system must never do?
- What failures worry you most?

Don't drip-feed these questions. One prompt, one answer. If the user provides a spec doc or design document instead, read that directly.

Step 2: Generate v0 Grader

Use OpenJudge's SimpleRubricsGenerator to create a zero-shot grader from the product description:

python
import asyncio
from openjudge.models.openai_chat_model import OpenAIChatModel
from openjudge.generator.simple_rubric.generator import (
    SimpleRubricsGenerator,
    SimpleRubricsGeneratorConfig,
)
from openjudge.runner.grading_runner import GradingRunner

# OpenAIChatModel reads OPENAI_API_KEY / OPENAI_BASE_URL from the environment.
# For Aliyun DashScope (Bailian): set OPENAI_BASE_URL to
# https://dashscope.aliyuncs.com/compatible-mode/v1 and OPENAI_API_KEY to your key.
model = OpenAIChatModel(model="qwen-plus")  # or "gpt-4o", etc.

config = SimpleRubricsGeneratorConfig(
    grader_name="Initial Quality Grader",
    model=model,
    task_description="<summarize from the interview>",
    scenario="<usage context from interview>",
    min_score=0,
    max_score=1,
)

generator = SimpleRubricsGenerator(config)
grader = await generator.generate(
    dataset=[],
    sample_queries=[
        "<example query 1 from interview>",
        "<example query 2 from interview>",
        "<example query 3 from interview>",
    ],
)

Why zero-shot instead of asking the user to write criteria? At this stage, the user doesn't know what "good" means operationally. The generator produces a reasonable starting point. The user refines it after seeing v0 results.

Step 3: Synthesize Eval Inputs

Generate 30 test inputs with stratification. Use 3 different prompt templates for diversity:

Template 1: "Generate a typical {domain} query for a {user_type}"
Template 2: "Create an ambiguous {domain} query where intent is unclear"
Template 3: "Generate an edge-case {domain} query that's unusual but realistic"

Target distribution:

  • 60% common/typical queries (18 inputs)
  • 30% boundary/ambiguous queries (9 inputs)
  • 10% edge-case/unusual queries (3 inputs)

Critical: Generate inputs ONLY. Never generate labels. The labels come from running the actual system and getting human judgments.

python
# The dataset format for GradingRunner
dataset = [
    {
        "query": "What's the status of my order #12345?",
        "response": "<will be filled by running the system>",
    },
    # ... 30 inputs
]

Step 4: Run v0 Evaluation

Plug the generated grader into GradingRunner:

python
from openjudge.runner.grading_runner import GradingRunner
from openjudge.graders.schema import GraderScore, GraderError

runner = GradingRunner(
    grader_configs={"v0_quality": grader},
    max_concurrency=8,
)

results = await runner.arun(dataset)

scores = [r.score for r in results["v0_quality"] if isinstance(r, GraderScore)]
errors = [r for r in results["v0_quality"] if isinstance(r, GraderError)]
print(f"V0 Results: avg={sum(scores)/len(scores):.2f}, errors={len(errors)}")

Step 5: Output Roadmap

The v0 grader is uncalibrated — you don't know its TPR/TNR yet. Give the user an exact path to trustworthiness:

Your v0 evaluation is ready. Here's the path to a calibrated system:

Phase 1 (now): Run the v0 grader on 30 inputs to get a baseline.
  → The grader is UNCALIBRATED. Treat scores as directional, not definitive.

Phase 2 (1-2 weeks): Collect 50 human-labeled examples (25 pass + 25 fail).
  → For each system output, have a human mark pass/fail against the criterion.
  → Store labels in labels/<grader_name>.jsonl

Phase 3: When you have 50 labels, run 03-align-human to:
  → Measure TPR/TNR of the v0 grader
  → Detect biases (position, verbosity, self-enhancement)
  → Get a human-reduction roadmap

Phase 4: When TPR >= 0.8 and TNR >= 0.8:
  → The grader is calibrated and can be used as a production gate
Show full SKILL.md (243 more words)Show less

Quick Mode vs Deep Mode

  • Quick mode (default): Steps 1-5 above. 30 minutes to v0. Use when stakes=low or when exploring.
  • Deep mode: If the user has 20+ labeled examples, use IterativeRubricsGenerator instead of SimpleRubricsGenerator for data-driven grader creation:
python
from openjudge.generator.iterative_rubric.generator import (
    IterativeRubricsGenerator,
    IterativePointwiseRubricsGeneratorConfig,
)

config = IterativePointwiseRubricsGeneratorConfig(
    grader_name="Data-Driven Grader",
    model=model,
    task_description="<from interview>",
    min_score=0, max_score=1,
    max_epochs=3,
    batch_size=10,
)
generator = IterativeRubricsGenerator(config)
grader = await generator.generate(dataset=labeled_data)  # 20+ labeled examples

Red Flags — STOP and Re-evaluate

  • "I'll generate both inputs and labels with the LLM to save time" → STOP. LLM-generating labels creates a self-consistency loop. TPR will look great until you test on real data, then it collapses.
  • "The v0 grader looks good, let's deploy it as a gate" → STOP. Uncalibrated graders have unknown TPR/TNR. They might pass everything or fail everything.
  • "I'll skip the roadmap, the user knows what to do next" → STOP. The roadmap IS the deliverable. Without it, bootstrap just produces an untrustworthy grader.

Common Mistakes

  • Over-interviewing. One prompt with 4 questions. Don't ask follow-ups unless the answers are genuinely unclear.
  • Too many principles in v0. SimpleRubricsGenerator works best with a focused task description. Don't try to evaluate 10 dimensions in v0 — start with the 2-3 most important ones.
  • Skipping stratification in synthetic inputs. If all 30 inputs are typical queries, you'll never see how the system handles edge cases.
  • Presenting v0 scores as truth. Always prefix v0 results with "UNVERIFIED — these scores are directional only."

Next Skills

After 08-bootstrap:

  • 03-align-human: Once 50 human labels are collected, calibrate the grader.
  • 01-eval-design: If you want a properly stratified dataset beyond the v0 30 inputs.
  • 02-metric-design: If you need multiple graders for different dimensions.

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/eval_pipeline/08-bootstrap of agentscope-ai/OpenJudge.

Open the folder on GitHubat commit d1e0642

Compare with similar skills

Bootstrap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bootstrap compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bootstrap this skillagentscope-ai/OpenJudge868—~2kAutomated safety check: PassApache-2.0
Agent BuildershareAI-lab/learn-claude-code78k6 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.8k15 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 6 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.8k GitHub starsUsed in 15 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from agentscope-ai/OpenJudge

All 19 skills in this repo
  • Align Human

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

    868 GitHub stars~3.1k tokensUpdated 27 days ago
    Auto-check passed
  • Prompt Regression

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

    868 GitHub stars~2.8k tokensUpdated 27 days ago
    Auto-check passed
  • RAG Eval

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

    868 GitHub stars~2.4k tokensUpdated 27 days ago
    Auto-check passed
  • 01 Auto Arena

    agentscope-ai/OpenJudge

    Automatically evaluate and compare multiple AI models or agents without pre-existing test data.

    868 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    868 GitHub stars~2.8k tokensUpdated 27 days ago
    Auto-check: warnings
  • 01 Graders And Pipeline

    agentscope-ai/OpenJudge

    Build custom LLM evaluation pipelines using the OpenJudge framework.

    868 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed

Questions about Bootstrap

What does Bootstrap do?

A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. Bootstrap is an agent skill from agentscope-ai/OpenJudge. Use when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch.

When should I use Bootstrap?

Bootstrap fits situations like: the user has nothing — no traces; no eval set — and needs to build a v0 evaluation from scratch; the user says I need to start evaluating my app but dont know where to begin; I want to set up eval for a new product.

How do I install Bootstrap in Claude Code?

Run `npx skills add agentscope-ai/OpenJudge --skill bootstrap -a claude-code`. Or copy the skill folder (skills/eval_pipeline/08-bootstrap in agentscope-ai/OpenJudge) into .claude/skills/bootstrap in your project. Claude Code loads it when a task matches its description.

How do I install Bootstrap in Codex?

Run `npx skills add agentscope-ai/OpenJudge --skill bootstrap -a codex`. Or copy the skill folder (skills/eval_pipeline/08-bootstrap in agentscope-ai/OpenJudge) into .agents/skills/bootstrap in your project. Codex loads it when a task matches its description.

Can I use Bootstrap in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill bootstrap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bootstrap, .gemini/skills/bootstrap, .github/skills/bootstrap and .opencode/skills/bootstrap in your project.

What does Bootstrap need to run?

Going by SKILL.md and its folder, Bootstrap needs the command-line tools its instructions call (pip) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Bootstrap access the network?

SKILL.md names 1 domain. In commands or code: dashscope.aliyuncs.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Bootstrap safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bootstrap use?

Bootstrap is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bootstrap use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bootstrap?

Skills that share tags, products or a category with Bootstrap: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bootstrap?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 868 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.

Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.