Agent skill

Align Human

by agentscope-ai in agentscope-ai/OpenJudge

A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

Apache-2.0Auto-check passedBusiness, Finance & HR

Install Align Human

skills CLI
$ npx skills add agentscope-ai/OpenJudge --skill align-human -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/OpenJudge align-human --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval_pipeline/03-align-human .claude/skills/align-human && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
align-human
GitHub stars
868
Token cost
~3.1k tokens
SKILL.md length
1,010 words
Files
2 (incl. scripts)
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

  • Works in 7 steps: Load and Pair Data → Measure TPR/TNR → Calculate Agreement → …
  • The user has a judge/grader and human-labeled data
  • SKILL.md covers When to Activate, Checklist, Fast path: run the bundled… and Step 1: Load and Pair Data, plus 8 more sections
  • Runs Python scripts from its folder; calls python

What it does

Align Human is an agent skill from agentscope-ai/OpenJudge. Use when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic evaluation can replace human review, or build a human-reduction roadmap. Also use when the user mentions calibration, TPR/TNR, judge validation, inter-rater agreement, Cohen's kappa, bias detection, or "is my automatic evaluation trustworthy." Merges the calibrate and align functions into one skill.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/calibration.py`).

It sits in Business, Finance & HR, covering Performance reviews. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.

When your agent uses it

  • The user has a judge/grader and human-labeled data
  • Wants to measure how well the judge agrees with humans
  • Detect systematic biases
  • Determine whether automatic evaluation can replace human review

Example prompts

  • “s kappa, bias detection, or”
  • “/align-human”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Load and Pair Data
  2. Measure TPR/TNR
  3. Calculate Agreement
  4. Five Bias Detection Checks
  5. Disagreement Pattern Analysis
  6. Human-Reduction Roadmap
  7. Confirmation and Output

What it can do on your machine

Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Align Human loads about 3.1k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 1,010 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 1,010 words, ~3,068 tokens.

Download SKILL.mdSave it as .claude/skills/align-human/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
align-human
description
Use when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic evaluation can replace human review, or build a human-reduction roadmap. Also use when the user mentions calibration, TPR/TNR, judge validation, inter-rater agreement, Cohen's kappa, bias detection, or "is my automatic evaluation trustworthy." Merges the calibrate and align functions into one skill.
<HARD-GATE>
NO calibrated:true WITHOUT TPR >= 0.8 AND TNR >= 0.8 AND boundary stratum TPR >= 0.6 AND TNR >= 0.6 AND n_dev >= 10 per class AND test-set drop < 10%.
NO human_reduction_phase >= 2 WITHOUT kappa >= 0.6 AND boundary kappa >= 0.6.
NO alignment conclusion WITHOUT all 5 bias checks completed.
</HARD-GATE>

Align Human

Measure whether your automatic judge agrees with human judgment, detect where and why they disagree, and build a roadmap to reduce human review over time.

When to Activate

  • You have a working judge/grader and 50+ human-labeled examples
  • You want to know if the judge is trustworthy enough to replace human review
  • You've noticed the judge's decisions being overturned by humans
  • You're preparing to deploy an evaluation as a production gate

Checklist

You MUST create a task for each item and complete them in order:

  1. Load paired data — match judge verdicts with human labels
  2. Measure TPR/TNR — confusion matrix + per-stratum breakdown
  3. Calculate agreement — Cohen's kappa, Gwet's AC1, systematic bias
  4. Run bias detection — 5 systematic bias checks
  5. Analyze disagreements — cluster patterns + diagnose root causes
  6. Build human-reduction roadmap — 4-phase transition plan
  7. Confirm and record — one confirmation, then write results

Fast path: run the bundled script

Don't hand-write the calibration statistics — that is exactly where subtle bugs hide. Run the bundled, tested script (scripts/calibration.py, standard library only, no OpenJudge dependency):

bash
python scripts/calibration.py --pairs pairs.jsonl                 # one paired file, OR
python scripts/calibration.py --verdicts verdicts.jsonl --labels labels.jsonl --stratum-key difficulty

Paired rows look like {"id","judge":"pass|fail","human":"pass|fail","stratum"?} (judge/human may also be 1/0). It prints the confusion matrix, TPR/TNR/F1 with bootstrap 95% CIs, Cohen's kappa, Gwet's AC1 (auto-flags the kappa paradox), directional bias, per-stratum TPR/TNR, and the calibration gate verdict (calibrated / not_calibrated / insufficient_evidence; exit code 0 only if calibrated). --json for machine output, --self-test to verify it.

Always report and interpret the actual numbers the script returns — TPR/TNR (with their 95% CIs), Cohen's kappa, Gwet's AC1, directional bias, per-stratum TPR/TNR, and the gate verdict — never just state that you ran it. If you don't yet have the paired verdicts/labels, say exactly what's missing (e.g. the judge verdicts file, or N more labels per class).

The Steps below explain what each number means and how to act on it — read them to interpret the script's output. The inline snippets are the reference behind the script; you normally just run the script rather than re-implementing it.

Step 1: Load and Pair Data

Human labels live in labels/<grader_name>.jsonl, one row per judged sample. Keep them separate from the dataset so they can be re-paired with any judge run:

python
{"id": "sample_017", "label": "pass",   # "pass"|"fail" (or 1/0); joins to a dataset row id
 "annotator": "alice", "rationale": "Order number matches context.",
 "timestamp": "2026-06-20T10:00:00Z", "schema_version": 1}

Match judge verdicts with human labels:

python
import json

# Load human labels and judge verdicts
labels = {item["id"]: item["label"] for item in json.load(open("labels.jsonl"))}
verdicts = json.load(open("runs/verdicts-dev.jsonl"))

# Pair them
paired = []
for v in verdicts:
    if v["id"] in labels:
        paired.append({
            "id": v["id"],
            "judge": v["verdict"],   # "pass" or "fail"
            "human": labels[v["id"]], # "pass" or "fail"
        })

print(f"Paired: {len(paired)}, Unmatched: {len(verdicts) - len(paired)}")

# Warn if severe class imbalance
pass_rate = sum(1 for p in paired if p["human"] == "pass") / len(paired)
if pass_rate > 0.8 or pass_rate < 0.2:
    print(f"WARNING: Human label pass rate is {pass_rate:.0%} — "
          "kappa may be paradoxically low. Use Gwet's AC1 as complement.")

Step 2: Measure TPR/TNR

Build a confusion matrix and compute per-stratum metrics:

python
from openjudge.analyzer.validation import (
    AccuracyAnalyzer, F1ScoreAnalyzer,
    FalsePositiveAnalyzer, FalseNegativeAnalyzer,
)

# Convert to OpenJudge-compatible dataset with labels
analysis_dataset = [
    {"query": p.get("query", ""), "response": p.get("response", ""),
     "label": 1 if p["human"] == "pass" else 0}
    for p in paired
]

# Binary grader results (1=pass, 0=fail)
grader_results = [
    GraderScore(name="judge", score=1.0 if p["judge"] == "pass" else 0.0, reason="")
    for p in paired
]

accuracy = AccuracyAnalyzer().analyze(analysis_dataset, grader_results, label_path="label")
f1 = F1ScoreAnalyzer().analyze(analysis_dataset, grader_results, label_path="label")
fpr = FalsePositiveAnalyzer().analyze(analysis_dataset, grader_results, label_path="label")
fnr = FalseNegativeAnalyzer().analyze(analysis_dataset, grader_results, label_path="label")

# TPR = 1 - FNR, TNR = 1 - FPR
tpr = 1 - fnr.false_negative_rate
tnr = 1 - fpr.false_positive_rate
print(f"TPR={tpr:.2f}, TNR={tnr:.2f}, F1={f1.f1_score:.2f}")

# Per-stratum breakdown if difficulty data exists
for stratum in ["easy", "boundary", "hard"]:
    stratum_data = [d for d in analysis_dataset
                    if d.get("metadata", {}).get("difficulty") == stratum]
    if len(stratum_data) >= 10:
        # Compute TPR/TNR per stratum
        ...

Why per-stratum matters: A judge with TPR=0.9 overall but TPR=0.5 on boundary cases is unreliable exactly where judgment matters most. This is the "progress illusion" (EMNLP 2025) — aggregate metrics hide stratum-level failure.

Bootstrap 95% CI

Use bootstrap resampling to quantify uncertainty:

python
import numpy as np

def bootstrap_ci(samples, metric_fn, n_iter=1000, ci=95):
    """Compute bootstrap confidence interval for a metric."""
    n = len(samples)
    values = []
    for _ in range(n_iter):
        idx = np.random.choice(n, n, replace=True)
        resampled = [samples[i] for i in idx]
        values.append(metric_fn(resampled))
    lower = np.percentile(values, (100 - ci) / 2)
    upper = np.percentile(values, 100 - (100 - ci) / 2)
    return np.mean(values), lower, upper

tpr_mean, tpr_low, tpr_high = bootstrap_ci(
    paired, lambda s: sum(1 for p in s if p["judge"] == "fail" and p["human"] == "fail")
                      / max(1, sum(1 for p in s if p["human"] == "fail"))
)

Step 3: Calculate Agreement

Cohen's Kappa (chance-corrected agreement)
kappa = (p_o - p_e) / (1 - p_e)
  • p_o: observed agreement rate
  • p_e: expected agreement by chance
  • = 0.8: substantial — judge is consistent with humans

  • 0.6-0.8: moderate — conditional trust, needs spot-checking
  • < 0.6: weak — cannot replace human judgment yet
Gwet's AC1 (robust to class imbalance)

When 90% of samples are "pass," kappa can be paradoxically low even with high agreement. Gwet's AC1 corrects for this. If kappa and AC1 differ by > 0.15, report both and note the class imbalance effect.

Systematic Bias
bias = P(judge=fail | human=pass) - P(judge=pass | human=fail)
  • bias > 0.1: judge is stricter than humans (over-flagging)
  • bias < -0.1: judge is more lenient than humans (under-flagging)
  • |bias| < 0.1: no significant directional bias
Show full SKILL.md (446 more words)Show less

Step 4: Five Bias Detection Checks

CheckWhat to look forHow to measure
Position biasDoes response order affect pairwise judgment?Swap A/B order, compare win rates. Diff > 0.05 = bias
Verbosity biasDo longer responses score higher?Pearson r between score and response length.
Self-enhancementDoes the judge favor its own model family?Check if judge model family = target model family. Same family = risk
Progress illusionDoes aggregate TPR hide boundary failure?Compare overall TPR vs boundary TPR. Gap > 0.2 = illusion
Label driftHave system outputs changed since labeling?Compare historical vs current pass rate. Shift > 0.15 = drift

Step 5: Disagreement Pattern Analysis

For samples where judge and human disagree, cluster them to find root causes:

  1. Extract disagreement samples — where judge verdict != human label
  2. Classify each disagreement:
    • judge_prompt_ambiguous: pass/fail definitions not clear for this case
    • human_inconsistent: multiple human annotators disagreed on this sample
    • task_inherently_subjective: the dimension is fundamentally subjective
    • label_error: human label appears wrong, judge's reasoning is more convincing
    • judge_too_strict: judge applies criteria more harshly than humans intend
    • judge_too_lenient: judge overlooks issues humans catch

Present the top 3 patterns with 3 exemplars each so the user can decide whether to refine the judge prompt or accept the disagreement as inherent noise.

Step 6: Human-Reduction Roadmap

PhaseConditionHuman roleJudge roleTrigger to advance
1: Advisorykappa < 0.6100% human reviewJudge is reference onlykappa >= 0.6
2: Assistedkappa >= 0.6Spot-check 20% of judgmentsJudge is primary screenerkappa >= 0.8, boundary >= 0.6
3: Auto-gatekappa >= 0.8Review only borderline + low-confidenceJudge is production gatekappa >= 0.9, all strata >= 0.8
4: Autonomouskappa >= 0.9Quarterly auditJudge runs independentlyContinuous monitoring

Step 7: Confirmation and Output

Present the alignment dashboard:

Alignment Results for [grader_name]:

TPR: 0.91  TNR: 0.88  F1: 0.90
Kappa: 0.87  AC1: 0.89  Bias: +0.03 (none)
95% CI (accuracy): [0.84, 0.93]

Per-stratum:
  easy:      TPR=0.96  TNR=0.94
  boundary:  TPR=0.82  TNR=0.79  ← weakest, monitor
  hard:      TPR=0.85  TNR=0.83

Bias checks:
  ✓ position:   no bias detected
  ✓ verbosity:  r=0.12 (clean)
  ✓ self-enh:   judge model != target model
  ✓ progress:   boundary gap 0.09 (acceptable)
  ✓ label drift: KL=0.08 (stable)

Phase: 3 (Auto-gate) — judge is calibrated and aligned

Disagreement patterns:
  - 12 samples: judge slightly stricter on multi-step queries
    Root cause: judge_prompt_ambiguous
    Recommendation: add a multi-step borderline example to few-shot

Recommendation: [calibrated + aligned | needs refinement | not ready]

Common Mistakes

  • Using raw accuracy instead of TPR/TNR. A judge that calls everything "pass" gets 90% accuracy when 90% of samples pass, but catches zero failures. TPR and TNR decompose accuracy into what actually matters.
  • Trusting kappa without checking class balance. With 95% pass rate, kappa can be 0.4 even with 95% raw agreement. Always report Gwet's AC1 alongside kappa.
  • Skipping per-stratum analysis. Aggregate TPR of 0.9 with boundary TPR of 0.5 means the judge is unreliable exactly where you need it most.
  • Not setting a stop condition for iteration. Refining the judge prompt has diminishing returns. After 3 iterations with no kappa improvement > 0.03, stop and collect more labeled data instead.
  • Forgetting judge model != target model constraint. Self-evaluation inflates TPR by 3-8%. Always verify these are different models.

Next Skills

After 03-align-human:

  • 04-eval-report: Generate a comprehensive report with maturity assessment.
  • 02-metric-design: If you need to redesign graders based on bias findings.
  • 01-eval-design: If disagreement patterns reveal dataset coverage gaps.

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/eval_pipeline/03-align-human of agentscope-ai/OpenJudge.

  • SKILL.md
  • scripts/calibration.py

Open the folder on GitHubat commit d1e0642

Compare with similar skills

Align Human next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Align Human compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Align Human this skillagentscope-ai/OpenJudge868—~3.1kAutomated safety check: PassApache-2.0
Wp Performance Reviewelvismdev/claude-wordpress-skills2351 repos~4.5kAutomated safety check: PassMIT
Performance ReportAffitor/affiliate-skills6991 repos~2.5kAutomated safety check: PassMIT
Run Mv Hoi Reconstructionnvidia-isaac/video_to_data850—~1.5kAutomated safety check: PassCustom licence
Company Analysiszhu1090093659/dsh-trading231—~4.2kAutomated safety check: PassCustom licence
Windbg Diagnostic Methodmicrosoft/win-dev-skills462—~1.9kAutomated safety check: PassMIT

Similar skills

  • Wp Performance Review

    elvismdev/claude-wordpress-skills

    WordPress performance code review and optimization analysis.

    235 GitHub starsUsed in 1 repo~4.5k tokens
    Business, Finance & HRAuto-check passed
  • Performance Report

    Affitor/affiliate-skills

    Generate affiliate performance reports with KPIs and recommendations.

    699 GitHub starsUsed in 1 repo~2.5k tokens
    Business, Finance & HRAuto-check passed
  • Run Mv Hoi Reconstruction

    nvidia-isaac/video_to_data

    Run and validate the repository-local multi-view camera calibration and human-object reconstruction pipelines.

    850 GitHub stars~1.5k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Company Analysis

    zhu1090093659/dsh-trading

    A skill your agent uses when the user wants to analyze a listed company, stock, business, or investment target; challenge or revise an existing company report; compare A/H or primary-listing/ADR…

    231 GitHub stars~4.2k tokensUpdated 5 days ago
    Business, Finance & HRAuto-check passed
  • Windbg Diagnostic Method

    microsoft/win-dev-skills

    Official

    Use with every WinDbg plugin investigation to apply evidence-first reasoning, confidence calibration, contrarian review, structured reporting, and deterministic validation.

    462 GitHub stars~1.9k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • AI Index

    mizchi/skills

    Method and tooling for measuring how AI-generated a piece of prose reads, in Japanese or English.

    356 GitHub stars~3.6k tokensUpdated 6 days ago
    Business, Finance & HRAuto-check passed

More from agentscope-ai/OpenJudge

All 19 skills in this repo
  • Prompt Regression

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

    868 GitHub stars~2.8k tokensUpdated 27 days ago
    Auto-check passed
  • RAG Eval

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

    868 GitHub stars~2.4k tokensUpdated 27 days ago
    Auto-check passed
  • 01 Auto Arena

    agentscope-ai/OpenJudge

    Automatically evaluate and compare multiple AI models or agents without pre-existing test data.

    868 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    868 GitHub stars~2.8k tokensUpdated 27 days ago
    Auto-check: warnings
  • 01 Graders And Pipeline

    agentscope-ai/OpenJudge

    Build custom LLM evaluation pipelines using the OpenJudge framework.

    868 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • 01 Paper Review

    agentscope-ai/OpenJudge

    Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline.

    868 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed

Questions about Align Human

What does Align Human do?

A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…. Align Human is an agent skill from agentscope-ai/OpenJudge. Use when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic evaluation can replace human review, or build a human-reduction roadmap.

When should I use Align Human?

Align Human fits situations like: the user has a judge/grader and human-labeled data; wants to measure how well the judge agrees with humans; detect systematic biases; determine whether automatic evaluation can replace human review.

How do I install Align Human in Claude Code?

Run `npx skills add agentscope-ai/OpenJudge --skill align-human -a claude-code`. Or copy the skill folder (skills/eval_pipeline/03-align-human in agentscope-ai/OpenJudge) into .claude/skills/align-human in your project. Claude Code loads it when a task matches its description.

How do I install Align Human in Codex?

Run `npx skills add agentscope-ai/OpenJudge --skill align-human -a codex`. Or copy the skill folder (skills/eval_pipeline/03-align-human in agentscope-ai/OpenJudge) into .agents/skills/align-human in your project. Codex loads it when a task matches its description.

Can I use Align Human in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill align-human -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/align-human, .gemini/skills/align-human, .github/skills/align-human and .opencode/skills/align-human in your project.

What does Align Human need to run?

Going by SKILL.md and its folder, Align Human needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Align Human access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Align Human safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Align Human use?

Align Human is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Align Human use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Align Human?

Skills that share tags, products or a category with Align Human: Wp Performance Review (elvismdev/claude-wordpress-skills, 235 stars), Performance Report (Affitor/affiliate-skills, 699 stars), Run Mv Hoi Reconstruction (nvidia-isaac/video_to_data, 850 stars) and Company Analysis (zhu1090093659/dsh-trading, 231 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Align Human?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 868 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.

Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.