Repository
agentscope-ai/OpenJudge agent skills
- skills
- 19
- GitHub stars
- 868
GitHub description: “OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards”
- Stars
- 868 (73 forks)
- Licence
- Apache-2.0
- Last push
- Sep 2026
- Created
- Jul 2025
- Homepage
- openjudge.me
- alignment
- reward
- reward-model
- rlhf
- agent
- evaluation
- grader
- llm
- agent-skills
- ai-agent
- skill-md
- skills
Install all skills
npx skills add agentscope-ai/OpenJudgeAdd --skill <name> for a single skill and -a <agent> to choose the agent (see the agent guides).
Skills in agentscope-ai/OpenJudge, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic… | agentscope-ai/ | 868 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 2 | A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline. | agentscope-ai/ | 868 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 3 | 3.RAG Eval A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues. | agentscope-ai/ | 868 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 4 | Automatically evaluate and compare multiple AI models or agents without pre-existing test data. | agentscope-ai/ | 868 | 1 repo | ~2.5k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 5 | A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a… | agentscope-ai/ | 868 | — | ~2.8k | Automated safety check: Warn | Apache-2.0 | 27 days ago |
| 6 | Build custom LLM evaluation pipelines using the OpenJudge framework. | agentscope-ai/ | 868 | 1 repo | ~1.3k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 7 | Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline. | agentscope-ai/ | 868 | 1 repo | ~2.4k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 8 | Verify a BibTeX file for hallucinated or fabricated references by cross-checking every entry against CrossRef, arXiv, and DBLP. | agentscope-ai/ | 868 | 1 repo | ~681 | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 9 | Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. | agentscope-ai/ | 868 | 1 repo | ~2.4k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 10 | 10.02 Rl Reward Build RL reward signals using the OpenJudge framework. An agent skill from agentscope-ai/OpenJudge. | agentscope-ai/ | 868 | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 11 | Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project. | agentscope-ai/ | 868 | 1 repo | ~5k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 12 | 12.Mmx CLI Generate text, images, video, speech, and music via the MiniMax AI platform. | agentscope-ai/ | 868 | 1 repo | ~655 | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 13 | A skill your agent uses when the user wants help with academic papers or citations but it's unclear which specific workflow fits — reviewing a paper, checking a BibTeX file for fake references, or… | agentscope-ai/ | 868 | — | ~976 | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 14 | A skill your agent uses when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom… | agentscope-ai/ | 868 | — | ~1k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 15 | 15.Bootstrap A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. | agentscope-ai/ | 868 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 16 | 16.Eval Report A skill your agent uses when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive… | agentscope-ai/ | 868 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 17 | 17.Meta Eval A skill your agent uses when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing at all. | agentscope-ai/ | 868 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 18 | Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks. | agentscope-ai/ | 868 | 1 repo | ~4.6k | Automated safety check: Warn | Apache-2.0 | 27 days ago |
| 19 | A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining… | agentscope-ai/ | 868 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
Questions, answered from the data.
What is the best skill in agentscope-ai/OpenJudge?
Align Human from agentscope-ai/OpenJudge ranks first of the 19 skills in agentscope-ai/OpenJudge listed here, with the highest score: its repository has 868 GitHub stars, its SKILL.md loads about 3.1k tokens and it passes the automated safety check with no findings. Next come Prompt Regression and RAG Eval.
Are the skills in agentscope-ai/OpenJudge official?
None yet. All 19 skills in agentscope-ai/OpenJudge listed here come from community repositories; a skill counts as official when the product's own GitHub organization publishes it.
How do I install all skills from agentscope-ai/OpenJudge?
Run npx skills add agentscope-ai/OpenJudge in your project: the open-source skills CLI installs the repository's skills into your coding agent's skills folder. To install a single skill, open its page here for the exact command.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.