Search
agentscope-ai/OpenJudge
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic… | agentscope-ai/ | 871 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 2 | A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline. | agentscope-ai/ | 871 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 3 | 3.RAG Eval A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues. | agentscope-ai/ | 871 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 4 | Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project. | agentscope-ai/ | 871 | 1 repo | ~5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 5 | A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a… | agentscope-ai/ | 871 | — | ~2.8k | Automated safety check: Warn | Apache-2.0 | 1 mo ago |
| 6 | Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks. | agentscope-ai/ | 871 | 1 repo | ~4.6k | Automated safety check: Warn | Apache-2.0 | 1 mo ago |
| 7 | A skill your agent uses when the user wants help with academic papers or citations but it's unclear which specific workflow fits — reviewing a paper, checking a BibTeX file for fake references, or… | agentscope-ai/ | 871 | — | ~976 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 8 | A skill your agent uses when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom… | agentscope-ai/ | 871 | — | ~1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 9 | Automatically evaluate and compare multiple AI models or agents without pre-existing test data. | agentscope-ai/ | 871 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 10 | Build custom LLM evaluation pipelines using the OpenJudge framework. | agentscope-ai/ | 871 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 11 | Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline. | agentscope-ai/ | 871 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 12 | Verify a BibTeX file for hallucinated or fabricated references by cross-checking every entry against CrossRef, arXiv, and DBLP. | agentscope-ai/ | 871 | — | ~681 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 13 | Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. | agentscope-ai/ | 871 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 14 | 14.02 Rl Reward Build RL reward signals using the OpenJudge framework. An agent skill from agentscope-ai/OpenJudge. | agentscope-ai/ | 871 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 15 | 15.Mmx CLI Generate text, images, video, speech, and music via the MiniMax AI platform. | agentscope-ai/ | 871 | — | ~655 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 16 | 16.Bootstrap A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. | agentscope-ai/ | 871 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 17 | 17.Eval Report A skill your agent uses when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive… | agentscope-ai/ | 871 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 18 | 18.Meta Eval A skill your agent uses when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing at all. | agentscope-ai/ | 871 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 19 | A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining… | agentscope-ai/ | 871 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |