Search

agentscope-ai/OpenJudge

19 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

agentscope-ai/OpenJudge871—~3.1kAutomated safety check: PassApache-2.01 mo ago
2

A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

agentscope-ai/OpenJudge871—~2.8kAutomated safety check: PassApache-2.01 mo ago
3

A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

agentscope-ai/OpenJudge871—~2.4kAutomated safety check: PassApache-2.01 mo ago
4

Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.

agentscope-ai/OpenJudge8711 repo~5kAutomated safety check: PassApache-2.01 mo ago
5

A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

agentscope-ai/OpenJudge871—~2.8kAutomated safety check: WarnApache-2.01 mo ago
6

Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.

agentscope-ai/OpenJudge8711 repo~4.6kAutomated safety check: WarnApache-2.01 mo ago
7

A skill your agent uses when the user wants help with academic papers or citations but it's unclear which specific workflow fits — reviewing a paper, checking a BibTeX file for fake references, or…

agentscope-ai/OpenJudge871—~976Automated safety check: PassApache-2.01 mo ago
8

A skill your agent uses when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom…

agentscope-ai/OpenJudge871—~1kAutomated safety check: PassApache-2.01 mo ago
9

Automatically evaluate and compare multiple AI models or agents without pre-existing test data.

agentscope-ai/OpenJudge871—~2.5kAutomated safety check: PassApache-2.01 mo ago
10

Build custom LLM evaluation pipelines using the OpenJudge framework.

agentscope-ai/OpenJudge871—~1.3kAutomated safety check: PassApache-2.01 mo ago
11

Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline.

agentscope-ai/OpenJudge871—~2.4kAutomated safety check: PassApache-2.01 mo ago
12

Verify a BibTeX file for hallucinated or fabricated references by cross-checking every entry against CrossRef, arXiv, and DBLP.

agentscope-ai/OpenJudge871—~681Automated safety check: PassApache-2.01 mo ago
13

Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP.

agentscope-ai/OpenJudge871—~2.4kAutomated safety check: PassApache-2.01 mo ago
14

Build RL reward signals using the OpenJudge framework. An agent skill from agentscope-ai/OpenJudge.

agentscope-ai/OpenJudge871—~1.6kAutomated safety check: PassApache-2.01 mo ago
15

Generate text, images, video, speech, and music via the MiniMax AI platform.

agentscope-ai/OpenJudge871—~655Automated safety check: PassApache-2.01 mo ago
16

A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch.

agentscope-ai/OpenJudge871—~2kAutomated safety check: PassApache-2.01 mo ago
17

A skill your agent uses when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive…

agentscope-ai/OpenJudge871—~2.4kAutomated safety check: PassApache-2.01 mo ago
18

A skill your agent uses when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing at all.

agentscope-ai/OpenJudge871—~2.5kAutomated safety check: PassApache-2.01 mo ago
19

A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining…

agentscope-ai/OpenJudge871—~5.1kAutomated safety check: PassApache-2.01 mo ago