Search

AI & LLM Engineering · agentscope-ai/OpenJudge

9 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

agentscope-ai/OpenJudge871—~2.8kAutomated safety check: PassApache-2.01 mo ago
2

A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

agentscope-ai/OpenJudge871—~2.4kAutomated safety check: PassApache-2.01 mo ago
3

Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.

agentscope-ai/OpenJudge8711 repo~5kAutomated safety check: PassApache-2.01 mo ago
4

Automatically evaluate and compare multiple AI models or agents without pre-existing test data.

agentscope-ai/OpenJudge871—~2.5kAutomated safety check: PassApache-2.01 mo ago
5

Build custom LLM evaluation pipelines using the OpenJudge framework.

agentscope-ai/OpenJudge871—~1.3kAutomated safety check: PassApache-2.01 mo ago
6

Build RL reward signals using the OpenJudge framework. An agent skill from agentscope-ai/OpenJudge.

agentscope-ai/OpenJudge871—~1.6kAutomated safety check: PassApache-2.01 mo ago
7

A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch.

agentscope-ai/OpenJudge871—~2kAutomated safety check: PassApache-2.01 mo ago
8

A skill your agent uses when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing at all.

agentscope-ai/OpenJudge871—~2.5kAutomated safety check: PassApache-2.01 mo ago
9

A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining…

agentscope-ai/OpenJudge871—~5.1kAutomated safety check: PassApache-2.01 mo ago