Search
AI & LLM Engineering · agentscope-ai/OpenJudge
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline. | agentscope-ai/ | 871 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 2 | 2.RAG Eval A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues. | agentscope-ai/ | 871 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 3 | Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project. | agentscope-ai/ | 871 | 1 repo | ~5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 4 | Automatically evaluate and compare multiple AI models or agents without pre-existing test data. | agentscope-ai/ | 871 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 5 | Build custom LLM evaluation pipelines using the OpenJudge framework. | agentscope-ai/ | 871 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 6 | Build RL reward signals using the OpenJudge framework. An agent skill from agentscope-ai/OpenJudge. | agentscope-ai/ | 871 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 7 | A skill your agent uses when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch. | agentscope-ai/ | 871 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 8 | A skill your agent uses when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines, or nothing at all. | agentscope-ai/ | 871 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 9 | A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining… | agentscope-ai/ | 871 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |