News Aggregator Skill
cclank/news-aggregator-skill
Comprehensive news aggregator that fetches, filters, and deeply analyzes real-time content from 44+ sources including Hacker News, Lobsters, Dev.to, GitHub, arXiv, Hugging Face Papers, AIHOT, TLDR…
A skill your agent uses when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive…
$ npx skills add agentscope-ai/OpenJudge --skill eval-report -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentscope-ai/OpenJudge eval-report --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval_pipeline/04-eval-report .claude/skills/eval-report && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "eval-report" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/04-eval-report into .claude/skills/eval-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-report", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/04-eval-reportType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentscope-ai/OpenJudge --skill eval-report -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentscope-ai/OpenJudge eval-report --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/eval_pipeline/04-eval-report .agents/skills/eval-report && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "eval-report" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/04-eval-report into .agents/skills/eval-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-report", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill eval-report -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentscope-ai/OpenJudge eval-report --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/eval_pipeline/04-eval-report .cursor/skills/eval-report && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "eval-report" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/04-eval-report into .cursor/skills/eval-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-report", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentscope-ai/OpenJudge.git --path skills/eval_pipeline/04-eval-report--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentscope-ai/OpenJudge --skill eval-report -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentscope-ai/OpenJudge eval-report --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/eval_pipeline/04-eval-report .gemini/skills/eval-report && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "eval-report" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/04-eval-report into .gemini/skills/eval-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-report", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentscope-ai/OpenJudge eval-reportInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentscope-ai/OpenJudge --skill eval-report -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/eval_pipeline/04-eval-report .github/skills/eval-report && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "eval-report" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/04-eval-report into .github/skills/eval-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-report", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill eval-report -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentscope-ai/OpenJudge eval-report --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/eval_pipeline/04-eval-report .opencode/skills/eval-report && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "eval-report" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/eval_pipeline/04-eval-report into .opencode/skills/eval-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "eval-report", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
eval-reportA skill your agent uses when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive…
Eval Report is an agent skill from agentscope-ai/OpenJudge. Use when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive summary. Also use when the user mentions eval health check, evaluation audit, ship readiness, evaluation maturity, or "how good is my evaluation system itself." This is a read-only analysis skill.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Writing & Content, covering Summarization. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Eval Report loads about 2.4k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 819 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 819 words, ~2,437 tokens.
.claude/skills/eval-report/SKILL.md (or your agent's skills folder).<HARD-GATE>
NO recommendation WITHOUT statistical evidence backing it.
NO "system ready" declaration WITHOUT all calibrated judges passing AND all production gates green.
NO trend analysis WITHOUT at least 2 data points in history.
</HARD-GATE>
Synthesize everything from your evaluation journey into a comprehensive report. This skill is read-only — it analyzes what exists, doesn't create new graders or datasets.
You MUST create a task for each item and complete them in order:
Read eval-design.md and all runs/ directories. Build a timeline:
Timeline:
2026-04-15 01-eval-design → 5 failure modes → 3 dimensions from 200 traces
2026-04-18 02-metric-design → 4 graders configured (2 LLM + 1 rule + 1 executable)
2026-04-25 (evaluation run) → 90-sample stratified dataset scored
2026-05-01 03-align-human → 2 judges Phase 3, 1 Phase 2, 1 Phase 1 (TPR/TNR + kappa)
2026-05-10 07-redteam → safety audit not yet runReport key metrics:
Rate the evaluation system across 5 dimensions:
| Dimension | L1 (Initial) | L2 (Developing) | L3 (Established) | L4 (Optimizing) |
|---|---|---|---|---|
| Failure Discovery | No systematic analysis | Failure modes identified | Coverage validated with stratification | Continuous triage from production |
| Judge Quality | v0 uncalibrated only | Some calibrated (TPR/TNR measured) | All calibrated with CI | Calibrated + aligned with humans |
| Label Coverage | < 50 labels | 50-200 labels | 200+ stratified labels | Coverage audit passed, drift monitored |
| Safety Coverage | No redteaming | Ad-hoc redteam run | Systematic redteam with policy doc | Continuous redteam with sign-off |
| Human Alignment | No alignment data | Kappa measured for some judges | Kappa ≥ 0.8 for all judges | Human spot-check only, quarterly audit |
Scoring rule: The overall maturity level is the minimum across dimensions (weakest link principle). If 4 dimensions are L3 but Safety is L1, the system is L1.
Find themes confirmed by multiple skills. Example:
Find where skills disagree. These are the most valuable findings:
What hasn't been touched by any skill?
Which principle/grader has the lowest pass rate? Where are failures clustering?
Compute Jaccard similarity between principle pairs — when sample A fails on principle X, does it also fail on principle Y? Highly correlated pairs (Jaccard > 0.5) likely share a root cause.
Which difficulty stratum performs worst across all principles? If boundary stratum TPR < 0.7 for 3 of 4 principles, boundary discrimination is a systemic weakness.
For each weakness area, classify the root cause:
| Type | Definition | Key indicator |
|---|---|---|
| system_problem | The application itself performs poorly | Low pass rate + high judge-human agreement |
| metric_problem | The judge/eval is flawed | Low pass rate + low judge-human agreement |
| data_problem | The eval dataset isn't representative | 01-eval-design coverage shows thin strata OR label drift detected |
| unclear | Not enough evidence | Conflicting signals, need more data |
This classification is critical — fixing a metric problem by changing the system (or vice versa) wastes effort.
Generate P0/P1/P2 actions. Each must include: priority, concrete action, current state, target state, expected impact, and the skill to use.
🔴 P0 | Calibrate hallucination judge
Current: TPR=0.74 (below 0.8 threshold)
Target: TPR >= 0.8
Impact: Judge becomes usable as production gate
Use: 03-align-human
🔴 P0 | Add boundary samples for hallucination
Current: n=6 boundary samples (CI half-width ±18%)
Target: n >= 20 (CI narrows to ±10%)
Impact: Reliable per-stratum TPR measurement
Use: 01-eval-design
🟡 P1 | Align tone_consistency judge
Current: kappa=0.72, bias=-0.15 (lenient)
Target: kappa >= 0.8, |bias| < 0.1
Impact: Reduce false positive rate ~15%
Use: 03-align-human
🟢 P2 | Enable auto-gate for factuality judge
Current: kappa=0.87, TPR=0.92, TNR=0.88
Impact: Eliminate 90% of human review for this dimension
Use: 03-align-human (mark Phase 3)One page for non-technical stakeholders:
Executive Summary
=================
System: Customer support chatbot for e-commerce
Stakes: Production
Report Date: 2026-05-12
SHIP READINESS: Conditional
2 items must be resolved before production gate:
1. Hallucination judge TPR below threshold (0.74 < 0.8)
2. No safety/redteam evaluation has been run
TOP 3 RISKS:
1. Hallucination detection unreliable — severity: HIGH
The judge measuring whether the bot fabricates information itself has poor
recall (TPR=0.74), meaning ~26% of hallucinations go undetected.
Mitigation: Calibrate with more boundary labels (2-3 weeks).
2. Safety coverage missing — severity: MEDIUM
No jailbreak, injection, or PII leakage testing has been performed.
Mitigation: Run redteam skill this sprint (1-2 days).
3. Tone evaluation is biased lenient — severity: LOW
The tone judge systematically rates responses as better than humans do.
Mitigation: Refine judge prompt with borderline examples.
EVAL MATURITY: L2 (Developing) → Target L3 in 3-4 weeks
Strongest: Failure Discovery (L3)
Weakest: Safety Coverage (L1), Human Alignment (L2)
NEXT ACTIONS (this sprint):
[P0] Add boundary labels + recalibrate hallucination judge
[P0] Run initial redteam evaluation
[P1] Refine tone judge alignment| File | Content |
|---|---|
eval-design.md | eval_report: namespace (maturity, signals, trends, actions) |
runs/eval-report/<ts>/report.md | Full analysis report |
runs/eval-report/<ts>/executive-summary.md | One-page stakeholder summary |
Each per-metric run you synthesize should be a runs/<skill>/<ts>/results.json row of the form:
{"metric": "order_accuracy", "mean": 0.82, "ci_95": [0.78, 0.86], "n": 90,
"by_stratum": {"easy": 0.95, "boundary": 0.71, "adversarial": 0.60},
"verdict": "pass | fail | insufficient_evidence"}After 04-eval-report:
© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/eval_pipeline/04-eval-report of agentscope-ai/OpenJudge.
Open the folder on GitHubat commit d1e0642
Eval Report next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Eval Report this skillagentscope-ai/OpenJudge | 870 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| News Aggregator Skillcclank/news-aggregator-skill | 1.3k | — | ~2.1k | Automated safety check: Pass | None | |
| AI Daily Newsgeekjourneyx/ai-daily-skill | 235 | — | ~2.3k | Automated safety check: Pass | None | |
| Vss Search ArchiveNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| AnalyzeriBigQiang/feedgrab | 614 | — | ~1k | Automated safety check: Pass | MIT | |
| Reportmicrosoft/data-formulator | 18k | — | ~1.5k | Automated safety check: Pass | MIT |
cclank/news-aggregator-skill
Comprehensive news aggregator that fetches, filters, and deeply analyzes real-time content from 44+ sources including Hacker News, Lobsters, Dev.to, GitHub, arXiv, Hugging Face Papers, AIHOT, TLDR…
geekjourneyx/ai-daily-skill
Fetches AI news from smol.ai RSS and generates structured markdown with intelligent summarization and categorization.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when a user wants to search archived VSS video or ingest or delete a source for search.
iBigQiang/feedgrab
Content Analyzer — any content (URL, text, transcript) into structured analysis report with actionable insights.
microsoft/data-formulator
Turn an exploration (threads, findings, charts) into a single Markdown report — note, blog post, executive summary, KPI dashboard, slide brief, or multi-section analytical report, with embedded…
earlyaidopters/claudeclaw
Summarize the current conversation into a TLDR note and save it to your notes folder.
agentscope-ai/OpenJudge
A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…
agentscope-ai/OpenJudge
A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.
agentscope-ai/OpenJudge
A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.
agentscope-ai/OpenJudge
Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.
agentscope-ai/OpenJudge
A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…
agentscope-ai/OpenJudge
Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.
Categories
A skill your agent uses when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive…. Eval Report is an agent skill from agentscope-ai/OpenJudge. Use when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive summary.
Eval Report fits situations like: the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment; cross-skill signals; prioritized actions; an executive summary.
Run `npx skills add agentscope-ai/OpenJudge --skill eval-report -a claude-code`. Or copy the skill folder (skills/eval_pipeline/04-eval-report in agentscope-ai/OpenJudge) into .claude/skills/eval-report in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentscope-ai/OpenJudge --skill eval-report -a codex`. Or copy the skill folder (skills/eval_pipeline/04-eval-report in agentscope-ai/OpenJudge) into .agents/skills/eval-report in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill eval-report -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-report, .gemini/skills/eval-report, .github/skills/eval-report and .opencode/skills/eval-report in your project.
SKILL.md names no scripts, command-line tools or credentials: Eval Report is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Eval Report is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Eval Report: News Aggregator Skill (cclank/news-aggregator-skill, 1.3k stars), AI Daily News (geekjourneyx/ai-daily-skill, 235 stars), Vss Search Archive (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars) and Analyzer (iBigQiang/feedgrab, 614 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 870 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.
Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.