AI Agent Evaluation Benchmarking
sickn33/agentic-awesome-skills
Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.
Evaluate analysis results for quality and reliability. An agent skill from openJiuwen-ai/sciencediscovery.
$ npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery result-evaluator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/result-evaluator .claude/skills/result-evaluator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "result-evaluator" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/result-evaluator into .claude/skills/result-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "result-evaluator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/result-evaluatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery result-evaluator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/result-evaluator .agents/skills/result-evaluator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "result-evaluator" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/result-evaluator into .agents/skills/result-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "result-evaluator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery result-evaluator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/result-evaluator .cursor/skills/result-evaluator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "result-evaluator" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/result-evaluator into .cursor/skills/result-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "result-evaluator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/openJiuwen-ai/sciencediscovery.git --path skills/result-evaluator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery result-evaluator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/result-evaluator .gemini/skills/result-evaluator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "result-evaluator" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/result-evaluator into .gemini/skills/result-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "result-evaluator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install openJiuwen-ai/sciencediscovery result-evaluatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/result-evaluator .github/skills/result-evaluator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "result-evaluator" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/result-evaluator into .github/skills/result-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "result-evaluator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery result-evaluator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/result-evaluator .opencode/skills/result-evaluator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "result-evaluator" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/result-evaluator into .opencode/skills/result-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "result-evaluator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
result-evaluatorEvaluate analysis results for quality and reliability. An agent skill from openJiuwen-ai/sciencediscovery.
Result Evaluator is an agent skill from openJiuwen-ai/sciencediscovery. Evaluate analysis results for quality and reliability. Scores Accuracy, Completeness, Robustness, and Relevance (0-10), checks source reliability and methodology, audits statistical rigor, and decides ACCEPTANDPROCEED or REVISEANDRETRY. NOT for performing analysis or modifying results.
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: ScienceDiscovery is an all‑in‑one agentic workbench built specifically for scientific research. The licence is Apache-2.0.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ab1403f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
pypi.tuna.tsinghua.edu.cnFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Result Evaluator loads about 4.4k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 1,777 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from openJiuwen-ai/sciencediscovery at commit ab1403f, republished under its Apache-2.0 licence (© openJiuwen-ai). 1,777 words, ~4,442 tokens.
.claude/skills/result-evaluator/SKILL.md (or your agent's skills folder).This skill evaluates analysis results against predefined criteria and decides ACCEPT_AND_PROCEED or REVISE_AND_RETRY. It follows a 4-phase protocol: criterion alignment → multi-dimensional evaluation with Source Reliability hard gate → statistical methodology audit → overall assessment. Hallucination detected → immediate REVISE; any checklist dimension FAIL → mandatory REVISE (hard gate, overrides scoring).
Always load this skill when:
code-engineer — has just produced as a Result PackageACCEPT_AND_PROCEED vs REVISE_AND_RETRY (or CONDITIONAL) decision before the results are used downstream (e.g. fed into a report, shared with stakeholders, or acted on)This skill evaluates analysis results with methodology documentation. Accepted input formats:
From code-engineer (recommended upstream skill):
--output-file JSON ([{col: val, ...}]) or CSV/MD export — provides the numerical/tabular resultsFrom other sources: any structured results with accompanying methodology description. Minimum required: results data + method description + data source identification.
If methodology documentation or data traceability is missing, note the gap in evaluation and flag Source Reliability as PARTIALLY_RELIABLE.
If you need to install new Python packages, install them through the Tsinghua PyPI mirror for reliability:
pip install [python package] -i https://pypi.tuna.tsinghua.edu.cn/simpleIdentify the evaluation context:
Prerequisites: results must be available and parseable; methodology documentation and data traceability should be provided (evaluation quality degrades without them); criteria must be specified or inferable.
Map each result to an evaluation criterion. Flag UNMAPPED results and uncovered criteria. Infer criteria from context if missing (document as inferred).
Source Reliability Hard Gate (check first — hallucination → immediate REVISE_AND_RETRY, skip rest):
For computational-type results (from code-engineer and similar tools):
| Check | What to detect |
|---|---|
| Data traceability | Cited data sources exist (file/sheet/column match actual data, row counts consistent) |
| Method consistency | Stated methods match the actual code implementation |
| Fabrication | Invented statistics, untraceable numbers, results that cannot be reproduced from given code and data |
| Code-data alignment | Code actually references the claimed data files/variables, not different ones |
For research-type results (literature-based, citing external references):
| Check | What to detect |
|---|---|
| Data traceability | Cited data sources exist (file/sheet/field match actual data) |
| Reference validity | Citations have author+year+DOI/PubMed (not "studies show") |
| Identifier authenticity | Standard entity/gene/protein names (not self-created) |
| Method consistency | Stated methods match implementation |
| Fabrication | Invented statistics, fake references, untraceable results |
Verdict: RELIABLE (PASS) / PARTIALLY_RELIABLE (FAIL, continue) / UNRELIABLE (REVISE, stop).
Unified Evaluation Matrix — score each dimension 0-10; each dimension also has a PASS/FAIL threshold (score ≥5 → PASS, score <5 → FAIL):
| Dimension | 9-10 | 7-8 | 4-6 | 0-3 | PASS threshold |
|---|---|---|---|---|---|
| Accuracy | Correct, methods match | Minor errors | Significant errors | Fundamental errors | ≥5 |
| Completeness | Complete, no gaps | Minor gaps | Significant gaps | Major omissions | ≥5 |
| Robustness | Sound methods, assumptions verified | 1-2 concerns | 3-4 issues | Invalid methods | ≥5 |
| Relevance | Directly addresses question | Mostly relevant | Partially relevant | Irrelevant | ≥5 |
| Methodology | Justified, rigorous, reproducible | Adequate justification | Weak justification | No justification | ≥5 |
| Critical reflection | Assumptions stated, limitations discussed | Some reflection | Minimal reflection | No reflection | ≥5 |
Hard Gate Rule: any dimension FAIL (score <5) → mandatory REVISE_AND_RETRY, regardless of the average score. The scoring average determines the severity grading of the REVISE decision, not whether to REVISE.
Composite Quality Rating (applies only when all dimensions PASS):
| Average | Rating |
|---|---|
| ≥8.0 | ROBUST |
| 6.0-7.9 | ACCEPTABLE |
| 5.0-5.9 | NEEDS_IMPROVEMENT |
Modifiers from Phase 3 RISK items: ≥3 RISK items → downgrade 1 level.
Per-Result Decision (when all dimensions PASS):
| Average | Decision |
|---|---|
| ≥7.0 | ACCEPT_AND_PROCEED |
| 5.0-6.9 | CONDITIONAL — ACCEPT with stated limitations |
When any dimension FAIL: the decision is always REVISE_AND_RETRY. The severity is graded by how many dimensions FAIL and the average score of passing dimensions:
| Failure pattern | Severity |
|---|---|
| 1 dimension FAIL, avg of others ≥7 | MODERATE — targeted revision on failed dimension |
| 1-2 dimensions FAIL, avg of others 5-6.9 | SIGNIFICANT — broader revision needed |
| ≥3 dimensions FAIL, or all passing dims <5 | CRITICAL — fundamental re-approach required |
| Item | YES | NO → RISK |
|---|---|---|
| Multiple testing / FDR | Method documented (Bonferroni, BH) | False positives likely |
| Model assumption verification | Tested with documented results | Model may be invalid |
| Confounder control | Known confounders included, justified | Spurious associations |
| Sample size / power | Power analysis conducted | Underpowered — false negatives |
| Batch effect / heterogeneity | Correction applied if multi-source | Batch confounded |
| Outlier / missing data | Strategy documented | Biased results |
| Reproducibility | Code provided, executable | Unverifiable results |
Domain priorities: Biology → batch, confounders, multiple testing; Chemistry → reproducibility, assumptions; Materials → sample size, uncertainty; Finance → assumptions, confounders, outlier handling.
For each NO: record RISK, assess severity (H/M/L), include in guidance if ≥MEDIUM.
Output evaluation results per the Output Schema below.
Every evaluation must produce the following structure:
{
"verdict": "ACCEPT_AND_PROCEED | CONDITIONAL | REVISE_AND_RETRY",
"severity": "MODERATE | SIGNIFICANT | CRITICAL",
"quality_rating": "ROBUST | ACCEPTABLE | NEEDS_IMPROVEMENT",
"source_reliability": "RELIABLE | PARTIALLY_RELIABLE | UNRELIABLE",
"dimension_scores": {
"accuracy": 0-10,
"completeness": 0-10,
"robustness": 0-10,
"relevance": 0-10,
"methodology": 0-10,
"critical_reflection": 0-10
},
"dimension_status": {
"accuracy": "PASS | FAIL",
"completeness": "PASS | FAIL",
"robustness": "PASS | FAIL",
"relevance": "PASS | FAIL",
"methodology": "PASS | FAIL",
"critical_reflection": "PASS | FAIL"
},
"risk_items": [
{"item": "description", "severity": "H | M | L"}
],
"revision_guidance": ["top 3 prioritized action items"],
"limitations": ["accepted weaknesses, if CONDITIONAL"]
}When presenting results to the user, format as a readable summary — not raw JSON. Highlight the verdict, failed dimensions (if any), and revision guidance (if REVISE).
| Domain | Key criteria | Score 9-10 | Score 0-3 |
|---|---|---|---|
| General — Data integrity | Missing values, duplicates, schema match | Clean data, transformations documented | Unchecked data quality |
| General — Calculation correctness | Formula verification, edge cases | Verified with test cases, edge cases handled | Unverified formulas |
| General — Output clarity | Labels, units, formatting | Clear labels, correct units, formatted tables | Ambiguous labels, missing units |
| Biology — Design validity | Controls, randomization, blinding | Proper controls + blinding documented | No controls |
| Biology — Statistical significance | p-values, correction, effect size | Corrected p-values + effect sizes + CI | Uncorrected only |
| Biology — Reproducibility | Protocol + code + data | Full protocol + code + raw data | No protocol, no code |
| Biology — Clinical relevance | Translational applicability | Clear relevance with limitations | Overgeneralized |
| Chemistry — Reaction reproducibility | Conditions, yields | Full conditions + error margins | Incomplete conditions |
| Chemistry — Characterization | Analytical methods coverage | NMR, XRD, MS, elemental all reported | Missing key methods |
| Chemistry — Computational validation | Theory-experiment agreement | Agreement within error, sensitivity tested | No comparison |
| Chemistry — Safety | Hazards, scalability | Safety documented, scalability assessed | No safety info |
| Materials — Measurement rigor | Standards, uncertainty | ASTM/ISO standards, uncertainty reported | Ad-hoc, no uncertainty |
| Materials — Sample prep | Reproducible synthesis, batch tracking | Reproducible with batch tracking | Single batch, no docs |
| Materials — Structure-property | Causal mechanism | Mechanistic link validated | Correlation without mechanism |
| Materials — Engineering applicability | Real-world constraints | Practical limits + failure modes assessed | Ideal conditions only |
| Finance — Risk-adjusted returns | Sharpe, drawdown, tail risk | Full risk metrics + tail risk | Raw returns only |
| Finance — Assumption validity | Distributional assumptions | Tested + regime detection | Assumed normality |
| Finance — Backtesting integrity | Out-of-sample, no leakage | Clean OOS, no data leakage | In-sample only, lookahead |
| Finance — Market microstructure | Costs, liquidity, slippage | Costs modeled, liquidity noted | Infinite liquidity assumed |
| Failure mode | Recovery |
|---|---|
| Results format mismatch | Attempt parse, mark CONDITIONAL, request re-format if REVISE |
| Criteria missing/vague | Infer from context, document as inferred |
| UNRELIABLE rating but ACCEPT | Flag contradiction: "ACCEPTED BUT RATED UNRELIABLE — verify" |
| No results to evaluate | Mark as evaluation failure |
| No code / unverifiable | Flag reproducibility NO with RISK, downgrade 1 level |
| Results irrelevant | Score Relevance ≤2 → FAIL → mandatory REVISE |
This skill works best in combination with code-engineer — load both for analysis tasks that require quality assurance. The typical workflow:
code-engineer performs the analysis and presents a Result Packageresult-evaluator evaluates the Result Package against quality criteriacode-engineer for re-analysisAnalysis task: "Correlate X and Y in dataset.csv and test statistical significance."
Structured data (--output-file JSON):
[
{"metric": "Pearson_r", "value": 0.8234},
{"metric": "p_value", "value": 0.0003},
{"metric": "sample_size", "value": 150}
]Methodology documentation (from conversation):
scipy.stats.pearsonr, pandasData traceability:
dataset.csv, columns X (float64, 148 non-null) and Y (float64, 150 non-null), 150 rows totalAnalysis code:
import pandas as pd
from scipy import stats
data = pd.read_csv('dataset.csv')
corr, p_value = stats.pearsonr(data['X'], data['Y'])
print(f"Pearson correlation: r={corr:.4f}, p={p_value:.6f}")Phase 1 — Criterion Alignment: Criteria inferred from task: statistical significance, method validity, data coverage. All three results map to criteria; no unmapped results or uncovered criteria.
Phase 2 — Source Reliability Hard Gate (computational-type):
pearsonr, which matches stated methoddata['X'] and data['Y'] from claimed filePhase 2 — Unified Evaluation Matrix:
| Dimension | Score | PASS/FAIL |
|---|---|---|
| Accuracy | 9 — r and p correct per method, sample size accurate | PASS |
| Completeness | 7 — includes r, p, n; missing confidence interval for r | PASS |
| Robustness | 6 — no normality test on X/Y before Pearson, 2 missing values in X not explained | PASS |
| Relevance | 9 — directly answers the correlation+significance question | PASS |
| Methodology | 7 — method justified (Pearson for continuous), but no assumption verification documented | PASS |
| Critical reflection | 5 — assumptions stated (continuous, linear) but limitations (outliers, non-linearity) not discussed | PASS |
All dimensions PASS → proceed to quality rating. Average: (9+7+6+9+7+5)/6 = 6.5 → ACCEPTABLE
Phase 3 — Statistical Methodology Audit:
RISK items: 3 (2 MEDIUM + 1 LOW). Modifier: ≥3 RISK → downgrade 1 level. ACCEPTABLE → NEEDS_IMPROVEMENT.
Phase 4 — Overall Assessment:
Verdict: REVISE_AND_RETRY
Severity: MODERATE
Quality Rating: NEEDS_IMPROVEMENT
Source Reliability: RELIABLE
Dimension Scores: Accuracy 9, Completeness 7, Robustness 6,
Relevance 9, Methodology 7, Critical Reflection 5
All dimensions: PASS
RISK Items:
- Model assumption verification: MEDIUM (normality not tested)
- Confounder control: MEDIUM (no confounders considered)
- Outlier/missing data: LOW (missing values not addressed)
Revision Guidance:
1. Test normality of X/Y before Pearson; use Spearman if non-normal
2. Document strategy for 2 missing values in X (drop or impute)
3. Identify and document potential confounders
Accepted Limitations: (none — REVISE)© openJiuwen-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/result-evaluator of openJiuwen-ai/sciencediscovery.
Open the folder on GitHubat commit ab1403f
Result Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Result Evaluator this skillopenJiuwen-ai/sciencediscovery | 159 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | |
| AI Agent Evaluation Benchmarkingsickn33/agentic-awesome-skills | 47k | 1 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 1 repos | ~8.1k | Automated safety check: Notes | MIT | |
| EvaluatorsArize-ai/phoenix | 12k | — | ~1.7k | Automated safety check: Pass | Custom licence | |
| Vss Evaluate Caption AccuracyNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | |
| LLM Evaluationdavila7/claude-code-templates | 33k | 12 repos | ~3.5k | Automated safety check: Pass | MIT |
sickn33/agentic-awesome-skills
Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Arize-ai/phoenix
Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.
NVIDIA-AI-Blueprints/video-search-and-summarization
Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
sickn33/agentic-awesome-skills
Evaluate agent behavior with versioned cases and explicit verifiers.
openJiuwen-ai/sciencediscovery
A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…
openJiuwen-ai/sciencediscovery
Operate GitCode issues, PRs, wikis, code/MR refs, and cached org templates.
openJiuwen-ai/sciencediscovery
Inspect a local PDB structure, summarize chains and residue composition, and identify protein atoms near a user-specified ligand or pocket center.
openJiuwen-ai/sciencediscovery
Prepare, launch, monitor, and summarize the real RFdiffusion to ProteinMPNN to Protenix antibody pipeline on a local or remote ScienceDiscovery Runner with sandboxed Ascend NPUs.
openJiuwen-ai/sciencediscovery
A skill your agent uses to orchestrate a multi-domain research team for literature/evidence research and data analysis.
openJiuwen-ai/sciencediscovery
A skill your agent uses when a research workflow needs verified academic source retrieval through literature-search MCP interfaces available in the current session before evidence extraction.
Evaluate analysis results for quality and reliability. An agent skill from openJiuwen-ai/sciencediscovery. Result Evaluator is an agent skill from openJiuwen-ai/sciencediscovery. Evaluate analysis results for quality and reliability.
Run `npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a claude-code`. Or copy the skill folder (skills/result-evaluator in openJiuwen-ai/sciencediscovery) into .claude/skills/result-evaluator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a codex`. Or copy the skill folder (skills/result-evaluator in openJiuwen-ai/sciencediscovery) into .agents/skills/result-evaluator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/result-evaluator, .gemini/skills/result-evaluator, .github/skills/result-evaluator and .opencode/skills/result-evaluator in your project.
Going by SKILL.md and its folder, Result Evaluator needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: pypi.tuna.tsinghua.edu.cn; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Result Evaluator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Result Evaluator: AI Agent Evaluation Benchmarking (sickn33/agentic-awesome-skills, 47k stars), Arize Evaluator (github/awesome-copilot, 40k stars), Evaluators (Arize-ai/phoenix, 12k stars) and Vss Evaluate Caption Accuracy (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
openJiuwen-ai (a GitHub organization) maintains it in openJiuwen-ai/sciencediscovery, which has 159 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 10, 2026.
Source: openJiuwen-ai/sciencediscovery on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.