Agent skill

Result Evaluator

by openJiuwen-ai in openJiuwen-ai/sciencediscovery

Evaluate analysis results for quality and reliability. An agent skill from openJiuwen-ai/sciencediscovery.

Apache-2.0Auto-check passed

Install Result Evaluator

skills CLI
$ npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openJiuwen-ai/sciencediscovery result-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/result-evaluator .claude/skills/result-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
result-evaluator
GitHub stars
159
Token cost
~4.4k tokens
SKILL.md length
1,777 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
Apache-2.0

At a glance

Evaluate analysis results for quality and reliability. An agent skill from openJiuwen-ai/sciencediscovery.

  • Works in 3 steps: Understand Evaluation Input → Execute 4-Phase Evaluation Protocol → Document Results
  • SKILL.md covers Overview, When to Use This Skill, Input Sources and Python Package Installation, plus 6 more sections
  • Calls pip; reaches pypi.tuna.tsinghua.edu.cn

What it does

Result Evaluator is an agent skill from openJiuwen-ai/sciencediscovery. Evaluate analysis results for quality and reliability. Scores Accuracy, Completeness, Robustness, and Relevance (0-10), checks source reliability and methodology, audits statistical rigor, and decides ACCEPTANDPROCEED or REVISEANDRETRY. NOT for performing analysis or modifying results.

Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: ScienceDiscovery is an all‑in‑one agentic workbench built specifically for scientific research. The licence is Apache-2.0.

Example prompts

  • “/result-evaluator”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Understand Evaluation Input
  2. Execute 4-Phase Evaluation Protocol
  3. Document Results

What it can do on your machine

Read from SKILL.md and the folder at commit ab1403f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • pypi.tuna.tsinghua.edu.cn

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Result Evaluator loads about 4.4k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 1,777 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openJiuwen-ai/sciencediscovery at commit ab1403f, republished under its Apache-2.0 licence (© openJiuwen-ai). 1,777 words, ~4,442 tokens.

Download SKILL.mdSave it as .claude/skills/result-evaluator/SKILL.md (or your agent's skills folder).
name
result-evaluator
description
Evaluate analysis results for quality and reliability. Scores Accuracy, Completeness, Robustness, and Relevance (0-10), checks source reliability and methodology, audits statistical rigor, and decides ACCEPT_AND_PROCEED or REVISE_AND_RETRY. NOT for performing analysis or modifying results.

Result Evaluator Skill

Overview

This skill evaluates analysis results against predefined criteria and decides ACCEPT_AND_PROCEED or REVISE_AND_RETRY. It follows a 4-phase protocol: criterion alignment → multi-dimensional evaluation with Source Reliability hard gate → statistical methodology audit → overall assessment. Hallucination detected → immediate REVISE; any checklist dimension FAIL → mandatory REVISE (hard gate, overrides scoring).

When to Use This Skill

Always load this skill when:

  • User asks to evaluate, audit, score, or quality-check analysis results that another skill — typically code-engineer — has just produced as a Result Package
  • User asks for an explicit ACCEPT_AND_PROCEED vs REVISE_AND_RETRY (or CONDITIONAL) decision before the results are used downstream (e.g. fed into a report, shared with stakeholders, or acted on)
  • User wants a Source Reliability check on computational or research-style results — to detect hallucinated numbers, fabricated statistics, invented citations, or code–data misalignment
  • User asks for a Statistical Methodology Audit covering multiple-testing correction, model-assumption verification, confounder control, sample-size/power, batch effects, outlier/missing-data handling, and reproducibility
  • User requests the multi-dimensional quality score (Accuracy / Completeness / Robustness / Relevance / Methodology / Critical Reflection, each 0–10) on a Result Package
  • User wants to know whether a result is reproducible from the supplied code and data, or whether the analysis should be re-run before being trusted

Input Sources

This skill evaluates analysis results with methodology documentation. Accepted input formats:

From code-engineer (recommended upstream skill):

  • Structured data: --output-file JSON ([{col: val, ...}]) or CSV/MD export — provides the numerical/tabular results
  • Methodology documentation: presented in conversation by the agent — includes libraries, statistical methods, method justification
  • Data traceability: source file names, sheet/column names, row counts, transformations applied
  • Analysis code: the complete code used to produce results (for reproducibility audit in Phase 3)

From other sources: any structured results with accompanying methodology description. Minimum required: results data + method description + data source identification.

If methodology documentation or data traceability is missing, note the gap in evaluation and flag Source Reliability as PARTIALLY_RELIABLE.

Python Package Installation

If you need to install new Python packages, install them through the Tsinghua PyPI mirror for reliability:

bash
pip install [python package] -i https://pypi.tuna.tsinghua.edu.cn/simple

Workflow

Step 1: Understand Evaluation Input

Identify the evaluation context:

  • Analysis results to evaluate: Structured output from code execution
  • Methodology documentation: How the results were produced (libraries, methods, code)
  • Data traceability: Source data identification (file names, column names, row counts)
  • Evaluation criteria: What aspects to evaluate and expected quality thresholds. Infer from context if missing (note limitation).
  • Analysis plan context: Domain, background information

Prerequisites: results must be available and parseable; methodology documentation and data traceability should be provided (evaluation quality degrades without them); criteria must be specified or inferable.

Step 2: Execute 4-Phase Evaluation Protocol
Phase 1 — Criterion Alignment

Map each result to an evaluation criterion. Flag UNMAPPED results and uncovered criteria. Infer criteria from context if missing (document as inferred).

Phase 2 — Per-Result Evaluation

Source Reliability Hard Gate (check first — hallucination → immediate REVISE_AND_RETRY, skip rest):

For computational-type results (from code-engineer and similar tools):

CheckWhat to detect
Data traceabilityCited data sources exist (file/sheet/column match actual data, row counts consistent)
Method consistencyStated methods match the actual code implementation
FabricationInvented statistics, untraceable numbers, results that cannot be reproduced from given code and data
Code-data alignmentCode actually references the claimed data files/variables, not different ones

For research-type results (literature-based, citing external references):

CheckWhat to detect
Data traceabilityCited data sources exist (file/sheet/field match actual data)
Reference validityCitations have author+year+DOI/PubMed (not "studies show")
Identifier authenticityStandard entity/gene/protein names (not self-created)
Method consistencyStated methods match implementation
FabricationInvented statistics, fake references, untraceable results

Verdict: RELIABLE (PASS) / PARTIALLY_RELIABLE (FAIL, continue) / UNRELIABLE (REVISE, stop).

Unified Evaluation Matrix — score each dimension 0-10; each dimension also has a PASS/FAIL threshold (score ≥5 → PASS, score <5 → FAIL):

Dimension9-107-84-60-3PASS threshold
AccuracyCorrect, methods matchMinor errorsSignificant errorsFundamental errors≥5
CompletenessComplete, no gapsMinor gapsSignificant gapsMajor omissions≥5
RobustnessSound methods, assumptions verified1-2 concerns3-4 issuesInvalid methods≥5
RelevanceDirectly addresses questionMostly relevantPartially relevantIrrelevant≥5
MethodologyJustified, rigorous, reproducibleAdequate justificationWeak justificationNo justification≥5
Critical reflectionAssumptions stated, limitations discussedSome reflectionMinimal reflectionNo reflection≥5

Hard Gate Rule: any dimension FAIL (score <5) → mandatory REVISE_AND_RETRY, regardless of the average score. The scoring average determines the severity grading of the REVISE decision, not whether to REVISE.

Composite Quality Rating (applies only when all dimensions PASS):

AverageRating
≥8.0ROBUST
6.0-7.9ACCEPTABLE
5.0-5.9NEEDS_IMPROVEMENT

Modifiers from Phase 3 RISK items: ≥3 RISK items → downgrade 1 level.

Per-Result Decision (when all dimensions PASS):

AverageDecision
≥7.0ACCEPT_AND_PROCEED
5.0-6.9CONDITIONAL — ACCEPT with stated limitations

When any dimension FAIL: the decision is always REVISE_AND_RETRY. The severity is graded by how many dimensions FAIL and the average score of passing dimensions:

Failure patternSeverity
1 dimension FAIL, avg of others ≥7MODERATE — targeted revision on failed dimension
1-2 dimensions FAIL, avg of others 5-6.9SIGNIFICANT — broader revision needed
≥3 dimensions FAIL, or all passing dims <5CRITICAL — fundamental re-approach required
Phase 3 — Statistical Methodology Quality Audit
ItemYESNO → RISK
Multiple testing / FDRMethod documented (Bonferroni, BH)False positives likely
Model assumption verificationTested with documented resultsModel may be invalid
Confounder controlKnown confounders included, justifiedSpurious associations
Sample size / powerPower analysis conductedUnderpowered — false negatives
Batch effect / heterogeneityCorrection applied if multi-sourceBatch confounded
Outlier / missing dataStrategy documentedBiased results
ReproducibilityCode provided, executableUnverifiable results

Domain priorities: Biology → batch, confounders, multiple testing; Chemistry → reproducibility, assumptions; Materials → sample size, uncertainty; Finance → assumptions, confounders, outlier handling.

For each NO: record RISK, assess severity (H/M/L), include in guidance if ≥MEDIUM.

Phase 4 — Overall Assessment
  1. Check Phase 2 hard gate: any dimension FAIL → REVISE_AND_RETRY (skip to step 4)
  2. If all PASS: compute average score → quality rating → apply RISK modifiers
  3. Final decision:
    • ROBUST → ACCEPT_AND_PROCEED
    • ACCEPTABLE → CONDITIONAL — ACCEPT with stated limitations
    • NEEDS_IMPROVEMENT → REVISE_AND_RETRY (MODERATE severity)
  4. If REVISE: prioritize guidance (FAIL dimensions > HIGH RISK > low passing scores), limit top 3 actionable items
Step 3: Document Results

Output evaluation results per the Output Schema below.

Output Schema

Every evaluation must produce the following structure:

json
{
  "verdict": "ACCEPT_AND_PROCEED | CONDITIONAL | REVISE_AND_RETRY",
  "severity": "MODERATE | SIGNIFICANT | CRITICAL",
  "quality_rating": "ROBUST | ACCEPTABLE | NEEDS_IMPROVEMENT",
  "source_reliability": "RELIABLE | PARTIALLY_RELIABLE | UNRELIABLE",
  "dimension_scores": {
    "accuracy": 0-10,
    "completeness": 0-10,
    "robustness": 0-10,
    "relevance": 0-10,
    "methodology": 0-10,
    "critical_reflection": 0-10
  },
  "dimension_status": {
    "accuracy": "PASS | FAIL",
    "completeness": "PASS | FAIL",
    "robustness": "PASS | FAIL",
    "relevance": "PASS | FAIL",
    "methodology": "PASS | FAIL",
    "critical_reflection": "PASS | FAIL"
  },
  "risk_items": [
    {"item": "description", "severity": "H | M | L"}
  ],
  "revision_guidance": ["top 3 prioritized action items"],
  "limitations": ["accepted weaknesses, if CONDITIONAL"]
}

When presenting results to the user, format as a readable summary — not raw JSON. Highlight the verdict, failed dimensions (if any), and revision guidance (if REVISE).

Show full SKILL.md (751 more words)Show less

Domain-Specific Evaluation Criteria

DomainKey criteriaScore 9-10Score 0-3
General — Data integrityMissing values, duplicates, schema matchClean data, transformations documentedUnchecked data quality
General — Calculation correctnessFormula verification, edge casesVerified with test cases, edge cases handledUnverified formulas
General — Output clarityLabels, units, formattingClear labels, correct units, formatted tablesAmbiguous labels, missing units
Biology — Design validityControls, randomization, blindingProper controls + blinding documentedNo controls
Biology — Statistical significancep-values, correction, effect sizeCorrected p-values + effect sizes + CIUncorrected only
Biology — ReproducibilityProtocol + code + dataFull protocol + code + raw dataNo protocol, no code
Biology — Clinical relevanceTranslational applicabilityClear relevance with limitationsOvergeneralized
Chemistry — Reaction reproducibilityConditions, yieldsFull conditions + error marginsIncomplete conditions
Chemistry — CharacterizationAnalytical methods coverageNMR, XRD, MS, elemental all reportedMissing key methods
Chemistry — Computational validationTheory-experiment agreementAgreement within error, sensitivity testedNo comparison
Chemistry — SafetyHazards, scalabilitySafety documented, scalability assessedNo safety info
Materials — Measurement rigorStandards, uncertaintyASTM/ISO standards, uncertainty reportedAd-hoc, no uncertainty
Materials — Sample prepReproducible synthesis, batch trackingReproducible with batch trackingSingle batch, no docs
Materials — Structure-propertyCausal mechanismMechanistic link validatedCorrelation without mechanism
Materials — Engineering applicabilityReal-world constraintsPractical limits + failure modes assessedIdeal conditions only
Finance — Risk-adjusted returnsSharpe, drawdown, tail riskFull risk metrics + tail riskRaw returns only
Finance — Assumption validityDistributional assumptionsTested + regime detectionAssumed normality
Finance — Backtesting integrityOut-of-sample, no leakageClean OOS, no data leakageIn-sample only, lookahead
Finance — Market microstructureCosts, liquidity, slippageCosts modeled, liquidity notedInfinite liquidity assumed

Error Handling

Failure modeRecovery
Results format mismatchAttempt parse, mark CONDITIONAL, request re-format if REVISE
Criteria missing/vagueInfer from context, document as inferred
UNRELIABLE rating but ACCEPTFlag contradiction: "ACCEPTED BUT RATED UNRELIABLE — verify"
No results to evaluateMark as evaluation failure
No code / unverifiableFlag reproducibility NO with RISK, downgrade 1 level
Results irrelevantScore Relevance ≤2 → FAIL → mandatory REVISE

Skill Pairing

This skill works best in combination with code-engineer — load both for analysis tasks that require quality assurance. The typical workflow:

  1. code-engineer performs the analysis and presents a Result Package
  2. result-evaluator evaluates the Result Package against quality criteria
  3. If REVISE_AND_RETRY: feed guidance back to code-engineer for re-analysis

Complete Example

Input (Result Package from code-engineer)

Analysis task: "Correlate X and Y in dataset.csv and test statistical significance."

Structured data (--output-file JSON):

json
[
  {"metric": "Pearson_r", "value": 0.8234},
  {"metric": "p_value", "value": 0.0003},
  {"metric": "sample_size", "value": 150}
]

Methodology documentation (from conversation):

  • Libraries: scipy.stats.pearsonr, pandas
  • Method: Pearson correlation with two-tailed test
  • Justification: X and Y are continuous variables, Pearson is appropriate for linear association

Data traceability:

  • Source: dataset.csv, columns X (float64, 148 non-null) and Y (float64, 150 non-null), 150 rows total

Analysis code:

python
import pandas as pd
from scipy import stats
data = pd.read_csv('dataset.csv')
corr, p_value = stats.pearsonr(data['X'], data['Y'])
print(f"Pearson correlation: r={corr:.4f}, p={p_value:.6f}")
Evaluation Walkthrough

Phase 1 — Criterion Alignment: Criteria inferred from task: statistical significance, method validity, data coverage. All three results map to criteria; no unmapped results or uncovered criteria.

Phase 2 — Source Reliability Hard Gate (computational-type):

  • Data traceability: PASS — columns X and Y exist in dataset.csv, 150 rows matches code
  • Method consistency: PASS — code uses pearsonr, which matches stated method
  • Fabrication: PASS — r=0.8234 and p=0.0003 are reproducible from given code and data
  • Code-data alignment: PASS — code references data['X'] and data['Y'] from claimed file
  • Verdict: RELIABLE

Phase 2 — Unified Evaluation Matrix:

DimensionScorePASS/FAIL
Accuracy9 — r and p correct per method, sample size accuratePASS
Completeness7 — includes r, p, n; missing confidence interval for rPASS
Robustness6 — no normality test on X/Y before Pearson, 2 missing values in X not explainedPASS
Relevance9 — directly answers the correlation+significance questionPASS
Methodology7 — method justified (Pearson for continuous), but no assumption verification documentedPASS
Critical reflection5 — assumptions stated (continuous, linear) but limitations (outliers, non-linearity) not discussedPASS

All dimensions PASS → proceed to quality rating. Average: (9+7+6+9+7+5)/6 = 6.5 → ACCEPTABLE

Phase 3 — Statistical Methodology Audit:

  • Multiple testing: N/A (single test) → skip
  • Model assumption verification: NO — normality of X/Y not tested → RISK (MEDIUM)
  • Confounder control: NO — no confounders considered → RISK (MEDIUM)
  • Sample size/power: YES — n=150, effect size r=0.82 provides adequate power
  • Outlier/missing data: NO — 2 missing in X not addressed → RISK (LOW)
  • Reproducibility: YES — full code provided

RISK items: 3 (2 MEDIUM + 1 LOW). Modifier: ≥3 RISK → downgrade 1 level. ACCEPTABLE → NEEDS_IMPROVEMENT.

Phase 4 — Overall Assessment:

  • All dimensions PASS, but 3 RISK items downgrade from ACCEPTABLE to NEEDS_IMPROVEMENT
  • Decision: REVISE_AND_RETRY (MODERATE severity — one dimension at 5, RISK items addressable)
  • Revision guidance (top 3, prioritized):
    1. Test normality of X and Y before using Pearson (assumption verification)
    2. Address 2 missing values in column X (outlier/missing data strategy)
    3. Consider potential confounders and document them
Output
Verdict: REVISE_AND_RETRY
Severity: MODERATE
Quality Rating: NEEDS_IMPROVEMENT
Source Reliability: RELIABLE

Dimension Scores: Accuracy 9, Completeness 7, Robustness 6,
  Relevance 9, Methodology 7, Critical Reflection 5
All dimensions: PASS

RISK Items:
  - Model assumption verification: MEDIUM (normality not tested)
  - Confounder control: MEDIUM (no confounders considered)
  - Outlier/missing data: LOW (missing values not addressed)

Revision Guidance:
  1. Test normality of X/Y before Pearson; use Spearman if non-normal
  2. Document strategy for 2 missing values in X (drop or impute)
  3. Identify and document potential confounders

Accepted Limitations: (none — REVISE)

© openJiuwen-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/result-evaluator of openJiuwen-ai/sciencediscovery.

Open the folder on GitHubat commit ab1403f

Compare with similar skills

Result Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Result Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Result Evaluator this skillopenJiuwen-ai/sciencediscovery159—~4.4kAutomated safety check: PassApache-2.0
AI Agent Evaluation Benchmarkingsickn33/agentic-awesome-skills47k1 repos~1.3kAutomated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k1 repos~8.1kAutomated safety check: NotesMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence
Vss Evaluate Caption AccuracyNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~2.1kAutomated safety check: NotesApache-2.0
LLM Evaluationdavila7/claude-code-templates33k12 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • AI Agent Evaluation Benchmarking

    sickn33/agentic-awesome-skills

    Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.

    47k GitHub starsUsed in 1 repo~1.3k tokens
    Agent WorkflowsAuto-check passed
  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated yesterday
    EducationAuto-check passed
  • Vss Evaluate Caption Accuracy

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

    1.9k GitHub stars~2.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    33k GitHub starsUsed in 12 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed

More from openJiuwen-ai/sciencediscovery

All 23 skills in this repo
  • Code Engineer

    openJiuwen-ai/sciencediscovery

    A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…

    162 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Gitcode

    openJiuwen-ai/sciencediscovery

    Operate GitCode issues, PRs, wikis, code/MR refs, and cached org templates.

    162 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Structure Pocket Inspection

    openJiuwen-ai/sciencediscovery

    Inspect a local PDB structure, summarize chains and residue composition, and identify protein atoms near a user-specified ligand or pocket center.

    162 GitHub stars~624 tokensUpdated today
    Auto-check passed
  • Antibody Design

    openJiuwen-ai/sciencediscovery

    Prepare, launch, monitor, and summarize the real RFdiffusion to ProteinMPNN to Protenix antibody pipeline on a local or remote ScienceDiscovery Runner with sandboxed Ascend NPUs.

    162 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Science Research Team

    openJiuwen-ai/sciencediscovery

    A skill your agent uses to orchestrate a multi-domain research team for literature/evidence research and data analysis.

    162 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Literature Searcher

    openJiuwen-ai/sciencediscovery

    A skill your agent uses when a research workflow needs verified academic source retrieval through literature-search MCP interfaces available in the current session before evidence extraction.

    162 GitHub stars~5.8k tokensUpdated today
    Auto-check passed

Questions about Result Evaluator

What does Result Evaluator do?

Evaluate analysis results for quality and reliability. An agent skill from openJiuwen-ai/sciencediscovery. Result Evaluator is an agent skill from openJiuwen-ai/sciencediscovery. Evaluate analysis results for quality and reliability.

How do I install Result Evaluator in Claude Code?

Run `npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a claude-code`. Or copy the skill folder (skills/result-evaluator in openJiuwen-ai/sciencediscovery) into .claude/skills/result-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Result Evaluator in Codex?

Run `npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a codex`. Or copy the skill folder (skills/result-evaluator in openJiuwen-ai/sciencediscovery) into .agents/skills/result-evaluator in your project. Codex loads it when a task matches its description.

Can I use Result Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openJiuwen-ai/sciencediscovery --skill result-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/result-evaluator, .gemini/skills/result-evaluator, .github/skills/result-evaluator and .opencode/skills/result-evaluator in your project.

What does Result Evaluator need to run?

Going by SKILL.md and its folder, Result Evaluator needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Result Evaluator access the network?

SKILL.md names 1 domain. In commands or code: pypi.tuna.tsinghua.edu.cn; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Result Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Result Evaluator use?

Result Evaluator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Result Evaluator use?

About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Result Evaluator?

Skills that share tags, products or a category with Result Evaluator: AI Agent Evaluation Benchmarking (sickn33/agentic-awesome-skills, 47k stars), Arize Evaluator (github/awesome-copilot, 40k stars), Evaluators (Arize-ai/phoenix, 12k stars) and Vss Evaluate Caption Accuracy (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Result Evaluator?

openJiuwen-ai (a GitHub organization) maintains it in openJiuwen-ai/sciencediscovery, which has 159 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 10, 2026.

Source: openJiuwen-ai/sciencediscovery on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.