Agent skill

Scholar Evaluation

by spacering-net in spacering-net/codeg

Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and…

MITAuto-check passedResearch & Science

Install Scholar Evaluation

skills CLI
$ npx skills add spacering-net/codeg --skill scholar-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install spacering-net/codeg scholar-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/spacering-net/codeg.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src-tauri/science/skills/scholar-evaluation .claude/skills/scholar-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scholar-evaluation
GitHub stars
3.8k
Used in
12 other repos
Token cost
~3.2k tokens
SKILL.md length
1,335 words
Files
5 (incl. scripts, references)
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and…

  • Works in 6 steps: Initial Assessment and Scope Definition → Dimension-Based Evaluation → Scoring and Rating → …
  • Tasks that involve Peer review
  • SKILL.md covers Overview, When to Use This Skill, Visual Enhancement with… and Evaluation Workflow, plus 6 more sections
  • Runs Python scripts from its folder; calls python

What it does

Scholar Evaluation is an agent skill from spacering-net/codeg. Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and writing with quantitative scoring and actionable feedback.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/evaluation_framework.md`, `scripts/calculate_scores.py` and `scripts/generate_schematic.py`).

It sits in Research & Science, covering Peer review and Data visualization. The repository describes itself as: Collaborative multi-agent AI coding workspace: aggregate sessions from Claude Code, Codex, OpenCode, Pi, Grok Build, etc. Desktop app, self-hosted server, or Docker. The licence is MIT.

When your agent uses it

  • Tasks that involve Peer review
  • Tasks that involve Data visualization

Example prompts

  • “/scholar-evaluation”

Requirements

  • Python 3
  • A credential in OPENROUTER_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Initial Assessment and Scope Definition
  2. Dimension-Based Evaluation
  3. Scoring and Rating
  4. Synthesize Overall Assessment
  5. Provide Actionable Feedback
  6. Contextual Considerations

What it can do on your machine

Read from SKILL.md and the folder at commit 592131c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scholar Evaluation loads about 3.2k tokens when it runs, and up to ~8.1k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 1,335 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from spacering-net/codeg at commit 592131c, republished under its MIT licence (© spacering-net). 1,335 words, ~3,159 tokens.

Download SKILL.mdSave it as .claude/skills/scholar-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
scholar-evaluation
description
Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and writing with quantitative scoring and actionable feedback.
license
MIT license
metadata.version
1.1
metadata.skill-author
K-Dense Inc.

Scholar Evaluation

Overview

Apply the ScholarEval framework to systematically evaluate scholarly and research work. This skill provides structured evaluation methodology based on peer-reviewed research assessment criteria, enabling comprehensive analysis of academic papers, research proposals, literature reviews, and scholarly writing across multiple quality dimensions.

When to Use This Skill

Use this skill when:

  • Evaluating research papers for quality and rigor
  • Assessing literature review comprehensiveness and quality
  • Reviewing research methodology design
  • Scoring data analysis approaches
  • Evaluating scholarly writing and presentation
  • Providing structured feedback on academic work
  • Benchmarking research quality against established criteria
  • Assessing publication readiness for target venues
  • Providing quantitative evaluation to complement qualitative peer review

Visual Enhancement with Scientific Schematics

When creating documents with this skill, always consider adding scientific diagrams and schematics to enhance visual communication.

If your document does not already contain schematics or diagrams:

  • Use the scientific-schematics skill to generate AI-powered publication-quality diagrams
  • Simply describe your desired diagram in natural language
  • Nano Banana Pro will automatically generate, review, and refine the schematic

For new documents: Scientific schematics should be generated by default to visually represent key concepts, workflows, architectures, or relationships described in the text.

How to generate schematics:

bash
python scripts/generate_schematic.py "your diagram description" -o figures/output.png

The AI will automatically:

  • Create publication-quality images with proper formatting
  • Review and refine through multiple iterations
  • Ensure accessibility (colorblind-friendly, high contrast)
  • Save outputs in the figures/ directory

When to add schematics:

  • Evaluation framework diagrams
  • Quality assessment criteria decision trees
  • Scholarly workflow visualizations
  • Assessment methodology flowcharts
  • Scoring rubric visualizations
  • Evaluation process diagrams
  • Any complex concept that benefits from visualization

For detailed guidance on creating schematics, refer to the scientific-schematics skill documentation.


Evaluation Workflow

Step 1: Initial Assessment and Scope Definition

Begin by identifying the type of scholarly work being evaluated and the evaluation scope:

Work Types:

  • Full research paper (empirical, theoretical, or review)
  • Research proposal or protocol
  • Literature review (systematic, narrative, or scoping)
  • Thesis or dissertation chapter
  • Conference abstract or short paper

Evaluation Scope:

  • Comprehensive (all dimensions)
  • Targeted (specific aspects like methodology or writing)
  • Comparative (benchmarking against other work)

Ask the user to clarify if the scope is ambiguous.

Step 2: Dimension-Based Evaluation

Systematically evaluate the work across the ScholarEval dimensions. For each applicable dimension, assess quality, identify strengths and weaknesses, and provide scores where appropriate.

Refer to references/evaluation_framework.md for detailed criteria and rubrics for each dimension.

Core Evaluation Dimensions:

  1. Problem Formulation & Research Questions

    • Clarity and specificity of research questions
    • Theoretical or practical significance
    • Feasibility and scope appropriateness
    • Novelty and contribution potential
  2. Literature Review

    • Comprehensiveness of coverage
    • Critical synthesis vs. mere summarization
    • Identification of research gaps
    • Currency and relevance of sources
    • Proper contextualization
  3. Methodology & Research Design

    • Appropriateness for research questions
    • Rigor and validity
    • Reproducibility and transparency
    • Ethical considerations
    • Limitations acknowledgment
  4. Data Collection & Sources

    • Quality and appropriateness of data
    • Sample size and representativeness
    • Data collection procedures
    • Source credibility and reliability
  5. Analysis & Interpretation

    • Appropriateness of analytical methods
    • Rigor of analysis
    • Logical coherence
    • Alternative explanations considered
    • Results-claims alignment
  6. Results & Findings

    • Clarity of presentation
    • Statistical or qualitative rigor
    • Visualization quality
    • Interpretation accuracy
    • Implications discussion
  7. Scholarly Writing & Presentation

    • Clarity and organization
    • Academic tone and style
    • Grammar and mechanics
    • Logical flow
    • Accessibility to target audience
  8. Citations & References

    • Citation completeness
    • Source quality and appropriateness
    • Citation accuracy
    • Balance of perspectives
    • Adherence to citation standards
Step 3: Scoring and Rating

For each evaluated dimension, provide:

Qualitative Assessment:

  • Key strengths (2-3 specific points)
  • Areas for improvement (2-3 specific points)
  • Critical issues (if any)

Quantitative Scoring (Optional): Use a 5-point scale where applicable:

  • 5: Excellent - Exemplary quality, publishable in top venues
  • 4: Good - Strong quality with minor improvements needed
  • 3: Adequate - Acceptable quality with notable areas for improvement
  • 2: Needs Improvement - Significant revisions required
  • 1: Poor - Fundamental issues requiring major revision

To calculate aggregate scores programmatically, use scripts/calculate_scores.py.

Step 4: Synthesize Overall Assessment

Provide an integrated evaluation summary:

  1. Overall Quality Assessment - Holistic judgment of the work's scholarly merit
  2. Major Strengths - 3-5 key strengths across dimensions
  3. Critical Weaknesses - 3-5 primary areas requiring attention
  4. Priority Recommendations - Ranked list of improvements by impact
  5. Publication Readiness (if applicable) - Assessment of suitability for target venues
Step 5: Provide Actionable Feedback

Transform evaluation findings into constructive, actionable feedback:

Feedback Structure:

  • Specific - Reference exact sections, paragraphs, or page numbers
  • Actionable - Provide concrete suggestions for improvement
  • Prioritized - Rank recommendations by importance and feasibility
  • Balanced - Acknowledge strengths while addressing weaknesses
  • Evidence-based - Ground feedback in evaluation criteria

Feedback Format Options:

  • Structured report with dimension-by-dimension analysis
  • Annotated comments mapped to specific document sections
  • Executive summary with key findings and recommendations
  • Comparative analysis against benchmark standards
Step 6: Contextual Considerations

Adjust evaluation approach based on:

Stage of Development:

  • Early draft: Focus on conceptual and structural issues
  • Advanced draft: Focus on refinement and polish
  • Final submission: Comprehensive quality check

Purpose and Venue:

  • Journal article: High standards for rigor and contribution
  • Conference paper: Balance novelty with presentation clarity
  • Student work: Educational feedback with developmental focus
  • Grant proposal: Emphasis on feasibility and impact

Discipline-Specific Norms:

  • STEM fields: Emphasis on reproducibility and statistical rigor
  • Social sciences: Balance quantitative and qualitative standards
  • Humanities: Focus on argumentation and scholarly interpretation
Show full SKILL.md (501 more words)Show less

Resources

references/evaluation_framework.md

Detailed evaluation criteria, rubrics, and quality indicators for each ScholarEval dimension. Load this reference when conducting evaluations to access specific assessment guidelines and scoring rubrics.

Search patterns for quick access:

  • "Problem Formulation criteria"
  • "Literature Review rubric"
  • "Methodology assessment"
  • "Data quality indicators"
  • "Analysis rigor standards"
  • "Writing quality checklist"
scripts/calculate_scores.py

Python script for calculating aggregate evaluation scores from dimension-level ratings. Supports weighted averaging, threshold analysis, and score visualization.

Usage:

bash
python scripts/calculate_scores.py --scores <dimension_scores.json> --output <report.txt>

Best Practices

  1. Maintain Objectivity - Base evaluations on established criteria, not personal preferences
  2. Be Comprehensive - Evaluate all applicable dimensions systematically
  3. Provide Evidence - Support assessments with specific examples from the work
  4. Stay Constructive - Frame weaknesses as opportunities for improvement
  5. Consider Context - Adjust expectations based on work stage and purpose
  6. Document Rationale - Explain the reasoning behind assessments and scores
  7. Encourage Strengths - Explicitly acknowledge what the work does well
  8. Prioritize Feedback - Focus on high-impact improvements first

Example Evaluation Workflow

User Request: "Evaluate this research paper on machine learning for drug discovery"

Response Process:

  1. Identify work type (empirical research paper) and scope (comprehensive evaluation)
  2. Load references/evaluation_framework.md for detailed criteria
  3. Systematically assess each dimension:
    • Problem formulation: Clear research question about ML model performance
    • Literature review: Comprehensive coverage of recent ML and drug discovery work
    • Methodology: Appropriate deep learning architecture with validation procedures
    • [Continue through all dimensions...]
  4. Calculate dimension scores and overall assessment
  5. Synthesize findings into structured report highlighting:
    • Strong methodology and reproducible code
    • Needs more diverse dataset evaluation
    • Writing could improve clarity in results section
  6. Provide prioritized recommendations with specific suggestions

Integration with Scientific Writer

This skill integrates seamlessly with the scientific writer workflow:

After Paper Generation:

  • Use Scholar Evaluation as an alternative or complement to peer review
  • Generate SCHOLAR_EVALUATION.md alongside PEER_REVIEW.md
  • Provide quantitative scores to track improvement across revisions

During Revision:

  • Re-evaluate specific dimensions after addressing feedback
  • Track score improvements over multiple versions
  • Identify persistent weaknesses requiring attention

Publication Preparation:

  • Assess readiness for target journal/conference
  • Identify gaps before submission
  • Benchmark against publication standards

Notes

  • Evaluation rigor should match the work's purpose and stage
  • Some dimensions may not apply to all work types (e.g., data collection for purely theoretical papers)
  • Cultural and disciplinary differences in scholarly norms should be considered
  • This framework complements, not replaces, domain-specific expertise
  • Use in combination with peer-review skill for comprehensive assessment

Citation

This skill is based on the ScholarEval framework introduced in:

Moussa, H. N., Da Silva, P. Q., Adu-Ampratwum, D., East, A., Lu, Z., Puccetti, N., Xue, M., Sun, H., Majumder, B. P., & Kumar, S. (2025). ScholarEval: Research Idea Evaluation Grounded in Literature. arXiv preprint arXiv:2510.16234. https://arxiv.org/abs/2510.16234

Abstract: ScholarEval is a retrieval augmented evaluation framework that assesses research ideas based on two fundamental criteria: soundness (the empirical validity of proposed methods based on existing literature) and contribution (the degree of advancement made by the idea across different dimensions relative to prior research). The framework achieves significantly higher coverage of expert-annotated evaluation points and is consistently preferred over baseline systems in terms of evaluation actionability, depth, and evidence support.

© spacering-net, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in src-tauri/science/skills/scholar-evaluation of spacering-net/codeg.

  • SKILL.md
  • references/evaluation_framework.md
  • scripts/calculate_scores.py
  • scripts/generate_schematic.py
  • scripts/generate_schematic_ai.py

Open the folder on GitHubat commit 592131c

Used in 12 other repositories

We found 20 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 12 other GitHub owners. This page covers the copy in spacering-net/codeg, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Scholar Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scholar Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scholar Evaluation this skillspacering-net/codeg3.8k12 repos~3.2kAutomated safety check: PassMIT
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k23 repos~5.9kAutomated safety check: NotesMIT
Nature-Style Scientific FiguresYuan1z0825/nature-skills46k—~2.9kAutomated safety check: PassApache-2.0
Academic Paper Writing PipelineImbad0202/academic-research-skills51k—~16kAutomated safety check: PassCustom licence
LLM Counciltenfoldmarc/llm-council-skill8192 repos~4.2kAutomated safety check: PassNone
Academic Paper ReviewerImbad0202/academic-research-skills51k—~11kAutomated safety check: PassCustom licence

Similar skills

  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 23 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Nature-Style Scientific Figures

    Yuan1z0825/nature-skills

    Creates, revises, audits and exports manuscript-ready scientific figures in Python or R, and routes AI-generated graphical abstracts to a separate workflow.

    46k GitHub stars~2.9k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Academic Paper Writing Pipeline

    Imbad0202/academic-research-skills

    Runs a 12-agent pipeline that plans, drafts, cites, reviews and formats academic papers, with modes for revision, rebuttals, abstracts and citation checks.

    51k GitHub stars~16k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed
  • LLM Council

    tenfoldmarc/llm-council-skill

    Run any question, idea, or decision through a council of 5 AI advisors who independently analyze it, peer-review each other anonymously, and synthesize a final verdict.

    819 GitHub starsUsed in 2 repos~4.2k tokens
    Research & ScienceAuto-check passed
  • Academic Paper Reviewer

    Imbad0202/academic-research-skills

    Simulates a journal peer review of a manuscript with a five-seat reviewer panel, an editorial synthesizer and several review modes.

    51k GitHub stars~11k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed
  • Peer Review

    K-Dense-AI/claude-scientific-writer

    Prepare evidence-bounded, constructive peer-review drafts and structured manuscript assessments.

    2.4k GitHub starsUsed in 2 repos~3.1k tokens
    Research & ScienceAuto-check: notes

More from spacering-net/codeg

All 9 skills in this repo
  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Auto-check passed
  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Auto-check: notes
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.8k GitHub starsUsed in 18 repos~5.9k tokens
    Auto-check: notes
  • Statistical Analysis

    spacering-net/codeg

    Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting.

    3.8k GitHub starsUsed in 4 repos~5k tokens
    Auto-check passed
  • Statistical Power

    spacering-net/codeg

    Sample-size and statistical power calculations for planning studies.

    3.8k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check: notes
  • Scientific Schematics

    spacering-net/codeg

    Create publication-quality scientific diagrams using Nano Banana 2 AI with smart iterative refinement.

    3.8k GitHub starsUsed in 11 repos~5.9k tokens
    Auto-check: notes

Questions about Scholar Evaluation

What does Scholar Evaluation do?

Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and…. Scholar Evaluation is an agent skill from spacering-net/codeg. Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and writing with quantitative scoring and actionable feedback.

When should I use Scholar Evaluation?

Scholar Evaluation fits situations like: tasks that involve Peer review; tasks that involve Data visualization.

How do I install Scholar Evaluation in Claude Code?

Run `npx skills add spacering-net/codeg --skill scholar-evaluation -a claude-code`. Or copy the skill folder (src-tauri/science/skills/scholar-evaluation in spacering-net/codeg) into .claude/skills/scholar-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Scholar Evaluation in Codex?

Run `npx skills add spacering-net/codeg --skill scholar-evaluation -a codex`. Or copy the skill folder (src-tauri/science/skills/scholar-evaluation in spacering-net/codeg) into .agents/skills/scholar-evaluation in your project. Codex loads it when a task matches its description.

Can I use Scholar Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add spacering-net/codeg --skill scholar-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scholar-evaluation, .gemini/skills/scholar-evaluation, .github/skills/scholar-evaluation and .opencode/skills/scholar-evaluation in your project.

What does Scholar Evaluation need to run?

Going by SKILL.md and its folder, Scholar Evaluation needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3; A credential in OPENROUTER_API_KEY.

Does Scholar Evaluation access the network?

SKILL.md names 1 domain. As links in the text: arxiv.org. This is read from the text; nothing was executed.

Is Scholar Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Scholar Evaluation use?

Scholar Evaluation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scholar Evaluation use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5k tokens, read only when the agent opens those files.

What are the alternatives to Scholar Evaluation?

Skills that share tags, products or a category with Scholar Evaluation: Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Nature-Style Scientific Figures (Yuan1z0825/nature-skills, 46k stars), Academic Paper Writing Pipeline (Imbad0202/academic-research-skills, 51k stars) and LLM Council (tenfoldmarc/llm-council-skill, 819 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scholar Evaluation?

spacering-net (a GitHub organization) maintains it in spacering-net/codeg, which has 3,848 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 7, 2026.

Source: spacering-net/codeg on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.