Agent skill

General Fair Data Review

by learningmatter-mit in learningmatter-mit/AtomisticSkills

Review a manuscript or code repository for FAIR data compliance (Findable, Accessible, Interoperable, Reusable), producing a structured report with pass/fail per principle and actionable remediation…

MITAuto-check passedDocuments & Office

Install General Fair Data Review

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill general-fair-data-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills general-fair-data-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/general-fair-data-review .claude/skills/general-fair-data-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
general-fair-data-review
GitHub stars
176
Token cost
~2.3k tokens
SKILL.md length
630 words
Files
1
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Review a manuscript or code repository for FAIR data compliance (Findable, Accessible, Interoperable, Reusable), producing a structured report with pass/fail per principle and actionable remediation…

  • Works in 6 steps: Identify Review Scope → Read the Manuscript and Data… → Inspect the Data/Code Repository → …
  • Documents & Office work in your project
  • SKILL.md covers Goal, Prerequisites, Instructions and Document-Specific Workflows, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

General Fair Data Review is an agent skill from learningmatter-mit/AtomisticSkills. Review a manuscript or code repository for FAIR data compliance (Findable, Accessible, Interoperable, Reusable), producing a structured report with pass/fail per principle and actionable remediation steps.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office. The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

When your agent uses it

  • Documents & Office work in your project

Example prompts

  • “/general-fair-data-review”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Identify Review Scope
  2. Read the Manuscript and Data Availability Statement
  3. Inspect the Data/Code Repository
  4. Score Each FAIR Sub-Principle
  5. Atomistic/Computational Science Specific Checks
  6. Generate the Structured Review Report

What it can do on your machine

Read from SKILL.md and the folder at commit 6257444. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doi.org
    • rsc.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

General Fair Data Review loads about 2.3k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 630 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 6257444, republished under its MIT licence (© learningmatter-mit). 630 words, ~2,256 tokens.

Download SKILL.mdSave it as .claude/skills/general-fair-data-review/SKILL.md (or your agent's skills folder).
name
general-fair-data-review
description
Review a manuscript or code repository for FAIR data compliance (Findable, Accessible, Interoperable, Reusable), producing a structured report with pass/fail per principle and actionable remediation steps.
metadata.category
general

General FAIR Data Review

Goal

Assess whether a manuscript submission or standalone code/data repository satisfies the FAIR Guiding Principles (Wilkinson et al., Sci. Data 2016). The output is a structured reviewer report — analogous to a peer-review report — that scores each FAIR sub-principle, identifies gaps, and provides concrete remediation steps the authors can act on before publication.

This skill is complementary to general-peer-review, which focuses on scientific methodology. Run both in sequence for a complete review.


Prerequisites

  • A manuscript (PDF or markdown) and/or a link to a code/data repository (GitHub, Zenodo, Figshare, etc.)
  • Read access to any supplementary files, Data Availability Statements (DAS), or README files provided by the authors

Instructions

1. Identify Review Scope

Determine what artifacts are under review. Three modes exist:

ModeInputFocus
Manuscript + data/codePDF + repo URLFull FAIR review
Manuscript onlyPDFDAS quality, metadata richness, identifier presence
Code/data repo onlyRepo URL / directoryRepository-level FAIR compliance

State the mode explicitly at the start of the review report.


2. Read the Manuscript and Data Availability Statement

Load the manuscript. Locate and extract:

  • The Data Availability Statement (DAS) — usually a dedicated section near the end.
  • Any Code Availability Statement.
  • All data/code repository URLs or DOIs mentioned.

If no DAS exists, flag immediately as a Critical Finding (fails F4, A1, R1.1).


3. Inspect the Data/Code Repository

For each repository URL found, check the following. If no repository exists, mark all sub-principles below as Fail.

Repository inspection checklist:
- Does a persistent identifier (DOI, Handle) exist? → F1
- Is metadata present and rich (title, authors, description, keywords, license)? → F2, R1
- Does the metadata explicitly reference the dataset/code identifier? → F3
- Is the repository indexed in a searchable resource (Zenodo, Figshare, OSF, etc.)? → F4
- Can the data/code be accessed via a standard protocol (HTTP/HTTPS, FTP)? → A1
- Is the protocol open and free (no proprietary portal login required)? → A1.1
- If restricted, is there a documented access procedure? → A1.2
- Does metadata remain accessible even if data is removed? → A2
- Are standard, community-recognized formats used (CIF, JSON, CSV, HDF5, not .xlsx or proprietary)? → I1
- Are domain ontologies or controlled vocabularies used for metadata fields? → I2
- Are cross-references to related datasets or publications included? → I3
- Is a clear, machine-readable license present (CC-BY, MIT, Apache 2.0, etc.)? → R1.1
- Is provenance documented (how data was generated, software versions, parameters)? → R1.2
- Do files conform to domain community standards (e.g., CIF for crystal structures, SMILES for molecules, HDF5 for trajectories)? → R1.3

4. Score Each FAIR Sub-Principle

For every sub-principle (F1–F4, A1–A2, I1–I3, R1–R1.3) assign:

  • Pass — requirement fully met
  • Partial — requirement partially met; improvement needed
  • Fail — requirement not met or artifact absent
  • N/A — not applicable to this submission type

5. Atomistic/Computational Science Specific Checks

In addition to the generic FAIR checklist, evaluate the following domain-specific criteria:

Structures & Trajectories

  • Crystal structures deposited as .cif (not as images or in supplementary PDF tables)
  • MD trajectories deposited in open formats (.xyz, .extxyz, .h5md, .lammpsdump) with a README specifying units, timestep, ensemble
  • Force field / MLIP checkpoints deposited with version, training set provenance, and validation metrics

Computational Parameters

  • DFT: INCAR/POTCAR/KPOINTS or equivalent included or described with exact values (not "similar to ref. X")
  • MLIP: model architecture, training hyperparameters, and train/val/test split recorded
  • MD: timestep, thermostat/barostat settings, equilibration protocol documented

Software Environment

  • environment.yml or requirements.txt present with pinned versions
  • Scripts runnable from the deposited repository without undocumented external dependencies

Show full SKILL.md (242 more words)Show less
6. Generate the Structured Review Report

Produce a report in the following format:

markdown
# FAIR Data Review Report

**Manuscript title:** [title]
**Review date:** [date]
**Reviewer:** AI FAIR Data Reviewer (general-fair-data-review skill)
**Review mode:** [Manuscript + data/code | Manuscript only | Code/data repo only]

---

## Summary

[2–4 sentences: overall FAIRness level, most critical gaps, overall recommendation: Ready / Minor Revisions / Major Revisions / Not Acceptable]

---

## FAIR Scorecard

| Principle | Sub-principle | Status | Evidence / Gap |
|-----------|--------------|--------|----------------|
| **Findable** | F1: Persistent identifier | Pass/Partial/Fail | ... |
| | F2: Rich metadata | | |
| | F3: Metadata references data ID | | |
| | F4: Indexed in searchable resource | | |
| **Accessible** | A1: Retrievable via standard protocol | | |
| | A1.1: Protocol open and free | | |
| | A1.2: Auth procedure documented | | |
| | A2: Metadata accessible if data removed | | |
| **Interoperable** | I1: Formal/shared knowledge representation | | |
| | I2: FAIR vocabularies used | | |
| | I3: Qualified references to other data | | |
| **Reusable** | R1: Rich, accurate, relevant attributes | | |
| | R1.1: Clear data usage license | | |
| | R1.2: Detailed provenance | | |
| | R1.3: Domain community standards met | | |

---

## Major Concerns

[Number sequentially. For each: state issue → why problematic → actionable fix.]

1. **[Issue title]**
   - *Problem:* ...
   - *Impact:* ...
   - *Fix:* ...

---

## Minor Concerns

- ...

---

## Atomistic/Computational Specific Findings

[Report on structure formats, trajectory deposits, software environments, parameter completeness.]

---

## Questions for Authors

1. ...

---

## Recommended Repositories (if none provided)

If no repository was deposited, suggest domain-appropriate options:

| Data type | Recommended repository |
|-----------|----------------------|
| Crystal structures | CCDC, ICSD, Materials Cloud, Zenodo |
| Molecular dynamics trajectories | Materials Cloud, Zenodo, NOMAD |
| ML models / checkpoints | Hugging Face, Zenodo, MACE-Models |
| General datasets | Zenodo, Figshare, Dryad |
| Code | GitHub + Zenodo DOI via Zenodo GitHub integration |

Document-Specific Workflows

Manuscript Review (PDF)

[!WARNING] Read the PDF text directly. Do not assume data exists unless a DOI or repository URL is explicitly present in the manuscript body or supplement.

Steps:

  1. Extract DAS and Code Availability Statement verbatim.
  2. Resolve any DOIs or URLs found.
  3. If DOIs resolve to a live repository, proceed with repository inspection (Step 3).
  4. If only "data available upon request" is stated: flag as Fail for F1, F4, A1, R1.1 — this does not meet FAIR standards.
Code/Data Repository Review Only

Steps:

  1. Read top-level README.md.
  2. Check for LICENSE, environment.yml/requirements.txt, CITATION.cff.
  3. Inspect directory structure for data files and their formats.
  4. Check metadata on the repository platform (Zenodo record, GitHub About section, etc.).

Constraints

  • Scope: Focus strictly on data/code FAIRness. Scientific methodology critique belongs in general-peer-review.
  • Tone: Objective and constructive. Every Fail must include a concrete, actionable fix.
  • "Data available upon request": Always flag as non-FAIR. Not acceptable per RSC Digital Discovery and FAIR principles.
  • Proprietary formats: .xlsx, .mat, Gaussian .chk, VASP WAVECAR without open alternatives are I1/R1.3 failures.
  • License absence: Unlicensed ≠ open. Always flag missing licenses as R1.1 Fail.

References

  • Wilkinson, M. D. et al., "The FAIR Guiding Principles for scientific data management and stewardship", Sci. Data 3, 160018 (2016). DOI: 10.1038/sdata.2016.18
  • RSC Digital Discovery Data Review Guidelines. rsc.org/publishing

See Also


Author: Magdalena Lederbauer Contact: GitHub @mlederbauer

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/general-fair-data-review of learningmatter-mit/AtomisticSkills.

Open the folder on GitHubat commit 6257444

Compare with similar skills

General Fair Data Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

General Fair Data Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
General Fair Data Review this skilllearningmatter-mit/AtomisticSkills176—~2.3kAutomated safety check: PassMIT
Research Writingalfonso0512/research-writing-skill4871 repos~818Automated safety check: PassMIT
Paper WritingMLNLP-World/Paper-Writing-Tips4.7k—~630Automated safety check: PassNone
PaperjurySpark-To-Paper-Skills/paperjury1.2k—~5.3kAutomated safety check: PassMIT
Literature Surveyai4s-research/ai4s-skills2372 repos~2kAutomated safety check: PassMIT
Venue TemplatesK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: PassMIT

Similar skills

  • Research Writing

    alfonso0512/research-writing-skill

    科研论文写作助手,提供 30 个 Prompt 模板覆盖论文写作全流程. An agent skill from alfonso0512/research-writing-skill.

    487 GitHub starsUsed in 1 repo~818 tokens
    Documents & OfficeAuto-check passed
  • Paper Writing

    MLNLP-World/Paper-Writing-Tips

    学术论文写作检查与优化助手。基于 MLNLP-World 社区整理的论文写作技巧,帮助检查和优化学术论文。Use when: (1) 检查论文 LaTeX 格式和排版, (2) 优化公式符号使用, (3) 改进图表设计, (4) 润色英文学术表达, (5) 检查参考文献格式, (6) 投稿前终稿检查, (7) 用户询问论文写作技巧或规范。

    4.7k GitHub stars~630 tokensUpdated 12 days ago
    Documents & OfficeAuto-check passed
  • Paperjury

    Spark-To-Paper-Skills/paperjury

    Three modes for CS-conference papers (CVPR/ICCV/ECCV vision, ACL/EMNLP/NAACL NLP, ICLR/NeurIPS/ICML/AAAI ML).

    1.2k GitHub stars~5.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Literature Survey

    ai4s-research/ai4s-skills

    A skill your agent uses when the user wants a comprehensive literature survey on a specific research topic.

    237 GitHub starsUsed in 2 repos~2k tokens
    Documents & OfficeAuto-check passed
  • Venue Templates

    K-Dense-AI/claude-scientific-writer

    Prepare journal manuscripts, conference papers, research posters, and grant documents using venue-specific formatting guidance and bundled LaTeX scaffolds.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    Documents & OfficeAuto-check passed
  • Press Conference Revision Evidence Bound

    lensback940701/Evidence-Bound-Press-Conference-Revision-Skill

    Diagnose and revise defensive academic writing while preserving claim ceilings, evidence status, scope conditions, rival explanations, and conceptual hierarchy.

    255 GitHub stars~2.2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    176 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    176 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    176 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    176 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    176 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    176 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about General Fair Data Review

What does General Fair Data Review do?

Review a manuscript or code repository for FAIR data compliance (Findable, Accessible, Interoperable, Reusable), producing a structured report with pass/fail per principle and actionable remediation…. General Fair Data Review is an agent skill from learningmatter-mit/AtomisticSkills. Review a manuscript or code repository for FAIR data compliance (Findable, Accessible, Interoperable, Reusable), producing a structured report with pass/fail per principle and actionable remediation steps.

When should I use General Fair Data Review?

General Fair Data Review fits situations like: documents & Office work in your project.

How do I install General Fair Data Review in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill general-fair-data-review -a claude-code`. Or copy the skill folder (skills/general-fair-data-review in learningmatter-mit/AtomisticSkills) into .claude/skills/general-fair-data-review in your project. Claude Code loads it when a task matches its description.

How do I install General Fair Data Review in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill general-fair-data-review -a codex`. Or copy the skill folder (skills/general-fair-data-review in learningmatter-mit/AtomisticSkills) into .agents/skills/general-fair-data-review in your project. Codex loads it when a task matches its description.

Can I use General Fair Data Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill general-fair-data-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/general-fair-data-review, .gemini/skills/general-fair-data-review, .github/skills/general-fair-data-review and .opencode/skills/general-fair-data-review in your project.

What does General Fair Data Review need to run?

SKILL.md names no scripts, command-line tools or credentials: General Fair Data Review is instructions for the agent only.

Does General Fair Data Review access the network?

SKILL.md names 3 domains. As links in the text: doi.org, rsc.org and github.com. This is read from the text; nothing was executed.

Is General Fair Data Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does General Fair Data Review use?

General Fair Data Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does General Fair Data Review use?

About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to General Fair Data Review?

Skills that share tags, products or a category with General Fair Data Review: Research Writing (alfonso0512/research-writing-skill, 487 stars), Paper Writing (MLNLP-World/Paper-Writing-Tips, 4.7k stars), Paperjury (Spark-To-Paper-Skills/paperjury, 1.2k stars) and Literature Survey (ai4s-research/ai4s-skills, 237 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains General Fair Data Review?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 176 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 7, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.