Agent skill

Arxiv Preflight

by aipoch in aipoch/medical-research-skills

Run a submission-readiness preflight on a manuscript before arXiv upload.

MITAuto-check passedResearch & Science

Install Arxiv Preflight

skills CLI
$ npx skills add aipoch/medical-research-skills --skill arxiv-preflight -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills arxiv-preflight --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'awesome-med-research-skills/Academic Writing/arxiv-preflight' .claude/skills/arxiv-preflight && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
arxiv-preflight
GitHub stars
2k
Token cost
~1.8k tokens
SKILL.md length
830 words
Files
24 (incl. scripts, references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

Run a submission-readiness preflight on a manuscript before arXiv upload.

  • Works in 8 steps: Identify inputs. Ask for or detect:… → Extract manuscript text. Run… → Hard-red-line scan. Run… → …
  • The user is preparing an arXiv submission
  • SKILL.md covers Workflow, Risk levels, Decision rule and Rules, plus 2 more sections
  • Asks to check a paper before uploading

What it does

Arxiv Preflight is an agent skill from aipoch/medical-research-skills. Run a submission-readiness preflight on a manuscript before arXiv upload. Use when the user is preparing an arXiv submission, asks to check a paper before uploading, mentions hallucinated or fake references, leftover LLM meta-comments / prompts in text, placeholder data (TODO, TBD, XX%), AI-use disclosure, scholarly integrity, research integrity, or arXiv moderation risk — even if they don't say "preflight". Also trigger on phrases like "check my paper before arXiv", "verify my references", "scan for AI…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 29 other files, including scripts and reference files (for example `eval_report_arxiv-preflight_result.json` and `evals/evals.json`).

It sits in Research & Science, covering Academic paper search and LaTeX. It works with arXiv and LaTeX. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • The user is preparing an arXiv submission
  • Asks to check a paper before uploading
  • Mentions hallucinated
  • Fake references

Example prompts

  • “t say”
  • “. Also trigger on phrases like”
  • “verify my references”
  • “/arxiv-preflight”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Identify inputs. Ask for or detect: LaTeX project directory, single PDF, BibTeX file, supplementary figures/tables, or experiment data…
  2. Extract manuscript text. Run scripts/extract_manuscript_text.py on the input. It merges \input/\include, strips LaTeX commands while…
  3. Hard-red-line scan. Run scripts/scan_ai_artifacts.py against the extracted text. It flags LLM meta-comments, prompt residue, placeholder…
  4. Verify references. Run scripts/verify_references.py on .bib (or extracted reference list). It performs (a) structural BibTeX checks, (b)…
  5. Cross-check claims, numbers, and figure/table references. Use manuscript.json to review extracted numbers, labels, and references. This…
  6. AI-use disclosure check. scripts/scan_ai_artifacts.py includes a limited disclosure heuristic for substantive LLM use. For nuanced cases…
  7. arXiv moderation pre-check. Do an agent-assisted scan for non-research content (opinion/proposal/course-project framing without research…
  8. Generate report. Run scripts/generate_preflight_report.py to merge all JSON outputs into a single Markdown report. Format and decision…

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Arxiv Preflight loads about 1.8k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 156 tokens; SKILL.md has 830 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~156
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 830 words, ~1,844 tokens.

Download SKILL.mdSave it as .claude/skills/arxiv-preflight/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.
name
arxiv-preflight
description
Run a submission-readiness preflight on a manuscript before arXiv upload. Use when the user is preparing an arXiv submission, asks to check a paper before uploading, mentions hallucinated or fake references, leftover LLM meta-comments / prompts in text, placeholder data (TODO, TBD, XX%), AI-use disclosure, scholarly integrity, research integrity, or arXiv moderation risk — even if they don't say "preflight". Also trigger on phrases like "check my paper before arXiv", "verify my references", "scan for AI artifacts", "scan for LLM residue", "is my submission ready", or "review .tex/.bib before submit".
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

arXiv Preflight

Submission-readiness review for a manuscript before arXiv upload. The goal is not to judge whether a paper was AI-written from style. The goal is to find concrete, locatable, reviewable evidence of unchecked LLM output, hallucinated references, placeholder data, unsupported claims, and arXiv-policy risks — and to hand the author a fix list.

Anchor on the actual arXiv stance: authors may use generative AI tools but are fully responsible for the content and must disclose significant use per field norms; AI tools cannot be listed as authors. See references/arxiv_policy_notes.md for the exact language.

This skill is intentionally platform-neutral. It should work for any AI coding/review agent that can read files and run local Python scripts; it does not require OpenAI-specific UI metadata.

Requires Python 3.9+.

Workflow

  1. Identify inputs. Ask for or detect: LaTeX project directory, single PDF, BibTeX file, supplementary figures/tables, or experiment data. Prefer LaTeX source over rendered PDF — PDF text extraction loses structure and produces noisier reference checks.

  2. Extract manuscript text. Run scripts/extract_manuscript_text.py on the input. It merges \input/\include, strips LaTeX commands while preserving section/figure/table/citation keys, and falls back to PDF text extraction. Output: manuscript.json with sections, figures, tables, cite_keys, numbers, and warnings. Treat unresolved or skipped includes as review risks.

  3. Hard-red-line scan. Run scripts/scan_ai_artifacts.py against the extracted text. It flags LLM meta-comments, prompt residue, placeholder content, TODO/TBD/XX%, and AI-listed-as-author. Patterns live in references/ai_artifact_patterns.md — extend the list when the user reports new failure modes. Any hit is a BLOCKER unless the user explicitly justifies it.

  4. Verify references. Run scripts/verify_references.py on .bib (or extracted reference list). It performs (a) structural BibTeX checks, (b) external metadata lookup against Crossref / arXiv / OpenAlex / Semantic Scholar (in that order), (c) cite-key vs. \cite reconciliation. Grading and API strategy: references/reference_verification.md. If network is unavailable, mark this section INCOMPLETE rather than skipping silently.

  5. Cross-check claims, numbers, and figure/table references. Use manuscript.json to review extracted numbers, labels, and references. This version exposes the raw structures for agent-assisted consistency review; it does not fully automate semantic claim verification. Surface candidate inconsistencies for the author to confirm, never rewrite their claims.

  6. AI-use disclosure check. scripts/scan_ai_artifacts.py includes a limited disclosure heuristic for substantive LLM use. For nuanced cases, compare detected AI-assistance signals against the manuscript's Acknowledgments / Methods / Ethics sections yourself. Recommend disclosure only when significant use is evident; do not push boilerplate disclosure on every paper. Calibration: references/arxiv_policy_notes.md.

  7. arXiv moderation pre-check. Do an agent-assisted scan for non-research content (opinion/proposal/course-project framing without research contribution), copyright risks (publisher PDFs, reviewer comments, unlicensed images, long third-party text), and salami/duplicate-submission signals. This version does not ship a dedicated moderation scanner. Do not predict whether arXiv will accept — only surface risk.

  8. Generate report. Run scripts/generate_preflight_report.py to merge all JSON outputs into a single Markdown report. Format and decision rules: references/report_template.md.

Show full SKILL.md (364 more words)Show less

Risk levels

  • BLOCKER — evidence that should stop submission: LLM meta-comment in body text, AI listed as author, reference that cannot be matched in any external database, unfilled placeholder in a results table.
  • HIGH — likely integrity or policy issue: DOI/title mismatch, cite-key pointing at wrong paper, a strong claim with no supporting number in the manuscript.
  • MEDIUM — inconsistency, missing metadata, unclear disclosure, broken \ref.
  • LOW — polish, formatting, capitalization.

Decision rule

  • Any BLOCKER → HOLD.
  • No blocker, any HIGH/MEDIUM → PASS_WITH_FIXES.
  • Only LOW or clean → PASS.

Never upgrade a finding's severity to be safe and never silently soften it. If you are unsure, mark it MEDIUM and explain.

Rules

  • Never call a paper AI-generated from writing style alone. Style is not evidence; meta-comments, fake citations, and placeholders are.
  • Every finding must cite a location. File + line number for .tex/.bib; for PDF input, use the PDF file plus extracted-text line unless page-aware extraction is available. Include a verbatim short snippet (≤ 25 words). A finding without a location is a suggestion, not a finding — move it to "Recommended Cleanup".
  • Treat external-metadata mismatches as leads. A single field disagreement (e.g., page numbers) is MEDIUM. Title + authors + year all unmatched across all four databases is BLOCKER.
  • Network-incomplete is not network-pass. If reference lookups failed for connectivity reasons, the report's reference section is INCOMPLETE; the overall decision cannot be PASS.
  • Do not rewrite the author's claims. Provide fixes, questions, and locations. Leave the rewriting to the author.
  • Minimise false positives at the cost of recall. This tool runs before submission; a false BLOCKER is more costly than a missed LOW.

When to read which reference

  • references/arxiv_policy_notes.md — when calibrating disclosure recommendations or moderation findings.
  • references/ai_artifact_patterns.md — when extending or interpreting the artifact scanner's hits.
  • references/reference_verification.md — when interpreting citation lookup output, deciding API order, or handling rate-limit / network failures.
  • references/report_template.md — when writing or regenerating the final report.

MVP scope

This skill currently ships: text extraction, AI-artifact scan, BibTeX + DOI/arXiv-ID reference verification, extracted label/citation/number structures, and Markdown report generation. Out of scope for v1: full semantic peer review, full claim/result verification, automated moderation judgement, AI-writing probability scoring, automatic rewriting, automated arXiv upload, large-scale plagiarism detection. If the user asks for these, say so explicitly rather than approximating.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 23 other files (scripts, references) in awesome-med-research-skills/Academic Writing/arxiv-preflight of aipoch/medical-research-skills.

  • SKILL.md
  • eval_report_arxiv-preflight_result.json
  • evals/evals.json
  • evals/fixtures/clean/main.tex
  • evals/fixtures/clean/refs.bib
  • evals/fixtures/hallucinated_ref/main.tex
  • evals/fixtures/hallucinated_ref/refs.bib
  • evals/fixtures/llm_meta/main.tex
  • evals/fixtures/llm_meta/refs.bib
  • evals/fixtures/multifile/intro.tex
  • evals/fixtures/multifile/main.tex
  • evals/fixtures/multifile/method.tex
  • evals/fixtures/multifile/refs.bib
  • evals/fixtures/multifile/results.tex
  • references
  • … and 9 more

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Arxiv Preflight next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Arxiv Preflight compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Arxiv Preflight this skillaipoch/medical-research-skills2k—~1.8kAutomated safety check: PassMIT
Arxiv MCP Serverblazickjp/arxiv-mcp-server3.2k—~353Automated safety check: PassApache-2.0
Arxiv Paper Writerappautomaton/latex-arxiv-SKILL458—~2.3kAutomated safety check: PassMIT
Arxiv Paper Writeryunshenwuchuxun/latex-paper-skills266—~3.2kAutomated safety check: PassMIT
Paper OrchestraAr9av/PaperOrchestra6771 repos~3.5kAutomated safety check: PassCustom licence
Section Writing AgentAr9av/PaperOrchestra6771 repos~3.3kAutomated safety check: PassCustom licence

Similar skills

  • Arxiv MCP Server

    blazickjp/arxiv-mcp-server

    A skill your agent uses when finding, comparing, reading, or monitoring arXiv papers, including requests for abstracts, citation graphs, original LaTeX, section-level technical details, or…

    3.2k GitHub stars~353 tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Arxiv Paper Writer

    appautomaton/latex-arxiv-SKILL

    Write LaTeX ML/AI review articles for arXiv using the IEEEtran template and verified BibTeX citations.

    458 GitHub stars~2.3k tokensUpdated 26 days ago
    Research & ScienceAuto-check passed
  • Arxiv Paper Writer

    yunshenwuchuxun/latex-paper-skills

    Writes ML/AI review and survey papers for arXiv using the IEEEtran LaTeX template with verified BibTeX citations.

    266 GitHub stars~3.2k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Paper Orchestra

    Ar9av/PaperOrchestra

    Orchestrate the full PaperOrchestra (Song et al., 2026, arXiv:2604.05018) five-agent pipeline to turn unstructured research materials (idea, experimental log, LaTeX template, conference guidelines…

    677 GitHub starsUsed in 1 repo~3.5k tokens
    Research & ScienceAuto-check passed
  • Section Writing Agent

    Ar9av/PaperOrchestra

    Step 4 of the PaperOrchestra pipeline (arXiv:2604.05018). An agent skill from Ar9av/PaperOrchestra.

    677 GitHub starsUsed in 1 repo~3.3k tokens
    Research & ScienceAuto-check passed
  • Academic Research

    voidful/academic-skills

    Complete academic research skill suite covering the full pipeline: paper reading (read/explain papers with storytelling), idea generation (brainstorm research directions), experiment design (plan…

    134 GitHub stars~887 tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 22 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 22 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 22 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 22 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 22 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 22 days ago
    Auto-check passed

Works with

Questions about Arxiv Preflight

What does Arxiv Preflight do?

Run a submission-readiness preflight on a manuscript before arXiv upload. Arxiv Preflight is an agent skill from aipoch/medical-research-skills. Run a submission-readiness preflight on a manuscript before arXiv upload.

When should I use Arxiv Preflight?

Arxiv Preflight fits situations like: the user is preparing an arXiv submission; asks to check a paper before uploading; mentions hallucinated; fake references.

How do I install Arxiv Preflight in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill arxiv-preflight -a claude-code`. Or copy the skill folder (awesome-med-research-skills/Academic Writing/arxiv-preflight in aipoch/medical-research-skills) into .claude/skills/arxiv-preflight in your project. Claude Code loads it when a task matches its description.

How do I install Arxiv Preflight in Codex?

Run `npx skills add aipoch/medical-research-skills --skill arxiv-preflight -a codex`. Or copy the skill folder (awesome-med-research-skills/Academic Writing/arxiv-preflight in aipoch/medical-research-skills) into .agents/skills/arxiv-preflight in your project. Codex loads it when a task matches its description.

Can I use Arxiv Preflight in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill arxiv-preflight -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arxiv-preflight, .gemini/skills/arxiv-preflight, .github/skills/arxiv-preflight and .opencode/skills/arxiv-preflight in your project.

What does Arxiv Preflight need to run?

SKILL.md names no scripts, command-line tools or credentials: Arxiv Preflight is instructions for the agent only. Our summary lists: Python 3.

Does Arxiv Preflight access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Arxiv Preflight safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Arxiv Preflight use?

Arxiv Preflight is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Arxiv Preflight use?

About 1.8k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5k tokens, read only when the agent opens those files.

What are the alternatives to Arxiv Preflight?

Skills that share tags, products or a category with Arxiv Preflight: Arxiv MCP Server (blazickjp/arxiv-mcp-server, 3.2k stars), Arxiv Paper Writer (appautomaton/latex-arxiv-SKILL, 458 stars), Arxiv Paper Writer (yunshenwuchuxun/latex-paper-skills, 266 stars) and Paper Orchestra (Ar9av/PaperOrchestra, 677 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Arxiv Preflight?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,978 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.