Read arXiv Paper
karpathy/nanochat
Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.
A skill your agent uses when the user wants a paper audited for integrity issues — image misuse, numerical anomalies, logical gaps — and needs a reviewable evidence report.
$ npx skills add ai4s-research/ai4s-skills --skill integrity-auditor -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ai4s-research/ai4s-skills integrity-auditor --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ai4s-research/ai4s-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/integrity-auditor .claude/skills/integrity-auditor && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "integrity-auditor" agent skill from https://github.com/ai4s-research/ai4s-skills/tree/main/skills/integrity-auditor into .claude/skills/integrity-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integrity-auditor", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ai4s-research/ai4s-skills/tree/main/skills/integrity-auditorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ai4s-research/ai4s-skills --skill integrity-auditor -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ai4s-research/ai4s-skills integrity-auditor --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai4s-research/ai4s-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/integrity-auditor .agents/skills/integrity-auditor && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "integrity-auditor" agent skill from https://github.com/ai4s-research/ai4s-skills/tree/main/skills/integrity-auditor into .agents/skills/integrity-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integrity-auditor", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai4s-research/ai4s-skills --skill integrity-auditor -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ai4s-research/ai4s-skills integrity-auditor --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai4s-research/ai4s-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/integrity-auditor .cursor/skills/integrity-auditor && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "integrity-auditor" agent skill from https://github.com/ai4s-research/ai4s-skills/tree/main/skills/integrity-auditor into .cursor/skills/integrity-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integrity-auditor", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ai4s-research/ai4s-skills.git --path skills/integrity-auditor--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ai4s-research/ai4s-skills --skill integrity-auditor -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ai4s-research/ai4s-skills integrity-auditor --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai4s-research/ai4s-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/integrity-auditor .gemini/skills/integrity-auditor && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "integrity-auditor" agent skill from https://github.com/ai4s-research/ai4s-skills/tree/main/skills/integrity-auditor into .gemini/skills/integrity-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integrity-auditor", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ai4s-research/ai4s-skills integrity-auditorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ai4s-research/ai4s-skills --skill integrity-auditor -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ai4s-research/ai4s-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/integrity-auditor .github/skills/integrity-auditor && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "integrity-auditor" agent skill from https://github.com/ai4s-research/ai4s-skills/tree/main/skills/integrity-auditor into .github/skills/integrity-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integrity-auditor", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ai4s-research/ai4s-skills --skill integrity-auditor -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ai4s-research/ai4s-skills integrity-auditor --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ai4s-research/ai4s-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/integrity-auditor .opencode/skills/integrity-auditor && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "integrity-auditor" agent skill from https://github.com/ai4s-research/ai4s-skills/tree/main/skills/integrity-auditor into .opencode/skills/integrity-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "integrity-auditor", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
integrity-auditorA skill your agent uses when the user wants a paper audited for integrity issues — image misuse, numerical anomalies, logical gaps — and needs a reviewable evidence report.
Integrity Auditor is an agent skill from ai4s-research/ai4s-skills. Use when the user wants a paper audited for integrity issues — image misuse, numerical anomalies, logical gaps — and needs a reviewable evidence report. Works on external papers (PDF / DOI / arXiv) and on outputs from a local paper-writer run. Single-stage skill.
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files, including reference files (for example `forensics_tools/README.md`, `forensics_tools/bilingual_cn_geography.json` and `forensics_tools/channel_check.py`).
It sits in Research & Science, covering Academic paper search. It works with arXiv. The repository describes itself as: Open-source agent skills for AI for Science: topic exploration, literature survey, experiments, paper writing, and integrity audit — driven by any coding agent. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 744ab20. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
curlpython3pdftotextpdftoppmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Integrity Auditor loads about 5k tokens when it runs, and up to ~26k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 2,096 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ai4s-research/ai4s-skills at commit 744ab20, republished under its MIT licence (© ai4s-research). 2,096 words, ~4,974 tokens.
.claude/skills/integrity-auditor/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.Paper-integrity audit package. Single stage, full quality from the start. The agent reads each reference, then carries out three evidence tracks (image / numerical / logical) and produces a structured audit_report.md with Level 1–4 graded findings.
This skill ships no LLM SDK — it is the skill instructions, references, templates, and single-purpose forensics_tools/ only.
The substantive work is decomposed into reference playbooks under references/:
| Reference | Topic |
|---|---|
references/00-incremental-execution.md | how to do this without losing work: batches, persistence, resume — read first |
references/01-image-evidence.md | image evidence: panel split, dup detection, rotate/flip alignment, Western-blot continuity |
references/02-numerical-evidence.md | numerical evidence: n-consistency, mean/SD/SEM recompute, P-value sanity, decimal trail, Benford with caveats, deterministic-column-pair and last-digit chi-square sweepers, variance-reporting consistency |
references/02a-supplement-acquisition.md | publisher CDN routes: how to get hi-res figures and source-data XLSXs even when the article PDF is paywalled |
references/02b-ml-paper-arithmetic.md | ML / non-biology papers: arithmetic re-derivation of every quoted improvement against tabulated benchmark cells; leaderboard archive routes |
references/03-logical-evidence.md | logical evidence: conclusion-chain compression, missing controls, replication gap |
references/04-evidence-grading.md | 4-level finding grading + reviewable-evidence format (DOI / figure-id / pointer / transformation / requested raw data) |
references/05-quality-gate.md | self-check before delivery |
Also:
templates/audit_report.md — report skeleton the agent fills.forensics_tools/image_dup.py — perceptual-hash (dHash + aHash) duplicate detector for figure / panel PNGs. Single-purpose pure-Python utility (Pillow only). Catches untransformed dups.forensics_tools/image_dup_orb.py — ORB feature-matching duplicate detector with horizontal-flip augmentation. Catches transformed dups (rotation / flip / crop / brightness change) that perceptual hashing misses. Pair with image_dup.py: use phash first, escalate to ORB when phash distance is suspicious-but-inconclusive (16–60 range). Deps: OpenCV + NumPy.forensics_tools/panel_split.py — whitespace-gutter panel splitter. Pair with image_dup.py / image_dup_orb.py for cross-panel duplicate detection; whole-figure phash without panel splitting almost never finds anything.forensics_tools/channel_check.py — RGB channel-content classifier (DAPI / Flag / Merge / other) for fluorescence sub-images. Catches within-panel label swaps (e.g., a "DAPI" sub-image that is actually a Merge); cross-panel phash cannot catch this class.forensics_tools/decimal_match.py — cross-cell last-N-decimal matching sweeper for source-data XLSX. Detects fabrication where many distinct values share trailing decimal patterns (Kang Tiebang whistleblower class). Single-purpose pure-Python utility (openpyxl only). See references/02-numerical-evidence.md Check 1.5.forensics_tools/magnitude_consistency.py — supplement-text vs source-data XLSX unit/scale consistency. Catches unit-confusion (TWh vs GWh, mM vs µM, MHz vs Hz, etc.) and order-of-magnitude transcription errors via entity-overlap + literal-value matching across a generic SI-prefix-aware unit taxonomy covering energy / power / mass / length / area / volume / time / voltage / current / frequency / pressure / concentration / amount / force / dose / genomics-bp / CO2 / currency. Cross-family pairs (e.g., kV vs TWh) are automatically rejected. Pair with bilingual_cn_geography.json (or your own JSON map) for cross-language entity matching. Empirical baseline: Hu et al. 2026 Nature Tongyu county 1000× unit error. See references/02-numerical-evidence.md Check 1.6.forensics_tools/xlsx_aggregate_consistency.py — cross-XLSX same-quantity sum/row consistency. Detects when two source-data tables in the same paper purport to carry the same aggregate quantity but disagree by a small systematic margin (e.g., Hu et al. 2026 Nature MOESM3 vs MOESM6 1110.78 vs 1103.89 TWh, 0.62 percent diff, all 31 provinces same sign). Reuses the unit taxonomy from magnitude_consistency.py. Empirical baseline same paper Level 1 finding. See references/02-numerical-evidence.md Check 1.7.tests/smoketest.sh — < 30-second pre-commit gate. Compiles every script, runs every --help (catches argparse % bugs), and runs positive + negative controls for decimal_match, magnitude_consistency, and xlsx_aggregate_consistency. Run before every change.forensics_tools/README.md for the design rule that distinguishes utility scripts from forbidden "skeleton → enrich" orchestration, and for the recommended pipeline.Read the relevant reference before writing, not after. The full audit does not fit in a single turn — references/00-incremental-execution.md is the only execution mode that completes.
paper-writer.experiment-suite.literature-survey.Detect input mode:
| Mode | Trigger | Acquisition |
|---|---|---|
| PDF path | local *.pdf argument | use directly |
| DOI / arXiv ID | 10.xxxx/..., arXiv:NNNN.NNNNN | WebFetch landing page; record DOI / arXiv URL and abstract; download PDF if open-access |
| Local paper-writer slug | matches output/paper-writer/<slug>/latest/ | use slug directly; also read output/experiment-suite/<slug>/latest/ if present |
Compute slug:
TITLE_OR_ID="<input identifier>"
SLUG=$(python3 -c "import re,hashlib,sys; t=sys.argv[1]; n=re.sub(r'[\\s_]+','-',re.sub(r'[^\\w\\s-]','',t.lower().strip())).strip('-')[:40].rstrip('-'); h=hashlib.sha1(t.encode()).hexdigest()[:8]; print(f'{n}-{h}')" "$TITLE_OR_ID")
# For local-slug mode, set SLUG to the existing paper-writer slug directly.
TS=$(date +%Y-%m-%d_%H%M%S)
RUN=output/integrity-auditor/$SLUG/$TS
mkdir -p "$RUN/findings/image" "$RUN/findings/numerical" "$RUN/findings/logical"
ln -sfn "$TS" "output/integrity-auditor/$SLUG/latest"In commands below $RUN = output/integrity-auditor/<slug>/latest.
Open references/02a-supplement-acquisition.md first if the input is a DOI / arXiv ID / external PDF reference. Many publishers (Nature / Springer being the canonical example) gate the article body behind a paywall while serving the supplementary tables, source data, and hi-res figures on open CDN endpoints. Never conclude "paywalled, audit blocked" before trying the supplement / figure CDN routes.
Acquisition order (try every step; do not stop early):
curl -sL -A "Mozilla/5.0" "<article URL>" -o $RUN/paper.html. This page typically embeds the abstract, all main-figure captions, and direct CDN URLs for every supplementary file and high-resolution figure.grep -oE 'https?://[^"]*(MOESM|Fig[0-9]_HTML|mmc|MEDIA)\.(?:pdf|xlsx|docx|png)' $RUN/paper.html | sort -u. Download every distinct URL into $RUN/supplementary/ and $RUN/figures_hires/.curl -sL -A "Mozilla/5.0" "<correction URL>.pdf" -o $RUN/correction.pdf then file $RUN/correction.pdf to confirm it is a real PDF (not HTML auth wall).
3a. CRITICAL — the correction's own supplementary — fetch the correction's landing HTML, grep for its MOESM supplementary URLs (separate DOI namespace from the parent article), download every one. These often contain the pre-correction originals of replaced figures. See references/02a-supplement-acquisition.md for the exact recipe. Empirical baseline: the Wang Ping 2025 Nature audit only surfaced its image-track Level 4 finding because the correction's MOESM1_ESM.pdf carried the original (pre-correction) Fig 2 and Extended Data Figs 7, 10.curl -sL -A "Mozilla/5.0" "<article URL>.pdf" -o $RUN/paper.pdf then file $RUN/paper.pdf. If gated, record that fact in the manifest and continue with whatever the earlier steps acquired.results.json, data_contract.md, figures/manifest.json, bibliography.bib.Write $RUN/input_manifest.md listing every artefact actually acquired. At minimum it must enumerate what was downloaded and what was attempted-but-blocked.
Extract structured material:
# Text (every numeric claim, caption, figure reference will be greppable later)
pdftotext -layout "$PDF" "$RUN/paper.txt"
# Panels (one .png / .ppm per embedded raster image)
mkdir -p "$RUN/panels"
pdfimages -all "$PDF" "$RUN/panels/page"
N_RASTER=$(ls "$RUN/panels" 2>/dev/null | wc -l)
# Fallback for vector-figure papers (e.g., paper-writer outputs):
# pdfimages only extracts embedded raster images; a paper built from matplotlib
# vector PDFs through \includegraphics will yield zero panels here. In that case,
# render each page to a PNG so visual inspection is still possible.
if [ "$N_RASTER" -eq 0 ]; then
pdftoppm -r 150 "$PDF" "$RUN/panels/page" -png
fi
# For a local paper-writer slug audit, also pull the production-side figure PDFs directly
# (these are the originals, before LaTeX embedding):
if [ "$INPUT_MODE" = "slug" ]; then
mkdir -p "$RUN/figures_from_suite"
cp output/experiment-suite/$SLUG/latest/figures/*.pdf "$RUN/figures_from_suite/" 2>/dev/null
cp output/experiment-suite/$SLUG/latest/figures/manifest.json "$RUN/figures_from_suite/" 2>/dev/null
fi
ls "$RUN/panels" | wc -l # record in input_manifest.mdIf pdfimages / pdftotext / pdftoppm from poppler-utils is not installed, install via system package manager and retry. Do not proceed without either raster panels or page renderings — image evidence track depends on them.
If the article PDF is paywalled and the article HTML page yields no supplementary or figure CDN URLs and the Author Correction PDF is also gated, then and only then write _paywall_blocked.md per track. The Wang Ping 2025 Nature audit (output/integrity-auditor/wang-ping-hdac6-valine-10-1038-s41586-024-08248-5/latest/) demonstrates that the article body being gated says nothing about whether the supplementary source data is gated; the latter is almost always open and is where the substantive audit lives.
Open references/00-incremental-execution.md first. Then carry out the three tracks below across many turns, persisting state to $RUN/ after every batch.
Before invoking any sweeper or forensics tool, classify the paper. The substantive audit method differs by class — running biology sweepers on an ML paper wastes the run and pollutes the report with "swept clean" lines that are misleading (nothing was actually swept).
| Class | Signals | Image-track method | Numerical-track method |
|---|---|---|---|
| bio | microscopy / Western blot / fluorescence / source-data XLSX / supplementary MOESM*.xlsx | forensics_tools/image_dup.py, panel_split.py, channel_check.py per references/01-image-evidence.md | sweepers from references/02-numerical-evidence.md Checks 1–4 |
| ml | schematic architecture diagrams, training-curve line plots, benchmark-score tables, no source-data XLSX, leaderboard URLs in body | figure-vs-text consistency only (forensics tools inapplicable — schematic figures lack pixel-level forensic signal) | arithmetic re-derivation per references/02b-ml-paper-arithmetic.md; sweepers do not apply (per-cell n too small for chi-square; cells heterogeneous in scale) |
| other (theory / review / clinical) | no source data, no original figures | document scope-limit; usually only logical-track is informative | logical-track only; record "numerical track inapplicable" |
Quick triage heuristic (works in one line):
# If ≥ 1 XLSX in supplementary/, treat as bio; if arXiv ID + no XLSX + ≤ 5 figures, treat as ml.
N_XLSX=$(ls $RUN/supplementary/*.xlsx 2>/dev/null | wc -l)
[ "$N_XLSX" -ge 1 ] && PAPER_CLASS=bio || PAPER_CLASS=ml
echo "$PAPER_CLASS" > $RUN/paper_class.txtRecord the chosen class in $RUN/input_manifest.md so reviewers can see the methodological branch.
When a sweeper or forensics tool is inapplicable for the paper class, write findings/<track>/_inapplicable.md (NOT _clean.md) — see "Important rules" below for the semantic distinction. A _clean.md claims "I swept and found nothing"; an _inapplicable.md claims "the tool does not apply to this paper class; here is what was checked instead". Conflating the two understates audit honesty.
Open: references/01-image-evidence.md (for bio class). For each figure panel, check within-figure duplicates, cross-figure duplicates, transformations (rotate / flip / crop / brightness), and Western-blot background continuity. Each anomaly becomes one $RUN/findings/image/<short-id>.md written in the per-finding format described in references/04-evidence-grading.md. A "no issue" outcome is recorded as _clean.md.
For ml class: forensics tools do not apply (schematic figures lack pixel-level signal). Substitute check is figure-vs-text consistency. Record as findings/image/_inapplicable.md listing which tools were considered and what consistency check was substituted.
Open: references/02-numerical-evidence.md (for bio class) or references/02b-ml-paper-arithmetic.md (for ml class). Bio class: extract every numeric claim from paper.txt with its location (section / figure / table); recompute means / SDs / SEMs against source data when available; otherwise check internal consistency (does n=6 in the body match the figure caption?); run the four sweepers (Checks 1–4). ML class: arithmetic re-derive every quoted delta / average / improvement against the table cells (Checks 1–4 are inapplicable; record as _inapplicable.md and the arithmetic re-derivation as _clean.md or as one finding per mismatch).
Each mismatch (either class) becomes $RUN/findings/numerical/<short-id>.md. For a local paper-writer slug, also reconcile every number in the paper against output/experiment-suite/<slug>/latest/results.json — every reported metric must trace back to a summary or runs entry.
In addition (both classes), run Check 5 — variance-reporting consistency from references/02-numerical-evidence.md: scan every table caption for restart / seed / std / error-bar mentions; flag the case where some tables in the same paper report restart-averaged numbers and others do not, when the un-averaged tables carry sub-1-pp claims. This was the BERT NSP-overclaim finding mechanism.
Open: references/03-logical-evidence.md. Compress each headline claim into "A through B causes C" form. Check whether the experiments actually exercise A, B, and the A→B→C link with controls (positive / negative / rescue / dose / time). Each gap becomes $RUN/findings/logical/<short-id>.md.
Open: references/04-evidence-grading.md and templates/audit_report.md. For each finding, assign Level 1 (suspicious / could be benign), Level 2 (obvious anomaly), Level 3 (cannot resolve without raw data), Level 4 (high suspicion of misconduct). Copy the template to $RUN/audit_report.md and fill it section by section, citing each finding file by relative path. The report must include a "Requested raw data" section listing exactly what the author needs to supply for any Level 3 finding to be resolved.
Open: references/05-quality-gate.md. Verify: no editorial verdicts in the report, every finding has a reviewable artefact pointer, every Level ≥ 2 finding names the raw data needed to resolve it, the manifest matches what was actually examined.
Report:
output/integrity-auditor/<slug>/latest/input_manifest.mdoutput/integrity-auditor/<slug>/latest/findings/{image,numerical,logical}/*.mdoutput/integrity-auditor/<slug>/latest/audit_report.mdreferences/05-quality-gate.md.When the input is a local paper-writer slug, the auditor is read-only against:
output/paper-writer/<slug>/latest/paper/main.pdf — the paper under auditoutput/paper-writer/<slug>/latest/paper/bibliography.bib — citation provenanceoutput/experiment-suite/<slug>/latest/results.json — numerical ground truth + provenance.mode (every "measured" claim in the paper must be backed here)output/experiment-suite/<slug>/latest/data_contract.md — dataset binding (does the data actually exist; checksum recoverable)output/experiment-suite/<slug>/latest/figures/manifest.json — figures the paper is allowed to reference (basenames only)Never modify another skill's outputs. The audit is a third-party read.
import anthropic / import openai. The skill is its instructions + references + template only._clean.md — the track's tools/sweepers were applicable and were run, and produced no anomaly. List what was checked and what passed. Empirical baseline: Wang Ping numerical track Mode A swept all 14 XLSX sheets and found 5 hits — the other 9 would be _clean.md content if reported per-sheet._inapplicable.md — the track's tools do not apply to this paper class (e.g., biology image-dup on a pure ML schematic-figure paper). List which tools were considered and rejected, and what substitute check was used instead. Empirical baseline: BERT (arXiv:1810.04805) image track wrote _inapplicable.md; the substitute was figure-vs-text consistency._inapplicable.md situation "clean" overstates audit coverage; calling a genuine _clean.md "inapplicable" understates it.forensics_tools/ directory if a concrete pain point demands it — the anti-pattern rule is against "skeleton → enrich" pipeline orchestrators and LLM SDK imports, not against single-purpose tools. v1 ships without forensics_tools/; revisit when needed.© ai4s-research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 22 other files (references) in skills/integrity-auditor of ai4s-research/ai4s-skills.
Open the folder on GitHubat commit 744ab20
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ai4s-research/ai4s-skills, which our catalogue first saw on October 7, 2026.
Integrity Auditor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Integrity Auditor this skillai4s-research/ai4s-skills | 237 | 1 repos | ~5k | Automated safety check: Pass | MIT | |
| Read arXiv Paperkarpathy/nanochat | 59k | 1 repos | ~494 | Automated safety check: Pass | MIT | |
| Literature Reviewneflibata-feng/MyArxiv-Agent | 126 | 20 repos | ~5.9k | Automated safety check: Notes | MIT | |
| Openalex Databaseneflibata-feng/MyArxiv-Agent | 126 | 12 repos | ~3k | Automated safety check: Pass | Custom licence | |
| Citation ManagementK-Dense-AI/claude-scientific-writer | 2.4k | 2 repos | ~3.9k | Automated safety check: Notes | MIT | |
| Citation Managementneflibata-feng/MyArxiv-Agent | 126 | 19 repos | ~8.1k | Automated safety check: Notes | MIT |
karpathy/nanochat
Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.
neflibata-feng/MyArxiv-Agent
Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).
neflibata-feng/MyArxiv-Agent
Query and analyze scholarly literature using the OpenAlex database.
K-Dense-AI/claude-scientific-writer
Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.
neflibata-feng/MyArxiv-Agent
Comprehensive citation management for academic research. An agent skill from neflibata-feng/MyArxiv-Agent.
bytedance/deer-flow
Searches arXiv across many papers on one topic, extracts each paper's methodology and findings in parallel, and synthesizes a cited literature review.
ai4s-research/ai4s-skills
A skill your agent uses when the user has a research question and needs a complete experiment package — design document, runnable code, results (measured or simulated with honest provenance)…
ai4s-research/ai4s-skills
A skill your agent uses when the user wants a comprehensive literature survey on a specific research topic.
ai4s-research/ai4s-skills
A skill your agent uses when the user wants a complete, publication-grade research paper on a specific topic — produces 200+ real citations, 4–8 publication-grade figures, and 7 sections of…
ai4s-research/ai4s-skills
A skill your agent uses when the user has a vague research direction and wants to explore feasible specific topics.
ai4s-research/ai4s-skills
A skill your agent uses when the user wants an end-to-end AI4S research pipeline — broad direction or specific topic in, full research package out (exploration + literature survey + experiment +…
ai4s-research/ai4s-skills
Generate beautiful, high-resolution mindmaps from Markdown unordered lists.
Works with
Categories
A skill your agent uses when the user wants a paper audited for integrity issues — image misuse, numerical anomalies, logical gaps — and needs a reviewable evidence report. Integrity Auditor is an agent skill from ai4s-research/ai4s-skills. Use when the user wants a paper audited for integrity issues — image misuse, numerical anomalies, logical gaps — and needs a reviewable evidence report.
Integrity Auditor fits situations like: the user wants a paper audited for integrity issues — image misuse; numerical anomalies; logical gaps — and needs a reviewable evidence report.
Run `npx skills add ai4s-research/ai4s-skills --skill integrity-auditor -a claude-code`. Or copy the skill folder (skills/integrity-auditor in ai4s-research/ai4s-skills) into .claude/skills/integrity-auditor in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ai4s-research/ai4s-skills --skill integrity-auditor -a codex`. Or copy the skill folder (skills/integrity-auditor in ai4s-research/ai4s-skills) into .agents/skills/integrity-auditor in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai4s-research/ai4s-skills --skill integrity-auditor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/integrity-auditor, .gemini/skills/integrity-auditor, .github/skills/integrity-auditor and .opencode/skills/integrity-auditor in your project.
Going by SKILL.md and its folder, Integrity Auditor needs Python for the scripts in its folder and the command-line tools its instructions call (curl, python3, pdftotext and pdftoppm). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Integrity Auditor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Integrity Auditor: Read arXiv Paper (karpathy/nanochat, 59k stars), Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars), Openalex Database (neflibata-feng/MyArxiv-Agent, 126 stars) and Citation Management (K-Dense-AI/claude-scientific-writer, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ai4s-research (a GitHub organization) maintains it in ai4s-research/ai4s-skills, which has 237 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on July 28, 2026.
Source: ai4s-research/ai4s-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.