Agent skill

Scrub

by ZimoLiao in ZimoLiao/scholaraio

A skill your agent uses when incrementally reviewing and repairing low-quality metadata after enrich, especially non-standard documents that need title, author, or year correction while skipping…

MITAuto-check passed

Install Scrub

skills CLI
$ npx skills add ZimoLiao/scholaraio --skill scrub -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZimoLiao/scholaraio scrub --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZimoLiao/scholaraio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/scrub .claude/skills/scrub && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scrub
GitHub stars
577
Token cost
~2k tokens
SKILL.md length
887 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when incrementally reviewing and repairing low-quality metadata after enrich, especially non-standard documents that need title, author, or year correction while skipping…

  • Works in 6 steps: Find unreviewed candidates → Inspect one paper at a time → Repair conservatively → …
  • Incrementally reviewing and repairing low-quality metadata after enrich
  • SKILL.md covers When To Use, Workflow, Heuristics and Acceptance Standard
  • Calls python

What it does

Scrub is an agent skill from ZimoLiao/scholaraio. Use when incrementally reviewing and repairing low-quality metadata after enrich, especially non-standard documents that need title, author, or year correction while skipping already reviewed records via .scrubbed.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Scholar All-In-One: A research infrastructure for AI agents. The licence is MIT.

When your agent uses it

  • Incrementally reviewing and repairing low-quality metadata after enrich
  • Especially non-standard documents that need title
  • Year correction while skipping already reviewed records via .scrubbed

Example prompts

  • “/scrub”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Find unreviewed candidates
  2. Inspect one paper at a time
  3. Repair conservatively
  4. Handle directory renames correctly
  5. Mark reviewed papers
  6. Rebuild indexes once per batch

What it can do on your machine

Read from SKILL.md and the folder at commit 777628b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scrub loads about 2k tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 887 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~55
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ZimoLiao/scholaraio at commit 777628b, republished under its MIT licence (© ZimoLiao). 887 words, ~1,991 tokens.

Download SKILL.mdSave it as .claude/skills/scrub/SKILL.md (or your agent's skills folder).
name
scrub
description
Use when incrementally reviewing and repairing low-quality metadata after enrich, especially non-standard documents that need title, author, or year correction while skipping already reviewed records via .scrubbed.

Scrub Metadata

Use this skill when the library contains already-ingested papers whose metadata is still clearly low quality after ingest or enrich, especially for non-standard documents that MinerU or fallback parsers converted successfully but described poorly.

scrub is a review-and-repair workflow, not a blind batch rewrite. It should reuse existing ScholarAIO repair and rename primitives, and it should treat .scrubbed as the durable marker for "reviewed and currently acceptable."

When To Use

Use this skill when the user wants to:

  • clean bad metadata after enrich
  • repair placeholder or garbled titles
  • fix suspicious author names
  • fill in missing years when the paper content supports it
  • incrementally review a large library without reprocessing already-reviewed papers

Do not use this skill for:

  • normal ingest
  • DOI or citation-count refresh
  • paper-content enrichment such as TOC/L3 extraction
  • directory normalization when metadata is already trustworthy and rename alone is enough

Workflow

1. Find unreviewed candidates

Skip papers that already contain .scrubbed.

You can list suspicious, unreviewed papers with a Python helper that resolves papers_dir from the active ScholarAIO config:

bash
python - <<'PY'
from scholaraio.services.audit import list_scrub_suspects
from scholaraio.core.config import load_config

cfg = load_config()

for issue in list_scrub_suspects(cfg.papers_dir):
    print(f"{issue.paper_id}\t{issue.rule}\t{issue.message}")
PY

If the user asked for a broad quality pass, it is also reasonable to start with:

bash
scholaraio audit

Then narrow to papers that are both:

  • not already .scrubbed
  • obviously bad enough to justify manual review
2. Inspect one paper at a time

For candidates with readable metadata, inspect:

bash
scholaraio show "<paper-id>" --layer 1

Before changing anything, record the stable paper UUID shown in the L1 header as stable_id. repair preserves this UUID even when the directory name changes.

Then read the source text as needed:

bash
scholaraio show "<paper-id>" --layer 4

If the suspect is invalid_metadata because the directory only has paper.md and no readable meta.json, show --layer 1 may not work yet. In that case, inspect paper.md directly from the configured papers directory and repair by directory name:

bash
python - <<'PY'
from scholaraio.core.config import load_config

paper_id = "<paper-id>"
cfg = load_config()
print((cfg.papers_dir / paper_id / "paper.md").resolve())
PY

If the default show --layer 4 view is too long, resolve the actual paper.md path with the same identifier semantics as show / repair, then inspect only the needed slice:

bash
python - <<'PY'
from scholaraio.cli import _resolve_paper
from scholaraio.core.config import load_config

paper_id = "<paper-id>"
cfg = load_config()
print((_resolve_paper(paper_id, cfg) / "paper.md").resolve())
PY

If the head of the file is insufficient, inspect a larger section or search relevant phrases in the resolved paper.md.

Focus on extracting only the identity-critical metadata needed to make the paper usable:

  • real title or at least a concise, accurate keyword title
  • real first author or organization when clearly stated
  • publication year when clearly supported by the document
3. Repair conservatively

Use repair to update only the fields you can support from the source:

bash
scholaraio repair "<paper-id>" --title "Correct Title" --author "First Author" --year 2024 --no-api --dry-run

Then run the real repair:

bash
scholaraio repair "<paper-id>" --title "Correct Title" --author "First Author" --year 2024 --no-api

repair now preserves existing metadata and only overwrites the fields you explicitly update through the CLI. Existing journal, abstract, paper type, citation counts, IDs, TOC/L3 fields, and other enriched metadata stay in place unless you intentionally replace them.

In scrub mode, --no-api should be the default. These records are often low-quality documents or weakly identified items, and conservative local repair is safer than letting API matches overwrite title, author, or year.

Only drop --no-api when the user explicitly wants metadata refetch behavior and has checked that the identifier quality is strong enough to support it.

Only pass --doi when you are intentionally correcting or adding the DOI. If you omit --doi, repair preserves the existing DOI.

Decision policy:

  • Title quality is the top priority.
  • Authors and year should be repaired when evidence is strong.
  • Do not fabricate DOI, journal, or venue.
  • If author or year cannot be confirmed reliably, leave them unresolved rather than guessing.
Show full SKILL.md (335 more words)Show less
4. Handle directory renames correctly

scholaraio repair already rewrites meta.json and renames the paper directory immediately when title, author, year, or DOI changes. The rename is derived from the updated identity fields, while the rest of the metadata is preserved unless explicitly overwritten.

That means the original directory name may stop existing right after the real repair. Resolve the current directory from the stable UUID you recorded before editing:

bash
python - <<'PY'
from scholaraio.cli import _resolve_paper
from scholaraio.core.config import load_config

stable_id = "<uuid-from-layer-1>"
cfg = load_config()
print(_resolve_paper(stable_id, cfg).name)
PY

If you repaired papers through scholaraio repair, skip rename --all for those same records. repair already rewrites meta.json and renames the directory immediately, including collision suffixes when needed.

Only use rename --all for records whose meta.json you edited outside repair, or for older records you did not already rename in the current scrub pass:

bash
scholaraio rename --all

Because rename may change the directory path, always create the marker using the post-rename directory name.

5. Mark reviewed papers

Once a paper has been reviewed and is acceptable for current library use, create the marker:

bash
python - <<'PY'
from scholaraio.cli import _resolve_paper
from scholaraio.core.config import load_config
from scholaraio.stores.papers import mark_scrubbed

stable_id = "<uuid-from-layer-1>"
cfg = load_config()
paper_d = _resolve_paper(stable_id, cfg)
mark_scrubbed(paper_d)
print(f"marked {paper_d.name} as scrubbed")
PY

Only mark a paper when:

  • you have reviewed the record
  • the remaining metadata quality is acceptable
  • there is no known blocking issue that should force future re-review

.scrubbed means reviewed, not perfect.

6. Rebuild indexes once per batch

After finishing the batch:

bash
scholaraio pipeline reindex

This keeps search and registry state aligned with renamed or repaired records.

Heuristics

The most common scrub targets are:

  • placeholder titles such as Introduction, TLDR, Overview, Summary
  • garbled titles containing replacement characters like �
  • missing or suspicious author names such as Unknown
  • missing years or placeholder-style directory names like XXXX
  • malformed directory names created from bad metadata

These are candidate heuristics, not auto-rewrite authority. The paper content is the final source of truth.

Acceptance Standard

A scrubbed paper should be:

  • identifiable in the library
  • searchable by a meaningful title
  • attributed to a plausible first author or organization when known
  • assigned a real year when known
  • normalized into the standard directory naming scheme

If you cannot achieve that threshold from the source text, stop short of marking the paper and report the ambiguity to the user.

© ZimoLiao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/scrub of ZimoLiao/scholaraio.

Open the folder on GitHubat commit 777628b

Compare with similar skills

Scrub next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scrub compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scrub this skillZimoLiao/scholaraio577—~2kAutomated safety check: PassMIT
Openclaw Repair Sweepopenclaw/openclaw392k—~1.8kAutomated safety check: PassMIT
TDD Repairruvnet/ruflo74k—~1.6kAutomated safety check: NotesMIT
Scrub Issuepytorch/pytorch104k—~4.6kAutomated safety check: PassCustom licence
Incremental Implementationaddyosmani/agent-skills103k1 repos~2.3kAutomated safety check: PassMIT
DB Repairgarrytan/gbrain31k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Openclaw Repair Sweep

    openclaw/openclaw

    Run scoped OpenClaw issue/PR repair campaigns: coordinate workers, prove root causes, and land or close verified work under the requested authority.

    392k GitHub stars~1.8k tokensUpdated today
    DevelopmentAuto-check passed
  • TDD Repair

    ruvnet/ruflo

    Test-Driven Repair — given a failing test, spawn a bounded headless claude -p (Read/Edit/Bash only) that makes the test pass without modifying it.

    74k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check: notes
  • Scrub Issue

    pytorch/pytorch

    Fetch, analyze, reproduce, and minimize GitHub issue reproductions.

    104k GitHub stars~4.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Incremental Implementation

    addyosmani/agent-skills

    Delivers a change in thin vertical slices, each implemented, tested, verified and committed before the next, using vertical, contract-first or risk-first slicing.

    103k GitHub starsUsed in 1 repo~2.3k tokens
    Agent WorkflowsAuto-check passed
  • DB Repair

    garrytan/gbrain

    Auto-fix gbrain's Postgres access so the brain stays available.

    31k GitHub stars~1.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Detecting Bluetooth Low Energy Attacks

    mukul975/Anthropic-Cybersecurity-Skills

    Detects and analyzes Bluetooth Low Energy (BLE) security attacks including sniffing, replay attacks, GATT enumeration abuse, and Man-in-the-Middle interception.

    34k GitHub stars~3.4k tokensUpdated 1 mo ago
    SecurityAuto-check passed

More from ZimoLiao/scholaraio

All 43 skills in this repo
  • Document

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to create or inspect DOCX, PPTX, or XLSX files, generate a downloadable Office deliverable, or verify its structure and layout warnings with scholaraio…

    577 GitHub stars~674 tokensUpdated 14 days ago
    Auto-check passed
  • Academic Writing

    ZimoLiao/scholaraio

    A skill your agent uses when the user needs help choosing or organizing an academic-writing workflow by deliverable, stage, or format, including review articles, guided reading, paper sections, PPT…

    577 GitHub stars~827 tokensUpdated 14 days ago
    Auto-check passed
  • Arxiv

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to browse arXiv preprints, search arXiv directly, fetch a PDF by arXiv ID or URL, or send a preprint into the ScholarAIO ingest pipeline.

    577 GitHub stars~799 tokensUpdated 14 days ago
    Auto-check passed
  • Bioinformatics

    ZimoLiao/scholaraio

    A skill your agent uses when working on bioinformatics workflows such as alignment, variant calling, phylogenetics, or protein-structure analysis, especially across BLAST, minimap2, samtools…

    577 GitHub stars~1.4k tokensUpdated 14 days ago
    Auto-check passed
  • Citation Check

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to verify citations in AI-generated or human-written text against the local knowledge base and catch hallucinated, wrong, or missing references.

    577 GitHub stars~454 tokensUpdated 14 days ago
    Auto-check passed
  • Draw

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants diagrams, flowcharts, architecture visuals, data relationships, timelines, concept maps, Mermaid, Graphviz, drawio, or polished paper figures generated…

    577 GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check: notes

Questions about Scrub

What does Scrub do?

A skill your agent uses when incrementally reviewing and repairing low-quality metadata after enrich, especially non-standard documents that need title, author, or year correction while skipping…. Scrub is an agent skill from ZimoLiao/scholaraio.scrubbed.

When should I use Scrub?

Scrub fits situations like: incrementally reviewing and repairing low-quality metadata after enrich; especially non-standard documents that need title; year correction while skipping already reviewed records via .scrubbed.

How do I install Scrub in Claude Code?

Run `npx skills add ZimoLiao/scholaraio --skill scrub -a claude-code`. Or copy the skill folder (.claude/skills/scrub in ZimoLiao/scholaraio) into .claude/skills/scrub in your project. Claude Code loads it when a task matches its description.

How do I install Scrub in Codex?

Run `npx skills add ZimoLiao/scholaraio --skill scrub -a codex`. Or copy the skill folder (.claude/skills/scrub in ZimoLiao/scholaraio) into .agents/skills/scrub in your project. Codex loads it when a task matches its description.

Can I use Scrub in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZimoLiao/scholaraio --skill scrub -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scrub, .gemini/skills/scrub, .github/skills/scrub and .opencode/skills/scrub in your project.

What does Scrub need to run?

Going by SKILL.md and its folder, Scrub needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Scrub access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scrub safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scrub use?

Scrub is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scrub use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scrub?

Skills that share tags, products or a category with Scrub: Openclaw Repair Sweep (openclaw/openclaw, 392k stars), TDD Repair (ruvnet/ruflo, 74k stars), Scrub Issue (pytorch/pytorch, 104k stars) and Incremental Implementation (addyosmani/agent-skills, 103k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scrub?

ZimoLiao (a GitHub user) maintains it in ZimoLiao/scholaraio, which has 577 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 25, 2026.

Source: ZimoLiao/scholaraio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.