Agent skill

Quant Paper Extractor

by CamusGIT in CamusGIT/EvoQuant

Convert quantitative research report PDFs to markdown, then extract structured knowledge (paperId, title, year, source, keywords, tldr, abstract, strategy, method, experiment, result) into JSONL…

Apache-2.0Auto-check passedDocuments & Office

Install Quant Paper Extractor

skills CLI
$ npx skills add CamusGIT/EvoQuant --skill quant-paper-extractor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install CamusGIT/EvoQuant quant-paper-extractor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/CamusGIT/EvoQuant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/EvoQuant/skills/quant-paper-extractor .claude/skills/quant-paper-extractor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
quant-paper-extractor
GitHub stars
151
Token cost
~2.4k tokens
SKILL.md length
854 words
Files
10 (incl. scripts, references, assets)
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Convert quantitative research report PDFs to markdown, then extract structured knowledge (paperId, title, year, source, keywords, tldr, abstract, strategy, method, experiment, result) into JSONL…

  • Works in 3 steps: PDF → Markdown → Markdown → JSONL → Refresh derived artifacts & verify
  • : adding quant research PDFs to the knowledge base
  • SKILL.md covers Setup, Pre-conditions, Phase 1: PDF → Markdown and Phase 2: Markdown → JSONL, plus 6 more sections
  • Runs Python scripts from its folder; calls python and pip

What it does

Quant Paper Extractor is an agent skill from CamusGIT/EvoQuant. Convert quantitative research report PDFs to markdown, then extract structured knowledge (paperId, title, year, source, keywords, tldr, abstract, strategy, method, experiment, result) into JSONL paper cards. Writes into the repo papers directory (papers/raw, papers/markdown, papers/cards) and refreshes contextbrief.md + index.jsonl. Use when: adding quant research PDFs to the knowledge base, extracting structured data from reports. Do NOT use for: academic paper search (use paper-navigator), idea generation (use…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts, reference files and assets (for example `assets/extraction-prompt.md`, `assets/jsonl-record-template.json` and `references/error-handling.md`).

It sits in Documents & Office, covering Brainstorming, Markdown and Deep research. The repository describes itself as: EvoQuant is a self-evolving AI research agent specialized in quantitative investment research. It runs the full research loop autonomously. The licence is Apache-2.0.

When your agent uses it

  • : adding quant research PDFs to the knowledge base
  • Extracting structured data from reports
  • : academic paper search (use paper-navigator)
  • Idea generation (use research-ideation)

Example prompts

  • “/quant-paper-extractor”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): write_file, edit_file, read_file, think_tool, execute

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. PDF → Markdown
  2. Markdown → JSONL
  3. Refresh derived artifacts & verify

What it can do on your machine

Read from SKILL.md and the folder at commit ac1c4b8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • write_file
    • edit_file
    • read_file
    • think_tool
    • execute

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Quant Paper Extractor loads about 2.4k tokens when it runs, and up to ~6.7k if it reads all its reference files. Until then it costs about 151 tokens; SKILL.md has 854 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~151
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from CamusGIT/EvoQuant at commit ac1c4b8, republished under its Apache-2.0 licence (© CamusGIT). 854 words, ~2,367 tokens.

Download SKILL.mdSave it as .claude/skills/quant-paper-extractor/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
quant-paper-extractor
description
Convert quantitative research report PDFs to markdown, then extract structured knowledge (paperId, title, year, source, keywords, tldr, abstract, strategy, method, experiment, result) into JSONL paper cards. Writes into the repo papers directory (papers/raw, papers/markdown, papers/cards) and refreshes context_brief.md + index.jsonl. Use when: adding quant research PDFs to the knowledge base, extracting structured data from reports. Do NOT use for: academic paper search (use paper-navigator), idea generation (use research-ideation), literature surveys (use research-survey).
allowed-tools
write_file, edit_file, read_file, think_tool, execute
metadata.author
quant-research-team
metadata.version
1.0.0
metadata.tags
quant, pdf-extraction, structured-data, jsonl, batch-processing

Quant Paper Extractor

Batch-convert quantitative research report PDFs (量化研究研报) to structured JSONL paper cards. Two-phase pipeline writing into the repo papers library:

papers/raw/{paperId}.pdf
      │
      ▼ Phase 1: PDF → Markdown (pdf_to_markdown.py)
papers/markdown/{paperId}.md
      │
      ▼ Phase 2: Markdown → card (agent-driven extraction)
papers/cards/{paperId}.jsonl
      │
      ▼ Phase 3: refresh derived artifacts
papers/context_brief.md + papers/index.jsonl

Setup

Scripts at scripts/. Run via python scripts/<name>.py from this skill's directory. Dependencies: pip install -e ..

Pre-conditions

The papers directory lives at the repo root (papers/; override with EVOSCIENTIST_PAPERS_DIR). Resolve it once and reuse:

bash
PAPERS=$(python -c "from EvoQuant.papers.paths import resolve_papers_dir; print(resolve_papers_dir())")

Layout (all three layers share the paperId join key):

$PAPERS/raw/       ← PDFs, renamed to {paperId}.pdf (see Red Line 8)
$PAPERS/markdown/  ← auto-created; converted markdown
$PAPERS/cards/     ← auto-created; extracted JSONL cards
$PAPERS/manifest.jsonl  ← processing-state ledger

Legacy layouts are dead. rawpaper/, markdown/, wiki/ in a workspace are deprecated — if you meet them, point the user at python -m EvoQuant.papers.migrate instead of writing there.

markdown/ is auto-created by Phase 1; create the cards dir up front:

bash
mkdir -p "$PAPERS/cards"

Phase 1: PDF → Markdown

Run the conversion script (fully automated, no LLM needed):

bash
python scripts/pdf_to_markdown.py \
  --rawpaper-dir "$PAPERS/raw" \
  --markdown-dir "$PAPERS/markdown" \
  --manifest-path "$PAPERS/manifest.jsonl"

This script:

  1. Scans all .pdf files in the raw dir (any filename — the hash is the identity)
  2. Computes SHA-256 hash of each PDF's binary content → paperId
  3. Incremental skip: if markdown/{paperId}.md already exists, skip
  4. Extracts text via three-tier fallback:
    • pymupdf4llm.to_markdown() — native Markdown output (best quality)
    • pymupdf.open() → page.get_text() — plain text with page headers
    • pypdf.PdfReader() → page.extract_text() — last resort
  5. Hard-truncates at 120,000 characters
  6. Writes structured Markdown to markdown/{paperId}.md
  7. Updates manifest.jsonl with status

Read the script's stdout for per-file success/failure reports.

Phase 2: Markdown → JSONL

The agent (you) performs the extraction reasoning. The extract.py script prepares context and validates output.

Step 2.1: List unextracted markdowns
bash
python scripts/manifest.py list \
  --manifest-path "$PAPERS/manifest.jsonl" --status markdown_done

Also check markdown_short status files. For each file without a corresponding cards/{paperId}.jsonl:

Step 2.2: Prepare extraction context
bash
python scripts/extract.py prepare \
  --markdown-file "$PAPERS/markdown/{paperId}.md"

This outputs:

  • MODE: single-pass or MODE: two-pass
  • The markdown content (truncated to 120,000 chars)
  • Section locator with heading tags (for two-pass mode)
  • Field template from assets/jsonl-record-template.json
Step 2.3: Extract fields (agent-driven)

Read the prepare output, then use think_tool to reason through the extraction following the rules below.

Read references/field-definitions.md for detailed field specs, word limits, and evidence rules.

Read references/quant-report-structure.md for section-heading heuristics and terminology glossary.

For two-pass mode, read references/two-pass-extraction.md for the detailed protocol:

  • Pass 1 (first ~8,000 chars): Extract paperId, title, year, source, tldr, abstract, keywords
  • Pass 2 (targeted sections via section locator): Extract strategy, method, experiment, result
Step 2.4: Write the JSONL record

Write a single JSON line to $PAPERS/cards/{paperId}.jsonl via write_file (or python -c if write_file's sandbox can't reach the papers library root).

Each record must have exactly these 11 fields:

FieldTypeWord LimitDescription
paperIdstrN/ASHA-256 of PDF binary content
titlestr≤30title of the report
yearint4 digitsPublication year
sourcestr≤10Source organization
keywordslist[str]3-8Quant finance keywords from the document
tldrstr≤40One-sentence core finding
abstractstr≤150Concise summary: question + approach + conclusion.
strategystr≤300Strategy description + evidence citation
methodstr≤300Methodology + evidence citation
experimentstr≤300Experimental setup + evidence citation
resultstr≤200Key metrics + evidence citation
Step 2.5: Validate
bash
python scripts/extract.py validate \
  --record "$PAPERS/cards/{paperId}.jsonl"

If validation fails, fix the record and re-validate.

Step 2.6: Update manifest

After successful validation, the manifest is updated automatically. Alternatively, rebuild from filesystem:

bash
python scripts/manifest.py rebuild \
  --rawpaper-dir "$PAPERS/raw" \
  --markdown-dir "$PAPERS/markdown" \
  --wiki-dir "$PAPERS/cards" \
  --manifest-path "$PAPERS/manifest.jsonl"
Show full SKILL.md (370 more words)Show less

Phase 3: Refresh derived artifacts & verify

After all cards are written, refresh the derived index/brief so the runtime tools see the new papers, then validate the whole cards directory:

bash
python -m EvoQuant.papers.refresh
python scripts/extract.py validate \
  --wiki-dir "$PAPERS/cards" --manifest-path "$PAPERS/manifest.jsonl"

refresh is a full recompute over cards/ (milliseconds) — cheap to run after every card, and the derived files can never drift.

Then report:

Extraction complete:
  PDFs scanned:     N
  Markdown created: M (K skipped, already existed)
  JSONL created:    P (Q skipped, already existed)
  Errors:           E
  Warnings:         W (short markdown, empty fields, etc.)

The report must also tell the user: 语料已更新,paper 工具自新会话起可用 (tools mount at agent startup — a restart is required, this is expected).

Red Lines (always)

  1. No fabrication. Every extracted field must be grounded in the source text. If information is not explicitly found, return empty string. Do not infer.
  2. Shape check. Every JSONL record must have exactly 11 fields, all present and non-null. Empty strings are allowed but flagged.
  3. Word limits. Strictly enforced per field (see references/field-definitions.md).
  4. English only. All output fields must be in English. Translate Chinese source text.
  5. Incremental. Never re-process a PDF that already has a markdown file, or a markdown that already has a JSONL.
  6. Quote-or-zero. The strategy, method, experiment, and result fields must each include at least one [source: "..."] inline evidence citation. No citation → flag warning.
  7. If information is not explicitly found, return empty string. Do not infer.
  8. One join key. After a PDF is processed, rename it in $PAPERS/raw/ to {paperId}.pdf (the hash) — raw/, markdown/, cards/ must stay keyed identically. The original filename survives in manifest sourcePdf.
  9. Write only inside the papers library. Never create rawpaper/, markdown/, or wiki/ in a workspace — that layout is deprecated (see AGENT.md).

Error Handling

Read references/error-handling.md for full details. Summary:

  • PDF extraction fails → skip, status=pdf_error, continue
  • Markdown too short (< 500 chars) → warn, attempt extraction, flag as likely incomplete
  • Empty required field → retry once; if still empty, write "", flag in _warnings
  • Validation failure → log errors, do not overwrite, let agent decide
  • Partial failure → continue with next file, do not roll back

References

FileRead when
references/field-definitions.mdUnderstanding field types, word limits, and evidence rules
references/quant-report-structure.mdUnderstanding quant report structure and terminology
references/two-pass-extraction.mdLong document extraction protocol
references/error-handling.mdFailure modes and recovery

Assets

FileUse
assets/jsonl-record-template.jsonTemplate for a single JSONL record
assets/extraction-prompt.mdExtraction rules and prompt template

Hand off to

GoalSkill
Find academic paperspaper-navigator
Research ideationresearch-ideation

© CamusGIT, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references, assets) in EvoQuant/skills/quant-paper-extractor of CamusGIT/EvoQuant.

  • SKILL.md
  • assets/extraction-prompt.md
  • assets/jsonl-record-template.json
  • references/error-handling.md
  • references/field-definitions.md
  • references/quant-report-structure.md
  • references/two-pass-extraction.md
  • scripts/extract.py
  • scripts/manifest.py
  • scripts/pdf_to_markdown.py

Open the folder on GitHubat commit ac1c4b8

Compare with similar skills

Quant Paper Extractor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Quant Paper Extractor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Quant Paper Extractor this skillCamusGIT/EvoQuant151—~2.4kAutomated safety check: PassApache-2.0
Paper Interpretationdigoal/blog8.6k—~1.5kAutomated safety check: PassGPL-2.0
Paper LensYSQ-boop/paper-lens101—~1.3kAutomated safety check: PassApache-2.0
Textbook To Mddrpwchen/textbook-to-note104—~3.5kAutomated safety check: PassMIT
Fulltext RetrievalAperivue/medsci-skills329—~1.9kAutomated safety check: PassMIT
Adu Corrections PDFmikeOnBreeze/cc-crossbeam293—~1.5kAutomated safety check: PassMIT

Similar skills

  • 从论文 PDF 文件或论文 PDF URL 生成通俗易懂、图文并茂、带批判性评估的中文 Markdown 解读,并保存到当前项目的 markdown 目录。Use when the user asks to interpret,精读,解读,summarize,explain,analyze, or write an article from an academic paper PDF…

    8.6k GitHub stars~1.5k tokensUpdated 9 days ago
    Documents & OfficeAuto-check passed
  • Paper Lens

    YSQ-boop/paper-lens

    Read and critically analyze one academic paper from an arXiv URL/ID or a local PDF, producing a source-grounded Markdown report that can grow from a quick read into a reviewer-level deep review.

    101 GitHub stars~1.3k tokensUpdated 8 days ago
    Documents & OfficeAuto-check passed
  • Textbook To Md

    drpwchen/textbook-to-note

    Convert PDF/EPUB textbooks to searchable markdown files for an AI agent's own reference.

    104 GitHub stars~3.5k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Fulltext Retrieval

    Aperivue/medsci-skills

    A skill your agent uses when you need full-text PDFs for a list of DOIs, such as a meta-analysis screening set.

    329 GitHub stars~1.9k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Adu Corrections PDF

    mikeOnBreeze/cc-crossbeam

    Formats a draft corrections letter (markdown) into a professional PDF.

    293 GitHub stars~1.5k tokensUpdated 7 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes

More from CamusGIT/EvoQuant

  • Local Paper Navigator

    CamusGIT/EvoQuant

    Find and read papers from the local papers library (repo papers/, mounted at /papers/).

    151 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Quant Experiment Runtime

    CamusGIT/EvoQuant

    Quant research experiment executor: discover an offline source database under the workdir's code-repo, build a panel, run a Research Artifact's entry point to compute research-object values, and…

    151 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Paper Review

    CamusGIT/EvoQuant

    Guides self-review of YOUR OWN academic paper before submission with adversarial stress-testing.

    151 GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed
  • Research Ideation

    CamusGIT/EvoQuant

    Quant-focused research ideation pipeline: scope selection (3 stages) → anchor-first literature grounding → single-core idea generation → iterative refinement → ELO tournament ranking (Final =…

    151 GitHub stars~5.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Find Skills

    CamusGIT/EvoQuant

    Helps users discover agent skills from the open ecosystem. An agent skill from CamusGIT/EvoQuant.

    151 GitHub stars~595 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Quant Paper Extractor

What does Quant Paper Extractor do?

Convert quantitative research report PDFs to markdown, then extract structured knowledge (paperId, title, year, source, keywords, tldr, abstract, strategy, method, experiment, result) into JSONL…. Quant Paper Extractor is an agent skill from CamusGIT/EvoQuant. Convert quantitative research report PDFs to markdown, then extract structured knowledge (paperId, title, year, source, keywords, tldr, abstract, strategy, method, experiment, result) into JSONL paper cards.

When should I use Quant Paper Extractor?

Quant Paper Extractor fits situations like: : adding quant research PDFs to the knowledge base; extracting structured data from reports; : academic paper search (use paper-navigator); idea generation (use research-ideation).

How do I install Quant Paper Extractor in Claude Code?

Run `npx skills add CamusGIT/EvoQuant --skill quant-paper-extractor -a claude-code`. Or copy the skill folder (EvoQuant/skills/quant-paper-extractor in CamusGIT/EvoQuant) into .claude/skills/quant-paper-extractor in your project. Claude Code loads it when a task matches its description.

How do I install Quant Paper Extractor in Codex?

Run `npx skills add CamusGIT/EvoQuant --skill quant-paper-extractor -a codex`. Or copy the skill folder (EvoQuant/skills/quant-paper-extractor in CamusGIT/EvoQuant) into .agents/skills/quant-paper-extractor in your project. Codex loads it when a task matches its description.

Can I use Quant Paper Extractor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add CamusGIT/EvoQuant --skill quant-paper-extractor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/quant-paper-extractor, .gemini/skills/quant-paper-extractor, .github/skills/quant-paper-extractor and .opencode/skills/quant-paper-extractor in your project.

What does Quant Paper Extractor need to run?

Going by SKILL.md and its folder, Quant Paper Extractor needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: write_file, edit_file, read_file, think_tool, execute.

Does Quant Paper Extractor access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Quant Paper Extractor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Quant Paper Extractor use?

Quant Paper Extractor is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Quant Paper Extractor use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.3k tokens, read only when the agent opens those files.

What are the alternatives to Quant Paper Extractor?

Skills that share tags, products or a category with Quant Paper Extractor: Paper Interpretation (digoal/blog, 8.6k stars), Paper Lens (YSQ-boop/paper-lens, 101 stars), Textbook To Md (drpwchen/textbook-to-note, 104 stars) and Fulltext Retrieval (Aperivue/medsci-skills, 329 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Quant Paper Extractor?

CamusGIT (a GitHub user) maintains it in CamusGIT/EvoQuant, which has 151 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 2, 2026.

Source: CamusGIT/EvoQuant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.