Agent skill

Document Grounding

by gaotiexinqu in gaotiexinqu/OneResearchClaw

Convert a raw document into a structured grounding note for downstream research and summarization.

MITAuto-check passedDocuments & Office

Install Document Grounding

skills CLI
$ npx skills add gaotiexinqu/OneResearchClaw --skill document-grounding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gaotiexinqu/OneResearchClaw document-grounding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gaotiexinqu/OneResearchClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/document-grounding .claude/skills/document-grounding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
document-grounding
GitHub stars
450
Token cost
~2k tokens
SKILL.md length
763 words
Files
5 (incl. scripts)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Convert a raw document into a structured grounding note for downstream research and summarization.

  • Works in 5 steps: Run the extraction script through the… → Read → If extracted.md contains AssetRef… → …
  • Tasks that involve Summarization
  • SKILL.md covers When to Use, Input, Output Bundle and Required Agent Workflow, plus 8 more sections
  • Runs Python and Shell scripts from its folder

What it does

Document Grounding is an agent skill from gaotiexinqu/OneResearchClaw. Convert a raw document into a structured grounding note for downstream research and summarization.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `scripts/extract_simple.py`, `scripts/ground_document.py` and `scripts/lightweight_extract.py`).

It sits in Documents & Office, covering Summarization. It works with Microsoft Word. The repository describes itself as: Any research. One Claw. 🦞 From any materials to research with fully autonomous & skill-driven researcher. The licence is MIT.

When your agent uses it

  • Tasks that involve Summarization

Example prompts

  • “/document-grounding”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Run the extraction script through the existing run.sh entrypoint.
  2. Read
  3. If extracted.md contains AssetRef blocks, inspect the referenced files in assets/... before writing grounded.md.
  4. Write grounded.md as a real structured grounding note.
  5. Do not stop after merely confirming that the extraction bundle exists.

What it can do on your machine

Read from SKILL.md and the folder at commit 37e86c6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python and Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Document Grounding loads about 2k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 763 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from gaotiexinqu/OneResearchClaw at commit 37e86c6, republished under its MIT licence (© gaotiexinqu). 763 words, ~2,024 tokens.

Download SKILL.mdSave it as .claude/skills/document-grounding/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
document-grounding
description
Convert a raw document into a structured grounding note for downstream research and summarization.

Document Grounding

Convert a raw document into a structured grounding note.

This skill is for document grounding, not a narrative recap. It should produce a stable intermediate note that is easy to read and easy for downstream skills to use.

When to Use

Use this skill when:

  • the input is a single document file
  • the document may be a PDF, DOCX, Markdown, or TXT file
  • you need structured document notes before downstream research or summary work
  • the document may contain non-textual evidence such as tables, figures, diagrams, formulas, or code blocks

Do not use this skill when:

  • the task is to write a polished final report
  • the task is to perform external literature search
  • the task is to render output into PDF/DOCX
  • the input is a meeting transcript and you want meeting-specific grounding

Input

A single document file.

Supported first-stage formats:

  • .pdf
  • .docx
  • .md
  • .txt

The document may contain:

  • plain text
  • section headings
  • tables
  • figures / diagrams
  • formulas
  • code blocks
  • captions
  • layout / reading-order challenges

Output Bundle

For each input document, create one bundle directory:

text
data/grounded_notes/<type>-<doc_id>_<timestamp>/

where <type> is the file extension (e.g. pdf, docx, md, txt), <doc_id> is the sanitized filename without extension, and <timestamp> is the Beijing-time execution timestamp (format: YYYYMMDDHHMMSS).

For example:

text
data/grounded_notes/pdf-paper_name_20260410153022/
data/grounded_notes/docx-notes_001_20260410153100/
data/grounded_notes/md-project_readme_20260410153215/

Inside that bundle, the expected outputs are:

text
<bundle_dir>/
├─ ground_id.txt       # Ground ID for this unit (reused by all downstream stages)
├─ extracted.md
├─ extracted_meta.json
├─ asset_index.json
├─ grounded.md
└─ assets/
   ├─ tables/
   ├─ figures/
   └─ formulas/

The <ground_id> (e.g. pdf-paper_name_20260410153022) is the single stable identifier for the entire pipeline — all downstream directories (lit_inputs, lit_results, report_inputs, review_outputs, reports, final_outputs) reuse this same <ground_id>.

Important separation of responsibilities
  • ground_document.py is responsible for building the extraction bundle:
    • extracted.md
    • extracted_meta.json
    • asset_index.json
    • assets/...
  • The agent is responsible for reading those outputs and then writing the final:
    • grounded.md

grounded.md must be a real grounding note. It must not remain a placeholder scaffold.

Required Agent Workflow

  1. Run the extraction script through the existing run.sh entrypoint.
  2. Read:
    • extracted.md
    • extracted_meta.json
    • asset_index.json
  3. If extracted.md contains AssetRef blocks, inspect the referenced files in assets/... before writing grounded.md.
  4. Write grounded.md as a real structured grounding note.
  5. Do not stop after merely confirming that the extraction bundle exists.

AssetRef Rule

extracted.md may contain blocks such as:

markdown
[AssetRef]
type: figure
id: figure_001
path: assets/figures/figure_001.png
instruction: Inspect this asset before writing grounded.md if it is relevant to the document's claims, comparisons, or conclusions.
[/AssetRef]

and

markdown
[AssetRef]
type: table
id: table_001
path_md: assets/tables/table_001.md
path_csv: assets/tables/table_001.csv
instruction: Inspect this table before writing grounded.md if it contains key evidence, comparisons, or numerical results.
[/AssetRef]

These are not decorative markers. They indicate that important evidence may exist outside the plain extracted text. If an AssetRef appears relevant, the agent must inspect the referenced asset before final grounding.

Output Format

Return markdown with exactly these sections.

markdown
# Document Grounding

## 1. Main Topic / Purpose
[2–4 sentence statement of the document’s main topic and purpose.]

## 2. Main Points
- [One bullet per major sub-topic, contribution, or argument]

## 3. Key Findings / Claims
- [Only items explicitly stated or strongly supported by the document]

## 4. Constraints / Risks
- [Only constraints, limitations, caveats, or risks explicitly stated or strongly supported by the document]

## 4a. Important Non-Textual Elements
- [Key tables, figures, diagrams, formulas, or code blocks that materially affect interpretation]
- [Mention how they affect the reading of the document if relevant]

## 5. Unresolved Issues
- [Preserve uncertainty, unanswered questions, incomplete evidence, or limitations that remain open]

## 6. Suggested Next Steps
- [Concrete follow-up directions grounded in the document]
- [Do not invent owners, deadlines, or commitments]

## 7. Search Keywords

### Problem Keywords
- ...

### Method / Solution Keywords
- ...

### Domain / Constraint Keywords
- ...

Instructions

  • Do not write a chronological recap of the document.
  • Do not invent facts, conclusions, owners, deadlines, or commitments.
  • Do not turn suggestions, hypotheses, or future work into established conclusions.
  • Do not add external knowledge, acronym expansions, years, benchmark details, or metadata not explicitly stated or strongly evidenced in the document.
  • Remove low-signal boilerplate when it does not affect meaning.
  • Group related content into high-level sub-topics instead of page-by-page bullets.
  • Keep the output concise and structured.
  • Treat non-textual evidence as first-class evidence when relevant.
Show full SKILL.md (297 more words)Show less

Special Rules

Key Findings / Claims

Only include something here if it is explicitly stated or strongly supported by the document. If it is merely hinted, proposed, or speculative, do not present it as a confirmed finding.

Constraints / Risks

Only include constraints, limitations, or risks that are explicitly stated or strongly evidenced. Do not infer hidden constraints from weak hints.

Important Non-Textual Elements

Only mention tables, figures, formulas, diagrams, or code blocks that materially affect interpretation. Do not list every asset mechanically. If an asset was referenced in extracted.md through an AssetRef block and it appears relevant, inspect it before deciding whether to include it.

Suggested Next Steps

Only include follow-up actions or directions that are explicitly discussed or strongly implied by the document. Do not invent action owners, deadlines, or commitments.

Search Keywords

Use specific noun phrases that are useful for later search. Avoid generic terms such as:

  • project
  • document
  • update
  • issue
  • optimization
  • system

Handling Difficult Cases

  • If the document is repetitive, summarize repeated points once.
  • If the document is incomplete, paraphrase only when meaning is clear.
  • If sections conflict, keep the conflict under unresolved issues.
  • If the main topic is genuinely unclear, write: Unclear — multiple loosely related topics appear in the document.
  • For long documents, identify major section shifts first, then synthesize.
  • If non-textual assets are present but unclear, mention them cautiously rather than over-claiming their meaning.

Execution Rule

The task is not complete if grounded.md is missing, empty, or still contains placeholder scaffold text. The agent must overwrite or newly create grounded.md with a real grounding note based on:

  • extracted.md
  • extracted_meta.json
  • asset_index.json
  • any relevant files referenced from assets/...

Example Invocation

/document-grounding

Environment Variables

When running the extraction script, ensure the Docling model path is set:

bash
export DOCLING_ARTIFACTS_PATH="${PROJECT_ROOT}/models/docling"

Supported model path locations (checked in order):

  • ${PROJECT_ROOT}/models/docling
  • ${PROJECT_ROOT}/models/docling-project/docling-models
  • /root/.cache/docling/models

© gaotiexinqu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in .cursor/skills/document-grounding of gaotiexinqu/OneResearchClaw.

  • SKILL.md
  • scripts/extract_simple.py
  • scripts/ground_document.py
  • scripts/lightweight_extract.py
  • scripts/run.sh

Open the folder on GitHubat commit 37e86c6

Compare with similar skills

Document Grounding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Document Grounding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Document Grounding this skillgaotiexinqu/OneResearchClaw450—~2kAutomated safety check: PassMIT
Doc SummarizerBlackBeltTechnology/pi-agent-dashboard315—~1.2kAutomated safety check: PassMIT
Document Generationautomateyournetwork/netclaw676—~1.9kAutomated safety check: PassApache-2.0
Tech Distilleryaofeino1/tech-distiller144—~1.2kAutomated safety check: PassNone
Beautiful ArticleConardLi/garden-skills13k—~4.7kAutomated safety check: PassMIT
Industry Bid Document WriterGet00/BiaoShu-SKILL173—~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • Doc Summarizer

    BlackBeltTechnology/pi-agent-dashboard

    Summarize documents of any size: extract with the document-converter engine, chunk to fit context, fan out to subagents, then synthesize one unified summary.

    315 GitHub stars~1.2k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Document Generation

    automateyournetwork/netclaw

    Generate Word (.docx), Excel (.xlsx) and PowerPoint (.pptx) documents and fill existing PDF forms, from real NetClaw data, with per-element provenance and no fabrication.

    676 GitHub stars~1.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Tech Distiller

    yaofeino1/tech-distiller

    Collects technical documents from URLs, files, wikis or APIs, filters them by role or technical direction, and writes a structured report on architecture, algorithms and design decisions.

    144 GitHub stars~1.2k tokensUpdated 6 mo ago
    Knowledge ManagementAuto-check passed
  • Beautiful Article

    ConardLi/garden-skills

    Turns a URL, PDF, DOCX, Markdown file, text or screenshots into a designed, shareable single-file HTML article through a staged review workflow.

    13k GitHub stars~4.7k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Industry Bid Document Writer

    Get00/BiaoShu-SKILL

    Converts a tender document into Markdown, extracts scoring criteria and requirements, then drafts an industry-formatted technical bid as a Word file.

    173 GitHub stars~4.6k tokensUpdated 16 days ago
    Documents & OfficeAuto-check passed
  • Persian Writing

    ali2000hos/persian-writing

    Write natural, human-sounding Persian (Farsi) and build correct right-to-left Persian deliverables.

    368 GitHub stars~4.4k tokensUpdated 6 days ago
    Documents & OfficeAuto-check: warnings

More from gaotiexinqu/OneResearchClaw

All 15 skills in this repo
  • Remote Input

    gaotiexinqu/OneResearchClaw

    Download remote content (arxiv papers, YouTube videos, Bilibili videos) to local storage and route to downstream grounding pipeline.

    450 GitHub stars~2.8k tokensUpdated 5 mo ago
    Auto-check passed
  • Grounded Research Lit

    gaotiexinqu/OneResearchClaw

    Run focused literature and web research from a grounded note.

    450 GitHub stars~11k tokensUpdated 5 mo ago
    Auto-check passed
  • Archive Grounding

    gaotiexinqu/OneResearchClaw

    Unpack a ZIP archive, inventory its files, run the corresponding child grounding skill for each supported child file, and then write a real archive-level grounded.md.

    450 GitHub stars~2.1k tokensUpdated 5 mo ago
    Auto-check: notes
  • Meeting Audio Grounding

    gaotiexinqu/OneResearchClaw

    Convert a meeting audio file into a transcript bundle, then use meeting-grounding to produce structured meeting grounding outputs.

    450 GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Meeting Video Grounding

    gaotiexinqu/OneResearchClaw

    Convert a meeting video into an audio-first transcript bundle, then use meeting-grounding to produce structured meeting grounding outputs.

    450 GitHub stars~1.2k tokensUpdated 5 mo ago
    Auto-check passed
  • PPTX Grounding

    gaotiexinqu/OneResearchClaw

    Extract a structured evidence bundle from a .pptx deck, then write a real grounded.md from the bundle.

    450 GitHub stars~2.1k tokensUpdated 5 mo ago
    Auto-check: notes

Works with

Questions about Document Grounding

What does Document Grounding do?

Convert a raw document into a structured grounding note for downstream research and summarization. Document Grounding is an agent skill from gaotiexinqu/OneResearchClaw. Convert a raw document into a structured grounding note for downstream research and summarization.

When should I use Document Grounding?

Document Grounding fits situations like: tasks that involve Summarization.

How do I install Document Grounding in Claude Code?

Run `npx skills add gaotiexinqu/OneResearchClaw --skill document-grounding -a claude-code`. Or copy the skill folder (.cursor/skills/document-grounding in gaotiexinqu/OneResearchClaw) into .claude/skills/document-grounding in your project. Claude Code loads it when a task matches its description.

How do I install Document Grounding in Codex?

Run `npx skills add gaotiexinqu/OneResearchClaw --skill document-grounding -a codex`. Or copy the skill folder (.cursor/skills/document-grounding in gaotiexinqu/OneResearchClaw) into .agents/skills/document-grounding in your project. Codex loads it when a task matches its description.

Can I use Document Grounding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gaotiexinqu/OneResearchClaw --skill document-grounding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/document-grounding, .gemini/skills/document-grounding, .github/skills/document-grounding and .opencode/skills/document-grounding in your project.

What does Document Grounding need to run?

Going by SKILL.md and its folder, Document Grounding needs Python and a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does Document Grounding access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Document Grounding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Document Grounding use?

Document Grounding is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Document Grounding use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Document Grounding?

Skills that share tags, products or a category with Document Grounding: Doc Summarizer (BlackBeltTechnology/pi-agent-dashboard, 315 stars), Document Generation (automateyournetwork/netclaw, 676 stars), Tech Distiller (yaofeino1/tech-distiller, 144 stars) and Beautiful Article (ConardLi/garden-skills, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Document Grounding?

gaotiexinqu (a GitHub user) maintains it in gaotiexinqu/OneResearchClaw, which has 450 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on May 9, 2026.

Source: gaotiexinqu/OneResearchClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.