Agent skill

Tabular Document Review

by borghei in borghei/Claude-Skills

Extract structured data from multiple documents into comparison matrix with citations.

MITAuto-check passedLegal & Compliance

Install Tabular Document Review

skills CLI
$ npx skills add borghei/Claude-Skills --skill tabular-document-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills tabular-document-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/legal/tabular-document-review .claude/skills/tabular-document-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tabular-document-review
GitHub stars
881
Token cost
~3k tokens
SKILL.md length
1,015 words
Files
5 (incl. scripts, references)
Skills in repo
349
Repo updated
First seen
Licence
MIT

At a glance

Extract structured data from multiple documents into comparison matrix with citations.

  • Works in 2 steps: Document Discovery… → Extraction Aggregator…
  • Bulk document review
  • SKILL.md covers Overview, Table of Contents, Clarify First and Tools, plus 8 more sections
  • Runs Python scripts from its folder; calls python

What it does

Tabular Document Review is an agent skill from borghei/Claude-Skills. Extract structured data from multiple documents into comparison matrix with citations. Use for bulk document review.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/common_extraction_columns.md`, `references/extraction_methodology.md` and `scripts/document_discovery.py`).

It sits in Legal & Compliance, covering Contract review, Schema markup and Citation management. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Bulk document review
  • Tasks that involve Contract review
  • Tasks that involve Schema markup

Example prompts

  • “/tabular-document-review”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Document Discovery (scripts/document_discovery.py)
  2. Extraction Aggregator (scripts/extraction_aggregator.py)

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tabular Document Review loads about 3k tokens when it runs, and up to ~9.1k if it reads all its reference files. Until then it costs about 35 tokens; SKILL.md has 1,015 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 1,015 words, ~3,006 tokens.

Download SKILL.mdSave it as .claude/skills/tabular-document-review/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
tabular-document-review
description
Extract structured data from multiple documents into comparison matrix with citations. Use for bulk document review.
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
The Glass Room
metadata.category
legal
metadata.domain
document-review
metadata.updated
2026-04-10
metadata.tags
document-review, extraction, comparison-matrix, contracts, bulk-review

⚠️ EXPERIMENTAL — This skill is provided for educational and informational purposes only. It does NOT constitute legal advice. All responsibility for usage rests with the user. Consult qualified legal professionals before acting on any output.

Tabular Document Review Skill

Overview

Production-ready toolkit for extracting structured data from multiple legal documents into a comparison matrix with citations. Supports user-defined extraction columns, parallel processing with up to 10 agents, confidence scoring, and output in markdown table or structured JSON. Designed for legal teams performing bulk contract review, NDA comparison, employment agreement analysis, and lease review.

Table of Contents

Clarify First

Before the review, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • The extraction columns — define exactly what each cell holds; a vague "Date" matches dozens of dates, while "Effective Date" with guidance does not
  • Document set location and formats — drives discovery and how many parallel agents to allocate (ceil(N/10), max 10)
  • Document type — contracts, NDAs, employment, leases — selects the pre-defined column set and per-column extraction guidance

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the matrix.

Tools

1. Document Discovery (scripts/document_discovery.py)

Scan a directory for legal documents and generate an inventory manifest.

bash
python scripts/document_discovery.py /path/to/contracts

python scripts/document_discovery.py /path/to/ndas --types pdf,docx --json

python scripts/document_discovery.py /path/to/leases --types pdf,docx,txt,md --min-size 1024
2. Extraction Aggregator (scripts/extraction_aggregator.py)

Aggregate multiple extraction result JSONs into a unified comparison matrix.

bash
python scripts/extraction_aggregator.py \
  --results extraction_1.json extraction_2.json extraction_3.json

python scripts/extraction_aggregator.py \
  --results-dir ./extraction_results/ --json

python scripts/extraction_aggregator.py \
  --results-dir ./extraction_results/ \
  --format markdown \
  --output review_matrix.md

python scripts/extraction_aggregator.py \
  --results extraction_1.json extraction_2.json \
  --columns "Parties,Effective Date,Term,Governing Law"

Reference Guides

ReferencePurpose
references/extraction_methodology.mdDocument extraction best practices, JSON schema, agent prompts
references/common_extraction_columns.mdPre-defined column sets for contracts, NDAs, employment, leases

Workflows

5-Step Document Review Pipeline
StepActionToolOutput
1. Gather RequirementsDefine document folder, output filename, columns to extractManualColumn list, file path
2. Discover DocumentsScan directory for target documentsdocument_discovery.pyDocument manifest JSON
3. Process DocumentsExtract values per column with citations (parallel agents)AI agents (external)Per-document extraction JSONs
4. Collect ResultsAggregate extraction JSONs into unified matrixextraction_aggregator.pyConsolidated matrix
5. Generate OutputExport as markdown table or structured JSONextraction_aggregator.pyFinal deliverable
Parallel Processing Strategy
AgentsDocuments per AgentUse When
1All1-5 documents
2-3ceil(N/agents)6-15 documents
4-6ceil(N/agents)16-40 documents
7-10ceil(N/agents)41-100 documents
10 (max)ceil(N/10)100+ documents
Agent Prompt Template

Each agent receives a prompt structured as:

You are reviewing {count} legal documents. For each document, extract the
following columns:

{column_definitions}

For each value extracted:
1. Provide the exact value found
2. Include the page number (PDF) or section/paragraph (DOCX/MD)
3. Rate your confidence: HIGH (exact match), MEDIUM (inferred), LOW (uncertain)
4. If not found, record "NOT FOUND" with confidence LOW

Output as JSON per the extraction schema.
Confidence Scoring
LevelColor CodeDefinition
HIGHGreenExact value found with clear citation
MEDIUMYellowValue inferred from context; multiple possible interpretations
LOWRed / Not FoundValue uncertain or not found in document
Output Format

Sheet 1: Document Review

DocumentPartiesEffective DateTermGoverning Law...
contract_a.pdfAcme / Beta [p.1]2026-01-15 [p.2]3 years [p.3]Delaware [p.12]...
contract_b.pdfGamma / Delta [p.1]NOT FOUND2 years [p.4]New York [p.10]...

Sheet 2: Summary

MetricValue
Documents processed25
Columns extracted8
Average confidence87%
Not found rate12%

Extraction Scenarios

Contract Review
ColumnWhat to Extract
PartiesAll contracting parties with full legal names
Effective DateContract effective or execution date
TermDuration of the agreement
RenewalAuto-renewal terms and notice period
Governing LawJurisdiction governing the agreement
Liability CapMaximum liability amount or formula
IndemnificationIndemnification obligations and scope
IP OwnershipIntellectual property ownership provisions
Termination RightsTermination triggers and notice requirements
Data ProtectionData protection or privacy obligations
NDA Review
ColumnWhat to Extract
PartiesDisclosing and receiving parties
TypeMutual or one-way
Definition ScopeHow "confidential information" is defined
ExceptionsStandard exceptions to confidentiality
TermDuration of confidentiality obligations
SurvivalSurvival period after termination
Return/DestructionObligations on termination
RemediesAvailable remedies for breach
Show full SKILL.md (427 more words)Show less

Troubleshooting

ProblemCauseSolution
Discovery finds 0 documentsWrong path or file typesVerify path exists; check --types matches actual file extensions
Extraction JSONs have wrong schemaAgent prompt incompleteUse the extraction schema from extraction_methodology.md
Aggregator shows conflictsMultiple values for same cellReview source documents; aggregator marks conflicts for manual review
High "NOT FOUND" rateColumns too specific for document typeUse column definitions from common_extraction_columns.md; broaden definitions
Confidence all LOWAgent unable to locate valuesCheck column definitions are specific enough; verify document is readable
Aggregator crashes on large setToo many result files loaded at onceProcess in batches of 50 results; use --columns to limit output width
Markdown table misalignedLong values or special charactersUse --format json for machine processing; truncate long values
Missing citationsAgent did not include page/section referencesReinforce citation requirement in agent prompt; check extraction schema

Success Criteria

  • Extraction Coverage: 90%+ of defined columns populated across all documents
  • Confidence Distribution: 70%+ of extractions rated HIGH confidence
  • Citation Accuracy: Every extracted value includes verifiable page/section citation
  • Processing Speed: 50+ documents processed within 30 minutes using parallel agents
  • Matrix Completeness: Final matrix includes all documents and all columns with no orphan rows

Scope & Limitations

This skill covers:

  • Document inventory and discovery across PDF, DOCX, TXT, and MD formats
  • Aggregation of extraction results from parallel agent processing into unified matrix
  • Pre-defined column sets for contracts, NDAs, employment agreements, and leases
  • Confidence scoring and conflict detection for extracted values
  • Markdown and JSON output formats

This skill does NOT cover:

  • Actual document parsing or text extraction (requires external libraries or AI agents)
  • OCR processing for scanned documents
  • Excel/XLSX output generation (use JSON output and convert externally)
  • Automated legal analysis or risk assessment of extracted values
  • Document comparison or redlining between versions

Anti-Patterns

Anti-PatternWhy It FailsBetter Approach
Vague column definitions"Date" could match dozens of dates in a contractUse specific definitions: "Effective Date" with guidance on where to look
Skipping document discoveryUnknown document count leads to wrong agent allocationAlways run discovery first; use manifest for pipeline planning
Ignoring LOW confidence resultsMissing or uncertain data treated as factReview all LOW confidence cells manually; flag in final report
Processing 100+ docs with 1 agentSlow, context window overflow, quality degradationUse parallel processing: ceil(N/10) documents per agent, max 10 agents
No citation requirementCannot verify extracted values against sourceRequire page/section citation for every extraction; reject uncited values

Tool Reference

scripts/document_discovery.py

Scan directory for legal documents and generate inventory manifest.

usage: document_discovery.py [-h] [--json]
                              [--types TYPES]
                              [--min-size MIN_SIZE]
                              [--max-size MAX_SIZE]
                              directory

positional arguments:
  directory             Path to directory containing documents

options:
  -h, --help            Show help message and exit
  --json                Output in JSON format
  --types TYPES         Comma-separated file extensions to include
                        (default: pdf,docx,doc,txt,md,rtf)
  --min-size MIN_SIZE   Minimum file size in bytes (default: 0)
  --max-size MAX_SIZE   Maximum file size in bytes (default: no limit)
scripts/extraction_aggregator.py

Aggregate extraction results into unified comparison matrix.

usage: extraction_aggregator.py [-h] [--json]
                                 [--results RESULTS [RESULTS ...]]
                                 [--results-dir RESULTS_DIR]
                                 [--format {markdown,json}]
                                 [--columns COLUMNS]
                                 [--output OUTPUT]

options:
  -h, --help            Show help message and exit
  --json                Output in JSON format (alias for --format json)
  --results             One or more extraction result JSON files
  --results-dir         Directory containing extraction result JSON files
  --format              Output format: markdown table or JSON (default: markdown)
  --columns             Comma-separated column names to include (default: all)
  --output              Write output to file instead of stdout

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in legal/tabular-document-review of borghei/Claude-Skills.

  • SKILL.md
  • references/common_extraction_columns.md
  • references/extraction_methodology.md
  • scripts/document_discovery.py
  • scripts/extraction_aggregator.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Tabular Document Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tabular Document Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tabular Document Review this skillborghei/Claude-Skills881—~3kAutomated safety check: PassMIT
Contract SnapshotLegalQuants/lq-ai150—~1.8kAutomated safety check: PassApache-2.0
Legal Design Assessmentlawve-ai/awesome-legal-skills836—~3.7kAutomated safety check: PassCC-BY-4.0
Msa SnapshotLegalQuants/lq-ai150—~1.7kAutomated safety check: PassApache-2.0
Contract Output Formatterinfometa/workbuddyskills344—~1.6kAutomated safety check: PassNone
Contract Reviewevolsb/claude-legal-skill4621 repos~3.6kAutomated safety check: PassMIT

Similar skills

  • Contract Snapshot

    LegalQuants/lq-ai

    A skill your agent uses when the user wants to compare the same handful of terms across N contracts side-by-side in a grid — what is the term, survival period, carveouts, and governing law in each…

    150 GitHub stars~1.8k tokensUpdated today
    Legal & ComplianceAuto-check passed
  • Legal Design Assessment

    lawve-ai/awesome-legal-skills

    Audits a legal document against legal design principles. An agent skill from lawve-ai/awesome-legal-skills.

    836 GitHub stars~3.7k tokensUpdated 5 days ago
    Legal & ComplianceAuto-check passed
  • Msa Snapshot

    LegalQuants/lq-ai

    A skill your agent uses when the user wants to compare the substantive commercial terms across N master services agreements side-by-side — term and renewal, payment terms, limitation of liability…

    150 GitHub stars~1.7k tokensUpdated today
    Legal & ComplianceAuto-check passed
  • Contract Output Formatter

    infometa/workbuddyskills

    Format contract work products into the delivery shape that best matches the hit scenario (C1-C9) and the reading audience.

    344 GitHub stars~1.6k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Contract Review

    evolsb/claude-legal-skill

    Review legal contracts, NDAs, employment agreements, SaaS terms, and M&A documents.

    462 GitHub starsUsed in 1 repo~3.6k tokens
    Legal & ComplianceAuto-check passed
  • Commercial Legal Pl

    apiotrowski-afk/commercial-legal-pl

    Skill do analizy i tworzenia umów według polskiego prawa, ze szczególnym uwzględnieniem umów B2B, IP i IT (body leasing, NDA, wdrożenia, SaaS, przeniesienie praw autorskich, ugody).

    176 GitHub starsUsed in 1 repo~4.1k tokens
    Legal & ComplianceAuto-check passed

More from borghei/Claude-Skills

All 349 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    881 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    881 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    881 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    881 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Tabular Document Review

What does Tabular Document Review do?

Extract structured data from multiple documents into comparison matrix with citations. Tabular Document Review is an agent skill from borghei/Claude-Skills. Extract structured data from multiple documents into comparison matrix with citations.

When should I use Tabular Document Review?

Tabular Document Review fits situations like: bulk document review; tasks that involve Contract review; tasks that involve Schema markup.

How do I install Tabular Document Review in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill tabular-document-review -a claude-code`. Or copy the skill folder (legal/tabular-document-review in borghei/Claude-Skills) into .claude/skills/tabular-document-review in your project. Claude Code loads it when a task matches its description.

How do I install Tabular Document Review in Codex?

Run `npx skills add borghei/Claude-Skills --skill tabular-document-review -a codex`. Or copy the skill folder (legal/tabular-document-review in borghei/Claude-Skills) into .agents/skills/tabular-document-review in your project. Codex loads it when a task matches its description.

Can I use Tabular Document Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill tabular-document-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tabular-document-review, .gemini/skills/tabular-document-review, .github/skills/tabular-document-review and .opencode/skills/tabular-document-review in your project.

What does Tabular Document Review need to run?

Going by SKILL.md and its folder, Tabular Document Review needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Tabular Document Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tabular Document Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Tabular Document Review use?

Tabular Document Review is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tabular Document Review use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.1k tokens, read only when the agent opens those files.

What are the alternatives to Tabular Document Review?

Skills that share tags, products or a category with Tabular Document Review: Contract Snapshot (LegalQuants/lq-ai, 150 stars), Legal Design Assessment (lawve-ai/awesome-legal-skills, 836 stars), Msa Snapshot (LegalQuants/lq-ai, 150 stars) and Contract Output Formatter (infometa/workbuddyskills, 344 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tabular Document Review?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 881 GitHub stars. The repository holds 349 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.