Summarize documents of any size: extract with the document-converter engine, chunk to fit context, fan out to subagents, then synthesize one unified summary.

MITAuto-check passedDocuments & Office

Install Doc Summarizer

skills CLI
$ npx skills add BlackBeltTechnology/pi-agent-dashboard --skill doc-summarizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BlackBeltTechnology/pi-agent-dashboard doc-summarizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BlackBeltTechnology/pi-agent-dashboard.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/document-converter/.pi/skills/doc-summarizer .claude/skills/doc-summarizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
doc-summarizer
GitHub stars
315
Token cost
~1.2k tokens
SKILL.md length
354 words
Files
1
Skills in repo
66
Repo updated
First seen
Licence
MIT

At a glance

Summarize documents of any size: extract with the document-converter engine, chunk to fit context, fan out to subagents, then synthesize one unified summary.

  • Works in 2 steps: Extract to Markdown via the engine → Decide direct vs. chunked
  • Tasks that involve Summarization
  • SKILL.md covers Prerequisites, Step 1 — Extract to Markdown…, Step 2 — Decide direct vs.… and Step 3a — Direct summarization…, plus 4 more sections
  • Calls npm

What it does

Doc Summarizer is an agent skill from BlackBeltTechnology/pi-agent-dashboard. Summarize documents of any size: extract with the document-converter engine, chunk to fit context, fan out to subagents, then synthesize one unified summary. Handles PDF, DOCX, PPTX, XLSX, HTML, CSV, TXT, MD. Triggers: "summarize this document", "what's in this PDF", "give me a summary of these files", "extract key points from", "condense this document", "TL;DR of this file".

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Summarization, Word documents and PowerPoint presentations. It works with Microsoft Excel, Microsoft PowerPoint, Microsoft Word and Docker. The repository describes itself as: Real-time web dashboard for pi coding-agent sessions. Multi-session view, live chat mirroring, integrated terminal, diff viewer, pi-flows execution, and mobile-first remote… The licence is MIT.

When your agent uses it

  • Tasks that involve Summarization
  • Tasks that involve Word documents
  • Tasks that involve PowerPoint presentations

Example prompts

  • “summarize this document”
  • “s in this PDF”
  • “give me a summary of these files”
  • “/doc-summarizer”

Requirements

  • Python 3
  • Docker

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Extract to Markdown via the engine
  2. Decide direct vs. chunked

What it can do on your machine

Read from SKILL.md and the folder at commit 86e8e4d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Doc Summarizer loads about 1.2k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 354 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from BlackBeltTechnology/pi-agent-dashboard at commit 86e8e4d, republished under its MIT licence (© BlackBeltTechnology). 354 words, ~1,158 tokens.

Download SKILL.mdSave it as .claude/skills/doc-summarizer/SKILL.md (or your agent's skills folder).
name
doc-summarizer
description
Summarize documents of any size: extract with the document-converter engine, chunk to fit context, fan out to subagents, then synthesize one unified summary. Handles PDF, DOCX, PPTX, XLSX, HTML, CSV, TXT, MD. Triggers: "summarize this document", "what's in this PDF", "give me a summary of these files", "extract key points from", "condense this document", "TL;DR of this file".

Document Summarizer

Summarize documents of any size. Extraction goes through the document-converter engine facade (dc.convertToMarkdown) — the same Docker-quarantined engine the document-converter skill uses. There are NO host-side extractor scripts here; the facade is the only extraction surface. Chunking and synthesis are agent work.

Prerequisites

  • The document-converter package built and runnable: Docker available, image built (cd packages/document-converter && npm run build:image). See the document-converter SKILL for the full facade contract.
  • Nothing else. No pdftotext/pandoc/Python on the host — the engine owns all format handling inside Docker.

Step 1 — Extract to Markdown via the engine

Call the facade; never invoke Python, docling, or pdftotext directly.

ts
import { createDocumentConverter } from "@blackbelt-technology/pi-dashboard-document-converter";
const dc = createDocumentConverter({
  image: "pi-doc-engine:0.1.0",
  stagingDir: "/abs/staging",
  mounts: ["/docs"],   // files outside cwd need a root (mounts or workspaceRoot), else PATH_NOT_ALLOWED
});

const { output } = await dc.convertToMarkdown("<file_path>");              // digital PDF/DOCX/…
// scanned PDF: pass OCR explicitly
await dc.convertToMarkdown("<file_path>", { ocr: { mode: "force", lang: ["english"] } });

The result is a provenance-stamped .md in stagingDir. Read that file to get the document text. On failure the call rejects with DocConverterError (.code, .stderr) — surface UNSUPPORTED_FORMAT, PATH_NOT_ALLOWED, OCR_LANG_UNSUPPORTED, INGEST_FAILED, DOCKER_UNAVAILABLE rather than retrying blindly.

Step 2 — Decide direct vs. chunked

Measure the extracted Markdown:

  • < ~8,000 words (~10k tokens): summarize directly in the current context (Step 3a).
  • >= ~8,000 words: chunk and fan out (Step 3b).

Step 3a — Direct summarization (small documents)

Read the extracted .md and produce a summary using the output format below: title/subject, key points, entities, document type, language.

Show full SKILL.md (189 more words)Show less

Step 3b — Chunked summarization (large documents)

  1. Chunk. Split the extracted Markdown into context-friendly pieces (~3,000–4,000 tokens each). Prefer natural boundaries — headings, sections, page markers in the engine output — over blind character cuts. No script needed; split with judgment.

  2. Fan out. For each chunk launch a subagent (Agent tool, subagent_type: "general-purpose"), up to ~3–4 concurrent:

    Summarize this text chunk (chunk {i}/{total} of document '{filename}').
    Extract: key points, entities (people/orgs/dates/amounts), topics, and any
    conclusions or action items. Output as structured markdown.
    
    Text:
    {chunk_text}
  3. Merge. Collect chunk summaries, deduplicate entities and key points, and produce one unified summary in the output format. If the merged result is still > ~8,000 words, run one more summarization pass on it.

Batch summarization

For a directory or glob: extract each file via dc.convertToMarkdown (run a few in parallel), then apply the single-document workflow per file. Emit a table:

markdown
| # | File | Type | Language | Words | Key Topics | Summary |
|---|------|------|----------|-------|------------|---------|
| 1 | invoice.pdf | Invoice | EN | 450 | AcmeCorp, 2024Q4 | Quarterly invoice… |

Summary output format

markdown
## Summary: {document_name}

**Type**: {document_type}
**Language**: {language}
**Word Count**: {word_count}
**Date**: {detected_date or file_modified_date}

### Key Points
- Point 1
- Point 2

### Entities
- **People**: …
- **Organizations**: …
- **Dates**: …
- **Amounts**: …

### Brief Summary
{2-3 paragraph narrative summary}

Special cases

  • Scanned PDF, no text: the engine returns little/empty text on mode: auto. Re-run with ocr: { mode: "force", lang: [...] } (canonical language names).
  • Encrypted / unsupported / empty: surface the DocConverterError.code and .stderr; report metadata only.
  • Mixed-language: report the primary language, note others present.

© BlackBeltTechnology, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/document-converter/.pi/skills/doc-summarizer of BlackBeltTechnology/pi-agent-dashboard.

Open the folder on GitHubat commit 86e8e4d

Compare with similar skills

Doc Summarizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Doc Summarizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Doc Summarizer this skillBlackBeltTechnology/pi-agent-dashboard315—~1.2kAutomated safety check: PassMIT
Light File ReadingLight0305/Light-skills640—~4.1kAutomated safety check: PassMIT
Heavy File Ingestion Claude CodeNateBJones-Projects/OB14.7k—~614Automated safety check: PassCustom licence
Heavy File Ingestion Claude DesktopNateBJones-Projects/OB14.7k—~547Automated safety check: PassCustom licence
Document Converterwentorai/Research-Claw858—~1.3kAutomated safety check: PassCustom licence
Compdf Conversion CLILeoYeAI/openclaw-master-skills2.2k—~5.9kAutomated safety check: PassProprietary

Similar skills

  • Light File Reading

    Light0305/Light-skills

    Light 多格式文件深度理解常驻技能:强大地读 Word / PDF / PPTX / Excel / CSV / 图片 / 视频 / 代码 / 压缩包,不只提取文字,而是理解结构 / 图表 / 数据 / 格式要求 / 隐含意图,产结构化"理解笔记"五面 (结构逻辑·关键内容·格式约束·视觉风格·可复用)并映射到下游技能动作(这个文件→接下来能做什么)。

    640 GitHub stars~4.1k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Heavy File Ingestion Claude Code

    NateBJones-Projects/OB1

    Use in Claude Code when a user asks to read, analyze, summarize, or extract from a heavyweight file such as PDF, DOCX, PPTX, XLSX, CSV, or TSV.

    4.7k GitHub stars~614 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Use in Claude Desktop when a user asks to read, analyze, summarize, or extract from a heavyweight file such as PDF, DOCX, PPTX, XLSX, CSV, or TSV.

    4.7k GitHub stars~547 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Document Converter

    wentorai/Research-Claw

    Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Compdf Conversion CLI

    LeoYeAI/openclaw-master-skills

    MUST use for ANY PDF or image format conversion task — converting PDF and images (JPG/JPEG/PNG/BMP/TIFF/TIF/WEBP/JPEG2000) to 10 formats (Word, Excel, PPT, HTML, Image, TXT, JSON, Markdown, RTF…

    2.2k GitHub stars~5.9k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • To Markdown

    Mathews-Tom/armory

    Convert any file or URL to clean Markdown: PDF, DOCX, XLSX, PPTX, HTML, images (OCR), audio, CSV, YouTube.

    328 GitHub stars~2k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed

More from BlackBeltTechnology/pi-agent-dashboard

All 66 skills in this repo
  • Browser

    BlackBeltTechnology/pi-agent-dashboard

    Browser automation via the agent-browser CLI. An agent skill from BlackBeltTechnology/pi-agent-dashboard.

    315 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • CI Troubleshoot

    BlackBeltTechnology/pi-agent-dashboard

    Diagnose failed GitHub Actions runs for pi-agent-dashboard: the 11-file workflow taxonomy, affected-test selection, the release pipeline, known failure modes, and how to read gh run logs and…

    315 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Debug Dashboard

    BlackBeltTechnology/pi-agent-dashboard

    Diagnose problems in the running pi-agent-dashboard system: server.log, /api/health, bridge WebSocket connectivity, vitest triage, known-issue FAQ entries.

    315 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Implement

    BlackBeltTechnology/pi-agent-dashboard

    Disciplined implementation in pi-agent-dashboard: the rebuild matrix (extension→reload, server→restart, client→build+restart, openspec-apply→full rebuild) plus the project's code discipline rules.

    315 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Pi Dashboard

    BlackBeltTechnology/pi-agent-dashboard

    Monitor and control the pi-dashboard server. An agent skill from BlackBeltTechnology/pi-agent-dashboard.

    315 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Session To Guideline

    BlackBeltTechnology/pi-agent-dashboard

    Turn a pi session into a Markdown "how-we-did-it" collaboration guideline: reads the session's JSONL transcript and synthesizes a reusable playbook of which prompts worked, what had to be steered…

    315 GitHub stars~3.2k tokensUpdated today
    Auto-check passed

Questions about Doc Summarizer

What does Doc Summarizer do?

Summarize documents of any size: extract with the document-converter engine, chunk to fit context, fan out to subagents, then synthesize one unified summary. Doc Summarizer is an agent skill from BlackBeltTechnology/pi-agent-dashboard. Summarize documents of any size: extract with the document-converter engine, chunk to fit context, fan out to subagents, then synthesize one unified summary.

When should I use Doc Summarizer?

Doc Summarizer fits situations like: tasks that involve Summarization; tasks that involve Word documents; tasks that involve PowerPoint presentations.

How do I install Doc Summarizer in Claude Code?

Run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill doc-summarizer -a claude-code`. Or copy the skill folder (packages/document-converter/.pi/skills/doc-summarizer in BlackBeltTechnology/pi-agent-dashboard) into .claude/skills/doc-summarizer in your project. Claude Code loads it when a task matches its description.

How do I install Doc Summarizer in Codex?

Run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill doc-summarizer -a codex`. Or copy the skill folder (packages/document-converter/.pi/skills/doc-summarizer in BlackBeltTechnology/pi-agent-dashboard) into .agents/skills/doc-summarizer in your project. Codex loads it when a task matches its description.

Can I use Doc Summarizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill doc-summarizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doc-summarizer, .gemini/skills/doc-summarizer, .github/skills/doc-summarizer and .opencode/skills/doc-summarizer in your project.

What does Doc Summarizer need to run?

Going by SKILL.md and its folder, Doc Summarizer needs the command-line tools its instructions call (npm). Our summary lists: Python 3; Docker.

Does Doc Summarizer access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Doc Summarizer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Doc Summarizer use?

Doc Summarizer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Doc Summarizer use?

About 1.2k tokens (SKILL.md is roughly 4.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Doc Summarizer?

Skills that share tags, products or a category with Doc Summarizer: Light File Reading (Light0305/Light-skills, 640 stars), Heavy File Ingestion Claude Code (NateBJones-Projects/OB1, 4.7k stars), Heavy File Ingestion Claude Desktop (NateBJones-Projects/OB1, 4.7k stars) and Document Converter (wentorai/Research-Claw, 858 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Doc Summarizer?

BlackBeltTechnology (a GitHub organization) maintains it in BlackBeltTechnology/pi-agent-dashboard, which has 315 GitHub stars. The repository holds 66 skills in this directory. The repository was last updated on October 8, 2026.

Source: BlackBeltTechnology/pi-agent-dashboard on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.