Agent skill

Word Document Reader and Writer

by HKUDS in HKUDS/DeepTutor

Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

Apache-2.0Auto-check passedDocuments & Office

Install Word Document Reader and Writer

skills CLI
$ npx skills add HKUDS/DeepTutor --skill docx -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/DeepTutor docx --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/DeepTutor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/deeptutor/skills/builtin/docx .claude/skills/docx && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
docx
GitHub stars
41k
Token cost
~2.5k tokens
SKILL.md length
755 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

  • Extracting or summarizing text and tables from a .docx file
  • SKILL.md covers Runtime, Read / extract, Create and Edit existing (simple), plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Generating a report, letter or memo as Word with headings, tables, images and page numbers

What it does

The skill treats a .docx as a ZIP of XML parts and works in two tiers. The default tier uses python-docx, which is preinstalled, for creating, reading and simple edits. The advanced tier edits raw OOXML through zipfile, only for what python-docx cannot express: tracked changes, comments and edits that must preserve every untouched byte. Code runs as complete Python source in one exec call.

For reading, it notes that doc.paragraphs skips text in tables, headers, footers and text boxes, so those are iterated separately, and that tracked changes must be parsed from the XML because python-docx ignores insertions and deletions. For creating, the rules include using list styles instead of typed bullet characters, one paragraph per visual line and style names that match the template. It also covers page setup, headers, footers and page numbers added as field XML.

When editing, it writes a new output file unless you explicitly ask for an authorized replacement, and it follows the turn's workspace instructions for finding inputs and presenting the finished file. It is not meant for PDF, .xlsx, .pptx or Google Docs.

When your agent uses it

  • Extracting or summarizing text and tables from a .docx file
  • Generating a report, letter or memo as Word with headings, tables, images and page numbers
  • Finding and replacing text across a Word document
  • Applying tracked changes or comments to a contract draft

Example prompts

  • “Summarize the tables in ./contracts/vendor-agreement.docx.”
  • “Write a two-page project memo as a Word file with a table of contents and page numbers.”
  • “Replace every mention of Acme Corp with Northwind Ltd in ./drafts/proposal.docx and save a new copy.”
  • “Add tracked changes and comments to the payment terms section of the lease draft instead of editing it directly.”

Requirements

  • Python with python-docx

What it can do on your machine

Read from SKILL.md and the folder at commit f07029c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and xml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Word Document Reader and Writer loads about 2.5k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 755 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/DeepTutor at commit f07029c, republished under its Apache-2.0 licence (© HKUDS). 755 words, ~2,544 tokens.

Download SKILL.mdSave it as .claude/skills/docx/SKILL.md (or your agent's skills folder).
name
docx
description
Read, create, or edit Microsoft Word .docx files — extract/summarize text and tables, generate reports/letters/memos with headings, tables, images, TOC and page numbers, do find-and-replace, or apply tracked changes (redlines) and comments. Use whenever the user has a .docx or wants a Word deliverable. Not for PDF, .xlsx, .pptx, or Google Docs.
tags
tool, office
requires.sandbox
shell

DOCX (Microsoft Word)

A .docx is a ZIP of XML parts. Body text lives in word/document.xml. Two tiers:

  • Default — python-docx (preinstalled): create, read, and simple edits. Use for almost everything.
  • Advanced — raw OOXML via zipfile: only for what python-docx cannot express — tracked changes (redlines), comments, and exact-fidelity edits that must preserve every untouched byte. See Raw OOXML.

Runtime

Use exec with complete Python source (language: python). python-docx is preinstalled. Prefer creating, saving, reopening, and validating the deliverable in one call; later calls can revise the same relative filename. Follow the turn's User workspace instructions for locating inputs, output boundaries, and presenting the finished file. When editing, write a new output unless the user explicitly requested an authorized replacement.

Read / extract

python
from docx import Document

doc = Document("in.docx")
text = "\n".join(p.text for p in doc.paragraphs)  # body paragraphs
for tbl in doc.tables:  # tables
    for row in tbl.rows:
        print([c.text for c in row.cells])

doc.paragraphs skips text inside tables, headers/footers, and text boxes — iterate doc.tables and doc.sections[i].header/.footer for those. Each paragraph's style: p.style.name (e.g. "Heading 1").

To read tracked changes, parse the XML directly — python-docx ignores <w:ins>/<w:del>:

python
import zipfile, re

xml = zipfile.ZipFile("in.docx").read("word/document.xml").decode("utf-8")
# inserted text = <w:ins>…<w:t>…  deleted = <w:del>…<w:delText>…
print(re.findall(r"<w:t[^>]*>(.*?)</w:t>", xml))

Create

python
from docx import Document
from docx.shared import Pt, Inches, RGBColor
from docx.enum.text import WD_ALIGN_PARAGRAPH

doc = Document()  # default template page size
doc.add_heading("Quarterly Report", level=0)  # 0 = title; 1..9 = H1..H9
p = doc.add_paragraph("Intro paragraph. ")
run = p.add_run("Bold tail.")
run.bold = True
doc.add_paragraph("First item", style="List Bullet")  # real list style, never a "• " literal
doc.add_paragraph("Step one", style="List Number")

# Table — header row + data
tbl = doc.add_table(rows=1, cols=2)
tbl.style = "Light Grid Accent 1"
tbl.rows[0].cells[0].text, tbl.rows[0].cells[1].text = "Metric", "Value"
for k, v in [("Revenue", "1.2M"), ("Growth", "15%")]:
    c = tbl.add_row().cells
    c[0].text, c[1].text = k, v

doc.add_picture("chart.png", width=Inches(5))  # image, scaled to width
doc.add_page_break()
doc.save("out.docx")

# Validate immediately; later exec calls can also reopen this relative path.
check = Document("out.docx")
assert check.paragraphs, "generated DOCX has no paragraphs"
import zipfile

with zipfile.ZipFile("out.docx") as package:
    assert package.testzip() is None, "generated DOCX has a corrupt ZIP member"

Rules:

  • Never type bullet/number characters (•, 1.) into text — use style="List Bullet"/"List Number". Only list styles defined in the doc's template are available.
  • No \n inside a run — each visual line is its own add_paragraph.
  • Built-in style names must match the template (e.g. "Heading 1", "List Bullet"); a wrong name raises KeyError.
  • Units: Pt, Inches, Cm from docx.shared. Colors: RGBColor(0x1F,0x4E,0x79).
Page setup, headers/footers, page numbers
python
from docx.shared import Inches

sec = doc.sections[0]
sec.page_width, sec.page_height = Inches(8.5), Inches(11)  # Letter
sec.top_margin = sec.bottom_margin = Inches(1)
sec.header.paragraphs[0].text = "Confidential"

Page-number fields aren't in the python-docx API; inject the field XML into a footer run:

python
from docx.oxml.ns import qn
from docx.oxml import OxmlElement


def add_page_number(paragraph):
    for t in ("begin", "instr", "end"):
        r = OxmlElement("w:r")
        if t == "instr":
            fld = OxmlElement("w:instrText")
            fld.set(qn("xml:space"), "preserve")
            fld.text = "PAGE"
        else:
            fld = OxmlElement("w:fldChar")
            fld.set(qn("w:fldCharType"), t)
        r.append(fld)
        paragraph._p.append(r)


add_page_number(doc.sections[0].footer.paragraphs[0])

A clickable Table of Contents is also a field; Word shows "right-click → Update Field" until refreshed. Same pattern with instrText = TOC \o "1-3" \h \z \u.

Edit existing (simple)

python-docx preserves the rest of the document; mutate then save under a new name.

python
doc = Document("in.docx")
# Find-and-replace, keeping each run's formatting:
for p in doc.paragraphs:
    if "{{CLIENT}}" in p.text:
        for r in p.runs:
            r.text = r.text.replace("{{CLIENT}}", "Acme Co")
doc.save("out.docx")

Gotcha: Word splits text across runs, so a phrase may not live in one run.text even though p.text shows it whole. If the placeholder spans runs, set p.runs[0].text = p.text.replace(...) and clear the rest (for r in p.runs[1:]: r.text = "") — this collapses formatting to the first run, acceptable for plain placeholders. For exact-fidelity edits, use the raw-OOXML tier.

.doc → .docx and PDF export (LibreOffice, optional)

Legacy binary .doc can't be read by python-docx, and there is no built-in PDF export. Both need LibreOffice, which is optional and often absent. Only use it from exec; locate it with shutil.which("soffice"), keep its profile/conversion directory relative to the stable working directory, and clean conversion intermediates when finished. Never search for a desktop installation by absolute path and never write conversion files to /tmp. If it is absent, degrade with a clear note.

python
import shutil

soffice = shutil.which("soffice")
if soffice is None:
    print("soffice unavailable — ask the user for .docx or omit PDF export")
# If present, invoke it with subprocess.run([...], check=True) here, using only
# relative paths below this run directory, then validate the result immediately.

Never fetch or install authoring dependencies during a document task. If a declared runtime dependency is absent, report an incomplete deployment and stop.

Show full SKILL.md (319 more words)Show less

Raw OOXML (advanced)

Only when python-docx can't express it: tracked changes, comments, exact-fidelity edits. Workflow: read the XML part → edit it as text → re-zip every original member, rewriting only the changed part.

python
import zipfile

src, dst = "in.docx", "out.docx"
with zipfile.ZipFile(src) as z:
    xml = z.read("word/document.xml").decode("utf-8")
xml = xml.replace("OLD", "NEW")  # or splice tracked-change elements (below)
with zipfile.ZipFile(src) as zin, zipfile.ZipFile(dst, "w", zipfile.ZIP_DEFLATED) as zout:
    for item in zin.infolist():
        data = (
            xml.encode("utf-8") if item.filename == "word/document.xml" else zin.read(item.filename)
        )
        zout.writestr(item, data)

Critical traps (these silently corrupt the file or lose text):

  • xml:space="preserve" on any <w:t>/<w:delText> with leading/trailing whitespace, or Word strips the space.
  • Don't pretty-print into text nodes — added newlines/indent inside <w:t> become visible spaces. Edit the XML as a string; never reserialize the whole tree with indentation.
  • Keep parts consistent: new image/part → add its <Relationship> in word/_rels/document.xml.rels and a content type in [Content_Types].xml, or the doc opens "corrupt."
  • Unique IDs: every w:id on <w:ins>/<w:del>/comments must be unique in the file. w14:paraId/w16cid:durableId must be < 0x7FFFFFFF (8-digit hex).
  • Element order in <w:pPr>: pStyle, numPr, spacing, ind, jc, then rPr last.
Tracked changes (redlines)

Use a consistent author (default "Claude" unless the user names one) and ISO date. Replace the whole <w:r> with siblings — never nest change tags inside a run — and copy the original <w:rPr> into the new runs to keep formatting.

xml
<!-- change "30 days" → "60 days" -->
<w:r><w:t xml:space="preserve">The term is </w:t></w:r>
<w:del w:id="1" w:author="Claude" w:date="2026-01-01T00:00:00Z">
  <w:r><w:delText>30</w:delText></w:r></w:del>
<w:ins w:id="2" w:author="Claude" w:date="2026-01-01T00:00:00Z">
  <w:r><w:t>60</w:t></w:r></w:ins>
<w:r><w:t xml:space="preserve"> days.</w:t></w:r>
  • Inside <w:del> use <w:delText> (not <w:t>); inside <w:ins> never use <w:delText>.
  • Deleting a whole paragraph: also mark its paragraph mark — add <w:del .../> inside <w:pPr><w:rPr> — or accepting changes leaves an empty paragraph.
  • Reject another author's insertion: nest your <w:del> inside their <w:ins>. Restore their deletion: add a new <w:ins> after their <w:del> — never edit their tags.
Comments

Comments live in a separate word/comments.xml part (create it + its relationship in word/_rels/document.xml.rels + a content-type override if absent). In document.xml, the anchor markers <w:commentRangeStart w:id="N"/> and <w:commentRangeEnd w:id="N"/> are siblings of <w:r>, never inside one; follow the end marker with <w:r><w:rPr><w:rStyle w:val="CommentReference"/></w:rPr><w:commentReference w:id="N"/></w:r>. This is fiddly — verify the output opens in Word.

Verify before returning

Always confirm the file reopens cleanly — a silent corruption is the most common failure:

python
from docx import Document

d = Document("out.docx")
print(len(d.paragraphs), "paragraphs OK")

For raw-OOXML edits, run zipfile.ZipFile('out.docx').testzip() and well-formedness-check each edited XML part with lxml.etree.parse in the same exec Python call that saves out.docx.

© HKUDS, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in deeptutor/skills/builtin/docx of HKUDS/DeepTutor.

Open the folder on GitHubat commit f07029c

Compare with similar skills

Word Document Reader and Writer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Word Document Reader and Writer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Word Document Reader and Writer this skillHKUDS/DeepTutor41k—~2.5kAutomated safety check: PassApache-2.0
Markdown to Word Convertercat-xierluo/SuitAgent206—~559Automated safety check: PassMIT
BiSheng DOCX Builderdataelement/bisheng12k—~2.6kAutomated safety check: PassApache-2.0
Word DOCX ToolkitTokenRhythm/opensquilla7.1k—~1.7kAutomated safety check: PassApache-2.0
Industry Bid Document WriterGet00/BiaoShu-SKILL167—~4.6kAutomated safety check: PassApache-2.0
Word Document Creation And Editingpipeshub-ai/pipeshub-ai3.8k—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Markdown to Word Converter

    cat-xierluo/SuitAgent

    Converts Markdown files into Word documents formatted to Chinese typesetting conventions, with presets for academic, legal, report and book layouts.

    206 GitHub stars~559 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • BiSheng DOCX Builder

    dataelement/bisheng

    Builds or edits Word .docx documents inside BiSheng's code executor with python-docx, handling Chinese fonts, tables of contents, page numbers and official-document layout.

    12k GitHub stars~2.6k tokensUpdated 6 days ago
    Documents & OfficeAuto-check passed
  • Word DOCX Toolkit

    TokenRhythm/opensquilla

    Inspects, edits in place or creates Word .docx files with bundled Python scripts, keeping existing styles intact when content changes.

    7.1k GitHub stars~1.7k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Industry Bid Document Writer

    Get00/BiaoShu-SKILL

    Converts a tender document into Markdown, extracts scoring criteria and requirements, then drafts an industry-formatted technical bid as a Word file.

    167 GitHub stars~4.6k tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • Word Document Creation And Editing

    pipeshub-ai/pipeshub-ai

    Routes a Word document request to the right approach: a TypeScript library for new files, XML editing for existing ones, and plain reading only.

    3.8k GitHub stars~1.3k tokensUpdated today
    Documents & OfficeAuto-check passed
  • DOCX

    LeastBit/Claude_skills_zh-CN

    全面的文档创建、编辑和分析功能,支持修订追踪、批注、格式保留和文本提取。当 Claude 需要处理专业文档(.docx 文件)时使用:(1) 创建新文档,(2) 修改或编辑内容,(3) 处理修订追踪,(4) 添加批注,或其他任何文档任务

    588 GitHub stars~1.3k tokensUpdated 8 mo ago
    Documents & OfficeAuto-check: notes

More from HKUDS/DeepTutor

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.3k tokensUpdated 3 days ago
    Auto-check passed
  • Reads, extracts from, creates, merges, splits, watermarks, encrypts and fills PDF files with pdfplumber, pypdf and reportlab inside a Python sandbox.

    41k GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed
  • Explains how to design and write DeepTutor skills: the SKILL.md anatomy, a trigger-focused description, concise content and progressive disclosure into reference files.

    41k GitHub stars~840 tokensUpdated 3 days ago
    Auto-check passed
  • Excel Workbook Editor

    HKUDS/DeepTutor

    Reads, creates and edits Excel workbooks with openpyxl, including formulas, styles, charts and CSV or TSV tables, with advice on formula values openpyxl cannot compute.

    41k GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed

Questions about Word Document Reader and Writer

What does Word Document Reader and Writer do?

Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits. docx as a ZIP of XML parts and works in two tiers. The default tier uses python-docx, which is preinstalled, for creating, reading and simple edits.

When should I use Word Document Reader and Writer?

Word Document Reader and Writer fits situations like: extracting or summarizing text and tables from a .docx file; generating a report, letter or memo as Word with headings, tables, images and page numbers; finding and replacing text across a Word document; applying tracked changes or comments to a contract draft.

How do I install Word Document Reader and Writer in Claude Code?

Run `npx skills add HKUDS/DeepTutor --skill docx -a claude-code`. Or copy the skill folder (deeptutor/skills/builtin/docx in HKUDS/DeepTutor) into .claude/skills/docx in your project. Claude Code loads it when a task matches its description.

How do I install Word Document Reader and Writer in Codex?

Run `npx skills add HKUDS/DeepTutor --skill docx -a codex`. Or copy the skill folder (deeptutor/skills/builtin/docx in HKUDS/DeepTutor) into .agents/skills/docx in your project. Codex loads it when a task matches its description.

Can I use Word Document Reader and Writer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/DeepTutor --skill docx -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/docx, .gemini/skills/docx, .github/skills/docx and .opencode/skills/docx in your project.

What does Word Document Reader and Writer need to run?

SKILL.md names no scripts, command-line tools or credentials: Word Document Reader and Writer is instructions for the agent only. Our summary lists: Python with python-docx.

Does Word Document Reader and Writer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Word Document Reader and Writer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Word Document Reader and Writer use?

Word Document Reader and Writer is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Word Document Reader and Writer use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Word Document Reader and Writer?

Skills that share tags, products or a category with Word Document Reader and Writer: Markdown to Word Converter (cat-xierluo/SuitAgent, 206 stars), BiSheng DOCX Builder (dataelement/bisheng, 12k stars), Word DOCX Toolkit (TokenRhythm/opensquilla, 7.1k stars) and Industry Bid Document Writer (Get00/BiaoShu-SKILL, 167 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Word Document Reader and Writer?

HKUDS (a GitHub organization) maintains it in HKUDS/DeepTutor, which has 40,857 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 4, 2026.

Source: HKUDS/DeepTutor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.