Agent skill

Word Analysis

by OpenSenseNova in OpenSenseNova/SenseNova-Skills

Word (.docx/.doc) 文档全量解析。覆盖:正文/段落文本提取、表格数据提取、高亮/颜色格式读取、多文件汇总对比、嵌入图片转 caption。

MITAuto-check passedDocuments & Office

Install Word Analysis

skills CLI
$ npx skills add OpenSenseNova/SenseNova-Skills --skill word-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenSenseNova/SenseNova-Skills word-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenSenseNova/SenseNova-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sn-da-non-spreadsheet-analysis/capability/word-analysis .claude/skills/word-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
word-analysis
GitHub stars
5.7k
Token cost
~2.2k tokens
SKILL.md length
181 words
Files
1
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Word (.docx/.doc) 文档全量解析。覆盖:正文/段落文本提取、表格数据提取、高亮/颜色格式读取、多文件汇总对比、嵌入图片转 caption。

  • Tasks that involve Word documents
  • SKILL.md covers Environment, Core Method 1: Full Text…, Core Method 2: Table… and Core Method 3: Format-Aware…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Word Analysis is an agent skill from OpenSenseNova/SenseNova-Skills. Word (.docx/.doc) 文档全量解析。覆盖:正文/段落文本提取、表格数据提取、高亮/颜色格式读取、多文件汇总对比、嵌入图片转 caption。

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Word documents. It works with Microsoft Word. The repository describes itself as: Modular SenseNova skills for building AI-powered office assistants and productivity workflows. The licence is MIT.

When your agent uses it

  • Tasks that involve Word documents

Example prompts

  • “/word-analysis”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 7838651. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Word Analysis loads about 2.2k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 181 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenSenseNova/SenseNova-Skills at commit 7838651, republished under its MIT licence (© OpenSenseNova). 181 words, ~2,173 tokens.

Download SKILL.mdSave it as .claude/skills/word-analysis/SKILL.md (or your agent's skills folder).
name
word-analysis
description
Word (.docx/.doc) 文档全量解析。覆盖:正文/段落文本提取、表格数据提取、高亮/颜色格式读取、多文件汇总对比、嵌入图片转 caption。

Word Analysis — .docx / .doc

Environment

python
from docx import Document
import os

# python-docx is available; for .doc (old format) convert via libreoffice first
def load_doc(path):
    """Load .docx directly; convert .doc to .docx first if needed."""
    if path.lower().endswith('.doc'):
        import subprocess
        out_dir = os.path.dirname(path)
        subprocess.run(
            ['libreoffice', '--headless', '--convert-to', 'docx', '--outdir', out_dir, path],
            check=True, capture_output=True
        )
        path = path.rsplit('.', 1)[0] + '.docx'
    return Document(path)

Core Method 1: Full Text Extraction

python
def extract_full_text(doc_path):
    """Extract all text: paragraphs + table cells, in document order."""
    doc = load_doc(doc_path)
    lines = []

    # Iterate paragraphs and tables in body order
    from docx.oxml.ns import qn
    for block in doc.element.body:
        tag = block.tag.split('}')[-1]
        if tag == 'p':
            # Paragraph
            from docx.text.paragraph import Paragraph
            para = Paragraph(block, doc)
            text = para.text.strip()
            if text:
                lines.append(text)
        elif tag == 'tbl':
            # Table
            from docx.table import Table
            tbl = Table(block, doc)
            for row in tbl.rows:
                row_text = '\t'.join(cell.text.strip() for cell in row.cells)
                if row_text.strip():
                    lines.append(row_text)

    return '\n'.join(lines)

# Usage
text = extract_full_text("/mnt/data/doc.docx")
print(text[:2000])  # preview first 2000 chars

Core Method 2: Table Extraction (Structured)

python
import pandas as pd

def extract_all_tables(doc_path):
    """Extract all tables from a Word document as list of DataFrames."""
    doc = load_doc(doc_path)
    tables = []

    for i, tbl in enumerate(doc.tables):
        rows = []
        for row in tbl.rows:
            rows.append([cell.text.strip() for cell in row.cells])
        if not rows:
            continue
        # Use first row as header if it looks like a header
        df = pd.DataFrame(rows[1:], columns=rows[0]) if rows else pd.DataFrame()
        tables.append((i, df))
        print(f"Table {i}: {df.shape[0]} rows × {df.shape[1]} cols")
        print(df.head(3))

    return tables

# Usage
tables = extract_all_tables("/mnt/data/doc.docx")

Core Method 3: Format-Aware Extraction (Color / Highlight)

Some questions require reading cell background color or text highlight color (e.g., "标黄的行", "红色文字"). Use XML-level access:

python
from docx import Document
from docx.oxml.ns import qn
from lxml import etree

def get_paragraph_highlight(para):
    """Return highlight color name of first run, or None."""
    for run in para.runs:
        rPr = run._r.find(qn('w:rPr'))
        if rPr is not None:
            hl = rPr.find(qn('w:highlight'))
            if hl is not None:
                return hl.get(qn('w:val'))  # e.g. 'yellow', 'cyan', 'red'
    return None

def get_table_cell_shading(cell):
    """Return background color hex of a table cell, or None."""
    tcPr = cell._tc.find(qn('w:tcPr'))
    if tcPr is not None:
        shd = tcPr.find(qn('w:shd'))
        if shd is not None:
            return shd.get(qn('w:fill'))  # hex color, e.g. 'FFFF00'
    return None

# Example: find all highlighted paragraphs
def find_highlighted_rows(doc_path, color='yellow'):
    doc = load_doc(doc_path)
    highlighted = []
    for i, para in enumerate(doc.paragraphs):
        hl = get_paragraph_highlight(para)
        if hl == color or (color == 'yellow' and hl in ('yellow', 'FFFF00')):
            highlighted.append((i, para.text))
    return highlighted

# For table cells with yellow background:
def find_highlighted_table_cells(doc_path, fill_colors=('FFFF00', 'FFD700')):
    doc = load_doc(doc_path)
    results = []
    for t_idx, tbl in enumerate(doc.tables):
        for r_idx, row in enumerate(tbl.rows):
            for c_idx, cell in enumerate(row.cells):
                color = get_table_cell_shading(cell)
                if color and color.upper() in fill_colors:
                    results.append({
                        'table': t_idx, 'row': r_idx, 'col': c_idx,
                        'color': color, 'text': cell.text.strip()
                    })
    return results

Core Method 4: Multi-File Aggregation

When the user asks about "these files" or the input is a directory:

python
def process_all_docs(file_list, extractor_fn):
    """Apply extractor to all files and aggregate results."""
    all_results = []
    for path in file_list:
        print(f"\n=== Processing: {os.path.basename(path)} ===")
        try:
            result = extractor_fn(path)
            all_results.append({'file': os.path.basename(path), 'data': result})
        except Exception as e:
            print(f"  ERROR: {e}")
    return all_results

# Example: extract text from all .docx in a directory
doc_files = [f for f in all_files if f.lower().endswith(('.docx', '.doc'))]
results = process_all_docs(doc_files, extract_full_text)

Core Method 5: Embedded Images → Caption

When a Word doc contains embedded images (charts, screenshots):

python
import zipfile, io, subprocess, json

CAPTION = "/path/to/skills/sn-da-image-caption/scripts/caption.py"

def extract_and_caption_images(doc_path, prompt=None):
    """Extract all images from .docx and caption each one."""
    # .docx is a ZIP archive; images are in word/media/
    results = []
    with zipfile.ZipFile(doc_path, 'r') as z:
        media_files = [n for n in z.namelist() if n.startswith('word/media/')]
        for media in media_files:
            ext = os.path.splitext(media)[-1].lower()
            if ext not in ('.png', '.jpg', '.jpeg', '.gif', '.bmp', '.wmf', '.emf'):
                continue
            # Save to temp
            tmp_path = f"/tmp/{os.path.basename(media)}"
            with z.open(media) as src, open(tmp_path, 'wb') as dst:
                dst.write(src.read())
            # Caption
            cmd = ["python3", CAPTION, tmp_path, "--json"]
            if prompt:
                cmd += ["--prompt", prompt]
            r = subprocess.run(cmd, capture_output=True, text=True, timeout=60)
            if r.returncode == 0:
                desc = json.loads(r.stdout).get("description", "")
                results.append({'image': media, 'caption': desc})
                print(f"  {media}: {desc[:100]}...")
            else:
                print(f"  {media}: caption failed — {r.stderr[:80]}")
    return results

Common Patterns

Font/size check (字号检查)
python
from docx.shared import Pt

def check_font_sizes(doc_path):
    doc = load_doc(doc_path)
    issues = []
    for i, para in enumerate(doc.paragraphs):
        for run in para.runs:
            size = run.font.size
            size_pt = size.pt if size else None
            # Also check style-level font
            if size_pt is None:
                style_size = run.style.font.size if run.style else None
                size_pt = style_size.pt if style_size else None
            issues.append({'para': i, 'text': run.text[:30], 'size_pt': size_pt})
    return issues
Spell/grammar check (错别字)
  • Use full-text extraction, then search with string matching or pass to LLM for proofreading
  • Do NOT try to install hunspell or other spell-check tools
Keyword search (全文定位)
python
def find_keyword(doc_path, keyword):
    text = extract_full_text(doc_path)
    idx = text.find(keyword)
    if idx >= 0:
        context = text[max(0, idx-100):idx+200]
        print(f"Found '{keyword}' at pos {idx}:\n{context}")
    else:
        print(f"'{keyword}' not found. Try broader search.")
        # Try case-insensitive or partial match
        for kw in keyword.split():
            if kw in text:
                print(f"  Partial match for '{kw}'")

Pitfalls

PitfallFix
Only read doc.paragraphs, miss tablesUse the body-order iterator in Method 1
Single file when input is multi-fileCheck os.path.isdir(), iterate all
Highlighted cells not detectedUse XML-level w:shd / w:highlight (Method 3)
.doc format fails to openConvert to .docx via libreoffice (Method 0)
Embedded charts look emptyExtract images from ZIP, caption each (Method 5)
Font size is NoneCheck both run-level and style-level (Method for font check)

© OpenSenseNova, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/sn-da-non-spreadsheet-analysis/capability/word-analysis of OpenSenseNova/SenseNova-Skills.

Open the folder on GitHubat commit 7838651

Compare with similar skills

Word Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Word Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Word Analysis this skillOpenSenseNova/SenseNova-Skills5.7k—~2.2kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
DOCXrvdbreemen/OTGW-firmware20733 repos~4.3kAutomated safety check: PassProprietary
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0
Word Document Reader and WriterHKUDS/DeepTutor41k—~2.5kAutomated safety check: PassApache-2.0
GenOffice Document CLIgenspark-ai/genoffice9k—~19kAutomated safety check: PassApache-2.0

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • DOCX

    rvdbreemen/OTGW-firmware

    A skill your agent uses whenever the user wants to create, read, edit, or manipulate Word documents (.docx files).

    207 GitHub starsUsed in 33 repos~4.3k tokens
    Documents & OfficeAuto-check passed
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

    41k GitHub stars~2.5k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    9k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • BiSheng DOCX Builder

    dataelement/bisheng

    Builds or edits Word .docx documents inside BiSheng's code executor with python-docx, handling Chinese fonts, tables of contents, page numbers and official-document layout.

    12k GitHub stars~2.6k tokensUpdated today
    Documents & OfficeAuto-check passed

More from OpenSenseNova/SenseNova-Skills

All 36 skills in this repo
  • SN Motion HTML

    OpenSenseNova/SenseNova-Skills

    Builds HTML stories where one continuous camera journey advances with page progress, using researched structure, AI stills, Seedance video clips and browser QA.

    5.7k GitHub stars~2.2k tokensUpdated today
    Auto-check: notes
  • SenseNova PPT Fallback Tools

    OpenSenseNova/SenseNova-Skills

    Fallback scripts for web search, image search and download, and image generation that PPT skills use only when the host agent lacks or fails its own tools.

    5.7k GitHub stars~575 tokensUpdated today
    Auto-check: notes
  • SenseNova PPT Workbench

    OpenSenseNova/SenseNova-Skills

    Opens the PPT Workbench web editor for an existing SenseNova HTML slide deck so you can preview, inspect and visually edit it without regenerating.

    5.7k GitHub stars~2.5k tokensUpdated today
    Auto-check: notes
  • SenseNova PPT Creative Renderer

    OpenSenseNova/SenseNova-Skills

    Turns an approved slide outline into a full-page image for every slide, one 16:9 PNG per page, and optionally packages the set into a PPTX.

    5.7k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • SenseNova PPT Entry

    OpenSenseNova/SenseNova-Skills

    Entry point for SenseNova presentation generation: creates a task folder, picks depth, output format and design richness, and routes to the right PPT skill.

    5.7k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes
  • China Market Open Data Search

    OpenSenseNova/SenseNova-Skills

    Researches Chinese market, macro, trade, procurement, listed-company and regulatory information from free official sources that need no sign-up or API key.

    5.7k GitHub stars~954 tokensUpdated today
    Auto-check: notes

Works with

Questions about Word Analysis

What does Word Analysis do?

Word (.docx/.doc) 文档全量解析。覆盖:正文/段落文本提取、表格数据提取、高亮/颜色格式读取、多文件汇总对比、嵌入图片转 caption。. Word Analysis is an agent skill from OpenSenseNova/SenseNova-Skills.

When should I use Word Analysis?

Word Analysis fits situations like: tasks that involve Word documents.

How do I install Word Analysis in Claude Code?

Run `npx skills add OpenSenseNova/SenseNova-Skills --skill word-analysis -a claude-code`. Or copy the skill folder (skills/sn-da-non-spreadsheet-analysis/capability/word-analysis in OpenSenseNova/SenseNova-Skills) into .claude/skills/word-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Word Analysis in Codex?

Run `npx skills add OpenSenseNova/SenseNova-Skills --skill word-analysis -a codex`. Or copy the skill folder (skills/sn-da-non-spreadsheet-analysis/capability/word-analysis in OpenSenseNova/SenseNova-Skills) into .agents/skills/word-analysis in your project. Codex loads it when a task matches its description.

Can I use Word Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenSenseNova/SenseNova-Skills --skill word-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/word-analysis, .gemini/skills/word-analysis, .github/skills/word-analysis and .opencode/skills/word-analysis in your project.

What does Word Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Word Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Word Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Word Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Word Analysis use?

Word Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Word Analysis use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Word Analysis?

Skills that share tags, products or a category with Word Analysis: Markitdown (ImCa0/just-laws, 782 stars), DOCX (rvdbreemen/OTGW-firmware, 207 stars), Gzh Design (isjiamu/gzh-design-skill, 3.9k stars) and Word Document Reader and Writer (HKUDS/DeepTutor, 41k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Word Analysis?

OpenSenseNova (a GitHub organization) maintains it in OpenSenseNova/SenseNova-Skills, which has 5,747 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on October 9, 2026.

Source: OpenSenseNova/SenseNova-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.