Agent skill

DOCX Dual Parse

by HKUDS in HKUDS/OpenSpace

Extract text from DOCX files using shell or Python zipfile, with environment-aware fallback

MITAuto-check passedDocuments & Office

Install DOCX Dual Parse

skills CLI
$ npx skills add HKUDS/OpenSpace --skill docx-dual-parse -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace docx-dual-parse --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/docx-shell-parse-enhanced-4bba79 .claude/skills/docx-dual-parse && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
docx-dual-parse
GitHub stars
7.8k
Token cost
~1.8k tokens
SKILL.md length
359 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Extract text from DOCX files using shell or Python zipfile, with environment-aware fallback

  • Works in 5 steps: Always check file existence before parsing → Test method availability in the target… → Capture stderr for debugging failed… → …
  • Tasks that involve Word documents
  • SKILL.md covers When to Use, Core Technique, Environment Detection and Method A: Shell-Based Extraction, plus 7 more sections
  • Calls python3

What it does

DOCX Dual Parse is an agent skill from HKUDS/OpenSpace. Extract text from DOCX files using shell or Python zipfile, with environment-aware fallback

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering Word documents. It works with Microsoft Word and Python. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve Word documents

Example prompts

  • “/docx-dual-parse”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Always check file existence before parsing
  2. Test method availability in the target environment
  3. Capture stderr for debugging failed extractions
  4. Validate output is non-empty before proceeding
  5. Handle XML entity decoding if needed (sed can expand basic entities)

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

DOCX Dual Parse loads about 1.8k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 359 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 359 words, ~1,800 tokens.

Download SKILL.mdSave it as .claude/skills/docx-dual-parse/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
docx-dual-parse
description
Extract text from DOCX files using shell or Python zipfile, with environment-aware fallback

DOCX Dual-Method Text Extraction

Extract text from Microsoft Word (.docx) files using either shell commands or Python's zipfile module, automatically selecting the most reliable method for your environment.

When to Use

  • Need reliable DOCX text extraction in varying environments (containers, sandboxes, minimal images)
  • Python environment may lack python-docx but has standard library access
  • Shell utilities (unzip, sed) may be unavailable or restricted
  • Want environment-aware fallback without manual intervention

Core Technique

DOCX files are ZIP archives containing XML files. This skill provides two extraction methods:

Method A (Shell): unzip -p + sed for tag stripping Method B (Python): zipfile module for archive access + string parsing

Environment Detection

Before extraction, detect which method will work:

bash
# Quick shell method test
if unzip -v >/dev/null 2>&1; then
    echo "Shell method available"
else
    echo "Shell method unavailable, try Python"
fi
python
# Quick Python method test
python3 -c "import zipfile; print('Python method available')" 2>/dev/null

Method A: Shell-Based Extraction

Use when unzip and sed are available and the environment allows shell operations.

Step-by-Step Instructions

1. Verify the DOCX file exists

bash
ls -la document.docx

2. Extract raw XML content

bash
unzip -p document.docx word/document.xml

3. Strip XML tags from content

bash
unzip -p document.docx word/document.xml | sed -e 's/<[^>]*>//g'

4. Clean up whitespace (optional)

bash
unzip -p document.docx word/document.xml | \
  sed -e 's/<[^>]*>//g' | \
  sed -e 's/^[[:space:]]*//' -e 's/[[:space:]]*$//' | \
  sed -e '/^$/d'

5. Save extracted text to file

bash
unzip -p document.docx word/document.xml | \
  sed -e 's/<[^>]*>//g' > output.txt
Reusable Shell Function
bash
parse_docx_shell() {
    local file="$1"
    if [ ! -f "$file" ]; then
        echo "Error: File not found: $file" >&2
        return 1
    fi
    if ! command -v unzip >/dev/null 2>&1; then
        echo "Error: unzip not available" >&2
        return 1
    fi
    unzip -p "$file" word/document.xml 2>/dev/null | \
        sed -e 's/<[^>]*>//g' | \
        sed -e 's/^[[:space:]]*//' -e 's/[[:space:]]*$//' | \
        sed -e '/^$/d'
}

# Usage: parse_docx_shell document.docx

Method B: Python Zipfile Extraction

Use when shell method fails or Python environment is more reliable than shell.

Step-by-Step Instructions

1. Verify the DOCX file exists

bash
ls -la document.docx

2. Run Python extraction via run_shell

bash
run_shell 'python3 -c "
import zipfile
import re
with zipfile.ZipFile(\"document.docx\", \"r\") as z:
    content = z.read(\"word/document.xml\").decode(\"utf-8\")
    text = re.sub(r\"<[^>]*>\", \"\", content)
    lines = [l.strip() for l in text.splitlines() if l.strip()]
    for line in lines:
        print(line)
"'

3. Save to file by redirecting output

bash
run_shell 'python3 -c "
import zipfile
import re
with zipfile.ZipFile(\"document.docx\", \"r\") as z:
    content = z.read(\"word/document.xml\").decode(\"utf-8\")
    text = re.sub(r\"<[^>]*>\", \"\", content)
    lines = [l.strip() for l in text.splitlines() if l.strip()]
    with open(\"output.txt\", \"w\") as f:
        for line in lines:
            f.write(line + \"\\n\")
"'
Reusable Python Function (via run_shell)
bash
parse_docx_python() {
    local file="$1"
    local output="$2"
    if [ ! -f "$file" ]; then
        echo "Error: File not found: $file" >&2
        return 1
    fi
    run_shell "python3 -c \"
import zipfile
import re
import sys
try:
    with zipfile.ZipFile(\\'$file\\', \\'r\\') as z:
        content = z.read(\\'word/document.xml\\').decode(\\'utf-8\\')
        text = re.sub(r\\'<[^>]*>\\', \\'\\', content)
        lines = [l.strip() for l in text.splitlines() if l.strip()]
        for line in lines:
            print(line)
except Exception as e:
    print(f\\'Error: {e}\\', file=sys.stderr)
    sys.exit(1)
\""
}

# Usage: parse_docx_python document.docx
# Or to file: parse_docx_python document.docx > output.txt

Unified Dual-Method Function

Automatically tries shell first, falls back to Python if shell fails:

bash
parse_docx() {
    local file="$1"
    if [ ! -f "$file" ]; then
        echo "Error: File not found: $file" >&2
        return 1
    fi
    
    # Try shell method first
    if command -v unzip >/dev/null 2>&1; then
        result=$(unzip -p "$file" word/document.xml 2>/dev/null | \
            sed -e 's/<[^>]*>//g' | \
            sed -e 's/^[[:space:]]*//' -e 's/[[:space:]]*$//' | \
            sed -e '/^$/d')
        if [ -n "$result" ]; then
            echo "$result"
            return 0
        fi
    fi
    
    # Fallback to Python method
    python3 -c "
import zipfile
import re
import sys
try:
    with zipfile.ZipFile('$file', 'r') as z:
        content = z.read('word/document.xml').decode('utf-8')
        text = re.sub(r'<[^>]*>', '', content)
        lines = [l.strip() for l in text.splitlines() if l.strip()]
        for line in lines:
            print(line)
except Exception as e:
    print(f'Error: {e}', file=sys.stderr)
    sys.exit(1)
"
}

# Usage: parse_docx document.docx
Show full SKILL.md (139 more words)Show less

Alternative: Extract to Temporary Directory

For complex parsing needs or debugging:

Shell approach:

bash
tmpdir=$(mktemp -d)
unzip document.docx -d "$tmpdir"
cat "$tmpdir/word/document.xml" | sed -e 's/<[^>]*>//g'
rm -rf "$tmpdir"

Python approach:

bash
python3 -c "
import zipfile
import tempfile
import os
with zipfile.ZipFile('document.docx', 'r') as z:
    tmpdir = tempfile.mkdtemp()
    z.extractall(tmpdir)
    with open(os.path.join(tmpdir, 'word/document.xml')) as f:
        print(f.read())
"

Verification

Confirm extraction worked:

bash
# Check output has content
parse_docx document.docx | head -20

# Verify file was created (if saving to file)
ls -la output.txt
wc -l output.txt

Method Selection Guide

EnvironmentRecommended Method
Standard Linux with unzipShell (faster, simpler)
Container without unzipPython zipfile
Sandboxed executionPython via execute_code_sandbox or run_shell
Minimal/busybox imagesShell if unzip available
Unknown/restrictedUse unified parse_docx function

Limitations

  • Does not preserve formatting, images, or table structure
  • May include some residual XML entity references ( , etc.)
  • Works best for simple text extraction needs
  • DOCX must be a valid Office Open XML format
  • Protected/encrypted DOCX files require additional handling

Error Handling Tips

  1. Always check file existence before parsing
  2. Test method availability in the target environment
  3. Capture stderr for debugging failed extractions
  4. Validate output is non-empty before proceeding
  5. Handle XML entity decoding if needed (sed can expand basic entities)

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/docx-shell-parse-enhanced-4bba79 of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

DOCX Dual Parse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

DOCX Dual Parse compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
DOCX Dual Parse this skillHKUDS/OpenSpace7.8k—~1.8kAutomated safety check: PassMIT
Word Document Reader and WriterHKUDS/DeepTutor41k—~2.5kAutomated safety check: PassApache-2.0
Word DOCX ToolkitTokenRhythm/opensquilla7.1k—~1.7kAutomated safety check: PassApache-2.0
Markdown to Word Convertercat-xierluo/SuitAgent205—~559Automated safety check: PassMIT
Software Certificate SkillIvanCodesDev/software-certificate-skill156—~1.6kAutomated safety check: PassMIT
Doc Cleanernotoriouslab/doc-cleaner309—~712Automated safety check: PassMIT

Similar skills

  • Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

    41k GitHub stars~2.5k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Word DOCX Toolkit

    TokenRhythm/opensquilla

    Inspects, edits in place or creates Word .docx files with bundled Python scripts, keeping existing styles intact when content changes.

    7.1k GitHub stars~1.7k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markdown to Word Converter

    cat-xierluo/SuitAgent

    Converts Markdown files into Word documents formatted to Chinese typesetting conventions, with presets for academic, legal, report and book layouts.

    205 GitHub stars~559 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Software Certificate Skill

    IvanCodesDev/software-certificate-skill

    面向普通用户,从真实软件项目全自动生成中国软件著作权申请资料:一次收集登记事实,自动分析业务、选择可追溯源码、取得真实界面证据,生成申请表信息、规范黑白灰操作手册、代码前后30页或全部材料及真实 DOCX/PDF;内部验证、渲染、哈希与备份只进入系统临时运行区,项目最终仅保留正式资料。适配 Codex、Claude…

    156 GitHub stars~1.6k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Doc Cleaner

    notoriouslab/doc-cleaner

    Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

    309 GitHub stars~712 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 15 days ago
    Documents & OfficeAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.8k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.8k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.8k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.8k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about DOCX Dual Parse

What does DOCX Dual Parse do?

Extract text from DOCX files using shell or Python zipfile, with environment-aware fallback. DOCX Dual Parse is an agent skill from HKUDS/OpenSpace.

When should I use DOCX Dual Parse?

DOCX Dual Parse fits situations like: tasks that involve Word documents.

How do I install DOCX Dual Parse in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill docx-dual-parse -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/docx-shell-parse-enhanced-4bba79 in HKUDS/OpenSpace) into .claude/skills/docx-dual-parse in your project. Claude Code loads it when a task matches its description.

How do I install DOCX Dual Parse in Codex?

Run `npx skills add HKUDS/OpenSpace --skill docx-dual-parse -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/docx-shell-parse-enhanced-4bba79 in HKUDS/OpenSpace) into .agents/skills/docx-dual-parse in your project. Codex loads it when a task matches its description.

Can I use DOCX Dual Parse in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill docx-dual-parse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/docx-dual-parse, .gemini/skills/docx-dual-parse, .github/skills/docx-dual-parse and .opencode/skills/docx-dual-parse in your project.

What does DOCX Dual Parse need to run?

Going by SKILL.md and its folder, DOCX Dual Parse needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does DOCX Dual Parse access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is DOCX Dual Parse safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does DOCX Dual Parse use?

DOCX Dual Parse is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does DOCX Dual Parse use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to DOCX Dual Parse?

Skills that share tags, products or a category with DOCX Dual Parse: Word Document Reader and Writer (HKUDS/DeepTutor, 41k stars), Word DOCX Toolkit (TokenRhythm/opensquilla, 7.1k stars), Markdown to Word Converter (cat-xierluo/SuitAgent, 205 stars) and Software Certificate Skill (IvanCodesDev/software-certificate-skill, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains DOCX Dual Parse?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,750 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.