Agent skill

Excel Debug Extraction

by HKUDS in HKUDS/OpenSpace

Iteratively debug Excel structure with exploratory scripts before writing extraction logic

MITAuto-check passedDocuments & Office

Install Excel Debug Extraction

skills CLI
$ npx skills add HKUDS/OpenSpace --skill excel-debug-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace excel-debug-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/excel-debug-extraction .claude/skills/excel-debug-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
excel-debug-extraction
GitHub stars
7.7k
Token cost
~1.3k tokens
SKILL.md length
313 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Iteratively debug Excel structure with exploratory scripts before writing extraction logic

  • Works in 5 steps: Initial Structure Reconnaissance → Map Column Positions → Identify Row Patterns → …
  • Tasks that involve Excel spreadsheets
  • SKILL.md covers When to Use This Skill, Workflow Steps, Best Practices and Common Pitfalls to Avoid, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Excel Debug Extraction is an agent skill from HKUDS/OpenSpace. Iteratively debug Excel structure with exploratory scripts before writing extraction logic

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering Excel spreadsheets. It works with Microsoft Excel. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve Excel spreadsheets

Example prompts

  • “/excel-debug-extraction”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Initial Structure Reconnaissance
  2. Map Column Positions
  3. Identify Row Patterns
  4. Document Findings
  5. Write Extraction Logic

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Excel Debug Extraction loads about 1.3k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 313 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 313 words, ~1,308 tokens.

Download SKILL.mdSave it as .claude/skills/excel-debug-extraction/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
excel-debug-extraction
description
Iteratively debug Excel structure with exploratory scripts before writing extraction logic

Excel Debug-First Extraction Workflow

When working with poorly-structured, complex, or unfamiliar Excel files, use this iterative debugging approach to map the data layout before writing your final extraction logic.

When to Use This Skill

  • Excel files with inconsistent formatting or merged cells
  • Files received from external sources with unknown structure
  • Complex workbooks with multiple sheets and interdependencies
  • When initial parsing attempts fail or produce unexpected results

Workflow Steps

Step 1: Initial Structure Reconnaissance

Before writing extraction logic, create a debug script to explore the file structure:

python
# debug_structure.py
from openpyxl import load_workbook

wb = load_workbook('file.xlsx')
print(f"Sheets: {wb.sheetnames}")

for sheet_name in wb.sheetnames:
    ws = wb[sheet_name]
    print(f"\n=== Sheet: {sheet_name} ===")
    print(f"Dimensions: {ws.dimensions}")
    
    # Print first 10 rows to understand header structure
    for row in ws.iter_rows(min_row=1, max_row=10, values_only=True):
        print([str(cell)[:50] for cell in row])
Step 2: Map Column Positions

Identify where key data fields are located:

python
# debug_columns.py
from openpyxl import load_workbook

wb = load_workbook('file.xlsx')
ws = wb['Sheet1']

# Examine header row to find column indices
header_row = 1
column_map = {}

for col in ws.iter_cols(min_row=header_row, max_row=header_row):
    for cell in col:
        if cell.value:
            column_map[str(cell.value)] = cell.column_letter

print("Column mapping:", column_map)

# Sample data rows to verify structure
for row_num in range(2, min(6, ws.max_row + 1)):
    row_data = [ws.cell(row=row_num, column=col).value 
                for col in range(1, ws.max_column + 1)]
    print(f"Row {row_num}: {row_data}")
Step 3: Identify Row Patterns

Understand how data rows are structured (e.g., summary rows, detail rows, blank separators):

python
# debug_rows.py
from openpyxl import load_workbook

wb = load_workbook('file.xlsx')
ws = wb['Sheet1']

row_types = []
for row_num in range(1, min(30, ws.max_row + 1)):
    row_values = [ws.cell(row=row_num, column=col).value 
                  for col in range(1, ws.max_column + 1)]
    
    non_empty = sum(1 for v in row_values if v is not None and str(v).strip())
    
    # Classify row type
    if non_empty == 0:
        row_type = "blank"
    elif non_empty == 1:
        row_type = "summary/label"
    elif non_empty == ws.max_column:
        row_type = "full_data"
    else:
        row_type = "partial"
    
    row_types.append((row_num, row_type, row_values[:5]))

for rt in row_types:
    print(f"Row {rt[0]} ({rt[1]}): {rt[2]}")
Step 4: Document Findings

Before writing extraction logic, summarize:

  • Sheet names and their purposes
  • Header row location and column mappings
  • Data row patterns (which rows contain actual data vs. headers/summaries)
  • Any special formatting (merged cells, blank separators, grouping rows)
Step 5: Write Extraction Logic

Incorporate findings into your final processing script:

python
# extract_data.py
from openpyxl import load_workbook
import pandas as pd

wb = load_workbook('file.xlsx')
ws = wb['Sheet1']

# Use column mappings from debug phase
STORE_COL = 'B'  # Column 2
WEEK1_COL = 'D'  # Column 4
WEEK2_COL = 'E'  # Column 5

# Skip header rows and summary rows based on debug findings
data_rows = []
for row_num in range(5, ws.max_row + 1):  # Start after header based on debug
    # Skip summary/blank rows
    if ws.cell(row=row_num, column=2).value is None:
        continue
    if 'TOTAL' in str(ws.cell(row=row_num, column=2).value).upper():
        continue
    
    row_data = {
        'store': ws.cell(row=row_num, column=2).value,
        'week1': ws.cell(row=row_num, column=4).value,
        'week2': ws.cell(row=row_num, column=5).value,
    }
    data_rows.append(row_data)

df = pd.DataFrame(data_rows)
print(df.head())

Best Practices

  1. Always start with exploration - Never assume Excel structure matches expectations
  2. Save debug scripts - Keep them in your project for future reference and debugging
  3. Print generously - Use verbose output during exploration to catch edge cases
  4. Verify row-by-row - Don't assume all data rows follow the same pattern
  5. Handle merged cells - Check for merged cells that span multiple rows/columns

Common Pitfalls to Avoid

  • Assuming header is always row 1
  • Assuming all rows between first and last contain data
  • Not checking for hidden sheets or protected ranges
  • Ignoring cell formatting that indicates row type (bold, indentation)
  • Not handling None values or empty strings consistently

File Naming Convention

Use descriptive names for debug scripts:

  • debug_structure.py - Overall file/sheet structure
  • debug_columns.py - Column positions and headers
  • debug_rows.py - Row patterns and data boundaries
  • debug_values.py - Value patterns and edge cases

Keep debug scripts alongside your extraction script for maintainability.

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/excel-debug-extraction of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

Excel Debug Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Excel Debug Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Excel Debug Extraction this skillHKUDS/OpenSpace7.7k—~1.3kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Data Table Managern8n-io/n8n207k—~2.3kAutomated safety check: PassCustom licence
Docx4jplutext/docx4j2.4k—~2.5kAutomated safety check: PassNone
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Cyber Pptcrazyykhllc-bit/CyberPPT1.8k—~10kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Official

    Load before calling data-tables or parse-file. An agent skill from n8n-io/n8n.

    207k GitHub stars~2.3k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Docx4j

    plutext/docx4j

    A skill your agent uses when writing Java code that creates, reads or edits Word (.docx), PowerPoint (.pptx) or Excel (.xlsx) files with docx4j — including generating documents, editing existing…

    2.4k GitHub stars~2.5k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Cyber Ppt

    crazyykhllc-bit/CyberPPT

    当用户需要把 DOCX、PDF、TXT、XLSX、研究报告、业务材料或原始数据转成高密度、可编辑、咨询风格 PPTX 时使用;也适用于需要 SCR 论证、视觉风格探索、详细图表和渲染质检的 PPT。

    1.8k GitHub stars~10k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Markit

    shift-labs-ai/markit

    Convert files and URLs to Markdown. An agent skill from shift-labs-ai/markit.

    1.3k GitHub stars~299 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.7k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.7k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.7k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.7k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Excel Debug Extraction

What does Excel Debug Extraction do?

Iteratively debug Excel structure with exploratory scripts before writing extraction logic. Excel Debug Extraction is an agent skill from HKUDS/OpenSpace.

When should I use Excel Debug Extraction?

Excel Debug Extraction fits situations like: tasks that involve Excel spreadsheets.

How do I install Excel Debug Extraction in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill excel-debug-extraction -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/excel-debug-extraction in HKUDS/OpenSpace) into .claude/skills/excel-debug-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Excel Debug Extraction in Codex?

Run `npx skills add HKUDS/OpenSpace --skill excel-debug-extraction -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/excel-debug-extraction in HKUDS/OpenSpace) into .agents/skills/excel-debug-extraction in your project. Codex loads it when a task matches its description.

Can I use Excel Debug Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill excel-debug-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/excel-debug-extraction, .gemini/skills/excel-debug-extraction, .github/skills/excel-debug-extraction and .opencode/skills/excel-debug-extraction in your project.

What does Excel Debug Extraction need to run?

SKILL.md names no scripts, command-line tools or credentials: Excel Debug Extraction is instructions for the agent only. Our summary lists: Python 3.

Does Excel Debug Extraction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Excel Debug Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Excel Debug Extraction use?

Excel Debug Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Excel Debug Extraction use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Excel Debug Extraction?

Skills that share tags, products or a category with Excel Debug Extraction: Markitdown (ImCa0/just-laws, 781 stars), Data Table Manager (n8n-io/n8n, 207k stars), Docx4j (plutext/docx4j, 2.4k stars) and Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Excel Debug Extraction?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,749 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.