Agent skill

Irregular Excel Parsing

by HKUDS in HKUDS/OpenSpace

Handle Excel files with irregular headers, merged cells, and unknown header row positions using pattern-matching and index-based extraction.

MITAuto-check passedDocuments & Office

Install Irregular Excel Parsing

skills CLI
$ npx skills add HKUDS/OpenSpace --skill irregular-excel-parsing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace irregular-excel-parsing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/irregular-excel-parsing .claude/skills/irregular-excel-parsing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
irregular-excel-parsing
GitHub stars
7.8k
Token cost
~1.4k tokens
SKILL.md length
248 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Handle Excel files with irregular headers, merged cells, and unknown header row positions using pattern-matching and index-based extraction.

  • Works in 5 steps: Read Excel Without Headers → Scan Rows to Find Header Pattern → Extract and Clean Headers → …
  • Tasks that involve Excel spreadsheets
  • SKILL.md covers Step-by-Step Instructions, Complete Example Function, When to Use This Pattern and When NOT to Use, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Irregular Excel Parsing is an agent skill from HKUDS/OpenSpace. Handle Excel files with irregular headers, merged cells, and unknown header row positions using pattern-matching and index-based extraction.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering Excel spreadsheets. It works with Microsoft Excel. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve Excel spreadsheets

Example prompts

  • “/irregular-excel-parsing”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Read Excel Without Headers
  2. Scan Rows to Find Header Pattern
  3. Extract and Clean Headers
  4. Extract Data Rows
  5. Validate and Clean Data

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Irregular Excel Parsing loads about 1.4k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 248 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 248 words, ~1,407 tokens.

Download SKILL.mdSave it as .claude/skills/irregular-excel-parsing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
irregular-excel-parsing
description
Handle Excel files with irregular headers, merged cells, and unknown header row positions using pattern-matching and index-based extraction.

Irregular Excel File Parsing

Use this skill when you encounter Excel files where standard pandas.read_excel() fails due to:

  • Headers not in row 0 (6-9+ header rows common)
  • Merged cells in header area
  • Unknown or inconsistent header row positions
  • Multiple title/metadata rows before actual data

Step-by-Step Instructions

Step 1: Read Excel Without Headers

First, read the entire sheet with header=None to get raw data:

python
import pandas as pd

# Read all data without assuming header position
df_raw = pd.read_excel('filename.xlsx', sheet_name='Sheet1', header=None)
Step 2: Scan Rows to Find Header Pattern

Search for the actual header row by looking for distinctive patterns:

python
def find_header_row(df):
    """Find header row by pattern-matching common column identifiers."""
    header_patterns = [
        r'Store ID',
        r'ID\d{4}',  # ID followed by 4 digits
        r'Week \d+',
        r'Date',
        r'Store',
        r'Product',
        r'ID'
    ]
    
    for row_idx in range(len(df)):
        row_values = df.iloc[row_idx].astype(str).str.lower()
        for pattern in header_patterns:
            if row_values.str.contains(pattern, case=False, regex=True).any():
                return row_idx
    
    # Fallback: return first non-empty row
    for row_idx in range(len(df)):
        if df.iloc[row_idx].notna().sum() > 0:
            return row_idx
    
    return 0

header_row = find_header_row(df_raw)
Step 3: Extract and Clean Headers

Extract the header row and clean column names:

python
# Extract header row
headers = df_raw.iloc[header_row].tolist()

# Clean headers: convert to string, strip whitespace, handle NaN
clean_headers = []
for h in headers:
    if pd.isna(h) or str(h).strip() == '':
        clean_headers.append(f'col_{len(clean_headers)}')
    else:
        clean_headers.append(str(h).strip())

# Handle duplicate headers by adding suffix
from collections import Counter
header_counts = Counter(clean_headers)
final_headers = []
for h in clean_headers:
    if header_counts[h] > 1:
        final_headers.append(f"{h}_{header_counts[h]}")
        header_counts[h] -= 1
    else:
        final_headers.append(h)
Step 4: Extract Data Rows

Extract data starting from the row after headers:

python
# Get data rows (everything after header)
data_df = df_raw.iloc[header_row + 1:].copy()
data_df.columns = final_headers

# Reset index
data_df = data_df.reset_index(drop=True)

# Remove completely empty rows
data_df = data_df.dropna(how='all')
Step 5: Validate and Clean Data

Perform basic validation and type conversion:

python
# Identify ID columns and preserve as string
for col in data_df.columns:
    if 'id' in col.lower() or 'code' in col.lower():
        data_df[col] = data_df[col].astype(str).str.strip()

# Convert numeric columns
numeric_cols = data_df.select_dtypes(include=['float64', 'int64']).columns
for col in numeric_cols:
    data_df[col] = pd.to_numeric(data_df[col], errors='coerce')

# Remove rows with invalid critical data
if 'Store ID' in data_df.columns:
    data_df = data_df[data_df['Store ID'].notna() & (data_df['Store ID'] != '')]

Complete Example Function

python
def parse_irregular_excel(filepath, sheet_name=0):
    """Parse Excel file with unknown/irregular header structure."""
    import pandas as pd
    import re
    from collections import Counter
    
    # Step 1: Read raw
    df_raw = pd.read_excel(filepath, sheet_name=sheet_name, header=None)
    
    # Step 2: Find header row
    patterns = [r'Store ID', r'ID\d{4}', r'Week', r'Date', r'Store']
    header_row = 0
    for idx in range(min(15, len(df_raw))):  # Check first 15 rows
        row_str = ' '.join(df_raw.iloc[idx].astype(str))
        for pattern in patterns:
            if re.search(pattern, row_str, re.IGNORECASE):
                header_row = idx
                break
    
    # Step 3: Extract headers
    headers = [str(h).strip() if pd.notna(h) else f'col_{i}' 
               for i, h in enumerate(df_raw.iloc[header_row])]
    
    # Handle duplicates
    counts = Counter(headers)
    final_headers = []
    for h in headers:
        if counts[h] > 1:
            final_headers.append(f"{h}_{counts[h]}")
            counts[h] -= 1
        else:
            final_headers.append(h)
    
    # Step 4: Extract data
    data_df = df_raw.iloc[header_row + 1:].copy()
    data_df.columns = final_headers
    data_df = data_df.dropna(how='all').reset_index(drop=True)
    
    return data_df

When to Use This Pattern

  • ✅ Excel files with 6-9+ title/metadata rows before data
  • ✅ Merged cells in header area causing misalignment
  • ✅ Headers not in predictable positions
  • ✅ Multiple sheets with inconsistent structures

When NOT to Use

  • ❌ Standard Excel files with headers in row 0 (use read_excel() directly)
  • ❌ Files with consistent, known structure (use explicit header= parameter)
  • ❌ When you know exact header row position (specify it directly)

Tips

  1. Always inspect first: Use df_raw.head(20) to visualize structure before parsing
  2. Pattern flexibility: Adjust regex patterns based on your specific column naming conventions
  3. Handle merged cells: Merged cells often result in NaN values - fill strategically if needed
  4. Save for reuse: Once you determine the correct header row for a file type, cache this for future runs

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/irregular-excel-parsing of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

Irregular Excel Parsing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Irregular Excel Parsing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Irregular Excel Parsing this skillHKUDS/OpenSpace7.8k—~1.4kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Data Table Managern8n-io/n8n207k—~2.3kAutomated safety check: PassCustom licence
Docx4jplutext/docx4j2.4k—~2.5kAutomated safety check: PassNone
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Cyber Pptcrazyykhllc-bit/CyberPPT1.8k—~10kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Official

    Load before calling data-tables or parse-file. An agent skill from n8n-io/n8n.

    207k GitHub stars~2.3k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Docx4j

    plutext/docx4j

    A skill your agent uses when writing Java code that creates, reads or edits Word (.docx), PowerPoint (.pptx) or Excel (.xlsx) files with docx4j — including generating documents, editing existing…

    2.4k GitHub stars~2.5k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Cyber Ppt

    crazyykhllc-bit/CyberPPT

    当用户需要把 DOCX、PDF、TXT、XLSX、研究报告、业务材料或原始数据转成高密度、可编辑、咨询风格 PPTX 时使用;也适用于需要 SCR 论证、视觉风格探索、详细图表和渲染质检的 PPT。

    1.8k GitHub stars~10k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Jev SEO

    AgriciDaniel/jev-seo

    Full live SEO audit of any website from its homepage URL, powered by Jev (TypeSafe's System One model).

    543 GitHub stars~2.5k tokensUpdated 18 days ago
    Documents & OfficeAuto-check: notes

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.8k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.8k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.8k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.8k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Irregular Excel Parsing

What does Irregular Excel Parsing do?

Handle Excel files with irregular headers, merged cells, and unknown header row positions using pattern-matching and index-based extraction. Irregular Excel Parsing is an agent skill from HKUDS/OpenSpace. Handle Excel files with irregular headers, merged cells, and unknown header row positions using pattern-matching and index-based extraction.

When should I use Irregular Excel Parsing?

Irregular Excel Parsing fits situations like: tasks that involve Excel spreadsheets.

How do I install Irregular Excel Parsing in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill irregular-excel-parsing -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/irregular-excel-parsing in HKUDS/OpenSpace) into .claude/skills/irregular-excel-parsing in your project. Claude Code loads it when a task matches its description.

How do I install Irregular Excel Parsing in Codex?

Run `npx skills add HKUDS/OpenSpace --skill irregular-excel-parsing -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/irregular-excel-parsing in HKUDS/OpenSpace) into .agents/skills/irregular-excel-parsing in your project. Codex loads it when a task matches its description.

Can I use Irregular Excel Parsing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill irregular-excel-parsing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/irregular-excel-parsing, .gemini/skills/irregular-excel-parsing, .github/skills/irregular-excel-parsing and .opencode/skills/irregular-excel-parsing in your project.

What does Irregular Excel Parsing need to run?

SKILL.md names no scripts, command-line tools or credentials: Irregular Excel Parsing is instructions for the agent only. Our summary lists: Python 3.

Does Irregular Excel Parsing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Irregular Excel Parsing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Irregular Excel Parsing use?

Irregular Excel Parsing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Irregular Excel Parsing use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Irregular Excel Parsing?

Skills that share tags, products or a category with Irregular Excel Parsing: Markitdown (ImCa0/just-laws, 781 stars), Data Table Manager (n8n-io/n8n, 207k stars), Docx4j (plutext/docx4j, 2.4k stars) and Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Irregular Excel Parsing?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,754 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.