Agent skill

Spreadsheet Source Validated

by HKUDS in HKUDS/OpenSpace

Execute Python scripts for spreadsheet operations with mandatory source data validation and fallback protocols

MITAuto-check passedDocuments & Office

Install Spreadsheet Source Validated

skills CLI
$ npx skills add HKUDS/OpenSpace --skill spreadsheet-source-validated -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace spreadsheet-source-validated --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/spreadsheet-direct-python-enhanced-enhanced-2e6773 .claude/skills/spreadsheet-source-validated && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spreadsheet-source-validated
GitHub stars
7.7k
Token cost
~3.2k tokens
SKILL.md length
700 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Execute Python scripts for spreadsheet operations with mandatory source data validation and fallback protocols

  • Works in 3 steps: Verify Data Source Availability → Verify Data Integrity → External Data Source Verification
  • Tasks that involve Excel spreadsheets
  • SKILL.md covers Pre-Execution Data Validation, When to Use This Skill, Handling Inaccessible Data… and Why Direct Execution?, plus 5 more sections
  • Calls python3; reaches dataservices.epa.illinois.gov and epa.gov

What it does

Spreadsheet Source Validated is an agent skill from HKUDS/OpenSpace. Execute Python scripts for spreadsheet operations with mandatory source data validation and fallback protocols

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering Excel spreadsheets. It works with Python. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve Excel spreadsheets

Example prompts

  • “/spreadsheet-source-validated”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Verify Data Source Availability
  2. Verify Data Integrity
  3. External Data Source Verification

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • dataservices.epa.illinois.gov
    • epa.gov
    • waterdata.usgs.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Spreadsheet Source Validated loads about 3.2k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 700 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 700 words, ~3,163 tokens.

Download SKILL.mdSave it as .claude/skills/spreadsheet-source-validated/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
spreadsheet-source-validated
description
Execute Python scripts for spreadsheet operations with mandatory source data validation and fallback protocols

Validated Python Execution for Spreadsheet Tasks

Pre-Execution Data Validation

CRITICAL: Before attempting any spreadsheet operations, validate that your source data is accessible and complete.

Step 1: Verify Data Source Availability
python
import os
import sys
from pathlib import Path

def validate_source_data(source_paths):
    """Validate all required data sources before processing."""
    missing = []
    inaccessible = []
    
    for path in source_paths:
        p = Path(path)
        if not p.exists():
            missing.append(str(path))
        elif not os.access(path, os.R_OK):
            inaccessible.append(str(path))
    
    if missing:
        print(f"ERROR: Missing data sources: {', '.join(missing)}", file=sys.stderr)
        return False
    if inaccessible:
        print(f"ERROR: Inaccessible data sources: {', '.join(inaccessible)}", file=sys.stderr)
        return False
    
    print(f"VALIDATED: {len(source_paths)} source(s) available")
    return True

# Usage
sources = ['input.xlsx', 'reference_data.csv']
if not validate_source_data(sources):
    sys.exit(1)
Step 2: Verify Data Integrity
python
import pandas as pd
from openpyxl import load_workbook

def validate_spreadsheet_integrity(file_path, required_sheets=None, required_columns=None):
    """Check spreadsheet structure before processing."""
    try:
        # Check file is readable
        wb = load_workbook(file_path, read_only=True)
        
        # Verify required sheets exist
        if required_sheets:
            missing_sheets = [s for s in required_sheets if s not in wb.sheetnames]
            if missing_sheets:
                print(f"ERROR: Missing sheets: {missing_sheets}", file=sys.stderr)
                wb.close()
                return False
        
        # Verify required columns (sample first sheet)
        if required_columns:
            ws = wb.active
            headers = [cell.value for cell in ws[1]]
            missing_cols = [c for c in required_columns if c not in headers]
            if missing_cols:
                print(f"ERROR: Missing columns: {missing_cols}", file=sys.stderr)
                wb.close()
                return False
        
        wb.close()
        print(f"VALIDATED: Spreadsheet structure OK")
        return True
        
    except Exception as e:
        print(f"ERROR: Cannot validate spreadsheet: {str(e)}", file=sys.stderr)
        return False
Step 3: External Data Source Verification

For data retrieved from APIs, websites, or external services:

python
import requests
from urllib3.exceptions import SSLError, MaxRetryError

def validate_external_source(url, timeout=30, max_retries=2):
    """Verify external data source is accessible."""
    for attempt in range(max_retries + 1):
        try:
            response = requests.head(url, timeout=timeout, allow_redirects=True)
            if response.status_code == 200:
                print(f"VALIDATED: External source accessible ({url})")
                return True
            else:
                print(f"WARNING: External source returned {response.status_code}", file=sys.stderr)
        except (SSLError, MaxRetryError, requests.ConnectionError) as e:
            if attempt == max_retries:
                print(f"ERROR: External source inaccessible after {max_retries + 1} attempts: {url}", file=sys.stderr)
                return False
            print(f"RETRY {attempt + 1}/{max_retries}: {str(e)}", file=sys.stderr)
    
    return False

When to Use This Skill

Use validated Python execution for spreadsheet operations when:

  • Source data must be verified before processing begins
  • Reading or writing complex Excel files with multiple sheets
  • External data sources (APIs, websites) must be accessed first
  • Applying formulas, formatting, or data transformations
  • Working with openpyxl, pandas, or similar libraries
  • The operation involves multiple steps that could exceed agent step limits
  • You need precise control over error handling and debugging
  • Fallback data sources should be identified if primary sources fail

Handling Inaccessible Data Sources

Protocol for Failed Data Access
  1. Document the failure with specific error messages
  2. Identify alternative sources (see Alternative Sources section below)
  3. Report the blockage clearly before attempting workarounds
  4. Limit retry attempts to 3 before escalating
Alternative Data Source Identification

When primary data sources are inaccessible:

python
ALTERNATIVE_SOURCES = {
    'epa_water_data': [
        'https://dataservices.epa.illinois.gov/swap',  # Primary
        'https://www.epa.gov/safewater/data-and-reports',  # Federal fallback
        'https://waterdata.usgs.gov/nwis',  # USGS fallback
    ],
    'financial_data': [
        'internal_database.xlsx',  # Primary
        'backup_financial_data.csv',  # Local backup
        'request_from_stakeholder',  # Manual acquisition
    ]
}

def try_alternative_sources(source_key):
    """Iterate through alternative sources until one succeeds."""
    alternatives = ALTERNATIVE_SOURCES.get(source_key, [])
    
    for i, source in enumerate(alternatives):
        print(f"Attempting alternative {i + 1}/{len(alternatives)}: {source}")
        if source.startswith('http'):
            if validate_external_source(source):
                return source
        else:
            if Path(source).exists():
                print(f"SUCCESS: Alternative source found: {source}")
                return source
    
    print(f"ERROR: All alternatives exhausted for {source_key}", file=sys.stderr)
    return None
Error Reporting Protocol
python
def report_data_access_failure(source, error_type, alternatives_tried=0):
    """Standardized error reporting for data access failures."""
    error_report = {
        'timestamp': datetime.now().isoformat(),
        'source': source,
        'error_type': error_type,
        'alternatives_tried': alternatives_tried,
        'action_required': 'Manual data acquisition or source configuration update'
    }
    
    print("=" * 60, file=sys.stderr)
    print("DATA ACCESS FAILURE REPORT", file=sys.stderr)
    print("=" * 60, file=sys.stderr)
    for key, value in error_report.items():
        print(f"{key}: {value}", file=sys.stderr)
    print("=" * 60, file=sys.stderr)
    
    return error_report

Why Direct Execution?

The shell_agent tool can:

  • Hit maximum step limits on complex multi-step operations
  • Produce unexplained errors on formatting operations
  • Fail on intricate spreadsheet reads/writes due to iterative parsing
  • Fail to parse heredoc syntax correctly, causing 'unknown error' failures

Direct run_shell with Python is more reliable because it:

  • Executes in a single step with no iteration limits
  • Provides clearer, immediate error messages
  • Handles complex operations without step constraints
  • Gives full control over library imports and execution flow
  • Writing scripts to .py files first avoids shell_agent parsing issues with heredocs

How to Use

bash
# Step 1: Write validation script
cat > validate_sources.py << 'EOF'
import sys
from pathlib import Path

sources = ['input.xlsx', 'config.json']
missing = [s for s in sources if not Path(s).exists()]

if missing:
    print(f"BLOCKED: Missing sources: {missing}", file=sys.stderr)
    sys.exit(1)

print("VALIDATED: All sources available")
EOF

# Step 2: Run validation
python3 validate_sources.py || exit 1

# Step 3: Write processing script
cat > process_spreadsheet.py << 'EOF'
import openpyxl
# Your spreadsheet code here
EOF

# Step 4: Execute processing
python3 process_spreadsheet.py

# Step 5: Clean up (optional)
rm validate_sources.py
Complete Workflow Example
python
#!/usr/bin/env python3
"""
Complete workflow: Validate -> Process -> Report
"""
import sys
from pathlib import Path
from datetime import datetime
from openpyxl import load_workbook

def main():
    # PHASE 1: Validation
    print(f"[{datetime.now().isoformat()}] Starting validation...")
    
    source_file = 'input_data.xlsx'
    if not Path(source_file).exists():
        print(f"ERROR: Source file not found: {source_file}", file=sys.stderr)
        # Check for alternatives
        for alt in ['backup_input.xlsx', 'data_backup.csv']:
            if Path(alt).exists():
                print(f"FALLBACK: Using alternative: {alt}")
                source_file = alt
                break
        else:
            print("ERROR: No alternative sources available", file=sys.stderr)
            sys.exit(1)
    
    # PHASE 2: Processing
    print(f"[{datetime.now().isoformat()}] Processing {source_file}...")
    try:
        wb = load_workbook(source_file)
        ws = wb.active
        
        # Your operations here
        for row in ws.iter_rows(min_row=2):
            pass  # Process data
        
        wb.save('output.xlsx')
        print(f"SUCCESS: Processed {ws.max_row - 1} rows")
        
    except Exception as e:
        print(f"ERROR: Processing failed: {str(e)}", file=sys.stderr)
        sys.exit(1)
    
    # PHASE 3: Cleanup
    print(f"[{datetime.now().isoformat()}] Complete")

if __name__ == '__main__':
    main()

Best Practices

  1. ALWAYS validate sources first before any spreadsheet operations
  2. Prefer file-based execution for complex scripts: write to .py file first, then execute via run_shell
  3. Identify 2-3 alternative sources for critical data before starting
  4. Import only needed libraries to reduce execution time
  5. Print clear success/error messages for debugging
  6. Save intermediate results for complex multi-step transformations
  7. Test with small data before scaling to large spreadsheets
  8. Use pandas for data manipulation and openpyxl for formatting when both are needed
  9. Clean up temporary script files after execution if they won't be reused
  10. Document all access failures with timestamps and error details

When NOT to Use This Skill

  • Simple single-cell reads/writes (use shell_agent or basic commands)
  • Operations that require interactive user input
  • Tasks where you need the agent to iteratively refine the approach
  • When source data is guaranteed available (skip validation overhead)
Show full SKILL.md (264 more words)Show less

Common Libraries

LibraryBest For
openpyxlReading/writing .xlsx files, formatting, formulas
pandasData manipulation, analysis, merging datasets
xlrdReading older .xls files (read-only)
xlsxwriterCreating new .xlsx files with advanced formatting
requestsValidating external API/data sources
pathlibCross-platform file path validation

Troubleshooting

Issue: Heredoc syntax fails with 'unknown error' when using shell_agent

  • Solution: Write the Python script to a .py file first, then execute it with python3 script.py. This pattern is significantly more reliable than inline heredoc execution when shell_agent is the executor.

Issue: Source data validation fails

  • Solution:
    1. Check file paths are absolute or relative to working directory
    2. Verify file permissions with ls -la
    3. Try alternative sources from your fallback list
    4. Report the failure with complete error details before proceeding

Issue: External data source inaccessible (SSL/proxy errors)

  • Solution:
    1. Try HTTP instead of HTTPS if appropriate
    2. Disable SSL verification temporarily: requests.get(url, verify=False)
    3. Check proxy settings in environment variables
    4. After 3 failed attempts, switch to alternative source or report blockage
    5. Do NOT exhaust 30+ iterations on a single inaccessible source

Issue: FileNotFoundError

  • Solution: Verify the file path is absolute or relative to the working directory; check for alternative backup files

Issue: PermissionError

  • Solution: Ensure the file is not open in another application; check file permissions

Issue: MemoryError on large files

  • Solution: Process data in chunks using pandas chunksize parameter

Issue: Formatting not applying

  • Solution: Ensure you're modifying cell styles before saving, and use .copy() for style objects

Issue: Data validation passes but processing fails

  • Solution: Add more detailed integrity checks (column types, value ranges, row counts)

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/spreadsheet-direct-python-enhanced-enhanced-2e6773 of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

Spreadsheet Source Validated next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spreadsheet Source Validated compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spreadsheet Source Validated this skillHKUDS/OpenSpace7.7k—~3.2kAutomated safety check: PassMIT
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Doc Cleanernotoriouslab/doc-cleaner309—~712Automated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
XLSXzzhonglei/GeoCode-Release186—~3.1kAutomated safety check: PassMIT
Python Bridgetmustier/pi-for-excel434—~820Automated safety check: PassMIT

Similar skills

  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Doc Cleaner

    notoriouslab/doc-cleaner

    Convert PDF, DOCX, XLSX, and text files to clean, structured Markdown.

    309 GitHub stars~712 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • XLSX

    zzhonglei/GeoCode-Release

    Create, edit, analyze, or convert Excel spreadsheets (.xlsx, .xlsm) where the workbook file is the primary deliverable.

    186 GitHub stars~3.1k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • Python Bridge

    tmustier/pi-for-excel

    Native Python execution via the local Python bridge. An agent skill from tmustier/pi-for-excel.

    434 GitHub stars~820 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Office To Md

    shuyu-labs/WebCode

    Convert Office documents (Word, Excel, PowerPoint, PDF) to Markdown format.

    278 GitHub stars~1k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check: notes

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.7k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.7k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.7k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.7k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Spreadsheet Source Validated

What does Spreadsheet Source Validated do?

Execute Python scripts for spreadsheet operations with mandatory source data validation and fallback protocols. Spreadsheet Source Validated is an agent skill from HKUDS/OpenSpace.

When should I use Spreadsheet Source Validated?

Spreadsheet Source Validated fits situations like: tasks that involve Excel spreadsheets.

How do I install Spreadsheet Source Validated in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill spreadsheet-source-validated -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/spreadsheet-direct-python-enhanced-enhanced-2e6773 in HKUDS/OpenSpace) into .claude/skills/spreadsheet-source-validated in your project. Claude Code loads it when a task matches its description.

How do I install Spreadsheet Source Validated in Codex?

Run `npx skills add HKUDS/OpenSpace --skill spreadsheet-source-validated -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/spreadsheet-direct-python-enhanced-enhanced-2e6773 in HKUDS/OpenSpace) into .agents/skills/spreadsheet-source-validated in your project. Codex loads it when a task matches its description.

Can I use Spreadsheet Source Validated in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill spreadsheet-source-validated -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spreadsheet-source-validated, .gemini/skills/spreadsheet-source-validated, .github/skills/spreadsheet-source-validated and .opencode/skills/spreadsheet-source-validated in your project.

What does Spreadsheet Source Validated need to run?

Going by SKILL.md and its folder, Spreadsheet Source Validated needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Spreadsheet Source Validated access the network?

SKILL.md names 3 domains. In commands or code: dataservices.epa.illinois.gov, epa.gov and waterdata.usgs.gov; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Spreadsheet Source Validated safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Spreadsheet Source Validated use?

Spreadsheet Source Validated is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Spreadsheet Source Validated use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Spreadsheet Source Validated?

Skills that share tags, products or a category with Spreadsheet Source Validated: Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Doc Cleaner (notoriouslab/doc-cleaner, 309 stars), Mineru (Nebutra/MinerU-Skill, 122 stars) and XLSX (zzhonglei/GeoCode-Release, 186 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spreadsheet Source Validated?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,743 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.