Agent skill

Sn Da Non Spreadsheet Analysis

by OpenSenseNova in OpenSenseNova/SenseNova-Skills

Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。遇到以下任一情况就主动使用本 skill:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 /…

MITAuto-check passedDocuments & Office

Install Sn Da Non Spreadsheet Analysis

skills CLI
$ npx skills add OpenSenseNova/SenseNova-Skills --skill sn-da-non-spreadsheet-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenSenseNova/SenseNova-Skills sn-da-non-spreadsheet-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenSenseNova/SenseNova-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sn-da-non-spreadsheet-analysis .claude/skills/sn-da-non-spreadsheet-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sn-da-non-spreadsheet-analysis
GitHub stars
5.7k
Token cost
~1.2k tokens
SKILL.md length
312 words
Files
2
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。遇到以下任一情况就主动使用本 skill:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 /…

  • Works in 4 steps: Identify file type and input scope → Load sub-skill by format → Parse and extract → …
  • Tasks that involve Excel spreadsheets
  • SKILL.md covers Workflow, Universal Rules, Caption Script (for… and Available sub-skills
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Sn Da Non Spreadsheet Analysis is an agent skill from OpenSenseNova/SenseNova-Skills. Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。遇到以下任一情况就主动使用本 skill:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 / 幻灯片分析 / 发票提取 / 合同分析 / 文档统计 / 错别字 / 语病 / 字号检查 / 简历分析 / 多文档对比;③任务涉及从文档中提取表格、数值、图表、格式(颜色/高亮/字号)、组织架构、时间线等结构化信息。仅不用于:Excel/CSV 数据分析(使用 sn-da-excel-workflow)、纯图片分析(使用 sn-da-image-caption)。

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering Excel spreadsheets, PowerPoint presentations and PDF. It works with Microsoft Excel, Microsoft PowerPoint and Microsoft Word. The repository describes itself as: Modular SenseNova skills for building AI-powered office assistants and productivity workflows. The licence is MIT.

When your agent uses it

  • Tasks that involve Excel spreadsheets
  • Tasks that involve PowerPoint presentations
  • Tasks that involve PDF

Example prompts

  • “/sn-da-non-spreadsheet-analysis”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Identify file type and input scope
  2. Load sub-skill by format
  3. Parse and extract
  4. Answer with verification

What it can do on your machine

Read from SKILL.md and the folder at commit 7838651. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sn Da Non Spreadsheet Analysis loads about 1.2k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 312 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenSenseNova/SenseNova-Skills at commit 7838651, republished under its MIT licence (© OpenSenseNova). 312 words, ~1,207 tokens.

Download SKILL.mdSave it as .claude/skills/sn-da-non-spreadsheet-analysis/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
sn-da-non-spreadsheet-analysis
description
Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。**遇到以下任一情况就主动使用本 skill**:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 / 幻灯片分析 / 发票提取 / 合同分析 / 文档统计 / 错别字 / 语病 / 字号检查 / 简历分析 / 多文档对比;③任务涉及从文档中提取表格、数值、图表、格式(颜色/高亮/字号)、组织架构、时间线等结构化信息。仅不用于:Excel/CSV 数据分析(使用 sn-da-excel-workflow)、纯图片分析(使用 sn-da-image-caption)。

Document Analysis Skill — Word / PDF / PPT

End-to-end workflow for Word, PDF, and PPT document parsing. Each format has specific parsing pitfalls — follow the format-specific sub-skill exactly.


Workflow

Step 0 — Identify file type and input scope
python
import os

input_path = "/mnt/data/..."  # from user

# Detect single file vs directory (multi-file scenario)
if os.path.isdir(input_path):
    all_files = [
        os.path.join(input_path, f)
        for f in os.listdir(input_path)
        if f.lower().endswith(('.docx', '.doc', '.pdf', '.pptx', '.ppt'))
    ]
    print(f"Found {len(all_files)} documents: {all_files}")
else:
    all_files = [input_path]

# Route by extension
ext = os.path.splitext(all_files[0])[-1].lower()
print(f"File type: {ext}")

Critical rule: When input_path is a directory OR the user says "这些文件" / "所有文档", process every file and aggregate. Never stop at the first file.


Step 1 — Load sub-skill by format
ExtensionSub-skill to load
.docx / .doccapability/word-analysis/SKILL.md
.pdfcapability/pdf-analysis/SKILL.md
.pptx / .pptcapability/ppt-analysis/SKILL.md
read_file(path="<skills_root>/sn-da-non-spreadsheet-analysis/capability/<format>-analysis/SKILL.md")

Load only the sub-skill you need — do not load all three at once.


Step 2 — Parse and extract

Follow the sub-skill's extraction pattern. For all formats:

  • Full scan: iterate all pages/slides/paragraphs — never stop early
  • Table extraction: get every table, not just the first one
  • Image/chart detection: if a page/slide yields no text, treat it as image-based and call caption.py

Step 3 — Answer with verification

After extracting data, verify before answering:

python
# For count/statistics questions: spot-check 3-5 items
sample = result_list[:3]
print(f"Sample check: {sample}")
print(f"Total count: {len(result_list)}")

# For numeric calculations: print intermediate values
print(f"Max={max_val}, Min={min_val}, Range={max_val - min_val}")

# For unit-sensitive answers: always include the unit
print(f"Answer: {value} {unit}")  # e.g., "475 千港元" not just "475"

Universal Rules

MUST DO
  • Always iterate all pages/slides/paragraphs — for page in doc, for slide in prs.slides, for para in doc.paragraphs
  • When input is a directory: collect and process all matching files, then aggregate results
  • For scanned PDFs: detect empty text → call caption.py for OCR
  • For image-only slides: text extraction returns empty → render slide as PNG → call caption.py
  • For calculations: show intermediate values; confirm unit matches the question
NEVER DO
  • Do NOT use pytesseract or easyocr as primary OCR — they are not installed; use caption.py
  • Do NOT use PIL pixel analysis to infer chart values — use vision model caption instead
  • Do NOT stop at the first file, first page, or first table
  • Do NOT guess content from filenames — always parse the actual file
  • Do NOT output percentage when the question asks for absolute value (and vice versa)

Caption Script (for image/chart content in any document)

When a page, slide, or embedded image needs vision understanding, load the sn-da-image-caption skill first, then use its scripts/caption.py:

read_file(path="<skills_root>/sn-da-image-caption/SKILL.md")
python
import subprocess, json

CAPTION = "/path/to/skills/sn-da-image-caption/scripts/caption.py"

def caption_image(image_path, prompt=None):
    cmd = ["python3", CAPTION, image_path, "--json"]
    if prompt:
        cmd += ["--prompt", prompt]
    result = subprocess.run(cmd, capture_output=True, text=True, timeout=60)
    if result.returncode != 0:
        raise RuntimeError(f"caption failed: {result.stderr[:200]}")
    return json.loads(result.stdout)["description"]

# Example prompts by content type:
# Table:  "提取表格所有内容,Markdown 表格格式,保持行列结构,数值不四舍五入。"
# Chart:  "提取图表标题、坐标轴标签、每个数据点的数值。Markdown 表格输出。"
# Diagram: "描述所有节点和连接关系。"

Available sub-skills

sn-da-non-spreadsheet-analysis/capability/word-analysis/SKILL.md   — .docx/.doc
sn-da-non-spreadsheet-analysis/capability/pdf-analysis/SKILL.md    — .pdf
sn-da-non-spreadsheet-analysis/capability/ppt-analysis/SKILL.md    — .pptx/.ppt

© OpenSenseNova, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/sn-da-non-spreadsheet-analysis of OpenSenseNova/SenseNova-Skills.

  • SKILL.md
  • capability

Open the folder on GitHubat commit 7838651

Compare with similar skills

Sn Da Non Spreadsheet Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sn Da Non Spreadsheet Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sn Da Non Spreadsheet Analysis this skillOpenSenseNova/SenseNova-Skills5.7k—~1.2kAutomated safety check: PassMIT
Anydocmagnus919/agent-skills116—~3.8kAutomated safety check: NotesMIT
Documentszhongkaifu/TensorSharp559—~4.2kAutomated safety check: PassBSD-3-Clause
Office ArtifactsPrismer-AI/PrismerCloud1.6k—~2.6kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5952 repos~1.7kAutomated safety check: PassNone

Similar skills

  • Anydoc

    magnus919/agent-skills

    Convert Word (.doc/.docx/.docm), PowerPoint (.ppt/.pps/.pot/.pptx/.pptm/.ppsx/.ppsm), Excel (.xls/.xlsx/.xlsm/.xlsb), OpenDocument (.odt/.ods/.odp), RTF, EPUB, CSV, and PDF documents to clean…

    116 GitHub stars~3.8k tokensUpdated yesterday
    Documents & OfficeAuto-check: notes
  • Documents

    zhongkaifu/TensorSharp

    Read and write real documents on the device - PDF, XLSX, DOCX, PPTX and CSV.

    559 GitHub stars~4.2k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Office Artifacts

    Prismer-AI/PrismerCloud

    Generate real DOCX, PPTX, XLSX, PDF, CSV files using python-docx / python-pptx / openpyxl / reportlab by writing them into the dispatch artifacts dir, then explicitly deliver each one with cloud…

    1.6k GitHub stars~2.6k tokensUpdated 9 days ago
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    595 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Office Documents

    ZS520L/HanakoPro

    A skill your agent uses when the user asks to open, read, inspect, understand, summarize, analyze, extract tables/text from, modify, update, repair, split, merge, rotate, or convert information from…

    103 GitHub stars~1.6k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed

More from OpenSenseNova/SenseNova-Skills

All 36 skills in this repo
  • SN Motion HTML

    OpenSenseNova/SenseNova-Skills

    Builds HTML stories where one continuous camera journey advances with page progress, using researched structure, AI stills, Seedance video clips and browser QA.

    5.7k GitHub stars~2.2k tokensUpdated today
    Auto-check: notes
  • SenseNova PPT Fallback Tools

    OpenSenseNova/SenseNova-Skills

    Fallback scripts for web search, image search and download, and image generation that PPT skills use only when the host agent lacks or fails its own tools.

    5.7k GitHub stars~575 tokensUpdated today
    Auto-check: notes
  • SenseNova PPT Workbench

    OpenSenseNova/SenseNova-Skills

    Opens the PPT Workbench web editor for an existing SenseNova HTML slide deck so you can preview, inspect and visually edit it without regenerating.

    5.7k GitHub stars~2.5k tokensUpdated today
    Auto-check: notes
  • SenseNova PPT Creative Renderer

    OpenSenseNova/SenseNova-Skills

    Turns an approved slide outline into a full-page image for every slide, one 16:9 PNG per page, and optionally packages the set into a PPTX.

    5.7k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • SenseNova PPT Entry

    OpenSenseNova/SenseNova-Skills

    Entry point for SenseNova presentation generation: creates a task folder, picks depth, output format and design richness, and routes to the right PPT skill.

    5.7k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes
  • China Market Open Data Search

    OpenSenseNova/SenseNova-Skills

    Researches Chinese market, macro, trade, procurement, listed-company and regulatory information from free official sources that need no sign-up or API key.

    5.7k GitHub stars~954 tokensUpdated today
    Auto-check: notes

Questions about Sn Da Non Spreadsheet Analysis

What does Sn Da Non Spreadsheet Analysis do?

Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。遇到以下任一情况就主动使用本 skill:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 /…. Sn Da Non Spreadsheet Analysis is an agent skill from OpenSenseNova/SenseNova-Skills.

When should I use Sn Da Non Spreadsheet Analysis?

Sn Da Non Spreadsheet Analysis fits situations like: tasks that involve Excel spreadsheets; tasks that involve PowerPoint presentations; tasks that involve PDF.

How do I install Sn Da Non Spreadsheet Analysis in Claude Code?

Run `npx skills add OpenSenseNova/SenseNova-Skills --skill sn-da-non-spreadsheet-analysis -a claude-code`. Or copy the skill folder (skills/sn-da-non-spreadsheet-analysis in OpenSenseNova/SenseNova-Skills) into .claude/skills/sn-da-non-spreadsheet-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Sn Da Non Spreadsheet Analysis in Codex?

Run `npx skills add OpenSenseNova/SenseNova-Skills --skill sn-da-non-spreadsheet-analysis -a codex`. Or copy the skill folder (skills/sn-da-non-spreadsheet-analysis in OpenSenseNova/SenseNova-Skills) into .agents/skills/sn-da-non-spreadsheet-analysis in your project. Codex loads it when a task matches its description.

Can I use Sn Da Non Spreadsheet Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenSenseNova/SenseNova-Skills --skill sn-da-non-spreadsheet-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sn-da-non-spreadsheet-analysis, .gemini/skills/sn-da-non-spreadsheet-analysis, .github/skills/sn-da-non-spreadsheet-analysis and .opencode/skills/sn-da-non-spreadsheet-analysis in your project.

What does Sn Da Non Spreadsheet Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Sn Da Non Spreadsheet Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Sn Da Non Spreadsheet Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sn Da Non Spreadsheet Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sn Da Non Spreadsheet Analysis use?

Sn Da Non Spreadsheet Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sn Da Non Spreadsheet Analysis use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sn Da Non Spreadsheet Analysis?

Skills that share tags, products or a category with Sn Da Non Spreadsheet Analysis: Anydoc (magnus919/agent-skills, 116 stars), Documents (zhongkaifu/TensorSharp, 559 stars), Office Artifacts (Prismer-AI/PrismerCloud, 1.6k stars) and Markitdown (ImCa0/just-laws, 782 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sn Da Non Spreadsheet Analysis?

OpenSenseNova (a GitHub organization) maintains it in OpenSenseNova/SenseNova-Skills, which has 5,747 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on October 9, 2026.

Source: OpenSenseNova/SenseNova-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.