Agent skill

Ppt Analysis

by OpenSenseNova in OpenSenseNova/SenseNova-Skills

PPT (.pptx/.ppt) 全量解析。覆盖:所有 slide 文本/表格/图表提取、嵌入图片 caption、纯图片 slide 渲染识别、数据标签提取。

MITAuto-check passedDocuments & Office

Install Ppt Analysis

skills CLI
$ npx skills add OpenSenseNova/SenseNova-Skills --skill ppt-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenSenseNova/SenseNova-Skills ppt-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenSenseNova/SenseNova-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sn-da-non-spreadsheet-analysis/capability/ppt-analysis .claude/skills/ppt-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ppt-analysis
GitHub stars
5.7k
Token cost
~2.8k tokens
SKILL.md length
171 words
Files
1
Skills in repo
36
Repo updated
First seen
Licence
MIT

At a glance

PPT (.pptx/.ppt) 全量解析。覆盖:所有 slide 文本/表格/图表提取、嵌入图片 caption、纯图片 slide 渲染识别、数据标签提取。

  • Tasks that involve PowerPoint presentations
  • SKILL.md covers Environment, Core Method 1: Full Text…, Core Method 2: Table… and Core Method 3: Chart Data…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Slides and decks

What it does

Ppt Analysis is an agent skill from OpenSenseNova/SenseNova-Skills. PPT (.pptx/.ppt) 全量解析。覆盖:所有 slide 文本/表格/图表提取、嵌入图片 caption、纯图片 slide 渲染识别、数据标签提取。

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PowerPoint presentations and Slides and decks. It works with Microsoft PowerPoint. The repository describes itself as: Modular SenseNova skills for building AI-powered office assistants and productivity workflows. The licence is MIT.

When your agent uses it

  • Tasks that involve PowerPoint presentations
  • Tasks that involve Slides and decks

Example prompts

  • “/ppt-analysis”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 5abde96. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ppt Analysis loads about 2.8k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 171 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenSenseNova/SenseNova-Skills at commit 5abde96, republished under its MIT licence (© OpenSenseNova). 171 words, ~2,758 tokens.

Download SKILL.mdSave it as .claude/skills/ppt-analysis/SKILL.md (or your agent's skills folder).
name
ppt-analysis
description
PPT (.pptx/.ppt) 全量解析。覆盖:所有 slide 文本/表格/图表提取、嵌入图片 caption、纯图片 slide 渲染识别、数据标签提取。

PPT Analysis — .pptx / .ppt

Environment

python
from pptx import Presentation
from pptx.util import Inches
import os, subprocess, json

# python-pptx is available
# For .ppt (old binary format): convert via libreoffice
def load_pptx(path):
    if path.lower().endswith('.ppt'):
        import subprocess
        out_dir = os.path.dirname(path)
        subprocess.run(
            ['libreoffice', '--headless', '--convert-to', 'pptx', '--outdir', out_dir, path],
            check=True, capture_output=True
        )
        path = path.rsplit('.', 1)[0] + '.pptx'
    return Presentation(path), path

Core Method 1: Full Text Extraction (ALL slides)

python
def extract_all_slides_text(pptx_path):
    """
    Extract text from every slide: text frames, tables, chart titles.
    For slides with no extractable text, flag them for image captioning.
    """
    prs, _ = load_pptx(pptx_path)
    slides_data = []

    for slide_num, slide in enumerate(prs.slides, start=1):
        slide_texts = []
        has_text = False

        for shape in slide.shapes:
            # Text frame (most common)
            if shape.has_text_frame:
                for para in shape.text_frame.paragraphs:
                    text = para.text.strip()
                    if text:
                        slide_texts.append(text)
                        has_text = True

            # Table
            if shape.has_table:
                tbl = shape.table
                for row in tbl.rows:
                    row_text = '\t'.join(cell.text.strip() for cell in row.cells)
                    if row_text.strip():
                        slide_texts.append(row_text)
                        has_text = True

            # Chart title
            if shape.shape_type == 3:  # MSO_SHAPE_TYPE.CHART
                try:
                    if shape.chart.has_title:
                        title = shape.chart.chart_title.text_frame.text
                        slide_texts.append(f"[Chart: {title}]")
                        has_text = True
                except Exception:
                    pass

        slides_data.append({
            'slide': slide_num,
            'text': '\n'.join(slide_texts),
            'has_text': has_text,
            'needs_caption': not has_text  # flag image-only slides
        })

    print(f"Total slides: {len(slides_data)}")
    image_only = sum(1 for s in slides_data if s['needs_caption'])
    print(f"Slides with text: {len(slides_data) - image_only}, image-only: {image_only}")
    return slides_data

Core Method 2: Table Extraction (Structured)

python
import pandas as pd

def extract_pptx_tables(pptx_path):
    """Extract all tables from all slides as DataFrames."""
    prs, _ = load_pptx(pptx_path)
    all_tables = []

    for slide_num, slide in enumerate(prs.slides, start=1):
        for shape in slide.shapes:
            if not shape.has_table:
                continue
            tbl = shape.table
            rows = []
            for row in tbl.rows:
                rows.append([cell.text.strip() for cell in row.cells])

            if not rows:
                continue

            # Use first row as header
            try:
                df = pd.DataFrame(rows[1:], columns=rows[0])
            except Exception:
                df = pd.DataFrame(rows)

            all_tables.append({'slide': slide_num, 'df': df})
            print(f"  Slide {slide_num}: table {df.shape[0]}r × {df.shape[1]}c")
            print(df.head(3).to_string())

    return all_tables

Core Method 3: Chart Data Extraction

python-pptx can read Chart data when it's stored as embedded Excel data. If that fails, fall back to captioning the slide image.

python
def extract_chart_data(pptx_path):
    """
    Extract data series from Chart shapes.
    Returns list of {slide, chart_title, series_name, categories, values}.
    """
    prs, _ = load_pptx(pptx_path)
    charts = []

    for slide_num, slide in enumerate(prs.slides, start=1):
        for shape in slide.shapes:
            if shape.shape_type != 3:  # not a chart
                continue
            try:
                chart = shape.chart
                title = chart.chart_title.text_frame.text if chart.has_title else f"Chart_S{slide_num}"

                for plot in chart.plots:
                    for series in plot.series:
                        try:
                            categories = [str(pt.label) for pt in series.data_labels] if hasattr(series, 'data_labels') else []
                            values = [pt.value for pt in series.values] if hasattr(series, 'values') else []
                            # Alternative: use xChart data
                            if not values:
                                values = list(series.values)
                        except Exception as e:
                            values = []
                            categories = []

                        charts.append({
                            'slide': slide_num,
                            'chart_title': title,
                            'series': getattr(series, 'name', ''),
                            'categories': categories,
                            'values': values
                        })
            except Exception as e:
                print(f"  Slide {slide_num}: chart extraction failed ({e}) — will use caption")

    return charts

Core Method 4: Render Image-Only Slides → Caption

When a slide has no extractable text (pure image/screenshot slides):

python
import fitz  # PyMuPDF can also render PPTX via LibreOffice conversion

CAPTION = "/path/to/skills/sn-da-image-caption/scripts/caption.py"

def caption_image_slides(pptx_path, slides_data, prompt=None):
    """
    For slides flagged as 'needs_caption', render to PNG and caption.
    Uses LibreOffice to convert PPTX to PDF first, then renders pages.
    """
    image_slides = [s for s in slides_data if s['needs_caption']]
    if not image_slides:
        print("No image-only slides to caption.")
        return slides_data

    # Convert PPTX → PDF (preserves slide visuals)
    out_dir = "/tmp"
    r = subprocess.run(
        ['libreoffice', '--headless', '--convert-to', 'pdf', '--outdir', out_dir, pptx_path],
        capture_output=True, text=True
    )
    pdf_name = os.path.basename(pptx_path).rsplit('.', 1)[0] + '.pdf'
    pdf_path = os.path.join(out_dir, pdf_name)

    if not os.path.exists(pdf_path):
        print(f"LibreOffice conversion failed: {r.stderr[:200]}")
        return slides_data

    # Render each image-only slide
    doc = fitz.open(pdf_path)
    for s in image_slides:
        page_idx = s['slide'] - 1  # 0-indexed
        if page_idx >= len(doc):
            continue
        page = doc[page_idx]
        mat = fitz.Matrix(150/72, 150/72)
        pix = page.get_pixmap(matrix=mat)
        img_path = f"/tmp/slide_{s['slide']}.png"
        pix.save(img_path)

        # Caption the slide image
        cmd = ["python3", CAPTION, img_path, "--json"]
        p = prompt or "提取幻灯片中所有文字、数值和表格内容,保持结构,Markdown格式输出。"
        cmd += ["--prompt", p]
        cr = subprocess.run(cmd, capture_output=True, text=True, timeout=90)
        if cr.returncode == 0:
            desc = json.loads(cr.stdout).get("description", "")
            s['text'] = desc
            s['needs_caption'] = False
            print(f"  Slide {s['slide']}: captioned ({len(desc)} chars)")
        else:
            print(f"  Slide {s['slide']}: caption failed — {cr.stderr[:80]}")

    doc.close()
    return slides_data

Common Patterns

Keyword search across all slides
python
def find_in_pptx(pptx_path, keyword, slides_data=None):
    """Find keyword across all slides (after text extraction + captioning)."""
    if slides_data is None:
        slides_data = extract_all_slides_text(pptx_path)

    results = []
    for s in slides_data:
        if keyword in s.get('text', ''):
            idx = s['text'].find(keyword)
            context = s['text'][max(0, idx-100):idx+200]
            results.append({'slide': s['slide'], 'context': context})

    print(f"'{keyword}' found in {len(results)} slides: {[r['slide'] for r in results]}")
    return results
Time-line / process extraction from PPT
python
def extract_timeline(pptx_path, date_pattern=r'\d{4}[年/\-]\d{1,2}'):
    """Extract date-tagged events from slide text."""
    import re
    slides_data = extract_all_slides_text(pptx_path)
    events = []
    for s in slides_data:
        for line in s['text'].split('\n'):
            if re.search(date_pattern, line):
                events.append({'slide': s['slide'], 'event': line.strip()})
    return events
Statistics from PPT tables (e.g., 录用占比)
python
def compute_ratio_from_pptx_table(pptx_path, numerator_col, denominator_col):
    """Example: compute ratio = col_A / col_B for all rows."""
    tables = extract_pptx_tables(pptx_path)
    for item in tables:
        df = item['df']
        # Try to find columns (flexible matching)
        num_col = next((c for c in df.columns if numerator_col in c), None)
        den_col = next((c for c in df.columns if denominator_col in c), None)
        if num_col and den_col:
            df[num_col] = pd.to_numeric(df[num_col].str.replace('人', '').str.strip(), errors='coerce')
            df[den_col] = pd.to_numeric(df[den_col].str.replace('人', '').str.strip(), errors='coerce')
            df['ratio'] = (df[num_col] / df[den_col] * 100).round(0).astype(str) + '%'
            print(df[['slide' if 'slide' in df.columns else df.columns[0], num_col, den_col, 'ratio']].to_string())

Full Workflow Example

python
pptx_path = "/mnt/data/report.pptx"

# 1. Extract text from all slides
slides_data = extract_all_slides_text(pptx_path)

# 2. Caption image-only slides
slides_data = caption_image_slides(pptx_path, slides_data)

# 3. Combine all text for analysis
all_text = '\n\n'.join(
    f"[Slide {s['slide']}]\n{s['text']}"
    for s in slides_data if s.get('text')
)

# 4. Search or analyze
results = find_in_pptx(pptx_path, '录用占比', slides_data)

# 5. Extract tables if needed
tables = extract_pptx_tables(pptx_path)

Pitfalls

PitfallFix
Skip slides with no text → miss chart dataFlag needs_caption, render & caption (Method 4)
shape.chart.plots[0].series fails → no dataCatch exception, fall back to captioning the slide
Table columns misread (企业名 vs 岗位名)Print headers + first 3 rows before computing; verify column meaning
Only read first N slidesAlways for slide in prs.slides — no index limit
.ppt format → python-pptx can't openConvert to .pptx via libreoffice first
PPT has overlapping text boxes → garbled orderSort shapes by top-left position: sorted(slide.shapes, key=lambda s: (s.top, s.left))

© OpenSenseNova, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/sn-da-non-spreadsheet-analysis/capability/ppt-analysis of OpenSenseNova/SenseNova-Skills.

Open the folder on GitHubat commit 5abde96

Compare with similar skills

Ppt Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ppt Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ppt Analysis this skillOpenSenseNova/SenseNova-Skills5.7k—~2.8kAutomated safety check: PassMIT
Image To Editable Pptningzimu/image-to-editable-ppt-skill2.8k—~4.3kAutomated safety check: PassMIT
Slidesfcakyon/claude-codex-settings1.2k1 repos~1.1kAutomated safety check: PassMIT
Ppt Image FirstNyxTides/ppt-image-first1.2k—~1.6kAutomated safety check: PassApache-2.0
Gpt Image2 PptJuneYaooo/gpt-image2-ppt-skills1.3k—~8.9kAutomated safety check: NotesApache-2.0
HTML Slide To PPTXmucsbr/ppt-agent-workflow-san643—~1.1kAutomated safety check: PassNone

Similar skills

  • Image To Editable Ppt

    ningzimu/image-to-editable-ppt-skill

    Rebuild slide images, scanned or image-based PPT/PPTX files, and PDF decks into object-level editable PowerPoint (.pptx), preserving speaker notes when supplied.

    2.8k GitHub stars~4.3k tokensUpdated 21 days ago
    Documents & OfficeAuto-check passed
  • Slides

    fcakyon/claude-codex-settings

    Create and edit presentation slide decks (.pptx) with PptxGenJS, bundled layout helpers, and render/validation utilities.

    1.2k GitHub starsUsed in 1 repo~1.1k tokens
    Documents & OfficeAuto-check passed
  • Ppt Image First

    NyxTides/ppt-image-first

    Build presentation plans for PPT / slides / decks through a conversation-first workflow, then propose multiple visual directions with preview images before writing deck specs.

    1.2k GitHub stars~1.6k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Gpt Image2 Ppt

    JuneYaooo/gpt-image2-ppt-skills

    Generate visually striking PPT slides via OpenAI's gpt-image-2 -- use any style in styles/<collection/STYLEID.md or mimic a user-supplied .pptx template; outputs high-res slide PNGs and a 16:9 .pptx.

    1.3k GitHub stars~8.9k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check: notes
  • HTML Slide To PPTX

    mucsbr/ppt-agent-workflow-san

    Convert structured single-slide or small deck HTML files into editable PPTX slides with native text boxes, shapes, chips, arrows, and panels.

    643 GitHub stars~1.1k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Paper Deck

    zsyggg/paper-craft-skills

    将论文、技术文章或知识内容制作成高真实感的 AIGC 幻灯片。先做叙事结构和逐页视觉导演,再调用生图模型生成每一页 16:9 slide image,最后合成为 PPTX/PDF。适合论文汇报、组会、公开课、技术分享、商业化研究展示;当用户提到“论文PPT”“AI生成PPT”“不像AI的PPT”“高质感幻灯片”“逐页生图PPT”时使用。

    1.3k GitHub stars~1.4k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed

More from OpenSenseNova/SenseNova-Skills

All 36 skills in this repo
  • SenseNova PPT Workbench

    OpenSenseNova/SenseNova-Skills

    Opens the PPT Workbench web editor for an existing SenseNova HTML slide deck so you can preview, inspect and visually edit it without regenerating.

    5.7k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check: notes
  • SN Motion HTML

    OpenSenseNova/SenseNova-Skills

    Builds HTML stories where one continuous camera journey advances with page progress, using researched structure, AI stills, Seedance video clips and browser QA.

    5.7k GitHub stars~2.2k tokensUpdated 20 days ago
    Auto-check: notes
  • China Market Open Data Search

    OpenSenseNova/SenseNova-Skills

    Researches Chinese market, macro, trade, procurement, listed-company and regulatory information from free official sources that need no sign-up or API key.

    5.7k GitHub starsUsed in 1 repo~954 tokens
    Auto-check: notes
  • SenseNova PPT Fallback Tools

    OpenSenseNova/SenseNova-Skills

    Fallback scripts for web search, image search and download, and image generation that PPT skills use only when the host agent lacks or fails its own tools.

    5.7k GitHub stars~575 tokensUpdated 20 days ago
    Auto-check: notes
  • SenseNova PPT Creative Renderer

    OpenSenseNova/SenseNova-Skills

    Turns an approved slide outline into a full-page image for every slide, one 16:9 PNG per page, and optionally packages the set into a PPTX.

    5.7k GitHub stars~1.2k tokensUpdated 20 days ago
    Auto-check passed
  • SenseNova PPT Entry

    OpenSenseNova/SenseNova-Skills

    Entry point for SenseNova presentation generation: creates a task folder, picks depth, output format and design richness, and routes to the right PPT skill.

    5.7k GitHub stars~2.7k tokensUpdated 20 days ago
    Auto-check: notes

Questions about Ppt Analysis

What does Ppt Analysis do?

PPT (.pptx/.ppt) 全量解析。覆盖:所有 slide 文本/表格/图表提取、嵌入图片 caption、纯图片 slide 渲染识别、数据标签提取。. Ppt Analysis is an agent skill from OpenSenseNova/SenseNova-Skills.

When should I use Ppt Analysis?

Ppt Analysis fits situations like: tasks that involve PowerPoint presentations; tasks that involve Slides and decks.

How do I install Ppt Analysis in Claude Code?

Run `npx skills add OpenSenseNova/SenseNova-Skills --skill ppt-analysis -a claude-code`. Or copy the skill folder (skills/sn-da-non-spreadsheet-analysis/capability/ppt-analysis in OpenSenseNova/SenseNova-Skills) into .claude/skills/ppt-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Ppt Analysis in Codex?

Run `npx skills add OpenSenseNova/SenseNova-Skills --skill ppt-analysis -a codex`. Or copy the skill folder (skills/sn-da-non-spreadsheet-analysis/capability/ppt-analysis in OpenSenseNova/SenseNova-Skills) into .agents/skills/ppt-analysis in your project. Codex loads it when a task matches its description.

Can I use Ppt Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenSenseNova/SenseNova-Skills --skill ppt-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ppt-analysis, .gemini/skills/ppt-analysis, .github/skills/ppt-analysis and .opencode/skills/ppt-analysis in your project.

What does Ppt Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Ppt Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Ppt Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ppt Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ppt Analysis use?

Ppt Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ppt Analysis use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ppt Analysis?

Skills that share tags, products or a category with Ppt Analysis: Image To Editable Ppt (ningzimu/image-to-editable-ppt-skill, 2.8k stars), Slides (fcakyon/claude-codex-settings, 1.2k stars), Ppt Image First (NyxTides/ppt-image-first, 1.2k stars) and Gpt Image2 Ppt (JuneYaooo/gpt-image2-ppt-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ppt Analysis?

OpenSenseNova (a GitHub organization) maintains it in OpenSenseNova/SenseNova-Skills, which has 5,747 GitHub stars. The repository holds 36 skills in this directory. The repository was last updated on September 18, 2026.

Source: OpenSenseNova/SenseNova-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.