Agent skill

Office File Transform

by CherryHQ in CherryHQ/cherry-studio

Derives new files from a selected part of a spreadsheet, Word document, PDF or slide deck, such as a cell range, paragraph, page or slide, without ever modifying the source file.

AGPL-3.0Auto-check passedDocuments & Office

Install Office File Transform

skills CLI
$ npx skills add CherryHQ/cherry-studio --skill office-transform -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install CherryHQ/cherry-studio office-transform --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/CherryHQ/cherry-studio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/resources/skills/office-transform .claude/skills/office-transform && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
office-transform
GitHub stars
52k
Token cost
~4.7k tokens
SKILL.md length
2,542 words
Files
12 (incl. scripts, references)
Skills in repo
30
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Derives new files from a selected part of a spreadsheet, Word document, PDF or slide deck, such as a cell range, paragraph, page or slide, without ever modifying the source file.

  • Works in 2 steps: The source file is read-only. Every… → Anchors are structural. Selections…
  • Extracting a worksheet range to CSV or Markdown
  • SKILL.md covers When to use which tool, Input: selection references, Operations and Output conventions, plus 2 more sections
  • Runs Python scripts from its folder; calls uv and pip

What it does

Two invariants hold throughout. The source file is read-only, so every result is a new file and both scripts refuse to write to the source path or overwrite an existing file. Selections use the document's own coordinates, namely a worksheet with an A1 range, a body-level paragraph ordinal or a one-based page number, never screen or DOM positions. A selection can arrive as a fenced selection-ref block from a chat surface, and its anchor object is passed straight to the scripts.

Supported operations are extracting an xlsx range to csv, markdown or xlsx, extracting docx paragraphs and replacing text, extracting pdf pages, extracting and editing pptx slides, shapes and table cells, and making targeted xlsx cell edits. Before a patch-copy edit the agent checks the file's size and modification time and confirms the selection still points where it did. If the file changed since you selected, it tells you and asks for a new selection instead of silently re-anchoring.

Use it when you want a file out. For merely reading a document to summarize it or answer questions, the skill points to a separate to_markdown tool. The scripts office_extract.py and office_patch_copy.py, per-format modules and edit guides for docx and pptx are included.

When your agent uses it

  • Extracting a worksheet range to CSV or Markdown
  • Pulling selected pages out of a PDF into a new file
  • Editing text in a Word document or slide as a copy
  • Extracting specific slides, shapes or table cells from a deck

Example prompts

  • “Extract A1:C10 from Sheet1 of report.xlsx into a new CSV file.”
  • “Copy pages 3 and 4 of contract.pdf into a separate PDF.”
  • “Make a copy of deck.pptx with the title on slide 2 changed, leaving the original alone.”

Requirements

  • Python, run through uv

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. The source file is read-only. Every result is a new file; both scripts refuse to
  2. Anchors are structural. Selections address the document's own coordinates —

What it can do on your machine

Read from SKILL.md and the folder at commit dd0767e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 9 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Office File Transform loads about 4.7k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 147 tokens; SKILL.md has 2,542 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from CherryHQ/cherry-studio at commit dd0767e, republished under its AGPL-3.0 licence (© CherryHQ). 2,542 words, ~4,658 tokens.

Download SKILL.mdSave it as .claude/skills/office-transform/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
office-transform
description
Derive new Office files from structural selections without touching the original. Use when the user points at part of a spreadsheet, Word document, PDF, or PowerPoint deck (a worksheet range, paragraph, page, slide, shape, or a pasted selection-ref block) and wants it extracted, converted, or edited — the result is always a NEW file; the source file is never modified. Covers xlsx range extraction to csv/markdown/xlsx, docx paragraph extraction and text replacement, pdf page extraction, pptx slide/shape/table-cell extraction and editing, and targeted xlsx cell edits.
version
1.0.0

Office Transform

Turn a structural selection of an Office/PDF document into a new derived file.

Two invariants hold for every operation:

  1. The source file is read-only. Every result is a new file; both scripts refuse to write to the source path or overwrite an existing file. Never work around this with ad-hoc shell edits to the original.
  2. Anchors are structural. Selections address the document's own coordinates — worksheet + A1 range, body-level paragraph ordinal, one-based page number — never screen or DOM positions.

When to use which tool

  • Only need to read a document to summarize or answer questions → use mcp__cherry-tools__to_markdown instead (see the cherry-tool-guide skill); it is lossy but built for reading.
  • The user wants a file out — an extracted range, a converted fragment, an edited copy → this skill.

Input: selection references

Chat surfaces may hand you a fenced selection-ref block:

selection-ref
{"path": "/abs/report.xlsx", "anchor": {"format": "xlsx", "sheet": "Sheet1", "range": "A1:C10"}, "excerpt": "…", "fileStamp": {"size": 1024, "mtimeMs": 1700000000000}}

The anchor object is exactly what the scripts take via --anchor.

Freshness check. fileStamp.mtimeMs is milliseconds since the Unix epoch, floored to whole milliseconds. Do not compare it to stat's default output: stat -f %m (macOS) and stat -c %Y (GNU) give whole seconds, so multiplying by 1000 loses the sub-second part and never matches. Read whole milliseconds the way the app wrote them:

bash
uv run python -c "import os,sys;st=os.stat(sys.argv[1]);print(st.st_size, st.st_mtime_ns//1_000_000)" '/abs/report.xlsx'

Treat the file as changed when the size differs, or when the mtimes differ by more than 2 ms. The tolerance is small on purpose: both sides floor the same nanosecond timestamp to milliseconds, so they agree exactly, and the couple of milliseconds only covers the rounding of a double that can no longer represent current epoch milliseconds exactly. A wider window would hide real edits — a same-size change made within it would read as unchanged, which is why the anchor check below is not optional. On a change, tell the user the file changed since they selected and ask them to re-select — never silently re-anchor.

Anchor check. Before a patch-copy edit, also verify the anchor still points where the user thinks: extract the anchored region and compare its text with the reference's excerpt, normalizing both sides the same way (NFC, collapse each whitespace run to one space, trim). The two are not equal and are not meant to be — the excerpt is what the user selected, while the extract is the whole region an edit would replace. A selection sitting inside one region makes the excerpt the shorter side; a selection running past that region anchors to where it starts, which makes it the longer one. So the test is containment, whichever way round fits: the excerpt appears inside the extract, or the extract's tail is where the excerpt begins. Only when neither holds has the anchor stopped pointing at the selection — then stop and tell the user, never edit at a mismatched anchor and never go searching for a "close enough" location. Two cases need care because the two sides are not directly comparable as-is:

  • xlsx: the excerpt is tab-separated cells and newline-separated rows, while extraction writes csv or md. Compare cell values, not the raw file — read the csv with a csv reader and join the fields with single spaces before normalizing, so , and | never count as a difference. That join drops the cell boundaries on both sides, so a match proves the text is unchanged, not the grid — the same words re-split across cells still reads as a match. The freshness check above is what catches that; never skip it on a passing anchor check. Number formats are the second asymmetry: the excerpt holds what Excel displays, because that is what the user pointed at (1,234.50, 45.67%), while extraction writes the stored value (1234.5, 0.4567). Neither side is wrong, so a formatted cell never matches literally. Compare those cells by value, and leave them out of the joined strings so the containment test runs on the text cells alone. Dates, times and durations are the same asymmetry seen from the other side: extraction writes one fixed shape (2024-01-03, 2024-01-03 15:04:05, 26:30:00) rather than the cell's format, so treat a cell that differs only in date or time presentation as a format difference too. A difference that is only the number format is not a moved anchor, and stopping on one sends the user back to re-select a selection that never moved.
  • docx with charRange: the slice is only part of what patch-copy compares and replaces — it rewrites the whole paragraph. Extract the paragraph without charRange as well, and read "Edit docx" below before writing.

Users may also describe the region in words ("sheet 2, columns A through C"); build the anchor JSON yourself, confirming the worksheet name or paragraph if ambiguous.

Anchor shapes:

FormatAnchor
xlsx{"format":"xlsx","sheet":"Sheet1","range":"A1:C10"} (range may be one cell)
docx{"format":"docx","paragraph":3,"paraId":"502E8D33","charRange":[0,12]} (paraId optional = the paragraph's w14:paraId, resolved first when present; charRange optional; ordinal counts body-level paragraphs only, tables excluded)
pdf{"format":"pdf","page":3,"charRange":[0,120]} (charRange optional, applies to extracted text)
pptx{"format":"pptx","slide":2,"nodeId":"4","paragraph":0} or {"format":"pptx","slide":2,"nodeId":"7","tableCell":{"row":1,"col":0}} (slide is one-based; nodeId is the OOXML shape id — omit for the whole slide; paragraph and tableCell are optional, mutually exclusive, and only valid together with nodeId)

Operations

Scripts live in this skill's scripts/ directory; resolve paths relative to this skill folder. The library-edit recipes routed to below live in references/ beside them. Python dependencies are per-format and provided at invocation time via uv run --with <pkg> (the bundled-shell idiom — do not pip install globally).

Which route an edit takes follows from the library that can write the format: openpyxl drops charts and drawings on a round-trip, which is why xlsx edits go through patch-copy, while python-pptx keeps XML it does not understand, which is why pptx edits go through the library.

Extract — pull the anchored region into a new file

Always single-quote paths — real documents have spaces in their names (Q1 report.xlsx). If a path itself contains a single quote, close and reopen the quoting around it: '/abs/Bob'\''s deck.pptx'.

bash
uv run --with openpyxl python scripts/office_extract.py \
  --file '/abs/report.xlsx' \
  --anchor '{"format":"xlsx","sheet":"Sheet1","range":"A1:C10"}' \
  --out '/abs/report-q1-range.csv'

The output format is inferred from --out's extension:

SourceDependency (--with)Output formats
xlsxopenpyxlxlsx, csv, md
docx'python-docx>=1.1,<2'docx, txt, md
pdfpypdfpdf (page copy), txt, md
pptxpython-pptxtxt, md (slide, shape, paragraph, or table-cell text)

The docx pin is not optional. Patch-copy's expectText gate compares a paragraph read with python-docx against the same paragraph read by the script's own paragraph_text, which reproduces python-docx's Paragraph.text element for element. That equivalence was checked against 1.x; a release that changes what .text spells would make the gate refuse paragraphs nobody edited. Quote the specifier — > and < are redirects to a shell.

xlsx extraction reads computed values (data_only), so formula cells yield their last saved result. docx extraction to docx carries text only, not run styling.

Patch-copy — derive an edited copy, standard library only
bash
uv run python scripts/office_patch_copy.py \
  --file '/abs/report.xlsx' \
  --edits '{"format":"xlsx","sheet":"Sheet1","cells":{"B2":42,"C3":"hello"}}' \
  --out '/abs/report-updated.xlsx'

OOXML packages are ZIPs of XML parts. Patch-copy copies every part byte-for-byte and re-serializes only what an edit reaches — the target worksheet or word/document.xml, plus the workbook bookkeeping noted below for xlsx — so styles, charts, images, and macros in untouched parts survive exactly. Edit shapes:

  • {"format":"xlsx","sheet":"S","cells":{"B2":42,"C3":"text","D4":true}} — numbers, strings, and booleans; an existing formula in an edited cell is replaced by the value. Patch-copy does not recalculate, so a formula reading an edited cell keeps the value it last cached; every write sets fullCalcOnLoad in xl/workbook.xml so Excel recomputes those on open. Replacing a formula also drops xl/calcChain.xml (a recalculation cache Excel rebuilds), along with the content-type and relationship entries that would otherwise point at a part no longer there; keeping a chain entry for a cell that no longer has a formula makes Excel report the derived file as damaged. Cells in a shared, array, or data-table formula group are refused — the expression lives in one member and the others only reference it, so overwriting a member would strip the formula from cells you never named. Rewrite such a range with openpyxl. Coordinates outside the worksheet grid (past XFD or row 1048576) are refused too.
  • {"format":"docx","replacements":[{"paragraph":3,"text":"new text","paraId":"502E8D33","expectText":"old text"}]} — the paragraph keeps its paragraph style and the first run's character style; extra run-level styling within that one paragraph is flattened into the new text. text must be the complete new paragraph. The whole body paragraph is replaced, and charRange does not narrow that — feeding back a charRange slice as text silently discards the rest of the sentence. The output is exactly w:p > [w:pPr] + w:r > [w:rPr] + w:t, so a paragraph holding anything that shape cannot carry is refused, not silently stripped. That covers bookmarks, comment anchors and fields (their start/end can pair across paragraphs, and rewriting one half unbalances the document), images, embedded objects, footnote/endnote references, hyperlinks, tracked changes and moves, content controls, equations — and anything else not on the short allow-list, including elements from namespaces that did not exist when this was written. A dropped w:del would even accept a pending deletion on the user's behalf. A page or column break (<w:br w:type="page"/>) is refused for a quieter reason: it carries no characters, so the extract reads the text on either side of it as one string and the anchor check cannot see that the rewrite would delete it. A bare <w:br/> line break still passes. To edit such a paragraph, see "Edit docx" below — do not reach for Paragraph.text, which destroys exactly the same content, only silently. paraId (optional) is resolved before the ordinal; a disagreement between the two is an error, never a silent pick. expectText (optional but strongly recommended) is a hard gate: the target paragraph's current text must match it after whitespace normalization or the edit is refused. Take its value from an extract of the same paragraph taken without charRange — the gate compares the whole paragraph, and a selection-ref excerpt is truncated at 2000 chars and may span more than the edit target.
  • Any text written into a cell or paragraph must be storable in XML: control characters other than tab, newline and carriage return are refused. Text extracted from a deck can carry them (python-pptx maps a soft line break to \x0B), so strip them before feeding extracted text back in as an edit value. A docx paragraph refuses tab, newline and carriage return as well: WordprocessingML spells those <w:tab/> and <w:br/>, and the single w:t this rewrite emits cannot carry them. Extracting a paragraph that holds them gives you \t / \n, so editing that string and writing it back would delete the elements while reading identically — split the content across separate body paragraphs, or see "Edit docx" below.

Text comparisons on both sides of this skill use one normalization rule, identical to the renderer's normalizeSelectionText: NFC-normalize, collapse every whitespace run to a single space, trim the ends.

Show full SKILL.md (792 more words)Show less
Generate — write ad-hoc library code for new documents

Generation (a fresh deck, workbook, or document — from scratch or from data you just extracted) has no source file to protect, so there is no fixed script: write a short Python program against the matching library (python-pptx, openpyxl, python-docx) and run it via uv run --with <pkg> python. The two skill invariants still apply: write to a path nothing occupies yet — open it with "xb" rather than handing the library a name, since every one of these save() methods overwrites — and verify the output by reopening it before reporting success.

Edit docx — edit runs, never Paragraph.text

When patch-copy refuses a paragraph, python-docx can still edit it — but only run by run. Never assign paragraph.text = "...": it inflicts exactly the loss patch-copy refused, minus the refusal. save(path) overwrites, so open the destination with "xb".

Read references/docx-edit.md before writing any python-docx edit.

Edit pptx — use python-pptx, saving to a new path

pptx edits do not go through office_patch_copy.py. Never assign .text at paragraph or shape level: it rebuilds that subtree as one unformatted run. Edit runs, and open the destination with "xb".

Read references/pptx-edit.md before writing any python-pptx edit.

Output conventions

  • Name derived files after the source with an operation suffix: report.xlsx → report-updated.xlsx, report-q1-range.csv, spec-p3.txt.
  • Write into the session workspace (or where the user asked). Both scripts print the written path on success — report it to the user.
  • --file and --out must both be absolute; a relative path is refused rather than resolved against whatever working directory the shell happens to be in.
  • Output is staged and renamed on success, so a failed run leaves nothing behind and the same command can be retried at the same path.

Verify before reporting success

Always reopen the derived file with the matching reader and check the result, e.g.:

bash
uv run --with openpyxl python -c "
from openpyxl import load_workbook
ws = load_workbook('/abs/report-updated.xlsx', data_only=True)['Sheet1']
print(ws['B2'].value)"

data_only reads cached values and patch-copy does not recalculate, so a formula reading a cell you just wrote still reads back its old result. Check the cells you wrote, and tell the user the dependent totals are recomputed when Excel opens the file.

A written xlsx string cell that held a literal _xHHHH_ or a carriage return reads back through openpyxl in its escaped form (_x005F_xHHHH_, _x000D_) — see the ST_Xstring note under "Limits". That is the reader, not a verification failure.

If verification fails, say so and show the error — do not present an unverified file.

Limits

  • pptx edits go through python-pptx (see "Edit pptx"), not patch-copy; slide-copy into a new deck is not supported (python-pptx cannot clone slides) — say so when asked for it.
  • Patched xlsx string cells become inline strings (valid OOXML; Excel reads them fine). Their text is written as ST_Xstring, so a literal _xHHHH_ in the value is escaped to _x005F_xHHHH_ and a carriage return is stored as _x000D_ — that is what makes Excel show the characters that were written. openpyxl decodes neither, so a verification read of such a cell shows the escaped form; that is the reader, not the file.
  • Edited XML parts may lose insignificant serialization details (attribute quoting, empty-element form); namespace prefixes and untouched content are preserved.
  • A patched docx paragraph keeps its characters but not every break opportunity: w:softHyphen leaves no mark in the extract and is dropped, and w:noBreakHyphen comes back as a plain -, so what both lose is only where a line may break. w:sym is refused instead — its glyph is a character the reader can see, not a hyphenation hint.
  • xlsx extraction refuses ranges over 1,000,000 cells; both scripts refuse packages with more than 10,000 entries, an entry over 256 MiB uncompressed, or over 1 GiB total uncompressed — ask the user for a smaller selection or file instead of retrying.
  • xlsx extraction does not mask merge followers — the streaming reader cannot see the merge ranges. Excel and openpyxl clear the covered cells when a merge is made, so followers normally extract as empty, matching what the user sees; a file that kept hidden text under a merge (LibreOffice offers to) extracts that text as its own cell. If the anchor check fails inside a merged range, look at the cells before telling the user to re-select.
  • csv and md carry cell text verbatim, formula-looking text included (=SUM(A1), +1, -); a spreadsheet opening the csv would evaluate it, which matters only for a source the user does not trust — say so if they mean to open the csv in one. In the derived xlsx, text that openpyxl would re-type — a leading = into a live formula, an error code such as #N/A into an error value — is written back as quote-prefixed text; +1 and - are already stored as text and stay that way, and a cell that really holds an error stays an error.
  • Scanned/image-only PDFs yield no text (no OCR here).

© CherryHQ, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts, references) in resources/skills/office-transform of CherryHQ/cherry-studio.

  • SKILL.md
  • references/docx-edit.md
  • references/pptx-edit.md
  • scripts/office/__init__.py
  • scripts/office/common.py
  • scripts/office/docx.py
  • scripts/office/ooxml.py
  • scripts/office/pdf.py
  • scripts/office/pptx.py
  • scripts/office/xlsx.py
  • scripts/office_extract.py
  • scripts/office_patch_copy.py

Open the folder on GitHubat commit dd0767e

Compare with similar skills

Office File Transform next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Office File Transform compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Office File Transform this skillCherryHQ/cherry-studio52k—~4.7kAutomated safety check: PassAGPL-3.0
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
MinerU Document Readeropendatalab/MinerU81k—~9.4kAutomated safety check: WarnCustom licence
GenOffice Document CLIgenspark-ai/genoffice8.8k—~19kAutomated safety check: PassApache-2.0
Markitdownjimmc414/Kosmos5942 repos~1.7kAutomated safety check: PassNone
Documentszhongkaifu/TensorSharp553—~4.2kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • MinerU Document Reader

    opendatalab/MinerU

    Reads, OCRs, searches and cites local documents through the mineru CLI, covering PDF, images, Office files, EPUB, HTML and CSV.

    81k GitHub stars~9.4k tokensUpdated yesterday
    Documents & OfficeAuto-check: warnings
  • GenOffice Document CLI

    genspark-ai/genoffice

    Creates, converts, reads and edits real pptx, xlsx, docx and PDF files locally through the genoffice command line.

    8.8k GitHub stars~19k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Documents

    zhongkaifu/TensorSharp

    Read and write real documents on the device - PDF, XLSX, DOCX, PPTX and CSV.

    553 GitHub stars~4.2k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Markdown Exporter

    bowenliang123/markdown-exporter

    Convert Markdown text to DOCX, PPTX, XLSX, PDF, PNG, SVG, HTML, IPYNB, MD, CSV, JSON, JSONL, XML files, and extract code blocks in Markdown to Python, Bash,JS and etc files.

    271 GitHub starsUsed in 1 repo~5.3k tokens
    Documents & OfficeAuto-check passed

More from CherryHQ/cherry-studio

All 30 skills in this repo
  • GitHub Issue Creator

    CherryHQ/cherry-studio

    Creates GitHub issues for the current repository by choosing the matching issue template and following its format, with a permission check for engineering tasks.

    52k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Cherry Studio PR Review

    CherryHQ/cherry-studio

    Reviews Cherry Studio branches, pull requests, commits, files and docs against the project's own architecture, naming, API-boundary and UI rules, report-only by default.

    52k GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Cherry Studio Regression Tests

    CherryHQ/cherry-studio

    Runs Cherry Studio's critical-path regression suite as deterministic Playwright E2E tests through a GitHub workflow on macOS and Windows runners.

    52k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Antigravity CLI Runner

    CherryHQ/cherry-studio

    Runs the Antigravity CLI headlessly with the agy command to analyze a repository or carry out a coding task, then checks its JSON result and the diff.

    52k GitHub stars~531 tokensUpdated today
    Auto-check passed
  • Kimi Code Delegation

    CherryHQ/cherry-studio

    Delegates one bounded repository task to Kimi Code in non-interactive prompt mode and reads back the final result from its JSON event stream.

    52k GitHub stars~504 tokensUpdated today
    Auto-check passed
  • GitHub PR Creation

    CherryHQ/cherry-studio

    Creates or updates GitHub pull requests by reading the repository's PR template, filling every section and picking the right base branch under its release rules.

    52k GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Questions about Office File Transform

What does Office File Transform do?

Derives new files from a selected part of a spreadsheet, Word document, PDF or slide deck, such as a cell range, paragraph, page or slide, without ever modifying the source file. Two invariants hold throughout. The source file is read-only, so every result is a new file and both scripts refuse to write to the source path or overwrite an existing file.

When should I use Office File Transform?

Office File Transform fits situations like: extracting a worksheet range to CSV or Markdown; pulling selected pages out of a PDF into a new file; editing text in a Word document or slide as a copy; extracting specific slides, shapes or table cells from a deck.

How do I install Office File Transform in Claude Code?

Run `npx skills add CherryHQ/cherry-studio --skill office-transform -a claude-code`. Or copy the skill folder (resources/skills/office-transform in CherryHQ/cherry-studio) into .claude/skills/office-transform in your project. Claude Code loads it when a task matches its description.

How do I install Office File Transform in Codex?

Run `npx skills add CherryHQ/cherry-studio --skill office-transform -a codex`. Or copy the skill folder (resources/skills/office-transform in CherryHQ/cherry-studio) into .agents/skills/office-transform in your project. Codex loads it when a task matches its description.

Can I use Office File Transform in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add CherryHQ/cherry-studio --skill office-transform -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/office-transform, .gemini/skills/office-transform, .github/skills/office-transform and .opencode/skills/office-transform in your project.

What does Office File Transform need to run?

Going by SKILL.md and its folder, Office File Transform needs Python for the scripts in its folder and the command-line tools its instructions call (uv and pip). Our summary lists: Python, run through uv.

Does Office File Transform access the network?

SKILL.md contains no URLs. Its commands use uv and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Office File Transform safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Office File Transform use?

Office File Transform is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Office File Transform use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Office File Transform?

Skills that share tags, products or a category with Office File Transform: Markitdown (ImCa0/just-laws, 781 stars), MinerU Document Reader (opendatalab/MinerU, 81k stars), GenOffice Document CLI (genspark-ai/genoffice, 8.8k stars) and Markitdown (jimmc414/Kosmos, 594 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Office File Transform?

CherryHQ (a GitHub organization) maintains it in CherryHQ/cherry-studio, which has 52,407 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on October 7, 2026.

Source: CherryHQ/cherry-studio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.