Agent skill

Structured Extraction

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when text must become a typed, schema-conformant object you can trust — pulling fields into a fixed JSON shape, extracting line items as typed records, classifying into…

MITAuto-check passedDocuments & Office

Install Structured Extraction

skills CLI
$ npx skills add ericrisco/rsc-harness --skill structured-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness structured-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/structured-extraction .claude/skills/structured-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
structured-extraction
GitHub stars
156
Token cost
~3.6k tokens
SKILL.md length
1,582 words
Files
5 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when text must become a typed, schema-conformant object you can trust — pulling fields into a fixed JSON shape, extracting line items as typed records, classifying into…

  • Works in 6 steps: Native structured outputs ON (or a… → Every object has additionalProperties:… → Every maybe-absent field is nullable… → …
  • Text must become a typed
  • SKILL.md covers Default: native constrained…, Schema design rules (each…, The reliability ladder —… and Value validation the schema…, plus 4 more sections
  • Runs Shell scripts from its folder

What it does

Structured Extraction is an agent skill from ericrisco/rsc-harness. Use when text must become a typed, schema-conformant object you can trust — pulling fields into a fixed JSON shape, extracting line items as typed records, classifying into enums, and building the Pydantic or Zod model plus the validate-and-retry loop. Covers extractors that throw parse errors, leak markdown fences, or fabricate a value where the field is absent instead of returning null. NOT getting the text out of a PDF, scan or DOCX first (that is document-processing), NOT general prompt craft untied to a…

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/providers.md`).

It sits in Documents & Office, covering Document parsing, Word documents and Forms and validation. It works with Pydantic, Microsoft Word, Zod and OpenAI. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Text must become a typed
  • Schema-conformant object you can trust — pulling fields into a fixed JSON shape
  • Extracting line items as typed records
  • Classifying into enums

Example prompts

  • “/structured-extraction”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Native structured outputs ON (or a documented reason the provider has no native mode).
  2. Every object has additionalProperties: false; every property is in required.
  3. Every maybe-absent field is nullable with a "use null if not stated" instruction.
  4. Closed sets are enums; value rules (ranges/format/cross-field) live in validators, not the schema.
  5. A reask loop, if any, is bounded and logged.
  6. You can state, for each field, whether a wrong output would be caught by decoding (shape) or only by

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Structured Extraction loads about 3.6k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 144 tokens; SKILL.md has 1,582 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,582 words, ~3,600 tokens.

Download SKILL.mdSave it as .claude/skills/structured-extraction/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
structured-extraction
description
Use when text must become a typed, schema-conformant object you can trust — pulling fields into a fixed JSON shape, extracting line items as typed records, classifying into enums, and building the Pydantic or Zod model plus the validate-and-retry loop. Covers extractors that throw parse errors, leak markdown fences, or fabricate a value where the field is absent instead of returning null. NOT getting the text out of a PDF, scan or DOCX first (that is `document-processing`), NOT general prompt craft untied to a schema (that is `prompt-engineering`).
tags
structured-outputs, json-schema, pydantic, extraction, validation, llm
recommends
document-processing, prompt-engineering, llm-pipeline, agent-eval, rag, data-cleaning
origin
risco

Structured extraction — text in, a typed object you can trust out

The deliverable is a typed object that conforms to a schema you defined — not prose, not "roughly JSON." The whole skill rests on one distinction the rest of the file keeps returning to:

Native structured outputs make the JSON valid and typed. They never make the values correct.

Constrained decoding guarantees the model cannot emit a token that breaks your schema, so JSON.parse errors, missing keys, wrong types, and stray markdown fences disappear at the source. It does nothing to stop the model from putting a plausible-but-wrong email in a string field, snapping a fuzzy category to the wrong enum, or coercing "$1,200" into 1200.0 when the currency mattered. Owning both halves — the shape (decoding) and the values (validation) — is this skill. If you only do the first half you ship a database full of well-typed lies.

Boundary test (bytes vs. schema). If the input is a PDF, scan, DOCX, or HTML and the deliverable is the raw text/Markdown/cells of that document, that is upstream: document-processing produces the text, this skill turns that text into typed fields. If you're holding text and want it shaped, you're in the right place.

Current as of 2026-06-02: OpenAI Structured Outputs (strict: true json_schema), Anthropic Structured Outputs (GA since the 2025-11-14 public beta; output_config.format), and Instructor (built on Pydantic, ~3M downloads/month). Exact request/response shapes and the per-provider limit tables live in references/providers.md so this file stays lean.

Default: native constrained decoding when the provider has it

If the model and provider support native structured outputs, use them. This is not a tuning knob — it is the difference between ~100% schema conformance and hoping a regex catches the fence.

Bad — prompt-and-pray, then parse raw text:

python
resp = client.chat.completions.create(
    model="gpt-5.1",
    messages=[{"role": "user", "content": f"Return JSON with name and email:\n{text}"}],
)
data = json.loads(resp.choices[0].message.content)  # markdown fence / preamble / missing key -> crash

Good — OpenAI strict json_schema (Chat Completions):

python
resp = client.chat.completions.create(
    model="gpt-5.1",
    messages=[{"role": "user", "content": text}],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "contact",
            "strict": True,
            "schema": {
                "type": "object",
                "additionalProperties": False,
                "required": ["name", "email"],
                "properties": {
                    "name": {"type": "string", "description": "Full name as written."},
                    "email": {"type": ["string", "null"],
                              "description": "Email exactly as written, or null if none is stated."},
                },
            },
        },
    },
)
data = json.loads(resp.choices[0].message.content)  # now guaranteed valid + typed

On the OpenAI Responses API the same block moves under text.format instead of response_format. On Anthropic, the equivalent is output_config={"format": {"type": "json_schema", "schema": {...}}} on Claude Opus 4.5–4.8 / Sonnet 4.5+ / Haiku 4.5; Anthropic compiles your schema into a grammar and caches it for 24h, and the SDKs ship helpers (client.messages.parse(...) in Python, zodOutputFormat(schema) in TS). The older deprecated output_format param and the deprecated structured-outputs-2025-11-13 beta header still work in a transition window — do not write new code against them. Full shapes in references/providers.md.

The non-negotiable strict-schema rule (OpenAI and Anthropic both):

  • every object sets "additionalProperties": false;
  • every property is listed in required;
  • "optional" is expressed as a union with null ("type": ["string", "null"]), never by leaving the field out of required. Omitting it is the single most common strict-mode error.

Schema design rules (each prevents a specific failure)

You wantExpress it asBecause
A field that may be absentnullable union ["string","null"] + description: "...or null if not stated"A non-null type forces a value, so the model fabricates one. Nullable + instruction yields null instead.
A closed set of categoriesenum: ["open","pending","closed"]Free-text string drifts ("Open", "in progress", "closd"); an enum makes drift impossible to emit.
Many items of one kindone object schema + a top-level {"items": {"type":"array","items": <object>}} wrapperOne object per extraction unit keeps each record independently validatable; arrays of scalars lose field structure.
The model to read your intenta description on every propertyThe model reads field descriptions at decode time; "amount in cents, no currency symbol" beats a bare integer.
A number in a range / a regex / a length capleave it out of the schema; enforce in a post-decode validatorStrict modes reject or silently ignore minimum/maximum/minLength/maxLength/complex regex — see the unsupported-features table in references.
A deeply nested or recursive shapeflatten it, or split into two extractionsNative modes reject recursion and cap nesting/complexity; flat schemas decode reliably.

Keep schemas flat and shallow. If you find yourself nesting four levels deep or describing a tree, that is two extractions, not one heroic schema.

The reliability ladder — escalate only on failure, and cap it

Climb from the cheapest mechanism upward. Each rung catches what the rung below cannot; you stop at the first rung that holds for your data.

RungMechanismCatchesDoes NOT catchWhen you stop here
1Native constrained decodingparse errors, wrong types, missing keys, fenceswrong values, bad units, wrong enumshape+types only, latest single provider
2Pydantic / Zod validation after decodeout-of-range, bad format, cross-field contradictions, null-vs-absentnothing the model genuinely got wrongvalue rules you can express as code
3Bounded reask (Instructor or hand-rolled)semantic errors the model can fix when shown the validation messagesystematic model blind spotsresidual errors; cap retries (e.g. 2) and log every reask
4Human / log revieweverything still wrong after 3—high-stakes fields or low-confidence rows

Rung 1 is mandatory when available. Rung 2 is mandatory the moment any field has a value rule (a range, a format, a "must match the order date") — because rung 1 structurally cannot enforce values. Rungs 3 and 4 are opt-in. Never make rung 3 unbounded: a retry loop with no cap turns one bad document into an unbounded bill.

Value validation the schema can't enforce

This is the half native decoding leaves on the table. Validate values after you have a typed object.

Pydantic — value rules + normalization the schema can't carry:

python
from pydantic import BaseModel, field_validator

class Order(BaseModel):
    amount_cents: int
    discount_pct: float | None  # nullable: may be absent
    order_date: str             # we'll normalize to ISO

    @field_validator("discount_pct")
    @classmethod
    def pct_in_range(cls, v):
        if v is not None and not (0 <= v <= 100):
            raise ValueError("discount_pct must be between 0 and 100")
        return v

    @field_validator("amount_cents", mode="before")
    @classmethod
    def strip_currency(cls, v):
        if isinstance(v, str):  # "$1,200.00" -> 120000
            return int(round(float(v.replace("$", "").replace(",", "")) * 100))
        return v

The Zod equivalent uses .refine() for cross-field and range checks and .transform() for normalization. Three normalizations bite constantly: currency ("$1,200" vs 1200 vs 120000 cents — pick one and enforce it), dates (free text → ISO 8601, and decide what a missing year means), and enum snapping (the model rounds "kinda urgent" to urgent; validate that the snap was legitimate, or widen the enum).

Null vs. absent. A nullable field with a clear instruction is the entire fix for "the model invents an email." "email": {"type": ["string","null"], "description": "...or null if the text states no email"} plus a one-line system instruction ("use null for any field not present in the source; never guess"). If you make the field non-nullable, you have told the model to produce a value — it will.

Bounded reask with Instructor — failed validation is fed back to the model as an error message:

python
import instructor

client = instructor.from_provider("openai/gpt-5.1")
order = client.chat.completions.create(
    response_model=Order,            # your Pydantic model, validators and all
    max_retries=2,                   # BOUND it; each retry is another paid call
    messages=[{"role": "user", "content": text}],
)

On a validation failure Instructor reasks with the ValueError text, so @field_validator rules the model never saw in the schema still get enforced through the loop. Log every reask (count + reason): a quietly climbing reask rate is your early signal that a field's instruction or schema is wrong.

Show full SKILL.md (526 more words)Show less

Multi-provider: Instructor from_provider

When you want one Pydantic model to run across OpenAI, Anthropic, and local backends without rewriting per SDK, use Instructor's unified entrypoint:

python
client = instructor.from_provider("anthropic/claude-opus-4-8")  # or "openai/gpt-5.1", "ollama/llama3.3"

Reach for Instructor when you need provider portability or value-level validation with reask. Reach for the native SDK helper (client.messages.parse, zodOutputFormat) when you're on one provider and want the simplest path with the fewest dependencies. Both sit on the same native decoding underneath.

Scale and accuracy live next door

This skill is the single extraction node and its per-call validation loop. Two concerns are explicitly not here:

  • Running 10k documents — batching, idempotency, retries-across-calls, cost/latency budgeting, multi-step chains → llm-pipeline.
  • Measuring extraction quality offline — a golden set, precision/recall, a regression gate that fails CI when accuracy drops → agent-eval. This skill builds the extractor; that one scores it.
  • Answering questions over a corpus by retrieving chunks → rag.
  • Post-extraction tabular cleanup (dedupe rows, coerce columns, normalize categories across a whole dataset) → data-cleaning.
  • Improving a prompt that isn't bound to a schema → prompt-engineering.

(Some routed siblings may not be built in this collection yet; the routing decision still holds.)

Anti-patterns

BadWhy it bitesGood
json.loads(resp.text) on raw model outputmarkdown fence, chatty preamble, or a missing key crashes at runtimenative structured outputs; parse only a decoder-guaranteed string
Stripping ```json fences with a regextreats the symptom; the model can still drop a key or change a typeturn on native decoding — the fence never appears
"type": "string" on a field that's often absentforces a value, so the model fabricates a plausible wrong onenullable union ["string","null"] + "null if not stated"
Omitting an optional field from required (strict mode)OpenAI/Anthropic strict reject it — every property must be in requiredkeep it in required, make its type a union with null
minimum/maxLength/lookahead-regex inside a strict schemarejected or silently ignored — the constraint does nothingleave value rules out of the schema; enforce in a Pydantic/Zod validator
max_retries unbounded (or a while reask loop)one bad doc becomes an unbounded bill and a hung jobcap at 2–3, log each reask, route the rest to review
Deep/recursive schema in one callnative modes reject recursion and cap complexity → compile failureflatten, or split into multiple extractions
Trusting decoding to make values correctvalid+typed ≠ true; you ship well-formed wrong dataadd the rung-2 validation step for every value rule
Building on Anthropic output_format / structured-outputs-2025-11-13 headerdeprecated transition-window APIuse output_config={"format": {...}}
One giant array of scalars for "many things"loses per-item field structure and per-item validationone object schema per unit, wrapped in a top-level items array

Checklist before you ship an extractor

  1. Native structured outputs ON (or a documented reason the provider has no native mode).
  2. Every object has additionalProperties: false; every property is in required.
  3. Every maybe-absent field is nullable with a "use null if not stated" instruction.
  4. Closed sets are enums; value rules (ranges/format/cross-field) live in validators, not the schema.
  5. A reask loop, if any, is bounded and logged.
  6. You can state, for each field, whether a wrong output would be caught by decoding (shape) or only by validation (value) — and you have the validation for the value cases.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/structured-extraction of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/providers.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Structured Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Structured Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Structured Extraction this skillericrisco/rsc-harness156—~3.6kAutomated safety check: PassMIT
Lexoid Pythonoidlabs-com/Lexoid109—~3.5kAutomated safety check: NotesApache-2.0
DOCX ToolkitXiaomiMiMo/MiMo-Code14k—~2.4kAutomated safety check: PassApache-2.0
Office File Processxstongxue/best-skills2.9k—~1.8kAutomated safety check: PassProprietary
MineruNebutra/MinerU-Skill122—~1.4kAutomated safety check: PassMIT
Md2docxYorick-Ryu/deep-share144—~1.9kAutomated safety check: PassCustom licence

Similar skills

  • Lexoid Python

    oidlabs-com/Lexoid

    Parse and convert documents (PDFs, images, web pages, DOCX/XLSX/PPTX, audio) inside a Python program using the lexoid library.

    109 GitHub stars~3.5k tokensUpdated today
    Documents & OfficeAuto-check: notes
  • DOCX Toolkit

    XiaomiMiMo/MiMo-Code

    Produces, edits and reads Microsoft Word files through python-docx and lxml, with a decision table for picking the lightest workflow for a given task.

    14k GitHub stars~2.4k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Office File Process

    xstongxue/best-skills

    处理 Office 文档的一站式 skill:Word(.doc/.docx/.dotx)、Excel(.xls/.xlsx/.xlsm/.csv)、PowerPoint(.ppt/.pptx/.potx) 的创建、读取、编辑、提取、转换、校验。触发:『读取 word 文档』『提取 excel 内容』『看 ppt 讲了什么』、.doc 老格式打不开、生成/编辑 Word…

    2.9k GitHub stars~1.8k tokensUpdated 24 days ago
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~1.4k tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • Md2docx

    Yorick-Ryu/deep-share

    Convert Markdown to Word (DOCX) documents. An agent skill from Yorick-Ryu/deep-share.

    144 GitHub stars~1.9k tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Markdown Converter

    Team-Commonly/commonly

    Convert binary documents (PDF, DOCX, XLSX, PPTX, HTML, EPUB, images) to clean LLM-friendly Markdown using Microsoft's markitdown Python tool.

    1.4k GitHub stars~557 tokensUpdated today
    Documents & OfficeAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Questions about Structured Extraction

What does Structured Extraction do?

A skill your agent uses when text must become a typed, schema-conformant object you can trust — pulling fields into a fixed JSON shape, extracting line items as typed records, classifying into…. Structured Extraction is an agent skill from ericrisco/rsc-harness. Use when text must become a typed, schema-conformant object you can trust — pulling fields into a fixed JSON shape, extracting line items as typed records, classifying into enums, and building the Pydantic or Zod model plus the validate-and-retry loop.

When should I use Structured Extraction?

Structured Extraction fits situations like: text must become a typed; schema-conformant object you can trust — pulling fields into a fixed JSON shape; extracting line items as typed records; classifying into enums.

How do I install Structured Extraction in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill structured-extraction -a claude-code`. Or copy the skill folder (skills/structured-extraction in ericrisco/rsc-harness) into .claude/skills/structured-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Structured Extraction in Codex?

Run `npx skills add ericrisco/rsc-harness --skill structured-extraction -a codex`. Or copy the skill folder (skills/structured-extraction in ericrisco/rsc-harness) into .agents/skills/structured-extraction in your project. Codex loads it when a task matches its description.

Can I use Structured Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill structured-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/structured-extraction, .gemini/skills/structured-extraction, .github/skills/structured-extraction and .opencode/skills/structured-extraction in your project.

What does Structured Extraction need to run?

Going by SKILL.md and its folder, Structured Extraction needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does Structured Extraction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Structured Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Structured Extraction use?

Structured Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Structured Extraction use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Structured Extraction?

Skills that share tags, products or a category with Structured Extraction: Lexoid Python (oidlabs-com/Lexoid, 109 stars), DOCX Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), Office File Process (xstongxue/best-skills, 2.9k stars) and Mineru (Nebutra/MinerU-Skill, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Structured Extraction?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.