Agent skill

Structured Extraction

by bionic-gpt in bionic-gpt/bionic-gpt

Extract structured, source-located evidence from uploaded documents with Xberg for document analysis and validation.

Apache-2.0Auto-check passedDocuments & Office

Install Structured Extraction

skills CLI
$ npx skills add bionic-gpt/bionic-gpt --skill structured-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bionic-gpt/bionic-gpt structured-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bionic-gpt/bionic-gpt.git skills-src && mkdir -p .claude/skills && cp -r skills-src/crates/builtin-skills/skills/structured-extraction .claude/skills/structured-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
structured-extraction
GitHub stars
2.4k
Token cost
~579 tokens
SKILL.md length
241 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
Apache-2.0

At a glance

Extract structured, source-located evidence from uploaded documents with Xberg for document analysis and validation.

  • Works in 6 steps: Read the ## Current attachments section… → For each source, call… → In that Python call, serialize the… → …
  • Tasks that involve Document parsing
  • SKILL.md covers Workflow and Quality checks
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Structured Extraction is an agent skill from bionic-gpt/bionic-gpt. Extract structured, source-located evidence from uploaded documents with Xberg for document analysis and validation.

Its SKILL.md is about 580 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Document parsing. The repository describes itself as: Bionic is sovereign Agentic AI for the enterprise — Runs on-premise and can securely work with your sensitive data and systems. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Document parsing

Example prompts

  • “/structured-extraction”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read the ## Current attachments section in the system prompt. It gives each original filename plus its preprocessed Markdown and…
  2. For each source, call document_conversion_api_extractdocument directly with run_python only when the preprocessed Markdown is missing or a…
  3. In that Python call, serialize the complete result to /home/user/work/.json instead of printing it. For example
  4. Prefer reading or searching the preprocessed /home/user/attachments//content.md with read_file or run_bash commands such as grep; it…
  5. Preserve the source filename and every available location marker when reporting evidence. Keep extraction separate from interpretation.
  6. If conversion fails or a format is unsupported, report the exact failure and do not substitute guessed content.

What it can do on your machine

Read from SKILL.md and the folder at commit c155c00. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Structured Extraction loads about 579 tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 241 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~579

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from bionic-gpt/bionic-gpt at commit c155c00, republished under its Apache-2.0 licence (© bionic-gpt). 241 words, ~579 tokens.

Download SKILL.mdSave it as .claude/skills/structured-extraction/SKILL.md (or your agent's skills folder).
name
structured-extraction
description
Extract structured, source-located evidence from uploaded documents with Xberg for document analysis and validation.

Structured Extraction

Use this skill when a task depends on facts contained in uploaded documents or other files and those facts must be extracted before analysis.

Workflow

  1. Read the ## Current attachments section in the system prompt. It gives each original filename plus its preprocessed Markdown and original-file paths.
  2. For each source, call document_conversion_api_extractdocument directly with run_python only when the preprocessed Markdown is missing or a fresh extraction is explicitly needed. Pass the original path as file_path.
  3. In that Python call, serialize the complete result to /home/user/work/<source-name>.json instead of printing it. For example:
python
import json

result = document_conversion_api_extractdocument(
    file_path="/home/user/attachments/<attachment-id>/original/<source-file>"
)
with open("/home/user/work/<source-file>.json", "w") as extracted:
    extracted.write(json.dumps(result, ensure_ascii=False, indent=2))
print("Saved extraction to /home/user/work/<source-file>.json")
  1. Prefer reading or searching the preprocessed /home/user/attachments/<attachment-id>/content.md with read_file or run_bash commands such as grep; it remains available across later turns. Work files also persist without appearing as generated files in the conversation.
  2. Preserve the source filename and every available location marker when reporting evidence. Keep extraction separate from interpretation.
  3. If conversion fails or a format is unsupported, report the exact failure and do not substitute guessed content.

Xberg supports many document families, including PDF, Word and other office documents, spreadsheets, presentations, ebooks, email, archives, HTML/XML, structured text, images with OCR, and audio or video transcription. Let the conversion response determine whether a particular file is supported.

Quality checks

  • Every extracted claim has a source filename and location when available.
  • Tables retain their headers and row or column meaning.
  • Empty, unreadable, or missing content is marked as unavailable.
  • Extraction and interpretation remain separate.

© bionic-gpt, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in crates/builtin-skills/skills/structured-extraction of bionic-gpt/bionic-gpt.

Open the folder on GitHubat commit c155c00

Compare with similar skills

Structured Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Structured Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Structured Extraction this skillbionic-gpt/bionic-gpt2.4k—~579Automated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
DOCX ToolkitXiaomiMiMo/MiMo-Code14k—~2.4kAutomated safety check: PassApache-2.0
Huashu Markdown Publishing Pipelinealchaincyf/huashu-md-html908—~4.8kAutomated safety check: PassMIT
Markitdownjimmc414/Kosmos5942 repos~1.7kAutomated safety check: PassNone
Liteparsebastani-inc/atomic846—~1.4kAutomated safety check: PassMIT

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • DOCX Toolkit

    XiaomiMiMo/MiMo-Code

    Produces, edits and reads Microsoft Word files through python-docx and lxml, with a decision table for picking the lightest workflow for a given task.

    14k GitHub stars~2.4k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Huashu Markdown Publishing Pipeline

    alchaincyf/huashu-md-html

    Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.

    908 GitHub stars~4.8k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Liteparse

    bastani-inc/atomic

    A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…

    846 GitHub stars~1.4k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Beautiful Article

    ConardLi/garden-skills

    Turns a URL, PDF, DOCX, Markdown file, text or screenshots into a designed, shareable single-file HTML article through a staged review workflow.

    13k GitHub stars~4.7k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed

More from bionic-gpt/bionic-gpt

All 21 skills in this repo
  • Static Sites

    bionic-gpt/bionic-gpt

    Create or modify Bionic's generated marketing site, documentation, blog, course pages, and static assets.

    2.4k GitHub stars~920 tokensUpdated today
    Auto-check passed
  • Automationbench Eval

    bionic-gpt/bionic-gpt

    Load the repository's AutomationBench OpenAPI catalogue and run evaluations against the local simulator.

    2.4k GitHub stars~619 tokensUpdated today
    Auto-check passed
  • Ron Auth

    bionic-gpt/bionic-gpt

    Implement identity and authorization boundaries in Rust on Nails applications.

    2.4k GitHub stars~667 tokensUpdated today
    Auto-check passed
  • Ron Web Pages

    bionic-gpt/bionic-gpt

    Build server-rendered Dioxus pages using opinionated Rust on Nails patterns.

    2.4k GitHub stars~616 tokensUpdated today
    Auto-check passed
  • Document Generation

    bionic-gpt/bionic-gpt

    Create polished printable documents and PDFs, including forms, checklists, reports, briefs, comparisons, worksheets, task lists, and other operational documents.

    2.4k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Local Deployment

    bionic-gpt/bionic-gpt

    Deploy one or more locally built Bionic services into the user's k3d cluster.

    2.4k GitHub stars~540 tokensUpdated today
    Auto-check: notes

Questions about Structured Extraction

What does Structured Extraction do?

Extract structured, source-located evidence from uploaded documents with Xberg for document analysis and validation. Structured Extraction is an agent skill from bionic-gpt/bionic-gpt. Extract structured, source-located evidence from uploaded documents with Xberg for document analysis and validation.

When should I use Structured Extraction?

Structured Extraction fits situations like: tasks that involve Document parsing.

How do I install Structured Extraction in Claude Code?

Run `npx skills add bionic-gpt/bionic-gpt --skill structured-extraction -a claude-code`. Or copy the skill folder (crates/builtin-skills/skills/structured-extraction in bionic-gpt/bionic-gpt) into .claude/skills/structured-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Structured Extraction in Codex?

Run `npx skills add bionic-gpt/bionic-gpt --skill structured-extraction -a codex`. Or copy the skill folder (crates/builtin-skills/skills/structured-extraction in bionic-gpt/bionic-gpt) into .agents/skills/structured-extraction in your project. Codex loads it when a task matches its description.

Can I use Structured Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bionic-gpt/bionic-gpt --skill structured-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/structured-extraction, .gemini/skills/structured-extraction, .github/skills/structured-extraction and .opencode/skills/structured-extraction in your project.

What does Structured Extraction need to run?

SKILL.md names no scripts, command-line tools or credentials: Structured Extraction is instructions for the agent only. Our summary lists: Python 3.

Does Structured Extraction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Structured Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Structured Extraction use?

Structured Extraction is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Structured Extraction use?

About 579 tokens (SKILL.md is roughly 2.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Structured Extraction?

Skills that share tags, products or a category with Structured Extraction: Markitdown (ImCa0/just-laws, 781 stars), DOCX Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), Huashu Markdown Publishing Pipeline (alchaincyf/huashu-md-html, 908 stars) and Markitdown (jimmc414/Kosmos, 594 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Structured Extraction?

bionic-gpt (a GitHub organization) maintains it in bionic-gpt/bionic-gpt, which has 2,383 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 7, 2026.

Source: bionic-gpt/bionic-gpt on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.