Agent skill

Offloading Extraction

by hashgraph-online in hashgraph-online/awesome-codex-plugins

A skill your agent uses when the user wants to extract a document via the cloud rather than the local kreuzberg CLI.

Apache-2.0Auto-check passedBackend & APIs

Install Offloading Extraction

skills CLI
$ npx skills add hashgraph-online/awesome-codex-plugins --skill offloading-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hashgraph-online/awesome-codex-plugins offloading-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/kreuzberg-dev/plugins/plugins/kreuzberg-cloud/skills/offloading-extraction .claude/skills/offloading-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
offloading-extraction
GitHub stars
1.2k
Token cost
~1.4k tokens
SKILL.md length
340 words
Files
1
Skills in repo
686
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user wants to extract a document via the cloud rather than the local kreuzberg CLI.

  • Works in 3 steps: Base64 JSON (small files, <5 MB… → Multipart (binary, recommended for… → URL crawl
  • The user wants to extract a document via the cloud rather than the local kreuzberg CLI
  • SKILL.md covers When to reach for this, Endpoint, Three submission shapes and Response (202), plus 7 more sections
  • Calls curl; reaches api.kreuzberg.dev; needs KREUZBERG_API_KEY

What it does

Offloading Extraction is an agent skill from hashgraph-online/awesome-codex-plugins. Use when the user wants to extract a document via the cloud rather than the local kreuzberg CLI. Covers POST /v1/extract — JSON vs multipart bodies, URL crawls, options block, webhook attachment, and the async response shape.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Backend & APIs, covering Webhooks. The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is Apache-2.0.

When your agent uses it

  • The user wants to extract a document via the cloud rather than the local kreuzberg CLI
  • Tasks that involve Webhooks

Example prompts

  • “/offloading-extraction”

Requirements

  • Python 3
  • A credential in KREUZBERG_API_KEY

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Base64 JSON (small files, <5 MB recommended)
  2. Multipart (binary, recommended for anything over ~1 MB)
  3. URL crawl

What it can do on your machine

Read from SKILL.md and the folder at commit 78497e5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.kreuzberg.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • KREUZBERG_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Offloading Extraction loads about 1.4k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 340 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hashgraph-online/awesome-codex-plugins at commit 78497e5, republished under its Apache-2.0 licence (© hashgraph-online). 340 words, ~1,384 tokens.

Download SKILL.mdSave it as .claude/skills/offloading-extraction/SKILL.md (or your agent's skills folder).
name
offloading-extraction
description
Use when the user wants to extract a document via the cloud rather than the local kreuzberg CLI. Covers POST /v1/extract — JSON vs multipart bodies, URL crawls, options block, webhook attachment, and the async response shape.

Offloading extraction

POST /v1/extract is the single submit endpoint. It returns 202 Accepted with job_ids (extraction) and crawl_job_ids (URL crawls) — never the extraction result inline. Pair every submit with either a poll loop (tracking-cloud-jobs skill) or a webhook.

When to reach for this

  • File is on a remote URL.
  • File is on disk but the local kreuzberg CLI is not installed.
  • You want server-side parallelism for a batch.
  • The user wants webhook-delivered results to skip blocking.
  • File is larger than ~50 MB → use presigned-uploads instead — the base64 JSON body is too big.

Endpoint

text
POST https://api.kreuzberg.dev/v1/extract
Authorization: Bearer $KREUZBERG_API_KEY
Content-Type: application/json | multipart/form-data

Returns 202 Accepted with ExtractResponse.

Three submission shapes

bash
curl -X POST https://api.kreuzberg.dev/v1/extract \
  -H "Authorization: Bearer $KREUZBERG_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<JSON
{
  "documents": [
    {
      "filename": "invoice.pdf",
      "mime_type": "application/pdf",
      "data": "$(base64 -w0 invoice.pdf)"
    }
  ],
  "options": {
    "extraction_config": {
      "output_format": "markdown",
      "ocr": { "backend": "tesseract", "language": "eng" }
    }
  }
}
JSON
bash
curl -X POST https://api.kreuzberg.dev/v1/extract \
  -H "Authorization: Bearer $KREUZBERG_API_KEY" \
  -F "file=@invoice.pdf;type=application/pdf" \
  -F 'options={"extraction_config":{"output_format":"markdown"}};type=application/json'

Add a webhook part as a JSON string:

bash
  -F 'webhook={"url":"https://hooks.example.com/x","secret":"shh"};type=application/json'
3. URL crawl
bash
curl -X POST https://api.kreuzberg.dev/v1/extract \
  -H "Authorization: Bearer $KREUZBERG_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "urls": [{"url": "https://example.com/docs"}],
    "crawl_config": {"max_depth": 2, "max_pages": 50, "stay_on_domain": true},
    "webhook": {"url": "https://hooks.example.com/x"}
  }'

URL crawls return crawl_job_ids instead of (or alongside) job_ids.

Response (202)

json
{
  "job_ids": ["550e8400-e29b-41d4-a716-446655440000"],
  "crawl_job_ids": [],
  "status": "pending"
}

status is always pending at submit time; the per-job status is retrieved via GET /v1/jobs/{id}.

The options block

Shape mirrors the local ExtractionConfig:

json
{
  "extraction_config": {
    "output_format": "markdown",
    "ocr": { "backend": "tesseract", "language": "eng+deu" },
    "extract_tables": true,
    "extract_images": false,
    "chunking": { "max_chars": 4000, "overlap": 200 }
  }
}

Supported output_format values: markdown, text, json, djot, html. Default is markdown.

The webhook block

json
{
  "url": "https://hooks.example.com/x",
  "secret": "shared-secret-32-bytes-min",
  "metadata": { "request_id": "abc123", "user_id": "u_42" }
}

secret is the HMAC key used to sign the webhook payload — see tracking-cloud-jobs for verification. metadata is echoed back in the delivered payload, useful for correlating server-side requests.

TypeScript SDK

ts
import { KreuzbergCloud } from "@kreuzberg/cloud";
import { readFile } from "node:fs/promises";

const client = new KreuzbergCloud({ apiKey: process.env.KREUZBERG_API_KEY! });

const data = await readFile("invoice.pdf");
const job = await client.extract({
  file: { name: "invoice.pdf", data, mimeType: "application/pdf" },
  options: { extractionConfig: { outputFormat: "markdown" } },
});
console.log(job.id, job.status);

For submit + wait in one call:

ts
const result = await client.extractAndWait({
  file: { name: "invoice.pdf", data },
});
console.log(result.result?.content);

Python SDK

python
from pathlib import Path
from kreuzberg_cloud import KreuzbergCloud

with KreuzbergCloud(api_key=os.environ["KREUZBERG_API_KEY"]) as client:
    job = client.extract(file=Path("invoice.pdf"))
    print(job.id, job.status)

Submit + wait:

python
job = client.extract_and_wait(file=Path("invoice.pdf"))
print(job.result.content if job.result else job.status)

Batch submission

JSON: pass multiple entries in documents. Multipart: repeat the file part. SDKs expose extractBatch / extract_batch helpers that fan out correctly per platform (parallel HTTP for the async Python client, sequential for the sync one).

Errors

StatusCauseFix
400Empty documents and urlsProvide at least one.
400Bad MIME typeUse a real RFC 6838 type, e.g. application/pdf.
401Missing BearerSet Authorization header.
413Request body too largeSwitch to presigned uploads.
429Quota or rate limitBackoff; check quota_remaining via /v1/usage.

Next step

After every submit, hand off to the tracking-cloud-jobs skill — cloud extraction is asynchronous and the result is delivered via either polling or webhook callback. Never assume a result is ready immediately after the 202 response.

© hashgraph-online, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/kreuzberg-dev/plugins/plugins/kreuzberg-cloud/skills/offloading-extraction of hashgraph-online/awesome-codex-plugins.

Open the folder on GitHubat commit 78497e5

Compare with similar skills

Offloading Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Offloading Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Offloading Extraction this skillhashgraph-online/awesome-codex-plugins1.2k—~1.4kAutomated safety check: PassApache-2.0
Novu Design Workflownovuhq/novu40k—~2.6kAutomated safety check: PassCustom licence
Stripe Appsfossasia/eventyay1.7k2 repos~3.6kAutomated safety check: PassApache-2.0
Golivemikehasa/golive-skill1.2k—~13kAutomated safety check: NotesMIT
Dingtalk Messageagentscope-ai/ReMe3.6k—~1.6kAutomated safety check: PassApache-2.0
PR Review Provideryansongda/pay5.4k—~2.4kAutomated safety check: PassMIT

Similar skills

  • Design notification workflows the Novu way — choose channels, set severity, decide when a workflow is critical, configure digests, and route based on subscriber state.

    40k GitHub stars~2.6k tokensUpdated today
    Backend & APIsAuto-check passed
  • Stripe Apps

    fossasia/eventyay

    A skill your agent uses when building, modifying, or reviewing a Stripe App — or when the user describes something that implies one (e.g.

    1.7k GitHub starsUsed in 2 repos~3.6k tokens
    Backend & APIsAuto-check passed
  • Golive

    mikehasa/golive-skill

    Take an agent-written app from repo to live production on the user's OWN accounts, with providers they choose (hosting, database, auth, payments, email, domain/DNS).

    1.2k GitHub stars~13k tokensUpdated 4 days ago
    Backend & APIsAuto-check: notes
  • Dingtalk Message

    agentscope-ai/ReMe

    钉钉消息发送技能。支持企业内部机器人(批量单聊/群聊)和 Webhook 自定义机器人两种接入方式,支持多机器人管理,支持文本、Markdown、链接、ActionCard、FeedCard等多种消息类型。

    3.6k GitHub stars~1.6k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • PR Review Provider

    yansongda/pay

    A skill your agent uses when reviewing PRs that add or modify a payment Provider in yansongda/pay - covers plugin pipeline, multi-tenant safety, signature verification, docs, and naming conventions.

    5.4k GitHub stars~2.4k tokensUpdated 9 days ago
    Backend & APIsAuto-check passed
  • Stripe Best Practices

    kanchengw/cnllm

    Guides Stripe integration decisions — API selection (Checkout Sessions vs PaymentIntents), Connect platform setup (Accounts v2, controller properties), billing/subscriptions, Treasury financial…

    175 GitHub starsUsed in 3 repos~925 tokens
    Backend & APIsAuto-check passed

More from hashgraph-online/awesome-codex-plugins

All 686 skills in this repo
  • Anime Reaction Gif

    hashgraph-online/awesome-codex-plugins

    Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.

    1.2k GitHub stars~922 tokensUpdated today
    Auto-check passed
  • Calibredb

    hashgraph-online/awesome-codex-plugins

    Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).

    1.2k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Rust API Test Harness

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…

    1.2k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Art

    hashgraph-online/awesome-codex-plugins

    Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…

    1.2k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Game Balance Economy

    hashgraph-online/awesome-codex-plugins

    Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.

    1.2k GitHub stars~618 tokensUpdated today
    Auto-check passed
  • Manuscript Engagement Analytics

    hashgraph-online/awesome-codex-plugins

    Analyze nonfiction manuscripts for reader engagement signals, including heading-level word counts, slow starts, long slogs, weak takeaway titles, value pacing, beta-reader comment dropoff, and…

    1.2k GitHub stars~875 tokensUpdated today
    Auto-check passed

Categories

Questions about Offloading Extraction

What does Offloading Extraction do?

A skill your agent uses when the user wants to extract a document via the cloud rather than the local kreuzberg CLI. Offloading Extraction is an agent skill from hashgraph-online/awesome-codex-plugins. Use when the user wants to extract a document via the cloud rather than the local kreuzberg CLI.

When should I use Offloading Extraction?

Offloading Extraction fits situations like: the user wants to extract a document via the cloud rather than the local kreuzberg CLI; tasks that involve Webhooks.

How do I install Offloading Extraction in Claude Code?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill offloading-extraction -a claude-code`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg-cloud/skills/offloading-extraction in hashgraph-online/awesome-codex-plugins) into .claude/skills/offloading-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Offloading Extraction in Codex?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill offloading-extraction -a codex`. Or copy the skill folder (plugins/kreuzberg-dev/plugins/plugins/kreuzberg-cloud/skills/offloading-extraction in hashgraph-online/awesome-codex-plugins) into .agents/skills/offloading-extraction in your project. Codex loads it when a task matches its description.

Can I use Offloading Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill offloading-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/offloading-extraction, .gemini/skills/offloading-extraction, .github/skills/offloading-extraction and .opencode/skills/offloading-extraction in your project.

What does Offloading Extraction need to run?

Going by SKILL.md and its folder, Offloading Extraction needs the command-line tools its instructions call (curl) and credentials named KREUZBERG_API_KEY. Our summary lists: Python 3; A credential in KREUZBERG_API_KEY.

Does Offloading Extraction access the network?

SKILL.md names 1 domain. In commands or code: api.kreuzberg.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Offloading Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Offloading Extraction use?

Offloading Extraction is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Offloading Extraction use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Offloading Extraction?

Skills that share tags, products or a category with Offloading Extraction: Novu Design Workflow (novuhq/novu, 40k stars), Stripe Apps (fossasia/eventyay, 1.7k stars), Golive (mikehasa/golive-skill, 1.2k stars) and Dingtalk Message (agentscope-ai/ReMe, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Offloading Extraction?

hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,242 GitHub stars. The repository holds 686 skills in this directory. The repository was last updated on October 8, 2026.

Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.