Agent skill

Document Extract

by amd in amd/gaia

Enumerate every requested item in long text documents or transcripts, with fields and source evidence.

MITAuto-check passed

Install Document Extract

skills CLI
$ npx skills add amd/gaia --skill document-extract -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/gaia document-extract --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/hub/agents/gaia/python/gaia_agent/skills/document-extract .claude/skills/document-extract && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
document-extract
GitHub stars
1.6k
Token cost
~524 tokens
SKILL.md length
245 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Enumerate every requested item in long text documents or transcripts, with fields and source evidence.

  • Works in 5 steps: Identify each source text file and the… → Call extract_document_items once per… → If an extraction fails, report it as… → …
  • Exhaustive inventories such as every exercise
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Not ordinary summaries

What it does

Document Extract is an agent skill from amd/gaia. Enumerate every requested item in long text documents or transcripts, with fields and source evidence. Use for exhaustive inventories such as every exercise, action item, or finding; not ordinary summaries, general knowledge lists, or code symbols.

Its SKILL.md is about 520 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Build AI agents for your PC. The licence is MIT.

When your agent uses it

  • Exhaustive inventories such as every exercise
  • Not ordinary summaries
  • General knowledge lists

Example prompts

  • “/document-extract”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Identify each source text file and the requested item type and fields. Use
  2. Call extract_document_items once per source. It processes bounded chunks
  3. If an extraction fails, report it as incomplete. Do not silently switch to
  4. If the user requested a saved inventory, use save_extracted_items for the
  5. The framework appends the full retained inventory to your final answer.

What it can do on your machine

Read from SKILL.md and the folder at commit 6c3bb5c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Document Extract loads about 524 tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 245 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~524

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/gaia at commit 6c3bb5c, republished under its MIT licence (© amd). 245 words, ~524 tokens.

Download SKILL.mdSave it as .claude/skills/document-extract/SKILL.md (or your agent's skills folder).
name
document-extract
description
Enumerate every requested item in long text documents or transcripts, with fields and source evidence. Use for exhaustive inventories such as every exercise, action item, or finding; not ordinary summaries, general knowledge lists, or code symbols.
version
0.1.0
license
MIT

Exhaustive document extraction

  1. Identify each source text file and the requested item type and fields. Use the user's current files; remembered summaries can help locate them but do not establish current contents or completeness.
  2. Call extract_document_items once per source. It processes bounded chunks sequentially and retains source-backed entries. Do not replace this with a summary or manually calculated page offsets. When specifying fields in the tool call, pass fields: ["name", "owner", "deadline"] (the actual requested names).
  3. If an extraction fails, report it as incomplete. Do not silently switch to a shorter list, a guessed count, or a script that has not been run.
  4. If the user requested a saved inventory, use save_extracted_items for the exact destination. The exporter preserves all entries and provenance as JSON, CSV, or readable text, and the framework reads the file back and checks it; don't edit or regenerate it by hand.
  5. The framework appends the full retained inventory to your final answer. Briefly state limitations or requested analysis; do not repeat or re-summarize the list.

The existing memory system records the action, source fingerprint, coverage, count and final inventory unless the session is private. Use conversation search for follow-ups about previous work; re-extract if the source or requested fields change. Never treat a remembered summary as proof that all current items were extracted, and don't store a whole private transcript as a durable personal fact. A compact summary may supplement the inventory, not replace it.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in hub/agents/gaia/python/gaia_agent/skills/document-extract of amd/gaia.

Open the folder on GitHubat commit 6c3bb5c

Compare with similar skills

Document Extract next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Document Extract compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Document Extract this skillamd/gaia1.6k—~524Automated safety check: PassMIT
Create Pull Request with Work Item IDmakeplane/plane60k—~824Automated safety check: PassAGPL-3.0
Extractalirezarezvani/claude-skills28k—~1.4kAutomated safety check: PassMIT
Pull Requestspnpm/pnpm37k—~2.2kAutomated safety check: PassMIT
Create Pull Requestcline/cline70k1 repos~1.6kAutomated safety check: PassApache-2.0
Reviewing Pull RequestsTriliumNext/Trilium38k—~4.9kAutomated safety check: PassAGPL-3.0

Similar skills

  • Opens a pull request for the current branch using the repo's template, a work item ID in the title and a description filled in from the actual diff.

    60k GitHub stars~824 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Extract

    alirezarezvani/claude-skills

    Turn a proven pattern or debugging solution into a standalone reusable skill with SKILL.md, reference docs, and examples.

    28k GitHub stars~1.4k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Pull Requests

    pnpm/pnpm

    Take a change through a pull request in the pnpm repository — opening it, then staying with it after every push until CI is green and the review round is quiet.

    37k GitHub stars~2.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Opens a GitHub pull request from your current branch with the gh CLI, after reviewing the commits and diff and gathering the details the PR needs.

    70k GitHub starsUsed in 1 repo~1.6k tokens
    DevelopmentAuto-check passed
  • Reviewing Pull Requests

    TriliumNext/Trilium

    A skill your agent uses whenever a Trilium pull request is to be judged — "review PR N", "is

    38k GitHub stars~4.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Brand Extract

    nexu-io/open-design

    Extract a complete Brand Kit from a live website by driving the in-app browser.

    100k GitHub stars~3.1k tokensUpdated today
    Productivity & AutomationAuto-check passed

More from amd/gaia

All 44 skills in this repo
  • Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

    1.6k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.

    1.6k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.

    1.6k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.

    1.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Questions about Document Extract

What does Document Extract do?

Enumerate every requested item in long text documents or transcripts, with fields and source evidence. Document Extract is an agent skill from amd/gaia. Enumerate every requested item in long text documents or transcripts, with fields and source evidence.

When should I use Document Extract?

Document Extract fits situations like: exhaustive inventories such as every exercise; not ordinary summaries; general knowledge lists.

How do I install Document Extract in Claude Code?

Run `npx skills add amd/gaia --skill document-extract -a claude-code`. Or copy the skill folder (hub/agents/gaia/python/gaia_agent/skills/document-extract in amd/gaia) into .claude/skills/document-extract in your project. Claude Code loads it when a task matches its description.

How do I install Document Extract in Codex?

Run `npx skills add amd/gaia --skill document-extract -a codex`. Or copy the skill folder (hub/agents/gaia/python/gaia_agent/skills/document-extract in amd/gaia) into .agents/skills/document-extract in your project. Codex loads it when a task matches its description.

Can I use Document Extract in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill document-extract -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/document-extract, .gemini/skills/document-extract, .github/skills/document-extract and .opencode/skills/document-extract in your project.

What does Document Extract need to run?

SKILL.md names no scripts, command-line tools or credentials: Document Extract is instructions for the agent only.

Does Document Extract access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Document Extract safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Document Extract use?

Document Extract is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Document Extract use?

About 524 tokens (SKILL.md is roughly 2.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Document Extract?

Skills that share tags, products or a category with Document Extract: Create Pull Request with Work Item ID (makeplane/plane, 60k stars), Extract (alirezarezvani/claude-skills, 28k stars), Pull Requests (pnpm/pnpm, 37k stars) and Create Pull Request (cline/cline, 70k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Document Extract?

amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.