Agent skill

Resilient Context Extraction

by HKUDS in HKUDS/OpenSpace

Ensures agents extract data from context files with validation and fallback strategies before resorting to assumptions or external searches.

MITAuto-check passedDocuments & Office

Install Resilient Context Extraction

skills CLI
$ npx skills add HKUDS/OpenSpace --skill resilient-context-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace resilient-context-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/prioritize-context-data-enhanced .claude/skills/resilient-context-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
resilient-context-extraction
GitHub stars
7.8k
Token cost
~2.5k tokens
SKILL.md length
1,232 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Ensures agents extract data from context files with validation and fallback strategies before resorting to assumptions or external searches.

  • Works in 6 steps: Scan Context for Attachments → Evaluate Relevance → Extract Data with Validation → …
  • Documents & Office work in your project
  • SKILL.md covers Objective, Critical Rule, Workflow Steps and Checklist, plus 4 more sections
  • Calls pdftotext

What it does

Resilient Context Extraction is an agent skill from HKUDS/OpenSpace. Ensures agents extract data from context files with validation and fallback strategies before resorting to assumptions or external searches.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Documents & Office work in your project

Example prompts

  • “Use the resilient-context-extraction skill to ensure agents extract data from context files with validation and fallback strategies before resorting…”
  • “/resilient-context-extraction”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Scan Context for Attachments
  2. Evaluate Relevance
  3. Extract Data with Validation
  4. Apply Fallback Extraction Strategies
  5. Cite Source Explicitly
  6. Fallback to Search (Last Resort Only)

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pdftotext

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Resilient Context Extraction loads about 2.5k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 1,232 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 1,232 words, ~2,508 tokens.

Download SKILL.mdSave it as .claude/skills/resilient-context-extraction/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
resilient-context-extraction
description
Ensures agents extract data from context files with validation and fallback strategies before resorting to assumptions or external searches.
category
workflow

Resilient Context Extraction

Objective

Prevent data hallucination and inefficiency by mandating that agents inspect, validate, and fully extract data from provided reference files before attempting web searches or generating synthetic data. When extraction is incomplete, use fallback strategies before making assumptions.

Critical Rule

If a reference file is provided in the task context, it is the source of truth. Do not fabricate data or search the web for information that may exist within the provided attachments. If extraction appears incomplete, attempt alternative methods before proceeding.

Workflow Steps

1. Scan Context for Attachments

At the start of every task, explicitly list all files provided in the context window or attachment panel.

  • Check for spreadsheets (.xlsx, .csv), documents (.pdf, .docx, .pptx), or data dumps (.json, .txt, .xml, .yaml).
  • Note the filename, file size (if available), and inferred content type.
  • Record this list for later verification.
2. Evaluate Relevance

Determine if any provided file contains the data required to complete the task.

  • Match Keywords: Do filenames or expected column headers match task requirements?
  • Check Scope: Does the data cover the necessary timeframe or region?
  • Prioritize: Rank files by likelihood of containing required data.
3. Extract Data with Validation

If relevant files are found:

  • Read the file content using appropriate tools (e.g., read_file, pandas, pdf_reader).
  • Validate Extraction Completeness (NEW CRITICAL STEP):
    • Check if output appears truncated (e.g., sudden cutoff mid-sentence, character limits hit).
    • Compare expected data points vs. extracted data points (e.g., "Task mentions pricing tiers; did extraction include pricing numbers?").
    • Look for structural indicators of incompleteness (e.g., unclosed tables, missing document endings).
  • If extraction is incomplete or suspicious, proceed to Step 4 (Fallback Strategies) before using the data.
4. Apply Fallback Extraction Strategies

When initial extraction is incomplete or critical data is missing, attempt these strategies in order before making assumptions:

4a. Alternative Library/Tool Approach
  • For .docx: Try python-docx directly via execute_code_sandbox if read_file was truncated.
  • For .xlsx/.csv: Try pandas with explicit sheet selection or openpyxl directly.
  • For .pdf: Try pdfplumber, PyPDF2, or shell-based pdftotext if available.
  • For any format: Try reading as raw bytes/text via shell commands (cat, xxd, strings).
4b. Shell-Based Extraction
  • Use shell_agent or run_shell for format-specific tools:
    • unzip -p file.docx word/document.xml | xmllint --format - (extract raw docx XML)
    • in2csv file.xlsx (convert Excel to CSV via shell)
    • pdftotext file.pdf - (extract PDF text via command line)
  • This bypasses potential sandbox or library limitations in read_file.
4c. Targeted Re-Reading
  • If specific sections are missing (e.g., "pricing table on page 3"), attempt to re-read with focus on that region.
  • Use tools that support page/section ranges if available.
4d. Document What's Missing
  • Before any assumptions or external searches, explicitly document:
    • What data was expected but not found.
    • What extraction methods were attempted.
    • Why each method failed or what portions remain unavailable.
  • Example: "Attempted read_file on Pricing_email.docx (truncated at 980 chars). Attempted python-docx extraction via sandbox—pricing table in section 2 not present in file. Pricing numbers unavailable from provided context."
5. Cite Source Explicitly

When presenting data in the final output:

  • Explicitly state which file the data came from.
  • Example: "According to Massabama_active_listings.xlsx..."
  • If fallback methods were used, note this: "Extracted via shell-based pdftotext from report.pdf..."
  • This verifies to the user that real data was used, not hallucinated.
6. Fallback to Search (Last Resort Only)

If all extraction strategies fail to provide the specific data needed:

  • State clearly what was missing from the files.
  • Document all extraction methods attempted and why they failed.
  • Then proceed with web search or estimation.
  • Mark any non-file data clearly as "External Search" or "Estimated (file data unavailable)".
  • Never present estimated data as if it came from the file.

Checklist

  • Did I list all attached files with their types?
  • Did I open and read the relevant files?
  • Did I validate extraction completeness (check for truncation, missing sections)?
  • If extraction was incomplete, did I try at least one fallback strategy?
  • Did I explicitly document what data is unavailable before making assumptions?
  • Did I cite the file name (and extraction method if non-standard) in my output?
  • Did I avoid fabricating numbers that should have come from the file?
  • Is any external/estimated data clearly marked as such?
Show full SKILL.md (549 more words)Show less

Example Usage

Task: Create a report on active property listings. Context: Massabama_active_listings.xlsx is attached.

Incorrect Approach:

  • Ignore the Excel file.
  • Search web for "Massabama property listings".
  • Fabricate numbers based on search snippets.
  • Result: Inaccurate data, ignored source of truth.

Correct Approach:

  • Identify Massabama_active_listings.xlsx in context.
  • Load the Excel file via read_file or pandas.
  • Validate: Check row count matches expected scope, verify all columns present.
  • Count rows and sum prices directly from the sheet.
  • Output: "Based on the 50 entries in Massabama_active_listings.xlsx, the total value is..."
  • Result: Accurate data, verified source.

Task: Extract pricing tiers from Pricing_email.docx. Context: Pricing_email.docx attached, but read_file output cuts off mid-sentence.

Incorrect Approach:

  • Accept truncated extraction as complete.
  • Assume pricing tiers based on incomplete info.
  • Proceed with fabricated numbers.

Correct Approach:

  • Identify truncation: read_file output ends at 980 chars, mid-sentence.
  • Attempt fallback: Use execute_code_sandbox with python-docx to read full document.
  • If fallback also incomplete: Document "Pricing table in section 2 not extractable; attempted read_file (truncated) and python-docx (table structure not preserved)."
  • Only then: Either request clarification from user, or mark pricing as "Estimated—file data unavailable" if task requires proceeding.
  • Result: User informed of limitation; no false claim of file-sourced data.

Warnings

  • Hallucination Risk: Ignoring provided files OR accepting incomplete extractions without validation are primary causes of data fabrication.
  • Efficiency: Reading a local file is faster and more reliable than web scraping—but only if extraction is complete.
  • User Expectation: Users attach files expecting them to be used. Ignoring them or failing to extract fully is a failure of instruction following.
  • Truncation Blind Spot: Many tools have character/content limits. Always verify output completeness against expected data scope.

Tool-Specific Notes

ToolKnown LimitationsFallback Strategy
read_fileMay truncate at ~1000-5000 chars depending on formatUse execute_code_sandbox with format-specific library
execute_code_sandboxMay have missing dependencies or sandbox errorsUse shell_agent or run_shell for CLI tools
shell_agentSlower, but more flexible with system toolsUse for pdftotext, unzip XML extraction, in2csv, etc.

Recovery Protocol

If all extraction attempts fail:

  1. Stop and assess: Is this data critical to task completion?
  2. Document all failed attempts with specific error messages or observations.
  3. Communicate to user: "Critical data from [filename] could not be extracted despite [N] attempts using [methods]. Please clarify or provide alternative source."
  4. Proceed cautiously: If task must continue, clearly demarcate any estimated/synthesized data.

*** End Files *** Add File: examples/extraction_fallback.sh #!/bin/bash

Fallback extraction strategies for common file formats

Use when read_file produces incomplete results

extract_docx_raw() { local file="$1" # Extract raw XML from docx (docx is a zip archive) unzip -p "$file" word/document.xml 2>/dev/null | xmllint --format - 2>/dev/null }

extract_xlsx_to_csv() { local file="$1" # Convert Excel to CSV using in2csv (from csvkit) if command -v in2csv &>/dev/null; then in2csv "$file" 2>/dev/null else echo "in2csv not available; try python pandas approach" return 1 fi }

extract_pdf_text() { local file="$1" # Extract text from PDF using pdftotext if command -v pdftotext &>/dev/null; then pdftotext "$file" - 2>/dev/null else echo "pdftotext not available; try PyPDF2 via Python sandbox" return 1 fi }

extract_raw_strings() { local file="$1" # Extract printable strings from any binary file if command -v strings &>/dev/null; then strings "$file" 2>/dev/null | head -500 else echo "strings not available" return 1 fi }

Usage: ./extraction_fallback.sh <command> <file>

Commands: docx_raw, xlsx_csv, pdf_text, raw_strings

case "$1" in docx_raw) extract_docx_raw "$2" ;; xlsx_csv) extract_xlsx_to_csv "$2" ;; pdf_text) extract_pdf_text "$2" ;; raw_strings) extract_raw_strings "$2" ;; *) echo "Usage: $0 <docx_raw|xlsx_csv|pdf_text|raw_strings> <file>" exit 1 ;; esac

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/prioritize-context-data-enhanced of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

Resilient Context Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Resilient Context Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Resilient Context Extraction this skillHKUDS/OpenSpace7.8k—~2.5kAutomated safety check: PassMIT
Markdown Article FormatterJimLiu/baoyu-skills26k6 repos~3.5kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Obsidian MarkdownAtmosphere/atmosphere3.8k20 repos~1.3kAutomated safety check: PassApache-2.0
DOCXrvdbreemen/OTGW-firmware20733 repos~4.3kAutomated safety check: PassProprietary
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0

Similar skills

  • Markdown Article Formatter

    JimLiu/baoyu-skills

    Reformats plain text or Markdown articles with frontmatter, a title, a summary, headings, bold, lists and code blocks, and saves a separate formatted copy.

    26k GitHub starsUsed in 6 repos~3.5k tokens
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Obsidian Markdown

    Atmosphere/atmosphere

    Create and edit Obsidian Flavored Markdown with wikilinks, embeds, callouts, properties, and other Obsidian-specific syntax.

    3.8k GitHub starsUsed in 20 repos~1.3k tokens
    Documents & OfficeAuto-check passed
  • DOCX

    rvdbreemen/OTGW-firmware

    A skill your agent uses whenever the user wants to create, read, edit, or manipulate Word documents (.docx files).

    207 GitHub starsUsed in 33 repos~4.3k tokens
    Documents & OfficeAuto-check passed
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

    41k GitHub stars~2.5k tokensUpdated yesterday
    Documents & OfficeAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.8k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.8k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.8k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.8k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.8k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Resilient Context Extraction

What does Resilient Context Extraction do?

Ensures agents extract data from context files with validation and fallback strategies before resorting to assumptions or external searches. Resilient Context Extraction is an agent skill from HKUDS/OpenSpace. Ensures agents extract data from context files with validation and fallback strategies before resorting to assumptions or external searches.

When should I use Resilient Context Extraction?

Resilient Context Extraction fits situations like: documents & Office work in your project.

How do I install Resilient Context Extraction in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill resilient-context-extraction -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/prioritize-context-data-enhanced in HKUDS/OpenSpace) into .claude/skills/resilient-context-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Resilient Context Extraction in Codex?

Run `npx skills add HKUDS/OpenSpace --skill resilient-context-extraction -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/prioritize-context-data-enhanced in HKUDS/OpenSpace) into .agents/skills/resilient-context-extraction in your project. Codex loads it when a task matches its description.

Can I use Resilient Context Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill resilient-context-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/resilient-context-extraction, .gemini/skills/resilient-context-extraction, .github/skills/resilient-context-extraction and .opencode/skills/resilient-context-extraction in your project.

What does Resilient Context Extraction need to run?

Going by SKILL.md and its folder, Resilient Context Extraction needs the command-line tools its instructions call (pdftotext). Our summary lists: Python 3.

Does Resilient Context Extraction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Resilient Context Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Resilient Context Extraction use?

Resilient Context Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Resilient Context Extraction use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Resilient Context Extraction?

Skills that share tags, products or a category with Resilient Context Extraction: Markdown Article Formatter (JimLiu/baoyu-skills, 26k stars), Markitdown (ImCa0/just-laws, 782 stars), Obsidian Markdown (Atmosphere/atmosphere, 3.8k stars) and DOCX (rvdbreemen/OTGW-firmware, 207 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Resilient Context Extraction?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,750 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.