Agent skill

Prioritize Reference Files

by HKUDS in HKUDS/OpenSpace

Ensures agents read and use provided reference files before searching or fabricating data

MITAuto-check passedDocuments & Office

Install Prioritize Reference Files

skills CLI
$ npx skills add HKUDS/OpenSpace --skill prioritize-reference-files -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUDS/OpenSpace prioritize-reference-files --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUDS/OpenSpace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmarks/gdpval/skills/prioritize-reference-files .claude/skills/prioritize-reference-files && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
prioritize-reference-files
GitHub stars
7.7k
Token cost
~1.3k tokens
SKILL.md length
515 words
Files
2
Skills in repo
199
Repo updated
First seen
Licence
MIT

At a glance

Ensures agents read and use provided reference files before searching or fabricating data

  • Works in 5 steps: Scan Task Context for Reference Files → Read Reference Files First → Extract and Validate Data → …
  • Tasks that involve Web search
  • SKILL.md covers Purpose, Core Principle, Workflow and Anti-Patterns to Avoid, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Prioritize Reference Files is an agent skill from HKUDS/OpenSpace. Ensures agents read and use provided reference files before searching or fabricating data

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Documents & Office, covering Web search. The repository describes itself as: "OpenSpace: The Skill Management Layer for AI Agents" -- https://open-space.cloud/. The licence is MIT.

When your agent uses it

  • Tasks that involve Web search

Example prompts

  • “Use the prioritize-reference-files skill to ensure agents read and use provided reference files before searching or fabricating data”
  • “/prioritize-reference-files”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Scan Task Context for Reference Files
  2. Read Reference Files First
  3. Extract and Validate Data
  4. Use Reference Data as Primary Source
  5. Document Data Source

What it can do on your machine

Read from SKILL.md and the folder at commit 3827781. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Prioritize Reference Files loads about 1.3k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 515 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUDS/OpenSpace at commit 3827781, republished under its MIT licence (© HKUDS). 515 words, ~1,256 tokens.

Download SKILL.mdSave it as .claude/skills/prioritize-reference-files/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
prioritize-reference-files
description
Ensures agents read and use provided reference files before searching or fabricating data

Prioritize Reference Files

Purpose

This skill ensures that when reference files are provided in task context, you MUST read and extract data from them FIRST before attempting web searches or generating synthetic data. Ignoring available structured data leads to fabricated outputs and incorrect results.

Core Principle

Reference files in context > Web search > Data fabrication (never)

Workflow

Step 1: Scan Task Context for Reference Files

Before taking any action, identify all files provided in the task context:

  • Look for file attachments, uploads, or references in the task description
  • Common formats: .xlsx, .csv, .json, .pdf, .docx, .txt
  • Check for phrases like "attached", "provided", "reference file", "see file"
Step 2: Read Reference Files First

Use the appropriate tool to read each reference file:

python
# For Excel files
read_file(filetype="xlsx", file_path="path/to/file.xlsx")

# For CSV files  
read_file(filetype="csv", file_path="path/to/file.csv")

# For PDF files
read_file(filetype="pdf", file_path="path/to/file.pdf")

# For JSON files
read_file(filetype="json", file_path="path/to/file.json")

# For text files
read_file(filetype="txt", file_path="path/to/file.txt")

If read_file fails on .docx files (returns error, empty content, or 'unknown error'):

Fallback Approach 1: Direct zipfile/XML extraction via run_shell

bash
# .docx files are ZIP archives containing XML; extract document.xml directly
unzip -p path/to/file.docx word/document.xml | grep -oP '(?<=<w:t>)[^<]+' | tr '\n' ' '

Or for more complete extraction:

bash
mkdir -p /tmp/docx_extract && cd /tmp/docx_extract && unzip path/to/file.docx && cat word/document.xml

Fallback Approach 2: Use shell_agent for complex extraction If direct extraction fails, delegate to shell_agent:

shell_agent(task="Extract text content from path/to/file.docx using zipfile and XML parsing")

The agent will attempt multiple extraction methods and report results.

Fallback Approach 3: Verify extraction success before proceeding After any extraction method, confirm content was retrieved:

  • Check that output is non-empty and contains expected document structure
  • Look for familiar text from the document (headings, known phrases)
  • If extraction appears incomplete, try an alternative method
  • Document which method succeeded: "Content extracted using [method] after read_file failed"

Important: Never proceed to data fabrication if reference files exist but read_file fails. Always attempt at least one fallback extraction method first.

Step 3: Extract and Validate Data

After reading:

  1. Parse the content to understand the data structure
  2. Identify relevant fields/columns for your task
  3. Verify the data is complete and usable
  4. Note any limitations or missing information
Step 4: Use Reference Data as Primary Source
  • Base all outputs on the reference file data
  • Only supplement with web search if reference data is incomplete
  • NEVER fabricate data when reference files exist
  • If reference data is insufficient, explicitly state what's missing
Show full SKILL.md (186 more words)Show less
Step 5: Document Data Source

In your outputs, acknowledge the source:

  • "Data sourced from [filename]"
  • "Based on provided reference file: [filename]"
  • This confirms you used the actual provided data

Anti-Patterns to Avoid

❌ Ignoring reference files and searching the web instead ❌ Fabricating data when structured data is available ❌ Assuming file contents without reading them ❌ Using outdated web data when current reference files exist ❌ Giving up after read_file fails without trying fallback extraction methods

Example

Task Context: "Create a property listings report. See Massabama_active_listings.xlsx for current data."

Correct Approach:

1. Read Massabama_active_listings.xlsx first
2. Extract property addresses, prices, specifications
3. Generate report using actual listing data
4. Note: "Data sourced from Massabama_active_listings.xlsx"

Incorrect Approach:

1. Search web for "Massabama property listings"
2. Fabricate property data from search results
3. Create report with unverified/generated data

Verification Checklist

Before completing any task with reference files:

  • I have identified all reference files in the task context
  • I have read each reference file using appropriate tools
  • I have extracted relevant data from the files
  • My output is based on the reference file data
  • I have documented the data source in my output

When Web Search is Appropriate

Only search the web when:

  • Reference files don't exist for the required data
  • Reference files are incomplete AND supplemental info is needed
  • You need to verify or update time-sensitive information
  • You explicitly state what the reference files lack

© HKUDS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in benchmarks/gdpval/skills/prioritize-reference-files of HKUDS/OpenSpace.

  • SKILL.md
  • .skill_id

Open the folder on GitHubat commit 3827781

Compare with similar skills

Prioritize Reference Files next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Prioritize Reference Files compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Prioritize Reference Files this skillHKUDS/OpenSpace7.7k—~1.3kAutomated safety check: PassMIT
Paper Submissionbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~3.4kAutomated safety check: NotesCustom licence
Nemo RetrieverNVIDIA/skills3.5k—~583Automated safety check: PassApache-2.0
Jev SEOAgriciDaniel/jev-seo515—~2.5kAutomated safety check: NotesMIT
Harness Useautonomous-ai/Physical-AI-Operating-System381—~8kAutomated safety check: PassApache-2.0
IngestPrismer-AI/PrismerCloud1.6k—~934Automated safety check: PassMIT

Similar skills

  • Paper Submission

    brycewang-stanford/Auto-Empirical-Research-Skills

    Evaluate a paper's contribution novelty, identify best-fit SSCI journal fields and ABS star rating, and recommend 20 target journals.

    4.5k GitHub stars~3.4k tokensUpdated 2 days ago
    Documents & OfficeAuto-check: notes
  • Nemo Retriever

    NVIDIA/skills

    Official

    A skill your agent uses when searching, extracting, ingesting, or querying a document collection with the NeMo Retriever 26.8.1 CLI, including local LanceDB indexes and deployed Retriever services.

    3.5k GitHub stars~583 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Jev SEO

    AgriciDaniel/jev-seo

    Full live SEO audit of any website from its homepage URL, powered by Jev (TypeSafe's System One model).

    515 GitHub stars~2.5k tokensUpdated 15 days ago
    Documents & OfficeAuto-check: notes
  • Harness Use

    autonomous-ai/Physical-AI-Operating-System

    Delegate digital work to agents on the computer paired through Harness; discover Store packages and prepare an agent when needed.

    381 GitHub stars~8k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Ingest

    Prismer-AI/PrismerCloud

    Turn external web URLs into LLM-ready content — load + cache web pages (HQCC compression) and search the web.

    1.6k GitHub stars~934 tokensUpdated 7 days ago
    Documents & OfficeAuto-check passed
  • Skywork Document

    aiskillstore/marketplace

    Generate professional documents in multiple formats (docx, pdf, html, md) from scratch or based on user files.

    430 GitHub stars~2.5k tokensUpdated today
    Documents & OfficeAuto-check passed

More from HKUDS/OpenSpace

All 199 skills in this repo
  • Walks through producing a master audio track plus stems in Python, from checking a reference file and timing sections by BPM to effects, a zip archive and final verification.

    7.7k GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Handle cascading data retrieval tool failures by falling back to embedded knowledge generation

    7.7k GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Gives an agent a workaround when its code-execution sandbox keeps failing: save the Python script to a file and run it through the shell instead.

    7.7k GitHub stars~588 tokensUpdated 1 mo ago
    Auto-check passed
  • A recovery routine for agents whose sandboxed code runner keeps failing: save the Python script to disk, then run it through the shell and read the output.

    7.7k GitHub stars~652 tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback ladder for failed sandboxed code runs, plus the habit of fixing the working directory first so generated files land in the right place.

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Fallback workflow for executing Python code when executecodesandbox fails repeatedly

    7.7k GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Prioritize Reference Files

What does Prioritize Reference Files do?

Ensures agents read and use provided reference files before searching or fabricating data. Prioritize Reference Files is an agent skill from HKUDS/OpenSpace.

When should I use Prioritize Reference Files?

Prioritize Reference Files fits situations like: tasks that involve Web search.

How do I install Prioritize Reference Files in Claude Code?

Run `npx skills add HKUDS/OpenSpace --skill prioritize-reference-files -a claude-code`. Or copy the skill folder (benchmarks/gdpval/skills/prioritize-reference-files in HKUDS/OpenSpace) into .claude/skills/prioritize-reference-files in your project. Claude Code loads it when a task matches its description.

How do I install Prioritize Reference Files in Codex?

Run `npx skills add HKUDS/OpenSpace --skill prioritize-reference-files -a codex`. Or copy the skill folder (benchmarks/gdpval/skills/prioritize-reference-files in HKUDS/OpenSpace) into .agents/skills/prioritize-reference-files in your project. Codex loads it when a task matches its description.

Can I use Prioritize Reference Files in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUDS/OpenSpace --skill prioritize-reference-files -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/prioritize-reference-files, .gemini/skills/prioritize-reference-files, .github/skills/prioritize-reference-files and .opencode/skills/prioritize-reference-files in your project.

What does Prioritize Reference Files need to run?

SKILL.md names no scripts, command-line tools or credentials: Prioritize Reference Files is instructions for the agent only. Our summary lists: Python 3.

Does Prioritize Reference Files access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Prioritize Reference Files safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Prioritize Reference Files use?

Prioritize Reference Files is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Prioritize Reference Files use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Prioritize Reference Files?

Skills that share tags, products or a category with Prioritize Reference Files: Paper Submission (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars), Nemo Retriever (NVIDIA/skills, 3.5k stars), Jev SEO (AgriciDaniel/jev-seo, 515 stars) and Harness Use (autonomous-ai/Physical-AI-Operating-System, 381 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Prioritize Reference Files?

HKUDS (a GitHub organization) maintains it in HKUDS/OpenSpace, which has 7,743 GitHub stars. The repository holds 199 skills in this directory. The repository was last updated on August 12, 2026.

Source: HKUDS/OpenSpace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.