Agent skill

Firstdata

by MLT-OSS in MLT-OSS/FirstData

Find official portals, APIs, and download paths for authoritative primary data sources (governments, international organizations, research institutions, etc.).

MITAuto-check passedResearch & Science

Install Firstdata

skills CLI
$ npx skills add MLT-OSS/FirstData --skill firstdata -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install MLT-OSS/FirstData firstdata --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/MLT-OSS/FirstData.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/firstdata .claude/skills/firstdata && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
firstdata
GitHub stars
183
Token cost
~3.1k tokens
SKILL.md length
1,330 words
Files
3 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Find official portals, APIs, and download paths for authoritative primary data sources (governments, international organizations, research institutions, etc.).

  • Users need to know where to find this data from an official source
  • SKILL.md covers What FirstData Is, Capabilities, Typical Queries and Quick Start, plus 3 more sections
  • Calls npx; reaches firstdata.deepminer.com.cn; needs FIRSTDATA_API_KEY
  • Which source is more authoritative

What it does

Firstdata is an agent skill from MLT-OSS/FirstData. Find official portals, APIs, and download paths for authoritative primary data sources (governments, international organizations, research institutions, etc.). Use when users need to know "where to find this data from an official source", "which source is more authoritative", or "how to cite primary data". Covers the live FirstData catalog with authority comparison and site navigation guidance.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `mcp-tool-descriptions-draft.md` and `references/firstdata-register.md`).

It sits in Research & Science, covering MCP servers. It works with Model Context Protocol. The repository describes itself as: The World's Most Comprehensive, Authoritative, and Structured Open Source Data Source Knowledge Base. The licence is MIT.

When your agent uses it

  • Users need to know where to find this data from an official source
  • Which source is more authoritative
  • How to cite primary data

Example prompts

  • “where to find this data from an official source”
  • “which source is more authoritative”
  • “how to cite primary data”
  • “/firstdata”

Requirements

  • Node.js
  • A credential in FIRSTDATA_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit e4a6687. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • firstdata.deepminer.com.cn

    Also links to:

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FIRSTDATA_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Firstdata loads about 3.1k tokens when it runs, and up to ~4.3k if it reads all its reference files. Until then it costs about 102 tokens; SKILL.md has 1,330 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from MLT-OSS/FirstData at commit e4a6687, republished under its MIT licence (© MLT-OSS). 1,330 words, ~3,054 tokens.

Download SKILL.mdSave it as .claude/skills/firstdata/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
firstdata
description
Find official portals, APIs, and download paths for authoritative primary data sources (governments, international organizations, research institutions, etc.). Use when users need to know "where to find this data from an official source", "which source is more authoritative", or "how to cite primary data". Covers the live FirstData catalog with authority comparison and site navigation guidance.
version
0.0.2

FirstData

What FirstData Is

FirstData is the External Facts Context Layer for AI Agents — a purpose-built, authoritative collection of primary data sources that helps agents locate official origins rather than generating unverified answers.

It does not replace raw data — it acts as an "authoritative data navigator", taking vague user needs as input, recommending the most appropriate primary sources, and providing clear access paths, API information, and download methods so both users and agents can trace back to original evidence. The live catalog size is returned by the MCP get_status tool and should not be hard-coded in agent responses.

Coverage:

  • International organizations: World Bank, IMF, OECD, WHO, FAO, etc.
  • Chinese government agencies: PBC, National Bureau of Statistics, General Administration of Customs, CSRC, etc.
  • National official agencies: US, Canada, Japan, UK, Australia, etc.
  • Academic & research databases: NBER, Penn World Table, PubMed, etc.
  • Corporate disclosure & market platforms: stock exchange disclosure systems, listed company filings, etc.
  • Industry-specific databases: energy, finance, health, climate, legal & regulatory, etc.

When to use: When users need to find official data sources, compare source authority, obtain official URLs/APIs/download paths, or build evidence-chain workflows. FirstData is a source locator, not an answer generator — after receiving results, guide users back to original sources for verification rather than treating them as final answers.

Capabilities

1. Source Locator — Returns the top 3–5 most relevant sources with authority level, matching rationale, access URL, API documentation, and download methods.

2. Site Pathfinder — Provides step-by-step navigation from homepage to target data for complex official websites, including alternative paths and API access methods.

3. Evidence-Ready Workflows — Can be embedded into workflows requiring evidence chains: deep research, policy analysis, investment research, compliance auditing, fact-checking, etc.

Each data source includes structured metadata: authority level (government / international / research / market / commercial / other), access URL, API information, download formats, geographic scope, update frequency, access level, etc.

Typical Queries

Typical query scenarios when agents call FirstData via MCP:

User NeedQuery DirectionExpected Output
"Which official source should I cite for China's 2023 NEV export volume?"China Customs, National Bureau of StatisticsOfficial source + authority level + data page URL
"Where to download IPO prospectus for a Hong Kong-listed company?"HKEXnewsOfficial platform + step-by-step navigation
"World Bank vs IMF GDP data — which is better for academic citation?"World Bank WDI, IMF WEOSource comparison + authority differences + API docs
"Need global climate data with API access"NASA Earthdata, NOAA CDOData source + API docs + access methods
"Where is the official data for China's M2 money supply?"People's Bank of ChinaOfficial data portal + update frequency + historical coverage

Full project background and feature documentation: README

Quick Start

This skill connects to the FirstData MCP server (firstdata.deepminer.com.cn), the project's official hosted API endpoint. An API key (FIRSTDATA_API_KEY) is required for authentication.

If you already have FIRSTDATA_API_KEY set, configure the MCP connection:

bash
npx mcporter config add firstdata https://firstdata.deepminer.com.cn/mcp --header 'Authorization=Bearer ${FIRSTDATA_API_KEY}'

Or add manually to your MCP config:

json
{
  "mcpServers": {
    "firstdata": {
      "type": "streamable-http",
      "url": "https://firstdata.deepminer.com.cn/mcp",
      "headers": {
        "Authorization": "Bearer <FIRSTDATA_API_KEY>"
      }
    }
  }
}

If you don't have an API key, see firstdata-register.md for the registration process (two API calls to the FirstData server to obtain a JWT token).

Once connected, call get_status first to verify the MCP connection, inspect the active tool list, and read the current catalog snapshot metadata. Then browse the tool list and select the appropriate tool based on your needs.

MCP Tools Reference

The FirstData MCP server provides 6 tools. Below is a reference with usage guidelines, limitations, and examples.

Common Limitations (all tools)
  • Authentication required: All tools require a valid API key (JWT token) via Authorization: Bearer <token> header.
  • Daily call quota: API usage is subject to a per-token daily call quota. Quota varies by API key tier (trial accounts: 30 calls/day). MCP tool calls do not return remaining quota information. To check quota, use the Token verification API (POST /api/token/verify) which returns remaining_daily in the response — this is a separate HTTP call, not available through MCP tool invocation.
  • Network dependency: All tools make HTTP calls to the FirstData server (firstdata.deepminer.com.cn). Network latency and server availability affect response times.
Tool: get_status

Purpose: Check the MCP server version, currently registered tools, and the catalog snapshot available to the server.

Use it first when configuring a new Agent or diagnosing a connection. The response includes catalog source counts, generated timestamp when available, and whether the runtime data directories are present. It does not return credentials.

Tool: search_source

Purpose: Unified data source search tool supporting keyword search, structured filtering, pagination, and multiple output modes.

Limitations:

  • Maximum 200 results per query (limit parameter range: 1–200, default: 20).
  • ⚠️ Each keyword is matched as an independent substring — pass each search term as a separate array element. For example, use ["中国", "GDP"] (173 results) instead of ["中国 GDP"] (0 results). This is by design to preserve multi-word terms like "New Zealand" or "World Bank".
  • Keyword matching is substring-based, not semantic search. Keywords are matched against source metadata fields (name, description, tags, content).
  • The domain parameter uses substring matching, not exact enum matching (e.g., "finance" matches "public-finance", "finance", "financial-markets").
  • No boolean operators (AND/OR/NOT). Multiple keywords in the array are combined with OR logic (results matching any keyword are returned, deduplicated).
  • Response time: typically ~1 second.
Show full SKILL.md (491 more words)Show less
Tool: get_source

Purpose: Retrieve full details for specific data sources by their IDs.

Limitations:

  • Invalid source_id values do NOT cause an error response (isError: false). Instead, the result array includes {"id": "xxx", "error": "Not found"} for each invalid ID alongside valid results. Callers must check individual items for error fields rather than relying solely on isError.
  • No schema-level limit on the number of source_ids per request, but performance with large batches (50+) is unverified. As a practical guideline (not a hard limit), consider batching in groups of ~20.
  • The fields parameter filters returned fields; when omitted, all fields are returned.
Tool: ask_agent

Purpose: LLM-powered intelligent search agent for complex, cross-domain, or ambiguous queries that require multi-step reasoning.

Limitations:

  • Query length: 2–1,000 characters.
  • Maximum results: 1–20 (default: 5).
  • Non-idempotent: Same query may return different results across calls (LLM reasoning varies).
  • Response time: typically 2–8 seconds (involves LLM inference). May take longer (10–30+ seconds) when the agent triggers web_search for external information.
  • Internally uses LangChain ReAct agent with jq for local data queries plus optional web_search. The web search step is not user-controllable.
  • Use search_source instead for simple keyword matching or structured filtering — it is faster, deterministic, and cheaper.
Tool: get_access_guide

Purpose: Generate detailed access instructions for a specific data source using RAG (Retrieval-Augmented Generation).

Limitations:

  • Not all data sources have instruction libraries. If a source has no pre-built instructions, results will be empty or irrelevant.
  • Invalid source_id returns {"error": "数据源 xxx 不存在"}.
  • top_k range: 1–5 (default: 3).
  • Response time is highly variable: 3–20 seconds, depending on RAG retrieval complexity and server load.
  • Retrieval quality depends heavily on the specificity of the operation parameter. Vague descriptions yield lower-quality matches. Use specific action verbs and entity names (e.g., "查询2024年M2货币供应量数据" rather than "查数据").
Tool: report_feedback

Purpose: Submit user feedback to the development team when FirstData has a confirmed issue.

Limitations:

  • feedback_message length: 10–2,000 characters.
  • Non-idempotent: Duplicate calls create duplicate feedback entries. Do not retry on success.
  • Only use when a genuine issue is confirmed (missing source, incorrect data, broken functionality). Do not use as a general comment channel.

Examples:

# Example 1: Broken link
feedback_message="链接失效:数据源 china-pbc 的 data_url 返回 404,无法访问数据页面。检索关键词:中国货币供应量"

# Example 2: Outdated content
feedback_message="数据内容过时:数据源 worldbank-open-data 的 update_frequency 标注为 quarterly,但实际已超过 6 个月未更新"

Description Quality Guidelines

When adding or modifying MCP tool descriptions, follow these principles (based on MCP tool description quality research):

Core principle: "Write it right before writing it all" — Functionality accuracy (+11.6% impact) matters ~8× more than Conciseness (+1.5%).

6-dimension checklist (check all before submitting):

  • Purpose: Is the tool's function clearly stated in the first sentence?
  • Guidelines: Are usage scenarios and when-to-use / when-not-to-use rules included?
  • Examples: Are typical input/output examples provided?
  • Limitations: Are constraints, edge cases, and known limitations documented?
  • Parameters: Are all parameters described with types, ranges, and defaults?
  • Return Format: Is the response structure documented?

Community

FirstData is an open-source project — join us in building the External Facts Context Layer for AI Agents:

  • ⭐ Star the project to help more agents and developers discover it
  • 📝 Issue to report problems, suggest new data sources, or propose improvements
  • 🔀 PR to contribute code, data sources, or documentation improvements

© MLT-OSS, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/firstdata of MLT-OSS/FirstData.

  • SKILL.md
  • mcp-tool-descriptions-draft.md
  • references/firstdata-register.md

Open the folder on GitHubat commit e4a6687

Compare with similar skills

Firstdata next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Firstdata compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Firstdata this skillMLT-OSS/FirstData183—~3.1kAutomated safety check: PassMIT
Read GitHubAgentTeam-TaichuAI/ScienceClaw6702 repos~638Automated safety check: PassNone
Report Issue Frameworkcyanheads/pubmed-mcp-server155—~3.2kAutomated safety check: PassApache-2.0
Serply Search MCPsickn33/agentic-awesome-skills47k1 repos~1.5kAutomated safety check: PassMIT
Bgpt MCPClawBio/ClawBio1.2k1 repos~3.2kAutomated safety check: PassMIT
Setup MedsciAperivue/medsci-skills329—~960Automated safety check: PassMIT

Similar skills

  • Read GitHub

    AgentTeam-TaichuAI/ScienceClaw

    Read and search GitHub repository documentation via gitmcp.io MCP service.

    670 GitHub starsUsed in 2 repos~638 tokens
    Research & ScienceAuto-check passed
  • Report Issue Framework

    cyanheads/pubmed-mcp-server

    File a bug or feature request against @cyanheads/mcp-ts-core when you hit a framework issue.

    155 GitHub stars~3.2k tokensUpdated 4 days ago
    Research & ScienceAuto-check passed
  • Serply Search MCP

    sickn33/agentic-awesome-skills

    Search Google, Bing, Google News and Google Scholar, and read public pages, with the Serply MCP server.

    47k GitHub starsUsed in 1 repo~1.5k tokens
    Research & ScienceAuto-check passed
  • Bgpt MCP

    ClawBio/ClawBio

    Search scientific papers via the BGPT MCP server and retrieve structured experimental data — methods, results, conclusions, quality scores, and 25+ metadata fields per paper.

    1.2k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Setup Medsci

    Aperivue/medsci-skills

    A skill your agent uses when a skill fails for a missing tool or the environment needs checking.

    329 GitHub stars~960 tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • G1

    brycewang-stanford/Auto-Empirical-Research-Skills

    VS-Enhanced Journal Matcher with Journal Intelligence MCP — Real-time journal data pipeline with checkpoint-based human decisions.

    4.5k GitHub stars~3.7k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed

Questions about Firstdata

What does Firstdata do?

Find official portals, APIs, and download paths for authoritative primary data sources (governments, international organizations, research institutions, etc.). Firstdata is an agent skill from MLT-OSS/FirstData.).

When should I use Firstdata?

Firstdata fits situations like: users need to know where to find this data from an official source; which source is more authoritative; how to cite primary data.

How do I install Firstdata in Claude Code?

Run `npx skills add MLT-OSS/FirstData --skill firstdata -a claude-code`. Or copy the skill folder (skills/firstdata in MLT-OSS/FirstData) into .claude/skills/firstdata in your project. Claude Code loads it when a task matches its description.

How do I install Firstdata in Codex?

Run `npx skills add MLT-OSS/FirstData --skill firstdata -a codex`. Or copy the skill folder (skills/firstdata in MLT-OSS/FirstData) into .agents/skills/firstdata in your project. Codex loads it when a task matches its description.

Can I use Firstdata in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MLT-OSS/FirstData --skill firstdata -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/firstdata, .gemini/skills/firstdata, .github/skills/firstdata and .opencode/skills/firstdata in your project.

What does Firstdata need to run?

Going by SKILL.md and its folder, Firstdata needs the command-line tools its instructions call (npx) and credentials named FIRSTDATA_API_KEY. Our summary lists: Node.js; A credential in FIRSTDATA_API_KEY.

Does Firstdata access the network?

SKILL.md names 2 domains. In commands or code: firstdata.deepminer.com.cn; the agent is likely to contact it when it follows the instructions. As links in the text: arxiv.org. This is read from the text; nothing was executed.

Is Firstdata safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Firstdata use?

Firstdata is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Firstdata use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Firstdata?

Skills that share tags, products or a category with Firstdata: Read GitHub (AgentTeam-TaichuAI/ScienceClaw, 670 stars), Report Issue Framework (cyanheads/pubmed-mcp-server, 155 stars), Serply Search MCP (sickn33/agentic-awesome-skills, 47k stars) and Bgpt MCP (ClawBio/ClawBio, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Firstdata?

MLT-OSS (a GitHub organization) maintains it in MLT-OSS/FirstData, which has 183 GitHub stars. The repository was last updated on October 3, 2026.

Source: MLT-OSS/FirstData on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.