Agent skill

Fin Data Acquisition

by csmar432 in csmar432/finai-research

根据REFINEDDESIGN.md中的变量定义,自动获取所需数据并生成可执行的回归分析脚本(Python/Stata)。

MITAuto-check passedResearch & Science

Install Fin Data Acquisition

skills CLI
$ npx skills add csmar432/finai-research --skill fin-data-acquisition -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install csmar432/finai-research fin-data-acquisition --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/csmar432/finai-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fin-data-acquisition .claude/skills/fin-data-acquisition && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fin-data-acquisition
GitHub stars
109
Token cost
~2k tokens
SKILL.md length
66 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

根据REFINEDDESIGN.md中的变量定义,自动获取所需数据并生成可执行的回归分析脚本(Python/Stata)。

  • Works in 5 steps: 必须先运行数据源预检查 — 不跳过他 → 禁止静默Fallback — 模拟数据必须用户授权 → 每个数据操作必须记录溯源 — 包括来源、时间戳、行数 → …
  • Tasks that involve Econometrics and empirical research
  • SKILL.md covers 触发条件, 核心原则, 数据源Fallback链 and DataFetcher API, plus 6 more sections
  • Needs TUSHARE_TOKEN and EODHD_API_KEY

What it does

Fin Data Acquisition is an agent skill from csmar432/finai-research. 根据REFINEDDESIGN.md中的变量定义,自动获取所需数据并生成可执行的回归分析脚本(Python/Stata)。

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Econometrics and empirical research and Design tokens. It works with Python and Model Context Protocol. The repository describes itself as: Evidence-first AI workflow for economic and financial research: literature → identification → data → econometrics → verifiable LaTeX. 43 data sources, 58 method modules, 18 AI… The licence is MIT.

When your agent uses it

  • Tasks that involve Econometrics and empirical research
  • Tasks that involve Design tokens

Example prompts

  • “/fin-data-acquisition”

Requirements

  • Python 3
  • A credential in TUSHARE_TOKEN
  • A credential in EODHD_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. 必须先运行数据源预检查 — 不跳过他
  2. 禁止静默Fallback — 模拟数据必须用户授权
  3. 每个数据操作必须记录溯源 — 包括来源、时间戳、行数
  4. 检查点强制暂停 — 用户确认前不继续
  5. 失败时显示具体原因 — 而非笼统错误

What it can do on your machine

Read from SKILL.md and the folder at commit 47eebb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and stata).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • TUSHARE_TOKEN
    • EODHD_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fin Data Acquisition loads about 2k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 66 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from csmar432/finai-research at commit 47eebb7, republished under its MIT licence (© csmar432). 66 words, ~1,996 tokens.

Download SKILL.mdSave it as .claude/skills/fin-data-acquisition/SKILL.md (or your agent's skills folder).
name
fin-data-acquisition
description
根据REFINED_DESIGN.md中的变量定义,自动获取所需数据并生成可执行的回归分析脚本(Python/Stata)。
trigger
获取数据|数据获取|data acquisition|下载数据|数据准备
version
1.0.0
created
2026-06-13
tags
data, acquisition, mcp, python, stata, regression

fin-data-acquisition

根据REFINED_DESIGN.md中的变量定义,自动获取所需数据并生成可执行的回归分析脚本(Python/Stata)。

触发条件

  • 关键词: 获取数据 数据获取 data acquisition 下载数据 数据准备 实证数据
  • Skill语法: Skill: fin-data-acquisition
  • 前置条件: 已完成 REFINED_DESIGN.md (研究设计文档)

核心原则

禁止行为 (未经用户明确授权不得执行)
❌ 静默回退到模拟数据
❌ 自动生成虚假回归结果
❌ 在用户未同意情况下继续流水线使用模拟数据
❌ 跳过数据源预检查直接获取数据
数据源预检查 (强制执行)

在任何数据获取前,必须先运行数据源检查:

python
from scripts.data_source_checker import DataSourceChecker, DataRequirement

# 第一步:定义数据需求
requirements = [
    DataRequirement(
        name="financial_data",
        user_facing_name="A股财务数据",
        description="ROA、资产负债率、企业规模、研发投入",
        sources=["tushare", "wind", "csmar", "akshare"],
        required=True,
    ),
    DataRequirement(
        name="esg_data",
        user_facing_name="ESG评级",
        sources=["msci", "商道融绿", "华证"],
        required=False,
    ),
    DataRequirement(
        name="macro_data",
        user_facing_name="宏观数据",
        sources=["user-financial", "user-wb-data", "user-imf-data"],
        required=True,
    ),
]

# 第二步:运行数据源检查
checker = DataSourceChecker()
results = checker.check(requirements)

# 第三步:展示可用性报告
checker.print_report(results)

数据源Fallback链

每个数据类型都有明确的降级路径:

A股财务数据
tushare (需TUSHARE_TOKEN)
  ↓ 失败/无Token
wind (需Wind账号)
  ↓ 失败/无账号
csmar (需机构账号)
  ↓ 失败/无账号
akshare (免费,备选)
  ↓ 失败
手动下载 -> 询问用户
宏观数据
user-financial (akshare, 免费)
  ↓ 失败
user-wb-data (World Bank, 免费)
  ↓ 失败
user-imf-data (IMF, 免费)
  ↓ 失败
手动下载 -> 询问用户
美股数据
user-yfinance (免费)
  ↓ 失败
user-eodhd (需EODHD_API_KEY)
  ↓ 失败/无Key
手动下载 -> 询问用户
学术文献数据
user-openalex (免费)
  ↓ 失败
user-arxiv (免费)
  ↓ 失败
user-nber-wp (免费)
  ↓ 失败
手动检索 -> 询问用户

DataFetcher API

python
from scripts.research_framework import DataFetcher, ProvenanceTracker

# 初始化 (带数据溯源)
tracker = ProvenanceTracker(output_dir="data/provenance/")
fetcher = DataFetcher(output_dir="data/", tracker=tracker, verbose=True)

# ============ 面板数据获取 ============
df = fetcher.fetch_panel(
    tickers=["000001.SZ", "600000.SH"],
    years=["2018", "2019", "2020", "2021", "2022"],
    statements=["balance", "income", "cashflow"],
    include_sustainability=True,  # ESG数据
)

# ============ 财务报表获取 ============
fin = fetcher.fetch_financials("000001.SZ", "income")  # 利润表
fin = fetcher.fetch_financials("000001.SZ", "balance")  # 资产负债表
fin = fetcher.fetch_financials("000001.SZ", "cashflow")  # 现金流量表

# ============ 公司信息获取 ============
info = fetcher.fetch_ticker_info("000001.SZ")  # 股票基本信息

# ============ ESG/可持续发展数据 ============
sust = fetcher.fetch_sustainability("000001.SZ")  # ESG评级等

# ============ 宏观数据获取 ============
macro = fetcher.fetch_macro(indicator="gdp", country="CHN")

# ============ 融资融券数据 ============
margin = fetcher.fetch_margin(ts_code="000001.SZ", start_date="20180101")

# ============ 陆股通/港股通 ============
hgt = fetcher.fetch_hsgt_top10(date="20240101")

# ============ 分析师预测 ============
forecast = fetcher.fetch_consensus("000001.SZ")

MCP数据获取

直接调用MCP工具
python
# A股行情数据 (tushare)
server: user-tushare
tool: get_daily_quote
params: { "ts_code": "000001.SZ", "start_date": "20240101", "end_date": "20241231" }

# 财务报告 (tushare)
server: user-tushare
tool: get_financial_report
params: { "ts_code": "000001.SZ", "report_type": "income" }

# 美股数据 (yfinance)
server: user-yfinance
tool: get_yf_historical
params: { "ticker": "AAPL", "start_date": "2024-01-01", "end_date": "2024-12-31" }

# 宏观数据 (World Bank)
server: user-wb-data
tool: get_wb_indicator
params: { "country_code": "CHN", "indicator": "wb_gdp_usd" }

# 研报数据 (eastmoney)
server: user-eastmoney-reports
tool: get_research_report
params: { "ts_code": "000001.SZ", "max_results": 20 }

回归脚本生成

Python脚本模板
python
"""
{研究标题} — 回归分析脚本
Generated by fin-data-acquisition skill
Date: {date}
"""

import pandas as pd
import numpy as np
import statsmodels.api as sm
from scipy import stats
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')

# ============ 配置 ============
DATA_PATH = "data/processed/{dataset_name}.csv"
OUTPUT_DIR = "output/fin-experiments/"

# ============ 数据加载 ============
df = pd.read_csv(DATA_PATH)
print(f"样本量: {len(df)}, 时间范围: {df['year'].min()}-{df['year'].max()}")
print(f"处理组: {df['treat'].sum()}, 对照组: {(1-df['treat']).sum()}")

# ============ 描述性统计 ============
desc = df[['y_var', 'x_var', 'controls']].describe()
desc.to_csv(f"{OUTPUT_DIR}/descriptive_stats.csv")
print(desc)

# ============ 相关性矩阵 ============
corr = df[['y_var', 'x_var', 'controls']].corr()
sns.heatmap(corr, annot=True, cmap='RdBu_r', center=0)
plt.savefig(f"{OUTPUT_DIR}/correlation_matrix.pdf", dpi=300)
plt.close()

# ============ 基准回归 ============
# OLS
X = sm.add_constant(df[['x_var'] + ['controls']])
y = df['y_var']
model = sm.OLS(y, X).fit(cov_type='cluster', cov_kwds={'groups': df['firmid']})
print(model.summary())

# DID回归 (带双向固定效应)
from linearmodels.panel import PanelOLS
df = df.set_index(['firmid', 'year'])
mod = PanelOLS.from_formula('y_var ~ x_var + EntityEffects + TimeEffects', df)
res = mod.fit(cov_type='clustered', cluster_entity=True)
print(res.summary)

# ============ 保存结果 ============
results = {
    'coefficient': model.params,
    'std_error': model.bse,
    'p_value': model.pvalues,
    'r_squared': model.rsquared,
}
pd.DataFrame(results).to_csv(f"{OUTPUT_DIR}/regression_results.csv")
Stata脚本模板
stata
/*
{研究标题} — Stata回归分析
Generated by fin-data-acquisition skill
Date: {date}
*/

clear all
cd "data/processed/"

* 加载数据
import delimited "{dataset_name}.csv", clear

* 描述性统计
estpost summarize y_var x_var controls, detail
esttab using "descriptive_stats.tex", cells("mean sd min p25 p50 p75 max") replace

* 相关性矩阵
pwcorr y_var x_var controls, star(5) sig

* 基准回归 (OLS + 聚类标准误)
reg y_var x_var controls, vce(cluster firmid)

* 双向固定效应 (DID)
encode firmid, gen(firm)
encode year, gen(year_dum)
xtset firm year_dum
xtreg y_var x_var controls i.year_dum, fe vce(cluster firm)

* 平行趋势检验
gen pre1 = (year == treated_year - 1)
gen post0 = (year == treated_year)
gen post1 = (year == treated_year + 1)
reg y_var pre1 post0 post1 controls i.year_dum, vce(cluster firm)

* 安慰剂检验
xtreg y_var placebo_* controls i.year_dum, fe vce(cluster firm)

* 异质性分析
bysort group: xtreg y_var x_var controls i.year_dum, fe vce(cluster firm)

esttab using "regression_results.tex", b(4) se(4) star(* 0.1 ** 0.05 *** 0.01) replace

数据溯源追踪

每个数据获取操作都记录溯源信息:

python
from scripts.core.provenance import DataProvenance

provenance = DataProvenance()

# 记录数据获取
provenance.record_fetch(
    source="tushare",
    table="financial_statement",
    timestamp=datetime.now(),
    rows=len(df),
    fields=list(df.columns),
    query_params={"ts_code": "000001.SZ", "year": 2022},
)

# 记录数据转换
provenance.record_transform(
    input_tables=["financial_statement", "trading_data"],
    output_table="merged_panel",
    transformation="merge on [ts_code, year]",
    rows_before=[1000, 500],
    rows_after=800,
)

# 导出溯源报告
provenance.export("data/provenance/provenance_report.json")

Checkpoint (强制交互)

[CHECKPOINT] 数据源预检查完成。

可用性报告:
✅ A股财务数据: tushare 可用 (Token已配置)
✅ 宏观数据: user-financial 可用
⚠️ ESG数据: MSCI需账号 (可选数据)
❌ 陆股通数据: tushare当前版本不支持

问题数据:
- 陆股通成分股数据需要Wind账号或手动下载

请选择:
1. 授权使用模拟数据 (用于测试脚本)
2. 提供替代数据源
3. 更换研究变量/设计
4. 继续 (缺失数据将在后续处理)

依赖项

  • scripts/data_source_checker.py — 数据源预检查
  • scripts/research_framework/data_fetcher.py — 数据获取接口
  • scripts/core/provenance.py — 数据溯源追踪
  • scripts/research_framework/regression_engine.py — 回归引擎

约束

  1. 必须先运行数据源预检查 — 不跳过他
  2. 禁止静默Fallback — 模拟数据必须用户授权
  3. 每个数据操作必须记录溯源 — 包括来源、时间戳、行数
  4. 检查点强制暂停 — 用户确认前不继续
  5. 失败时显示具体原因 — 而非笼统错误

© csmar432, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/fin-data-acquisition of csmar432/finai-research.

Open the folder on GitHubat commit 47eebb7

Compare with similar skills

Fin Data Acquisition next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fin Data Acquisition compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fin Data Acquisition this skillcsmar432/finai-research109—~2kAutomated safety check: PassMIT
Aer Statspaibrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~3kAutomated safety check: PassCustom licence
NSFC Literature Review WriterHuiyuLi-2000/Chinese-Grant-Writer-Skills4341 repos~1.4kAutomated safety check: NotesMIT
Arxiv MCP Serverblazickjp/arxiv-mcp-server3.2k—~353Automated safety check: PassApache-2.0
Stata C Pluginsdylantmoore/stata-skill2911 repos~5.8kAutomated safety check: PassCustom licence
Stata AuditSepineTam/mcp-for-stata264—~1.2kAutomated safety check: PassAGPL-3.0

Similar skills

  • Aer Statspai

    brycewang-stanford/Auto-Empirical-Research-Skills

    A skill your agent uses when aer-identification has fixed the design, after methodology choice and before aer-robustness or aer-tables-figures, to run an AER-track analysis with StatsPAI — the…

    4.5k GitHub stars~3k tokensUpdated 4 days ago
    Research & ScienceAuto-check passed
  • NSFC Literature Review Writer

    HuiyuLi-2000/Chinese-Grant-Writer-Skills

    Writes the research-status literature review and critique section of an NSFC grant proposal, backed by a bundled multi-source literature search.

    434 GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check: notes
  • Arxiv MCP Server

    blazickjp/arxiv-mcp-server

    A skill your agent uses when finding, comparing, reading, or monitoring arXiv papers, including requests for abstracts, citation graphs, original LaTeX, section-level technical details, or…

    3.2k GitHub stars~353 tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Stata C Plugins

    dylantmoore/stata-skill

    Develop high-performance C/C++ plugins for Stata using the stplugin.h SDK.

    291 GitHub starsUsed in 1 repo~5.8k tokens
    Research & ScienceAuto-check passed
  • Stata Audit

    SepineTam/mcp-for-stata

    Inspect, validate, summarize, and render local Stata-MCP audit evidence under .statamcp.

    264 GitHub stars~1.2k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • CLI Anything Zotero

    PiaoyangGuohai1/cli-anything-zotero

    Full-featured CLI for Zotero reference management. An agent skill from PiaoyangGuohai1/cli-anything-zotero.

    138 GitHub stars~2.6k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed

More from csmar432/finai-research

All 15 skills in this repo
  • Fin Arch Diagram

    csmar432/finai-research

    生成研究/项目架构图、流程图、层次图(swimlane / processflow / hierarchytree)。适合 PPT 汇报、技术文档、综述插图。输出风格接近 draw.io,可选 graphviz(高质量)/ matplotlib(零依赖)双后端。

    109 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check passed
  • Fin Brief Generator

    csmar432/finai-research

    根据用户输入或已有研究输出(文献综述/想法报告/新颖性报告),自动生成或更新FINBRIEF.md,减少用户填写负担. An agent skill from csmar432/finai-research.

    109 GitHub stars~1.7k tokensUpdated 4 days ago
    Auto-check passed
  • Fin Experiment Design

    csmar432/finai-research

    经济金融实证方法设计。根据研究想法和REFINEDDESIGN.md,生成完整的实证研究设计方案,覆盖识别策略选择、样本构建、变量定义、稳健性检验清单和内生性处理方案。

    109 GitHub stars~4.2k tokensUpdated 4 days ago
    Auto-check passed
  • Fin Generate Idea

    csmar432/finai-research

    针对经济金融研究方向的创意生成与评估。生成8-12个可发表的研究idea,过滤后在数据可行的情况下进行小规模实证验证,输出排序后的研究想法报告。

    109 GitHub stars~2.5k tokensUpdated 4 days ago
    Auto-check passed
  • Fin Idea Discovery

    csmar432/finai-research

    经济金融研究的完整想法发现流程。从研究方向出发,经过文献综述、想法生成、新颖性验证、实证方法设计和数据获取,输出经过数据实证验证的可执行研究方案。

    109 GitHub stars~2.7k tokensUpdated 4 days ago
    Auto-check passed
  • Fin Lit Review

    csmar432/finai-research

    经济金融领域的系统性文献综述。整合 Semantic Scholar + ArXiv + OpenAlex + NBER 构建引文网络,识别研究缺口,生成结构化文献地图。

    109 GitHub stars~1.2k tokensUpdated 4 days ago
    Auto-check passed

Questions about Fin Data Acquisition

What does Fin Data Acquisition do?

根据REFINEDDESIGN.md中的变量定义,自动获取所需数据并生成可执行的回归分析脚本(Python/Stata)。. Fin Data Acquisition is an agent skill from csmar432/finai-research.

When should I use Fin Data Acquisition?

Fin Data Acquisition fits situations like: tasks that involve Econometrics and empirical research; tasks that involve Design tokens.

How do I install Fin Data Acquisition in Claude Code?

Run `npx skills add csmar432/finai-research --skill fin-data-acquisition -a claude-code`. Or copy the skill folder (.agents/skills/fin-data-acquisition in csmar432/finai-research) into .claude/skills/fin-data-acquisition in your project. Claude Code loads it when a task matches its description.

How do I install Fin Data Acquisition in Codex?

Run `npx skills add csmar432/finai-research --skill fin-data-acquisition -a codex`. Or copy the skill folder (.agents/skills/fin-data-acquisition in csmar432/finai-research) into .agents/skills/fin-data-acquisition in your project. Codex loads it when a task matches its description.

Can I use Fin Data Acquisition in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add csmar432/finai-research --skill fin-data-acquisition -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fin-data-acquisition, .gemini/skills/fin-data-acquisition, .github/skills/fin-data-acquisition and .opencode/skills/fin-data-acquisition in your project.

What does Fin Data Acquisition need to run?

Going by SKILL.md and its folder, Fin Data Acquisition needs credentials named TUSHARE_TOKEN and EODHD_API_KEY. Our summary lists: Python 3; A credential in TUSHARE_TOKEN; A credential in EODHD_API_KEY.

Does Fin Data Acquisition access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Fin Data Acquisition safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Fin Data Acquisition use?

Fin Data Acquisition is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fin Data Acquisition use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fin Data Acquisition?

Skills that share tags, products or a category with Fin Data Acquisition: Aer Statspai (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars), NSFC Literature Review Writer (HuiyuLi-2000/Chinese-Grant-Writer-Skills, 434 stars), Arxiv MCP Server (blazickjp/arxiv-mcp-server, 3.2k stars) and Stata C Plugins (dylantmoore/stata-skill, 291 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fin Data Acquisition?

csmar432 (a GitHub user) maintains it in csmar432/finai-research, which has 109 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 6, 2026.

Source: csmar432/finai-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.