A skill your agent uses when writing the data and sample section of an Economic-Research manuscript — naming databases, building variable-definition and descriptive-statistics tables, and leaving an…

MITAuto-check passedResearch & Science

Install Er Data Sample

skills CLI
$ npx skills add brycewang-stanford/Awesome-Journal-Skills --skill er-data-sample -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brycewang-stanford/Awesome-Journal-Skills er-data-sample --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brycewang-stanford/Awesome-Journal-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Economic-Research-Journal-Skills/skills/er-data-sample .claude/skills/er-data-sample && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
er-data-sample
GitHub stars
1.2k
Token cost
~1.1k tokens
SKILL.md length
176 words
Files
1
Skills in repo
2,387
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when writing the data and sample section of an Economic-Research manuscript — naming databases, building variable-definition and descriptive-statistics tables, and leaving an…

  • Writing the data and sample section of an Economic-Research manuscript — naming databases
  • SKILL.md covers 触发时机, 数据说明段落规范, 变量定义表规范 and 描述性统计表规范, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Building variable-definition and descriptive-statistics tables

What it does

Er Data Sample is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when writing the data and sample section of an Economic-Research manuscript — naming databases, building variable-definition and descriptive-statistics tables, and leaving an auditable sample-filtering trail to 发表级.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Statistics. The repository describes itself as: Journal-specific Claude Code/Codex skill packs covering mainstream journals — AER, QJE, Nature, Cell, 管理世界, 经济研究 & 200+ more — your fast track to getting published. | 覆盖主流期刊的… The licence is MIT.

When your agent uses it

  • Writing the data and sample section of an Economic-Research manuscript — naming databases
  • Building variable-definition and descriptive-statistics tables
  • Leaving an auditable sample-filtering trail to 发表级

Example prompts

  • “/er-data-sample”

What it can do on your machine

Read from SKILL.md and the folder at commit 932eb23. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are stata).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Er Data Sample loads about 1.1k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 176 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brycewang-stanford/Awesome-Journal-Skills at commit 932eb23, republished under its MIT licence (© brycewang-stanford). 176 words, ~1,117 tokens.

Download SKILL.mdSave it as .claude/skills/er-data-sample/SKILL.md (or your agent's skills folder).
name
er-data-sample
description
Use when writing the data and sample section of an Economic-Research manuscript — naming databases, building variable-definition and descriptive-statistics tables, and leaving an auditable sample-filtering trail to 发表级.

数据与样本(er-data-sample)

触发时机

  • 正文有「数据与样本」一节但只写了「数据来源于公开渠道」「样本为 A 股上市公司」一句话
  • 变量定义表用文字描述(「反映税负偏差」)而不给计算公式
  • 描述性统计表均值 / 极值看着别扭,但正文没解释
  • 样本从原始库到回归样本怎么筛的,自己都说不清,更别说审稿人复现
  • 审稿人质疑:核心变量这么度量合理吗?样本代表性?有没有选择偏误?

配套代码:resources/code/stata/01_clean.do(清洗 + 筛选留痕)、 resources/code/stata/02_descriptive.do(描述统计 + 变量表)。 样本筛选每一步须可在代码复现,呼应 er-reproducibility。

数据说明段落规范

「数据与样本」开头第一段约 200 字,固定四块:时间跨度 + 数据库(点名)+ 样本范围 + N;筛选标准;缩尾处理;多源合并键。模板:

本文使用 2008—2022 年中国 A 股上市公司年度数据,财务数据来自国泰安(CSMAR)
数据库,专利数据来自中国研究数据服务平台(CNRDS),城市层面变量取自《中国城市
统计年鉴》。样本筛选:(1)剔除金融业(证监会行业 J 门类);(2)剔除 ST、*ST 及
退市公司;(3)剔除核心变量缺失的观测;(4)剔除资产负债率大于 1 的异常样本。最终
得到 2,841 家公司、共 28,317 个公司—年度观测的非平衡面板。为消除极端值影响,对所有
连续变量在上下 1% 分位进行缩尾(winsorize)处理。多源数据以「股票代码 + 年份」为
键合并,公司与城市数据按公司注册城市代码匹配。
  • 数据库必须点名:国泰安CSMAR、Wind、CNRDS、中国工业企业数据库、中国海关数据库、全国税收调查、CHFS、CHARLS、CFPS。微观调查数据注明调查年份与抽样框。
  • 禁忌:写「数据来源于公开渠道」「相关数据库」。审稿人据此无法判断口径,等同没说。
  • 时间跨度给起止理由(如政策实施年、数据可得性截止年),不要只甩一个区间。

变量定义表规范

每个变量有且仅有一行;定义给计算公式而非文字描述;数据来源精确到数据库名。分四类排列:被解释变量 / 核心解释变量 / 控制变量 / 工具变量。

类别变量符号定义(计算公式)数据来源
被解释变量企业避税BTD=(税前会计利润−应纳税所得额)/ 期末总资产CSMAR 财务报表
核心解释变量税收执法强度Enforce=实际税负−预期税负(行业—地区回归残差)全国税收调查
控制变量企业规模Size=ln(期末总资产)CSMAR
控制变量资产负债率Lev=总负债 / 总资产CSMAR
工具变量政策冲击IV_reform=2002 年所得税分享改革后注册=1,否则=0作者手工整理
  • 写公式:=实际税负-预期税负,不写「反映税负偏差」这类描述。
  • 衍生变量注明上游字段或回归来源(如「行业—地区回归残差」),与 01_clean.do 的 gen 一一对应。
  • 表注列明:缩尾口径、单位、对数化的变量、虚拟变量取值含义。

描述性统计表规范

报告均值 / 标准差 / 最小值 / p25 / 中位数 / p75 / 最大值 / N;连续变量为缩尾后数值;变量顺序与定义表完全一致(一一呼应)。

  • 异常的均值 / 极值要在正文解释:例如核心解释变量均值接近 0(残差类变量正常)、某控制变量最大值偏高(已缩尾后仍高,说明行业特性)。
  • 虚拟变量报告均值即组占比;正文点出处理组 / 对照组样本比例是否失衡。
  • 若分组(处理 vs 对照、改革前 vs 后),加分组均值与差异检验,为识别铺垫。
  • N 与数据说明段落的最终观测数一致;若个别变量 N 偏小,说明缺失来源。

样本筛选留痕

每一步筛选可追溯、可在代码复现,正文给「漏斗」式交代,代码留痕呼应 er-reproducibility:

stata
* 01_clean.do —— 样本筛选漏斗,每步记录剩余观测数
use "$data/raw/csmar_firm.dta", clear
count                                          // 原始:512,043
drop if inlist(ind_code,"J")                   // 剔除金融业
drop if st_flag==1                             // 剔除 ST/*ST/退市
drop if missing(btd, enforce, size, lev)       // 剔除核心变量缺失
drop if lev>1 & !missing(lev)                  // 剔除资不抵债异常
winsor2 btd enforce size lev, cuts(1 99) replace  // 上下 1% 缩尾
count                                          // 最终:28,317
  • 正文交代是非平衡面板还是平衡面板,以及为何(强平衡会损失大量样本则说明)。
  • 多源合并报告匹配率:如「专利数据成功匹配 26,108 个观测,匹配率 92.2%,未匹配主要为当年无专利申请企业」。
  • 合并键与口径写清(股票代码 vs 公司全称 vs 统一社会信用代码),跨库口径不一致须说明清洗规则。

审稿人高频质疑预防

  • 度量合理性:核心变量为何这么算?给文献依据(如某算法源自 某作者,年份)+ 至少 1 个替代度量留作稳健性(呼应 er-robustness)。
  • 样本代表性:样本占总体比例、行业 / 地区 / 年份分布是否偏;若用子样本(如仅制造业),论证不损外部有效性或明确限定结论边界。
  • 选择偏误:筛选是否系统性排除某类企业(如剔除缺失值是否与被解释变量相关);必要时报告 Heckman 两步 / 与全样本的均值对比,预判而非等审稿人问。

必查清单

  • 数据说明段落点名具体数据库(不写「公开渠道」),含时间跨度起止理由 + 最终 N
  • 四块齐全:库 + 范围 + N、筛选标准、缩尾口径、合并键
  • 变量定义表每变量一行,定义为计算公式,来源精确到库名,四类分组
  • 描述统计含均值 / 标准差 / 分位数 / N,连续变量为缩尾后
  • 异常均值 / 极值在正文有解释
  • 描述统计变量顺序与定义表一致,N 与数据段落一致
  • 样本筛选漏斗可在 01_clean.do 复现,每步剩余观测数留痕
  • 交代非平衡 / 平衡面板,多源合并报告匹配率
  • 核心变量度量有文献依据 + 替代度量;样本代表性与选择偏误已预判

反模式

  • 数据来源含糊:「数据来源于公开渠道」「相关数据库」「Wind 等」——必须点名到具体库
  • 变量定义用文字描述(「反映企业创新水平」)而不给公式
  • 描述统计出现明显异常均值 / 极值却不解释,把疑虑留给审稿人
  • 筛选标准不透明:只说「经过筛选得到 N 个样本」,无漏斗、代码无法复现
  • 缩尾只说「做了处理」不给分位(1%/99%?5%/95%?)
  • 描述统计与变量定义表变量不对应、N 对不上
  • 在正文(理论、机制、稳健性各处)反复重新定义同一变量,口径还彼此打架

输出格式

【数据说明段落】四块齐全 / 缺:[库点名 / N / 筛选 / 缩尾 / 合并键]
【数据库点名】具体(CSMAR / CNRDS / ...)/ 含糊待改:[...]
【变量定义表】公式化且四类分组 / 问题:[某变量用描述/缺来源/缺类别]
【描述统计】合规(缩尾后, 含分位数)/ 异常未解释:[变量]
【表—文呼应】一致 / 不一致:[顺序 or N 对不上]
【筛选留痕】漏斗可复现 / 不透明:[缺步骤]
【面板与匹配】非平衡/平衡 已交代 + 匹配率 X% / 缺
【质疑预防】度量依据 / 代表性 / 选择偏误:[已备 / 待补]
【下一步】数据与样本扎实 → er-identification 落识别策略与主回归

参考

  • 著录采用著者—出版年制;正文引用如(范子英、田彬彬,2013),不使用 [1][2] / [J][M] 序号制。

附属资源

© brycewang-stanford, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in Economic-Research-Journal-Skills/skills/er-data-sample of brycewang-stanford/Awesome-Journal-Skills.

Open the folder on GitHubat commit 932eb23

Compare with similar skills

Er Data Sample next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Er Data Sample compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Er Data Sample this skillbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.1kAutomated safety check: PassMIT
Manuscript Statistics AuditYuan1z0825/nature-skills46k2 repos~2.1kAutomated safety check: PassApache-2.0
Math Modeling SolverLupynow/math-modeling-skills416—~2.1kAutomated safety check: PassMIT
JS Perf InvestigationSAP/project-foxhound1801 repos~4.1kAutomated safety check: PassGPL-3.0
Academic Paper Reproduction Methodologyxjtulyc/MedgeClaw6171 repos~1.3kAutomated safety check: PassNone
Data Scientistmagnus919/hermes-profiles278—~3.3kAutomated safety check: PassMIT

Similar skills

  • Manuscript Statistics Audit

    Yuan1z0825/nature-skills

    Audits or rewrites the statistical reporting in a manuscript: experimental units, replication, tests, uncertainty and figure legends, without inventing missing details.

    46k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Math Modeling Solver

    Lupynow/math-modeling-skills

    数学建模竞赛解题全流程指导。覆盖国赛(CUMCM)和美赛(MCM/ICM)全部题型(A-F),提供12种问题本质分析、95+场景模型决策矩阵、5本算法Cookbook、11本完整例题Playbook、22个Python+7个MATLAB可运行代码模板。与math-modeling-paper形成"解题→写作"配对。当用户提及建模思路、选什么模型、怎么建模、赛题求解、粘贴赛题文本、美赛/国赛题目分…

    416 GitHub stars~2.1k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • JS Perf Investigation

    SAP/project-foxhound

    Official

    Structured performance opportunity investigation for SpiderMonkey (the Firefox JavaScript engine).

    180 GitHub starsUsed in 1 repo~4.1k tokens
    Research & ScienceAuto-check passed
  • Six-phase process for reproducing a published paper's results from provided data, from variable mapping and sample filtering through regression tables and a written report.

    617 GitHub starsUsed in 1 repo~1.3k tokens
    Research & ScienceAuto-check passed
  • Data Scientist

    magnus919/hermes-profiles

    PhD-level expertise in data science, statistics, and machine learning.

    278 GitHub stars~3.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    32k GitHub starsUsed in 12 repos~4k tokens
    Research & ScienceAuto-check passed

More from brycewang-stanford/Awesome-Journal-Skills

All 2,387 skills in this repo
  • Aaag Data Analysis

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running and reporting the analysis for an Annals of the American Association of Geographers manuscript — spatial statistics and modeling, remote-sensing accuracy, or…

    1.2k GitHub stars~1.3k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Literature Positioning

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when positioning an Annals of the American Association of Geographers manuscript in the literature — engaging geographic scholarship across the relevant area and the…

    1.2k GitHub stars~1.3k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Rebuttal

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when responding to an Annals of the American Association of Geographers decision letter (major/minor revision) — building a point-by-point response to the subject editor and…

    1.2k GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Research Design

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when defending the research design of an Annals of the American Association of Geographers manuscript — spatial/quantitative analysis and GIScience, remote-sensing and…

    1.2k GitHub stars~1.4k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Review Process

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when you need to understand how the Annals of the American Association of Geographers evaluates a manuscript — double-anonymous review routed through a subject editor by…

    1.2k GitHub stars~1.3k tokensUpdated 10 days ago
    Auto-check passed
  • Aaag Submission

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when running the final pre-submission preflight for the Annals of the American Association of Geographers via ScholarOne Manuscripts — area/article-type selection…

    1.2k GitHub stars~1.6k tokensUpdated 10 days ago
    Auto-check passed

Questions about Er Data Sample

What does Er Data Sample do?

A skill your agent uses when writing the data and sample section of an Economic-Research manuscript — naming databases, building variable-definition and descriptive-statistics tables, and leaving an…. Er Data Sample is an agent skill from brycewang-stanford/Awesome-Journal-Skills. Use when writing the data and sample section of an Economic-Research manuscript — naming databases, building variable-definition and descriptive-statistics tables, and leaving an auditable sample-filtering trail to 发表级.

When should I use Er Data Sample?

Er Data Sample fits situations like: writing the data and sample section of an Economic-Research manuscript — naming databases; building variable-definition and descriptive-statistics tables; leaving an auditable sample-filtering trail to 发表级.

How do I install Er Data Sample in Claude Code?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill er-data-sample -a claude-code`. Or copy the skill folder (Economic-Research-Journal-Skills/skills/er-data-sample in brycewang-stanford/Awesome-Journal-Skills) into .claude/skills/er-data-sample in your project. Claude Code loads it when a task matches its description.

How do I install Er Data Sample in Codex?

Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill er-data-sample -a codex`. Or copy the skill folder (Economic-Research-Journal-Skills/skills/er-data-sample in brycewang-stanford/Awesome-Journal-Skills) into .agents/skills/er-data-sample in your project. Codex loads it when a task matches its description.

Can I use Er Data Sample in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill er-data-sample -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/er-data-sample, .gemini/skills/er-data-sample, .github/skills/er-data-sample and .opencode/skills/er-data-sample in your project.

What does Er Data Sample need to run?

SKILL.md names no scripts, command-line tools or credentials: Er Data Sample is instructions for the agent only.

Does Er Data Sample access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Er Data Sample safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Er Data Sample use?

Er Data Sample is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Er Data Sample use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Er Data Sample?

Skills that share tags, products or a category with Er Data Sample: Manuscript Statistics Audit (Yuan1z0825/nature-skills, 46k stars), Math Modeling Solver (Lupynow/math-modeling-skills, 416 stars), JS Perf Investigation (SAP/project-foxhound, 180 stars) and Academic Paper Reproduction Methodology (xjtulyc/MedgeClaw, 617 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Er Data Sample?

brycewang-stanford (a GitHub user) maintains it in brycewang-stanford/Awesome-Journal-Skills, which has 1,216 GitHub stars. The repository holds 2,387 skills in this directory. The repository was last updated on September 27, 2026.

Source: brycewang-stanford/Awesome-Journal-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.