Research Methodology
chekusu/wanman
Methodology for market research and data collection, ensuring data quality and source traceability
Finds usable public datasets, judges whether the data can support a research idea, and checks train and test splits for leakage before results are trusted.
SKILL.md written in Chinese; this summary is our English description.
$ npx skills add Light0305/Light-skills --skill light-data-engineering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Light0305/Light-skills light-data-engineering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/light-data-engineering .claude/skills/light-data-engineering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "light-data-engineering" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-data-engineering into .claude/skills/light-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-data-engineering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Light0305/Light-skills/tree/master/skills/light-data-engineeringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Light0305/Light-skills --skill light-data-engineering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Light0305/Light-skills light-data-engineering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/light-data-engineering .agents/skills/light-data-engineering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "light-data-engineering" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-data-engineering into .agents/skills/light-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-data-engineering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Light0305/Light-skills --skill light-data-engineering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Light0305/Light-skills light-data-engineering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/light-data-engineering .cursor/skills/light-data-engineering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "light-data-engineering" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-data-engineering into .cursor/skills/light-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-data-engineering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Light0305/Light-skills.git --path skills/light-data-engineering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Light0305/Light-skills --skill light-data-engineering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Light0305/Light-skills light-data-engineering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/light-data-engineering .gemini/skills/light-data-engineering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "light-data-engineering" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-data-engineering into .gemini/skills/light-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-data-engineering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Light0305/Light-skills light-data-engineeringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Light0305/Light-skills --skill light-data-engineering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/light-data-engineering .github/skills/light-data-engineering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "light-data-engineering" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-data-engineering into .github/skills/light-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-data-engineering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Light0305/Light-skills --skill light-data-engineering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Light0305/Light-skills light-data-engineering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/light-data-engineering .opencode/skills/light-data-engineering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "light-data-engineering" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-data-engineering into .opencode/skills/light-data-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-data-engineering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
light-data-engineeringFinds usable public datasets, judges whether the data can support a research idea, and checks train and test splits for leakage before results are trusted.
This is the second step of a multi-stage research pipeline, written in Chinese. It answers two questions before an idea is finalized: whether the available data is large and clean enough to support the study, and whether the way it was split hides leakage that would inflate results. Leaks such as normalizing before splitting, looking ahead in time series, entity overlap between train and test, and target-encoding leakage are treated as blocking failures.
For finding data, the agent first freezes the task, unit of observation, minimum size, licence and storage budget. It then uses dataset_intake.py to check public candidates from hosts such as HF, OpenML, UCI and Kaggle for access, licence, revision, size, splits and dataset card, flagging gated, tag-only licence or sensitive-label items for review. Rows are sampled and the revision and a SHA256 recorded before any full download. Further scripts cover a feasibility gate, a quality gate, drift checks, derived evaluation sets for generalization and sensitivity tests, and Croissant metadata export; a data card template and a resource map are included.
Both gates report to a checkpoint command, and a critical finding makes it exit with an error, while warnings such as a tight sample size or an awkward split do not block. When the data cannot support an idea, the skill sends the work back to the idea stage with the gap and a fix, and stops to ask you. The statistical-power screen is a rough rule of thumb, not a formal power analysis, and leakage detection is heuristic.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 6b44f57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 10 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Research Data Feasibility and Leakage Checks loads about 4.9k tokens when it runs, and up to ~7.7k if it reads all its reference files. Until then it costs about 153 tokens; SKILL.md has 1,078 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from Light0305/Light-skills at commit 6b44f57, republished under its MIT licence (© Light0305). 1,078 words, ~4,945 tokens.
.claude/skills/light-data-engineering/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.你是 Light 科研流水线的 DAG 第 2 节点。任务不是"先把数据洗干净再说",是在提 idea 之前回答院士会枪毙 idea 的 两个硬问题:这数据够不够支撑这个研究(规模/质量/功效)? 和 这套划分有没有藏着让结果虚高的数据泄漏? 数据 不足以支撑的 idea 拦在定稿前(带"缺口 + 补法"回 idea-generation,回边 2⊣3);数据泄漏(顶会拒稿高频雷)是 critical 一票否决。
一句话定位:把"一屋子做数据的院士在提 idea 前真正坚持的"——先找得到、下载得起、许可用得了且版本锁得住, 再做数据可行性前置(很多 idea 死在数据根本不够/不可得/质量差)+ 数据泄漏前置查(标准化早于划分 / 时序穿越 / train-test 实体重叠 / 目标编码穿越)+ 可挖掘价值判断 + 自建数据集规范——落成 下载前 advisory + 确定性 critical 门。深度对标真相源 =
docs/competitors/data-engineering.md(11 个真同类 + 机制锚 + 诚实边界)。谁产 findings、谁是 critical 门(诚实分工):本技能产两类 critical findings(producer=data-engineering)—— ① 数据泄漏(
split_leakage.py→leak_findings.json,HIGH=critical);② 数据可行性不足/idea-killing (data_feasibility_gate.py,功效粗筛 insufficient / 四问 insufficient = critical)。均被run_checkpoint --stage 2聚合 → critical fail exit 1。warn 不阻断:样本量偏紧、划分不合理(spec §4.2 口径)。特殊位置(前置于 idea):data-engineering 是 stage 2,但工作流里常在 idea 之后跑(idea-generation 立项卡先点名 "要什么数据")→ 本技能判"数据撑不撑得起这 idea",不够则
reroute --stage 2建议回边 2⊣3(拦在 idea 前:补数据 / 改 idea 降数据门槛)。这是 idea-generation 立项卡"数据可行性必答字段"的前置守门方。是横切常驻吗? 否。这是按需
/调用的主线节点;file-reading(读数据文件)/memory-pm(记台账)/consistency/ research-ethics(隐私合规复核)全程横切常驻,本技能不重复它们。
references/data-resource-map.md 的 intake 闭环。run_checkpoint --stage 2);数据泄漏 / idea-killing 不足 → critical fail 确定性阻断。derive_eval_set.py 出鲁棒性/泛化/敏感性评测集。每个动作先归类:该自己做(ACT)、该停下问用户(ASK)、还是绝不(NEVER)?
dataset_intake.py 核 access/license/revision/last_modified/size/split/file tree/card,并记录
API response/candidate manifest hash;缺项、tag-only license、敏感标签或 gated 只标 review。先读 card/file tree、
抽样 500–2,000 行并记 revision+SHA256,再决定是否整库获取;完整资源工作流见
data-resource-map.md。data_doctor.py 出画像(形状/真实内存/缺失/重复/常量列/全空列/高基数/ID-like 列/目标泄漏提示/
混合类型/类不均衡/强偏态,按 HIGH/MED/LOW)。先看一眼数据长什么样,再谈可行性。data_feasibility_gate.py 编排 sample_size_check(经验功效粗筛)+
data_feasibility(四问)→ 产 light.findings.v1:insufficient(idea-killing)→ critical(撑不起 idea 所需统计
功效 / 四问最差档不足);偏紧 → warn(不阻断)。critical → run_checkpoint --stage 2 exit 1 → reroute 建议 2⊣3。safe_split.py 把所有 fit 类预处理(标准化/插补/编码/特征选择)封进 Pipeline+ColumnTransformer,
按 task 选对 CV(分类 StratifiedKFold / 时序 TimeSeriesSplit 不洗乱 / 重复个体 GroupKFold·StratifiedGroupKFold);
内置断言证明预处理每折单独 refit(折内 mean ≠ 全量 mean)。时序给 --time-col 升序校验防乱序穿越;group 用
--group-clf/--group-reg 显式声明(不靠 nunique 猜)。split_leakage.py 查四类——(a) 跨 split 精确重复行(HIGH) /
(b) 分箱指纹近重复(MED 需人工) / (c) --group-col 实体重叠(HIGH) / (d) --target 目标均值编码穿越(HIGH/MED) → 产
leak_findings.json(light.findings.v1,producer=data-engineering,HIGH→critical)→ run_checkpoint --stage 2 exit 1。quality_gate.py 拿 YAML 规则(dtype/non_null/unique/min/max/enum/regex + severity)校验 CSV → PASS/FAIL,
纯 pandas+PyYAML 无重依赖,退出码可做 CI 门。drift_check(KS+PSI,PSI 为主 p 为辅)/ check_access_level(raw 数据流向
public 被阻断)/ croissant_export(出 Croissant JSON-LD 元数据)/ derive_eval_set(research-plan 回边的派生评测集)。| 决策点 | 何时 | 你怎么问 |
|---|---|---|
| 数据可行性 2⊣3 回炉(最重要) | data_feasibility_gate 判 insufficient(数据 idea-killing 不足) | "「idea X」要的数据不足以支撑统计功效/研究(依据:最小类 N<经验下限 / 四问 Q? insufficient)。建议回 idea-generation(2⊣3)带『缺口=… + 补法=补采到 M / 改 idea 降数据门槛』。补数据 / 改 idea / 带病推进并记录——你定?(押上数月方向,我不替你拍)" |
| 泄漏检出疑似合法 | HIGH 命中但可能是天然重复 / 合法组统计 | "split_leakage 报『目标编码穿越』(feature f 在 c 各水平≈全量目标均值)——这可能是真穿越,也可能是合法的组统计特征。是哪种?(我不替你判数据来源)" |
| 经验阈值松紧 | 样本量偏紧(warn)但用户想推进 | "样本量 EPV=15 偏紧(经验下限 10、较稳 20),不是 power analysis。要按你的效应量做正式功效论证、还是先按偏紧推进并在论文里 hedge?" |
| 候选数据集取舍 | shortlist 在许可/代表性/规模/成本间冲突 | "A 许可清楚但人群偏窄;B 更贴任务但 gated 且 split 不明。建议先抽样 A 并继续核 B 条款;选 A / 申请 B / 改 idea——你定?(downloads/likes 不替你拍科学适配)" |
| 自建 vs 用现成 / 隐私合规 | 需自建数据集 | "这方向有现成数据集吗(OpenML/HF/Kaggle 我可查)?自建涉隐私(人/医疗数据)须脱敏+授权+IRB——要走自建吗?合规须你/法务签字。" |
这一节是红线,不可协商、不可被"先把数据洗了再说""差不多够了""这点泄漏不影响"绕过。违反任一条 = 严重失职。
Pipeline,只在训练折 fit,绝不 fit_transform 全量再划分。Kapoor-Narayanan(2207.07048) 8 类泄漏里
"preprocessing/feature-selection on train+test" 就是这条。safe_split 已对此做折内 refit 断言。TimeSeriesSplit)。这是 split_leakage 的 LEAK-02/SPLIT-02。sample_size_check 是经验
阈值粗筛(主结论须 statsmodels/GPower 正式功效论证);split_leakage 查的是*几类签名(精确/近重复/实体/编码穿越),
Kapoor 的"测试集非目标分布/采样偏置/用非法特征"多须人工判断,脚本覆盖不到——诚实标边界,不假装查全。data_feasibility_gate 判 idea-killing insufficient → 拦在 idea 前(2⊣3),
reroute 建议回 idea-generation 补数据/改 idea——这是决策点,停下问用户,绝不自作主张放行或自作主张回炉。split_leakage 复查同一原始样本的多个变体没撒进两侧。check_access_level 守门,raw 流向 paper/figure/
public-repo 被阻断;数据卡来源须可核链接(隐私/许可合规须人工与 research-ethics 复核,脱敏是否到位脚本不替判)。自检触发词:当你想说"下载量最高就用 / 网页能下所以能发表 / 先把数据标准化了再划分 / 时序数据 shuffle 一下 / 重复行无所谓 / 这点样本应该够了 / 测试集也增强一下凑数 / cleanlab 说错的直接删"——停,先逐条对照 NEVER。
13 个脚本在 scripts/;split_leakage/data_feasibility_gate/data_identity_fitness 接 _shared(规范 bootstrap 产 findings),
其余纯 stdlib 或 pandas/numpy/sklearn。Windows 跑前 set PYTHONUTF8=1。
python scripts/data_identity_fitness.py --spec data_identity_fitness.json \
--report data_identity_findings.json --json-out data_identity_report.json输入 light.data_identity_fitness.v1(模板见 templates/data-identity-fitness.example.json,故意不完整,直接跑应 exit 1):锁定 as_of、dataset_id/version/source_locator/snapshot_at/raw_sha256,逐项核 license/consent/DUA/ethics_review,登记 raw→clean→split 衍生链,统一 TIME_CROSSOVER/GROUP_OVERLAP/ENTITY_OVERLAP/PREPROCESSING_BEFORE_SPLIT/TARGET_LEAKAGE/NEAR_DUPLICATE/AUGMENTATION_LEAK threat matrix,并给 measurement_quality/label_quality/missingness/sample_power/bias/staleness 适用性裁定。UNKNOWN 不是 pass:权限未知、split threat 未排除、stale 无 impact、DERIVED 无血缘、未来 snapshot_at/created_at/data_valid_at、存在 blocker 却声明 FIT 均 critical fail。decision=NOT_FIT/UNKNOWN 本身阻断推进;只有无 blocker 且限制已下沉时,才可 FIT_WITH_LIMITATIONS。
python scripts/dataset_intake.py --query "breast cancer" --limit 10 --sort downloads \
--report data_candidates.json
python scripts/dataset_intake.py --inspect scikit-learn/breast-cancer-wisconsin输出 light.data_candidates.v1,含 raw_response_sha256 与 candidate_manifest_sha256;缺
license/revision/last_modified/size/split/file tree/card、license 只来自 tag、gated/private 或命中医疗/隐私/人类等
sensitive tags → review,API 失败 →
UNAVAILABLE exit 2。它不产 light.findings.v1、不进入 STAGE_GATES;完整闭环见
references/data-resource-map.md。
# 规模够不够支撑 idea 所需统计功效(经验粗筛)+ 四问 → critical/warn findings:
python scripts/data_feasibility_gate.py --spec feasibility_spec.json --report feas_findings.json # insufficient → exit 1
# 交总控聚合(stage 2 确认点,critical fail → exit 1 确定性阻断):
python ../light-orchestrator/scripts/run_checkpoint.py --file .light/passport.yaml --stage 2 \
--findings feas_findings.json --write --ts 2026-06-18T11:00
# fail → 根因回炉建议(命中 ROUTES[2],建议 2⊣3:拦在 idea 前,只建议不执行,停下问用户):
python ../light-orchestrator/scripts/reroute.py --findings feas_findings.json --stage 2 \
--passport .light/passport.yaml
# 用户拍板回炉后落账:
python ../light-orchestrator/scripts/passport.py add-back-edge --to 3 --from 2 \
--root-cause "数据不足以支撑 idea 所需统计功效" --evidence-ptr "<reroute 给的指针>"feasibility_spec.json:{project, idea, sample{task,n,classes,features,positives,per_class}, feasibility{sufficiency, quality,feature_value}}(scale 缺省由 sample 自动回填)。spec 源自 idea-generation 立项卡的"数据可行性必答字段"。
# 四类泄漏审计 → leak_findings.json(HIGH=critical):
python scripts/split_leakage.py --train train.csv --test test.csv --group-col user_id --target y \
--out leak_audit.md --findings leak_findings.json # 任一 HIGH → exit 1
# 单文件带 split 列:--csv data.csv --split-col split
# 交总控聚合(critical fail → exit 1 阻断;泄漏在 stage 2 内修复,非跨阶段回边):
python ../light-orchestrator/scripts/run_checkpoint.py --file .light/passport.yaml --stage 2 \
--findings leak_findings.json --write --ts 2026-06-18T11:30python scripts/data_doctor.py --csv data.csv --target y --out report.md # 体检画像(先做)
python scripts/safe_split.py --csv data.csv --target y --task group --group-col user_id --group-clf # Pipeline+CV 折内 refit
python scripts/quality_gate.py --csv data.csv --rules rules.yaml --out gate.md # YAML 数据门禁(CI)
python scripts/sample_size_check.py --task clf --n 1200 --classes 3 --features 20 # 经验功效粗筛(非 power analysis)
python scripts/data_feasibility.py --project X --q1 ok:... --scale-json size.json --q4 ok:... --out data_feasibility.md
python scripts/drift_check.py --ref train.csv --cur test.csv --out drift.md # KS+PSI(PSI 为主 p 为辅)
python scripts/check_access_level.py --level raw --sink paper # raw→public 阻断
python scripts/croissant_export.py --in card_fields.json --out ds.croissant.json # Croissant 元数据
python scripts/derive_eval_set.py --base data.csv --spec derive_spec.json --outdir derived/ # research-plan 回边各脚本 --selftest/--help 即接口;资源闭环见 data-resource-map.md,
逐工具 API/已知坑见 references.md。
提 idea 之前先问四问:这 idea 要什么数据?规模/质量/标注够不够支撑统计显著?sample_size_check 给经验粗筛
(分类每类 ≥50 偏紧/≥100 较稳;回归 EPV 样本/特征 ≥10/≥20;二分类正例 EPV,Peduzzi 1996),data_feasibility 四问取
最差档。数据 idea-killing 不足 = critical 前置门,reroute 建议 2⊣3(拦在 idea 前,补数据/改 idea)——别让一个
数据撑不起的 idea 押上数月。
Kapoor-Narayanan(Patterns 2023, 2207.07048) survey 出 17 个领域 329 篇论文因泄漏结论过度乐观,给 8 类泄漏。
本技能查可机检的几类(对标 Deepchecks TrainTestSamplesMix/DateTrainTestLeakage*/IndexTrainTestLeakage):
safe_split 折内 refit 杜绝。TimeSeriesSplit + --time-col 升序校验。split_leakage --group-col,GroupKFold 防。split_leakage --target 查签名。
任一 HIGH→critical→exit 1。这是 stage 2 的 STAGE_GATES(leakage)。data_doctor 画像 + 四问 Q4:特征-目标关系是否真实(非 ID-like 误用、非目标泄漏)、有没有可建模的结构。data-centric
视角(DataPerf):改数据有时胜过堆模型。这是定性判断 + 画像启发,不是可比分数。
templates/annotation_guide.md(类目定义/边界规则/LLM 辅助+人工审核闭环/质检抽样率/IAA)+ assets/data_card_template.md
(对齐 Datasheets for Datasets / HF Dataset Card / Croissant:动机/构成/采集/标注/用途/分发/维护 + 偏差·隐私·访问分级·溯源)。
标注质量:IAA(sklearn cohen_kappa_score / statsmodels fleiss_kappa)评流程整体 + cleanlab 置信学习定位具体可疑样本
(人裁定 top-K 不全自动删)。隐私/许可合规须人工与 research-ethics 复核。
dataset_intake 报告有 raw_response_sha256 与 candidate_manifest_sha256 吗?抽样前记 URL,整库前记 SHA256 了吗?data_identity_fitness.py 吗?数据身份、权限链、衍生链、split threat matrix、fitness 与 staleness 都有 locator/hash/impact 吗?as_of、snapshot_at、split created_at、data_valid_at 是真实核验/冻结时间吗?没有预填未来日期吧?review/unresolved 吧?Pipeline 只训练折 fit 了吗?没在划分前 fit_transform 全量吧?TimeSeriesSplit 没洗乱吧?重复个体用 GroupKFold 防实体重叠了吗?split_leakage 查了四类泄漏吗?HIGH 命中是真污染还是疑似合法(停下问用户)?check_access_level)?数据来源可核、没编造 DOI 吧?真增量(v2 兑现,已 selftest):⓪ 数据身份/权限/血缘/split threat/fitness 统一契约(data_identity_fitness.py,Round 3 新增)——产 light.findings.v1,把 license/consent/DUA/ethics、raw SHA256、衍生链、split threat matrix、staleness impact 与 FIT/FIT_WITH_LIMITATIONS/NOT_FIT/UNKNOWN 对齐;UNKNOWN/禁止/未排除威胁、未来时间、NOT_FIT/UNKNOWN 裁定不会被冒充 pass。① 数据候选下载前 intake(dataset_intake.py,Round 2 新增/Round 3 续补)——HF 公开 API
元数据归一为 light.data_candidates.v1,核 access/license/revision/last_modified/size/split/file tree/card,记录
raw_response_sha256 与 candidate_manifest_sha256;缺项、tag-only license、敏感标签或 gated/private 诚实 review,不冒充 usable、
不扩大 critical 面。② 数据泄漏 critical 门 producer(split_leakage.py 港 v1,v2 修硬编码
../../_shared→规范 bootstrap + producer m02→data-engineering)——四类泄漏 → leak_findings.json(light.findings.v1,
HIGH=critical),被 run_checkpoint --stage 2 聚合 exit 1;输出名正是 STAGE_GATES[2] 引用的标准件。③ 数据可行性
前置 critical 门 producer(data_feasibility_gate.py,v2 净新增接线)——编排港来的 sample_size_check+data_feasibility
(v1 纯工具、零接 _shared,grep 实证)产 critical/warn findings,insufficient → reroute 命中 ROUTES[2] 建议 2⊣3
(拦在 idea 前)。④ 防 fit 穿越的 Pipeline+CV(safe_split 折内 refit 断言)。⑤ 零重依赖数据门(quality_gate 是
GX 哲学的轻量同构,纯 pandas+PyYAML)。
裸模型本就会的(不吹):"数据要先划分再标准化""注意别泄漏""样本量要够""数据集要写卡"——裸 Opus 都会。本技能价值 = ① 把防泄漏落成确定性机读门 + 折内 refit 断言(裸模型会嘴上说不泄漏、手上还是全量 fit);② 数据可行性前置于 idea 定稿 + 2⊣3 回边(裸模型不会"拦在 idea 前"喂回 idea 阶段);③ 机读 critical findings + 确定性阻断 + 根因回炉 (裸模型给口头结论,编排器读不了、阻断不了)。
诚实落后项(已知没做到):
metadata-ready 只表示字段
较齐,不表示任务适配、许可终判、无隐私/无泄漏。HF API 失败时 exit 2 + UNAVAILABLE,不用旧缓存冒充当天真值。split_leakage 查几类签名(精确/近重复/实体/目标编码穿越),≠"查全了所有泄漏";
近重复靠分箱指纹(巧合会误报,标 MED 需人工);目标编码穿越的"合法组统计"也可能命中(须人工核来源)。Kapoor 的
"测试集非目标分布/采样偏置/用非法特征"多须人工判断,脚本覆盖不到。sample_size_check 是领域经验下限粗筛(每类样本/EPV/检测实例),不替代
效应量+显著性+功效的正式论证(statsmodels/G*Power)。阈值经验默认、可调。data_feasibility_gate 只聚合判据 + 出 findings,不替你判"数据到底
够不够"——GIGO(输入的四问/规模参数错,结论就错)。data_doctor 是粗筛画像非完整
EDA;漂移/质量门用轻量自写实现(drift_check 纯 numpy 渐近 p、quality_gate 无 GX 重依赖),表达力不及 Deepchecks/GX/
ydata-profiling,重场景仍建议用专业库(references 有真实端点)。check_access_level 按声明判流向(真脱敏须人工+research-ethics);
croissant_export 出关键层(完整 spec 校验/Hub 上传留外部工具)。emit_artifacts.py 未港——其"标准工件名 + passport
登记"在 v2 归 memory-pm pm.py / orchestrator passport.py append-stage,不重造(标准工件名约定见本 SKILL「产出」)。
v1 的 code_assets/ 共享统计库(stats_tests/agreement)v2 未港,统计/κ 用 statsmodels/sklearn 直接做。标准产出工件:
data_feasibility.md(交 idea-generation/idea-critique,前置 2⊣3)·leak_findings.json(泄漏 critical 门)·quality_report.md/data_card.md(交 research-plan/experiment-coding 做实验)。落.light/,passport 登记交 memory-pm。
docs/competitors/data-engineering.md(11 个同类一手核 + 横切可借 + 超越点 + 诚实边界)references/data-resource-map.md(需求卡→多源发现→preflight→抽样→血缘→门→发布;access 分级)references.md(pandas/polars/DuckDB/Deepchecks/GX/ydata-profiling/cleanlab/OpenML/HF/Kaggle/数据增强 等)scripts/——各 --selftest/--help 即接口;split_leakage.py(泄漏 critical 门)· data_feasibility_gate.py(可行性 critical 门 + 2⊣3)是 findings 核心assets/data_card_template.md(datasheet)· templates/annotation_guide.md(标注规范+IAA)· examples/worked_example.md(山羊行为数据走查)· examples/rules.example.yaml · examples/derive_spec.example.json_shared/README.md(findings_schema · gate_runner · 规范 bootstrap)light-idea-generation(stage 3,立项卡"数据可行性必答字段",2⊣3 前置回边)· light-orchestrator/scripts/run_checkpoint.py(stage 2 聚合 critical fail→exit 1)· reroute.py(ROUTES[2] 建议 2⊣3)· research-plan(stage 5,派生评测集回边)© Light0305, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 22 other files (scripts, references, assets) in skills/light-data-engineering of Light0305/Light-skills.
Open the folder on GitHubat commit 6b44f57
Research Data Feasibility and Leakage Checks next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Research Data Feasibility and Leakage Checks this skillLight0305/Light-skills | 640 | — | ~4.9k | Automated safety check: Pass | MIT | |
| Research Methodologychekusu/wanman | 688 | — | ~533 | Automated safety check: Pass | Apache-2.0 | |
| Data Quality Frameworkswshobson/agents | 40k | 10 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Datalineage Summarygoogle/skills | 21k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Monte Carlo Context Detectionsickn33/agentic-awesome-skills | 47k | 1 repos | ~2.6k | Automated safety check: Warn | MIT | |
| Bio Metabolomics Normalization QcGPTomics/bioSkills | 1.2k | 1 repos | ~5k | Automated safety check: Pass | MIT |
chekusu/wanman
Methodology for market research and data collection, ensuring data quality and source traceability
wshobson/agents
Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.
google/skills
Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS.
sickn33/agentic-awesome-skills
Route data-related requests to the right Monte Carlo skill or workflow.
GPTomics/bioSkills
Designs QC, corrects signal drift, removes batch effects, filters features, normalizes samples, and imputes missing values for untargeted LC-MS/GC-MS metabolomics, framing each step as a measurement…
aws/agent-toolkit-for-aws
Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables.
Light0305/Light-skills
Verifies that every reference in a manuscript is real, correctly identified and actually supports its claim, and produces a citation registry for typesetting.
Light0305/Light-skills
Coordinates and recovers multi-stage Light research projects from a single passport file, with checkpoints, stale-work tracking and rerouting only when you approve.
Light0305/Light-skills
Builds an evidence-backed invention disclosure packet from a project or research result for attorney or patent-agent review, without giving legal advice.
Light0305/Light-skills
Audits, scaffolds and safely migrates research project folder structures, keeping existing repositories read-only until you approve exact moves from a plan.
Light0305/Light-skills
Prepares draft materials for a China software copyright registration from a real project: application worksheet, source deposit plan, operation manual and consistency checks.
Light0305/Light-skills
Evidence-based workflow for designing or modernizing a software system: current-state inventory, options, API and schema contracts, migration plans, ADRs and verification.
Categories
Finds usable public datasets, judges whether the data can support a research idea, and checks train and test splits for leakage before results are trusted. This is the second step of a multi-stage research pipeline, written in Chinese. It answers two questions before an idea is finalized: whether the available data is large and clean enough to support the study, and whether the way it was split hides leakage that would inflate results.
Research Data Feasibility and Leakage Checks fits situations like: choosing and downloading a public dataset while checking its licence, version and size; judging whether a dataset is large and clean enough to support a research idea; suspecting that training and test data overlap or that a split leaks information; designing a train, validation and test split or a cross-validation scheme.
Run `npx skills add Light0305/Light-skills --skill light-data-engineering -a claude-code`. Or copy the skill folder (skills/light-data-engineering in Light0305/Light-skills) into .claude/skills/light-data-engineering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Light0305/Light-skills --skill light-data-engineering -a codex`. Or copy the skill folder (skills/light-data-engineering in Light0305/Light-skills) into .agents/skills/light-data-engineering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Light0305/Light-skills --skill light-data-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/light-data-engineering, .gemini/skills/light-data-engineering, .github/skills/light-data-engineering and .opencode/skills/light-data-engineering in your project.
Going by SKILL.md and its folder, Research Data Feasibility and Leakage Checks needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python to run the bundled scripts; The surrounding research pipeline's checkpoint command for the blocking gates.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Research Data Feasibility and Leakage Checks is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Research Data Feasibility and Leakage Checks: Research Methodology (chekusu/wanman, 688 stars), Data Quality Frameworks (wshobson/agents, 40k stars), Datalineage Summary (google/skills, 21k stars) and Monte Carlo Context Detection (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Light0305 (a GitHub user) maintains it in Light0305/Light-skills, which has 640 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on July 6, 2026.
Source: Light0305/Light-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.