Lammps Deepmd
Hello-QM/catgo-LRG
Run LAMMPS molecular dynamics with DeePMD-kit machine learning potentials.
Light 科研主线第 7 步·结果分析:不描述好坏、解释「为什么」,把每条结论绑死到 claim + 证据强度,并防 p-hacking。
$ npx skills add Light0305/Light-skills --skill light-result-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Light0305/Light-skills light-result-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/light-result-analysis .claude/skills/light-result-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "light-result-analysis" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-result-analysis into .claude/skills/light-result-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-result-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Light0305/Light-skills/tree/master/skills/light-result-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Light0305/Light-skills --skill light-result-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Light0305/Light-skills light-result-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/light-result-analysis .agents/skills/light-result-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "light-result-analysis" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-result-analysis into .agents/skills/light-result-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-result-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Light0305/Light-skills --skill light-result-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Light0305/Light-skills light-result-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/light-result-analysis .cursor/skills/light-result-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "light-result-analysis" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-result-analysis into .cursor/skills/light-result-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-result-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Light0305/Light-skills.git --path skills/light-result-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Light0305/Light-skills --skill light-result-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Light0305/Light-skills light-result-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/light-result-analysis .gemini/skills/light-result-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "light-result-analysis" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-result-analysis into .gemini/skills/light-result-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-result-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Light0305/Light-skills light-result-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Light0305/Light-skills --skill light-result-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/light-result-analysis .github/skills/light-result-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "light-result-analysis" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-result-analysis into .github/skills/light-result-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-result-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Light0305/Light-skills --skill light-result-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Light0305/Light-skills light-result-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Light0305/Light-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/light-result-analysis .opencode/skills/light-result-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "light-result-analysis" agent skill from https://github.com/Light0305/Light-skills/tree/master/skills/light-result-analysis into .opencode/skills/light-result-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "light-result-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
light-result-analysisLight 科研主线第 7 步·结果分析:不描述好坏、解释「为什么」,把每条结论绑死到 claim + 证据强度,并防 p-hacking。
Light Result Analysis is an agent skill from Light0305/Light-skills. Light 科研主线第 7 步·结果分析:不描述好坏、解释「为什么」,把每条结论绑死到 claim + 证据强度,并防 p-hacking。 何时用:实验跑完要解读结果 / 问「这些数说明什么」/ 要做显著性检验 + 效应量 + 置信区间 + 多重比较校正 / 担心 p-hacking (多重比较不校正、选择性报告、HARKing) / 要给每条 claim 定证据强度供写作校准措辞 / 判结果支不支撑假设、可不可复现。 触发词:结果分析 / 解读数据 / 这些结果说明什么 / 显著性 / p 值 / 效应量 effect size / Cohen's d / 置信区间 CI / 多重比较 / BH-FDR / Bonferroni / 校正 / p-hacking / 选择性报告 / garden of forking paths / HARKing / 证据强度 / claim 证据绑定 / SHAP / 消融分析 / 切片分析 / 配对检验 / result analysis。 核心纪律:统计错误 / p-hacking = critical(spec…
Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files, including scripts and assets (for example `assets/result_analysis_report_template.md`, `examples/analysis_audit.example.json` and `examples/method_compatibility.example.json`).
It sits in Data & Analytics, covering Machine learning. The repository describes itself as: An AI workflow skill pack for research, competitions, and innovation projects. The licence is MIT.
9 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 6b44f57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 10 files in scripts/ (Python and R, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Light Result Analysis loads about 5.5k tokens when it runs. Until then it costs about 176 tokens; SKILL.md has 1,206 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from Light0305/Light-skills at commit 6b44f57, republished under its MIT licence (© Light0305). 1,206 words, ~5,461 tokens.
.claude/skills/light-result-analysis/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.你是 Light 科研流水线的 DAG 第 7 节点。任务不是「描述结果好不好」,是把执行出来的结果解释清「为什么」—— 哪些证明方法有效、哪些暴露问题、哪些异常要排查、哪些能成论文亮点——并把每条能写进论文的论断(claim)绑死到它的 统计证据 + 证据强度档,守住让结论不可信的红线:p-hacking(多重比较不校正 / 选择性报告 / HARKing / garden of forking paths)。统计错误/p-hacking = critical;过度解读、效应量缺失 = warn。显著性看 q 不看 p。
一句话定位:把「一屋子做实验的院士在看结果时真正死磕的」——这提升是统计显著还是噪声(效应量多大、CI 含不含 0、 多重比较校正没有)、换数据集/换种子还成立吗(稳健性、可复现)、每条 claim 配多强证据(强证据强措辞、弱证据 hedge、 不显著只能报「未见显著差异」)——落成确定性机读门 + critical findings + 证据强度档。 深度对标真相源 =
docs/competitors/result-analysis.md(Round 2:8 个真同类 SKILL + 机制锚 + 超越点 + 诚实边界);真实用户闭环见result-analysis-resource-map.md。谁产 findings、谁是 critical 门(诚实分工):本技能产统计严谨/证据强度 critical findings(producer=result-analysis,
stat_rigor_gate.py四 gate)——stat_validity(多重比较未校正/选择性报告→真重算 BH-FDR→critical)、hypothesis_support(假设被结果证否→critical)、reproducibility(多种子不稳→critical)被run_checkpoint --stage 7聚合 → critical fail exit 1;evidence_strength(证据档 + 过度解读/效应量缺失)= warn 不阻断 DAG(spec §4.2 口径)+ emitevidence_strength.json。与 research-ethics 的分工(evidence_contract 是桥):result-analysis 在 stage 7 定证据强度(产
evidence_strength.json: 每条 claim 的 q/效应量/CI → 证据档 strong/moderate/weak/none + 允许/禁止措辞);research-ethics 在 stage 8claim_evidence_bind查措辞是否超过证据(消费同一个evidence_strength.json)。本技能定强度、它查措辞,不重叠;_shared/evidence_contract是两者共用的桥。特殊位置(回炉发起方,与 experiment-coding 相反):experiment-coding 是 7→6 的回炉落点(被动接);result-analysis 是 7→5 + 7→6 两条回边的发起方(主动发)——判结果不支撑假设(findings 带「假设/支撑/效应」信号)→ 总控
reroute --stage 7建议 7→5 回 research-plan;判结果不可复现(带「种子/复现」信号)→ 建议 7→6 回 experiment-coding。这是本技能的 非线性核心:不是终点,是把结果送回上游修的枢纽。(p-hacking critical 则是 stage 7 内重做分析,reroute 给 manual。)是横切常驻吗? 否。这是按需
/调用的主线节点;file-reading / memory-pm / project-structure / consistency / research-ethics 全程横切常驻,本技能不重复它们。
run_manifest.md:多种子指标 + 产物路径)——主用法。evidence_strength.json)。每个动作先归类:该自己做(ACT)、该停下问用户(ASK)、还是绝不(NEVER)?
python scripts/analysis_plan_audit.py --spec analysis_audit.json --report analysis_audit_findings.json --json-out analysis_audit_full.json——核结果前 plan lock、统计单位/复杂设计、comparison family、
expected↔observed seed/fold/sample coverage,以及 raw result 的 hash/owner/time/run manifest/commit。method_compatibility.py 核对
domain_scope/input_modalities/task_types/requires_access/supported_dependence/labels;
已知不兼容 FAIL,条件缺失 UNRESOLVED,不得把“脚本能跑”误写成“方法适用”。python scripts/analyze_results.py results.csv --group method --metric acc f1——EDA(n/均值±std/中位/95%CI/
正态性)+ 按正态性与组数自动选检验(2 组正态→Welch t / 非正态→Mann-Whitney;≥3 组→先 Levene 方差齐性→ANOVA+Tukey
或 Welch-ANOVA / 非正态→Kruskal-Wallis)+ 每对 Cohen's d(Hedges 校正)+ BH-FDR 跨比较校正。共享种子/折加 --paired-by seed
走配对 t / Wilcoxon(功效更高)。给了 --paired-by 后,claim/evidence 只采用配对比较;独立样本结果仅留
advisory,不得生成 duplicate claims。--slice-by <col> 切片分析防聚合掩盖子群失败(小 n 切片自动标「待核查」)。python scripts/stat_rigor_gate.py --spec stat_spec.json --report stat_findings.json --evidence-out evidence_strength.json 编排 BH-FDR 真重算 + 消费 evidence_contract → 产 light.findings.v1:多重比较未校正/
选择性报告 / 假设证否 / 多种子不稳 → critical;过度解读/效应量缺失 → warn。critical → run_checkpoint --stage 7 exit 1。analyze_results.py --emit-claim-table(claim_evidence_table.md:每个比较↔检验/p/q/d/CI/n)+
--emit-evidence(evidence_strength.json:挂接 _shared/evidence_contract,q/效应量/CI→证据档+措辞档)。显著性一律以
BH-FDR 后 q 为准,不显著的比较标「不得声称更好」。significance_test.py(cohens_d/mean_diff_ci/bootstrap_ci/benjamini_hochberg/delong_two_auroc
比较同测试集两模型 AUROC 差是否显著)。只报 p 不报效应量 = 误用:p 小不代表差异大。python scripts/r_analysis_crosscheck.py --input results.csv --group method --metric acc --paired-by seed --out r_acc.csv。launcher 会找 RSCRIPT/PATH/Windows Program Files;base R 真算 paired t 或
independent Welch。R 不可用就明确返回 unavailable;复杂 mixed/repeated/nested 设计仍需专门模型。explain_shap.py(SHAP beeswarm/bar/waterfall,非因果,shap 缺失优雅降级)+
leakage_overfit_check.py(train/val/test gap + 特征-标签高相关泄漏 + 重复行)——指标好得反常先查泄漏。assets/result_analysis_report_template.md 每个发现写「现象→原因→证据→对论文的意义」+ 亮点/异常/待补实验清单。| 决策点 | 何时 | 你怎么问 |
|---|---|---|
| 回炉发起 7→5(不支撑假设)(最重要) | hypothesis_support 判主假设 grade=none | 「主假设 H1『新模块提升 acc』未被结果支撑:BH-FDR 后 q=0.2≥.05、CI=[-0.2,0.4] 含 0、效应量 d=0.1 过小。建议回 research-plan(7→5)重审假设/设计(也许 H 本就不成立,或需更强实验)。回炉带『哪条假设没撑住 + 对应效应量/CI』——还是带病推进 / 转已知局限?(方向你定,绝不 HARKing 删掉换个成功假设重报)」 |
| 回炉发起 7→6(不可复现) | reproducibility 判多种子 sign-flip/CV 过大 | 「结果不可复现:claim X 的效应跨 5 个种子 sign-flip([-0.31,+0.55]),换次跑结论会飘。建议回 experiment-coding(7→6)查种子覆盖 / 实现 bug,带『失败的复现证据』。还是多种子报均值±std 并 hedge?别把单次峰值当结论。」 |
| 查出 p-hacking | stat_validity 多重比较未校正/选择性报告 | 「扫到 5 个比较未做多重比较校正:4 个裸 p<.05 中 4 个经 BH-FDR 后 q≥.05(假阳性)。建议在 stage 7 内重做分析:对全部比较做 BH-FDR,显著性以校正后 q 为准。要我直接重算吗?校正后『显著』可能消失——那才是真值。」 |
| 证据弱却想强措辞 | evidence_strength 判 asserted_grade 强于实算档 | 「claim Y 证据档=weak(显著但小效应 d=0.2),但措辞写了『显著优于』。建议降到『在本实验中略优 / 初步提示』并加 hedge。强措辞会被 research-ethics 的 stage-8 措辞门拦。」 |
| 异常结果:排查 vs 当亮点 | 切片/某指标异常高或异常 | 「切片『夜间』acc 异常高(0.97 vs 整体 0.85)。可能是真亮点(该场景方法特别有效),也可能是泄漏/小样本(该切片 n=12)。建议先查泄漏 + 看 n 再定,别急着写进 contribution。先查哪个?」 |
这一节是红线,不可协商、不可被「结果好就行」「p<.05 就是显著」「先写强点好发」绕过。违反任一条 = 严重失职。
evidence_strength warn。stat_validity critical。evidence_strength.json 是跨技能「措辞不强于证据」的单一数据源,下游 paper-writing/research-ethics 据此卡。stat_validity critical。stat_rigor_gate 查「多重比较校没校正 / 假设统计上撑不撑得住 / 结果稳不稳」——
机检有边界,效应量解读、机制因果、外推性的终判仍需人 / 领域判断;SHAP 是模型关联非因果,绝不当因果证据。诚实标边界。seed/fold/sample 是否独立由设计决定,不由 CSV 行数决定;
repeated measures / clustered / nested CV / repeated holdout 不能靠简单 t 检验终判。comparison family 由结果前计划定义,
必须保留 stable family_id、planned/reported coverage;不得拆成多次调用规避校正。evidence_strength.json 当前只含统计强度与措辞上限,不等于完整 run provenance;
claim/metric 只能带 source file/run/commit 作为 consistency 候选,不得由 result-analysis 直接写入 canonical
.light/consistency。自检触发词:当你想说「p<.05 就够了别管校正 / 这个不显著但趋势对也算亮点 / 措辞强点好发 / 假设没撑住就换一个 / 跑了十个报最好那个 / SHAP 说明这个特征导致了结果」——停,八成踩了 NEVER 第 1/2/3/4/5/6 条,或漏了 ASK 的回炉/p-hacking/措辞决策。
当前 11 个 Python 脚本 + 1 个 R 脚本在 scripts/;stat_rigor_gate/result_card_gate/analyze_results/
analysis_plan_audit 接 _shared(规范 bootstrap),统计件优先复用 statsmodels/scipy,DeLong 港 v1
(statsmodels 没有)。R 路径只依赖 base R;高级 R 包逐项检测,不假装已装。Windows 跑前 set PYTHONUTF8=1。
python scripts/result_card_gate.py --spec result_card.json \
--report result_card_findings.json --json-out result_card_report.json输入 light.result_card.v1(模板见 templates/result-card.example.json,故意不完整,直接跑应 exit 1):每条 claim 必须绑定 as_of、target/analysis_set/missingness/provenance/assumption/guardrail_analysis/comparison_family/practical_threshold/effect/language;provenance 至少记录 source run IDs、run manifest/raw result/analysis code locator 与 SHA-256、computed_at、owner_skill;guardrail_analysis 必须显式说明是否适用,适用时绑定 guardrail evidence locator/SHA-256 和逐项 PASS/FAIL/WARN/UNKNOWN;locked_at/results_available_at/computed_at 不能来自未来,且 computed_at 不得早于结果可见时间;decision=REVISION_REQUIRED/UNKNOWN 本身阻断写作交接;decision_ledger 记录 PRE/POST 结果的排除、换指标、模型选择、切片、阈值等分析决策;sensitivity 明确 sensitivity vs supplementary。p=0.049/0.051 不许让叙事翻面:q≥.05 或 CI 含 0 只能写“未见显著差异/不确定”;guardrail FAIL/UNKNOWN 不得 CLAIM_READY;阈值附近必须披露敏感性;POST_RESULTS 的 EXCLUSION/METRIC/MODEL_CHOICE/SUBGROUP 不得冒充 CONFIRMATORY。
# 编排 BH-FDR 真重算 + 消费 evidence_contract → light.findings.v1 + evidence_strength.json:
python scripts/stat_rigor_gate.py --spec stat_spec.json --report stat_findings.json --evidence-out evidence_strength.json
# 交总控聚合(stage 7 确认点,critical fail → exit 1 确定性阻断):
python ../light-orchestrator/scripts/run_checkpoint.py --file .light/passport.yaml --stage 7 \
--findings stat_findings.json --write --ts 2026-06-20T10:00
# p-hacking critical = 在 stage 7 内重做分析(reroute 给 manual);不支撑假设/不可复现 = 发起回边(见 ③)。stat_spec.json:{project, claims[{claim_id,p,q_fdr?,effect_size,ci95,n,asserted_grade?,is_hypothesis?,hypothesis_id?,seeds?}], correction(none/bh/bonferroni), comparisons_run?, comparisons_reported?, results_csv?}。示例见 examples/stat_spec.example.json。
python scripts/analyze_results.py results.csv --group method --metric acc f1 \
--paired-by seed --slice-by subgroup --emit-claim-table --emit-evidence # 自动选检验+效应量+FDR+两工件
python scripts/significance_test.py --selftest # p/d/CI/FDR/DeLong 函数库(对齐 scipy/statsmodels)
python scripts/leakage_overfit_check.py --train tr.csv --test te.csv --target y # 泄漏/过拟合体检
python scripts/explain_shap.py # SHAP 三图(非因果;shap 缺失优雅降级 exit 0)# warn-only:缺计划锁/复杂设计/family/coverage/provenance 显式出现,但不扩大 critical 阻断面
python scripts/analysis_plan_audit.py --spec analysis_audit.json \
--report analysis_audit_findings.json --json-out analysis_audit_full.json
# 声明方法能力与真实研究条件;FAIL/UNRESOLVED 都不得直接用该方法终判
python scripts/method_compatibility.py --input examples/method_compatibility.example.json
python scripts/method_compatibility.py --selftest
# 同一真实两组比较走 base R;关键 p/mean_diff/effect 与 Python 交叉核数
python scripts/r_analysis_crosscheck.py --input results.csv --group method --metric acc \
--paired-by seed --out r_acc.csv# stat_rigor_gate 判"不支撑假设/不可复现" → findings 带信号词 → 总控 reroute 按 ROUTES[7] 分流:
python ../light-orchestrator/scripts/reroute.py --findings stat_findings.json --stage 7 --passport .light/passport.yaml
# 不支撑假设(信号 假设/支撑/效应)→ 建议 7→5 回 research-plan;不可复现(信号 种子/复现)→ 建议 7→6 回 experiment-coding。
# 用户拍板回炉后落一等回边(记在目标阶段,不破拓扑):
python ../light-orchestrator/scripts/passport.py add-back-edge --to 5 --from 7 \
--root-cause "结果不支撑假设(效应量小/CI 含 0)" --evidence-ptr "<reroute 给的指针>" # 或 --to 6(不可复现)
python ../light-orchestrator/scripts/passport.py validate --file .light/passport.yaml # 回边不破拓扑 → 仍 PASS各脚本 --selftest/--help 即接口;工具一手核(statsmodels multipletests / scipy.stats / pingouin / SHAP 非因果 /
Evidently v7 破坏式变更 / GRADE / Cohen's d·Cliff δ)详见 references.md。
逐层深入:描述(指标±CI vs baseline)→ 解释(归因到方法哪个组件,结合消融)→ 诊断(哪些证明创新有效、哪些暴露问题/矛盾) → 洞察(能成论文亮点的规律、意外发现、可解释性证据)→ 行动(哪些异常要排查、哪些结论需补实验、哪些不能过度声称)。
每条能写进论文的论断连到它的检验/p/q(FDR)/效应量/CI/n(claim_evidence_table.md);经 evidence_contract.grade_evidence
机械定档:strong(q<.01 且 |d|≥.5 且 CI 不含 0 且 n≥30)/ moderate(显著中效应)/ weak(显著但小效应或小样本)/
none(不显著或 CI 含 0)。下游据此校准措辞——这是 evidence_strength.json 的真实第一消费链(paper-writing 也消费)。
p 值三件套(p + 效应量 + CI);多重比较校正(BH-FDR 控 FDR / Bonferroni 控 FWER,显著性看 q);自动选检验(正态性 + 组数 + 方差齐性,配对设计用配对检验);效应量配检验类型(参数→Cohen's d+Hedges,非参/序数→Cliff's δ 更稳)。统计错误/p-hacking = critical(spec §4.2)。
这提升是统计显著还是噪声(看校正后 q,不看裸 p)?效应量多大(d/δ,不只看 p)?换数据集/换种子还成立吗(稳健性、可复现, 多种子报均值±std)?有没有 p-hacking(选择性报告 / 多重比较不校正 / HARKing / garden of forking paths)?——答不上来的,分析没到及格线。
result_card_gate.py 吗?每条 claim 的 target、analysis set、missingness、assumption、guardrail_analysis、family、practical threshold、effect、language 都绑在一张 result card 上了吗?as_of、locked_at、results_available_at、computed_at 都是真实已发生时间吗?分析没有早于结果可见、也没有未来预填吧?decision_ledger 写清 PRE/POST 结果的排除、换指标、模型选择、切片、阈值了吗?POST_RESULTS 决策没冒充 CONFIRMATORY 吧?evidence_strength.json 产了吗(每条 claim 有档 + 措辞上限,交 paper-writing/research-ethics)?unit_of_analysis、复杂设计、comparison family_id、planned/reported
coverage、raw result hash/run manifest/commit 都留痕了吗?--paired-by 时,claim/evidence 是否只有 paired 主比较、没有同一数据再产 independent duplicate claim?真增量(v2 兑现,已 selftest + E2E 实测):⓪ result card + analysis decision ledger gate(result_card_gate.py,Round 3 新增)——产 light.findings.v1,把 target、analysis set、missingness、provenance、assumptions、guardrail_analysis、comparison family、practical threshold、effect、language 与 PRE/POST 分析决策账本绑定;p≈0.05 的语言稳定性、非显著误释、结果后 confirmatory 伪装、sensitivity/supplementary 混淆、guardrail FAIL/UNKNOWN 后继续强 claim 均可机读阻断。Round 3 再补 raw-run provenance 与时间轴:每条 claim 必须带 source run IDs、run manifest/raw result/analysis code locator 与 SHA-256;locked_at/results_available_at/computed_at 不能未来预填,REVISION_REQUIRED/UNKNOWN 不能冒充 ready,防止孤立数字或未完成结果直接进入写作;Round3 续补 guardrail evidence 消费,防止 experiment-coding 的 guardrails.json 到写作前消失。① 统计严谨/证据强度 critical 门 producer(stat_rigor_gate.py,v2 净新增接线)——
编排 BH-FDR 真重算(裸 p 经校正后掉到 q≥.05 = 假阳性 → critical)+ 消费 _shared/evidence_contract 给每条 claim 定档 → 产
light.findings.v1(producer=result-analysis):多重比较未校正/选择性报告/假设证否/多种子不稳 → critical(对齐
STAGE_GATES[7]=[stat_validity,evidence_strength]),过度解读/效应量缺失 → warn,被 run_checkpoint --stage 7 聚合 exit 1。
② 本技能是 7→5 / 7→6 两条回边的发起方(与 experiment-coding 的「落点」相反)——hypothesis_support 判不支撑假设 → findings 带
「假设/支撑/效应」信号 → reroute 建议 7→5;reproducibility 判多种子 sign-flip/CV 过大 → 带「种子/复现」信号 → 建议 7→6;
p-hacking critical 刻意不带 5/6 信号 → reroute 给 manual(在 stage 7 内修),落点诚实(E2E 三线实测分流正确)。③ 港 v1
统计资产修 bootstrap/命名:stats_tests/significance_test/analyze_results 修硬编码 ../../../code_assets、../../_shared→规范
bootstrap,m06/金矿1/m07/a10→v2 技能名,source:m06:*→result-analysis:*。④ evidence_strength.json 是 result-analysis↔
paper-writing/research-ethics 的桥:本技能定证据强度、research-ethics(stage 8)查措辞是否超过证据,evidence_contract 共用、不重叠。
⑤ DeLong 相关 AUROC 比较港 v1(statsmodels/scipy 没有,自测对齐 sklearn roc_auc_score)。
裸模型本就会的(不吹):「做显著性检验」「报效应量别只报 p」「多重比较要校正」「措辞别太强」「SHAP 非因果」——裸 Opus 都会说,
近零增量。本技能价值 = ① 把 p-hacking/证据强度落成确定性 critical 机读门 + 确定性阻断(裸模型嘴上说「要校正」,手上还是报裸
p<.05 当显著;编排器读不了它的「嘴上说」,读 light.findings.v1 的 verdict);② BH-FDR 真重算判假阳性(不是提醒「记得校正」,是
真算出「这 4 个校正后 q≥.05」);③ claim↔证据档单一数据源(evidence_strength.json 跨 5 个下游技能卡措辞,裸模型每处各凭感觉);
④ 机读 findings + 根因回炉发起(result-analysis 是 7→5/7→6 回炉枢纽,裸模型无此非线性编排闭环)。
诚实落后项(已知没做到):
method_compatibility.py 能拒绝已知不兼容并暴露未知条件,但不会自动证明
第三方方法声明真实,也不证明统计识别、实现正确、结果可信或因果有效。stat_rigor_gate 查「多重比较校没校正 / 假设统计上撑不撑得住 / 结果稳不稳」——绝不
「证明了方法真有效 / 因果成立」。效应量解读、机制因果、外推性的终判仍需人 / 领域判断;机检给的是「统计上站不站得住」的必要条件。DataDriftPreset API
迁到 evidently.future——用前必核版本(铁律 2 一手核出的坑)。两者都不是 critical 门,只作洞察/参考。标准产出工件:结果分析报告(现象→原因→证据→对论文的意义)·
claim_evidence_table.md(claim↔证据)·evidence_strength.json(证据档→措辞档,交 paper-writing/research-ethics)·stat_findings.json(统计严谨/证据强度门)· 推荐图表清单(交 figure)。 亮点 → paper-writing 写作支撑;不支撑假设 → 回 research-plan(7→5);不可复现 → 回 experiment-coding(7→6);结论台账交 memory-pm。
docs/competitors/result-analysis.md(Round 2 真搜:
8 个真同类 SKILL + statsmodels/scipy/base R/规范等机制锚;每条带整仓 star、commit、路径、行号)result-analysis-resource-map.md(分析计划→raw runs→
设计感知统计→稳健性→claim/provenance→总控;含 Python/R 双路径与资源访问分级)references.mdscripts/——各 Python 脚本 --selftest/--help 即接口,R 脚本有 --selftest;
stat_rigor_gate.py(critical 门)、analysis_plan_audit.py(warn-only 计划/设计/provenance)、
method_compatibility.py(方法×数据×访问条件兼容门)、
analyze_results.py(设计感知检验+FDR+两工件)、r_analysis_crosscheck.py/.R(base-R 交叉核验)、
significance_test.py(p/d/CI/FDR/DeLong)、make_figs.py/explain_shap.py/leakage_overfit_check.pyassets/result_analysis_report_template.md(四段式 + 亮点/异常/待补实验 + 回炉判定)examples/worked_example.py(EDA→显著性→图→泄漏体检→报告)· examples/stat_spec.example.json(stat_rigor_gate 输入)examples/method_compatibility.example.json_shared/README.md(evidence_contract 核心消费 · findings_schema · gate_runner · 规范 bootstrap)light-experiment-coding(stage 6,run_manifest.md 多种子指标;7→6 回炉目标)·
light-research-plan(stage 5,7→5 回炉目标)· light-research-ethics(stage 8
claim_evidence_bind 消费 evidence_strength.json 查措辞)· run_checkpoint.py(stage 7 聚合 exit 1)·
reroute.py(ROUTES[7] 两条出边 7→5/7→6)· paper-writing(stage 8,据证据档校准措辞)© Light0305, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 21 other files (scripts, assets) in skills/light-result-analysis of Light0305/Light-skills.
Open the folder on GitHubat commit 6b44f57
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Light0305/Light-skills, which our catalogue first saw on October 7, 2026.
Light Result Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Light Result Analysis this skillLight0305/Light-skills | 641 | 1 repos | ~5.5k | Automated safety check: Pass | MIT | |
| Lammps DeepmdHello-QM/catgo-LRG | 205 | 1 repos | ~1k | Automated safety check: Pass | AGPL-3.0 | |
| scikit-survival Time-to-Event Modelingdavila7/claude-code-templates | 32k | 12 repos | ~3.7k | Automated safety check: Pass | MIT | |
| Neuropixels AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~5k | Automated safety check: Pass | MIT | |
| Molfeatdavila7/claude-code-templates | 32k | 10 repos | ~3.7k | Automated safety check: Pass | MIT | |
| IcmlnanoAgentTeam/research-claw | 293 | — | ~2.4k | Automated safety check: Pass | MIT |
Hello-QM/catgo-LRG
Run LAMMPS molecular dynamics with DeePMD-kit machine learning potentials.
davila7/claude-code-templates
Fits and evaluates survival models with scikit-survival: Cox models, Random Survival Forests, boosting, survival SVMs, concordance index, Brier score and competing risks.
K-Dense-AI/scientific-agent-skills
Analyzes Neuropixels extracellular recordings end-to-end with SpikeInterface.
davila7/claude-code-templates
Molecular featurization for ML (100+ featurizers). An agent skill from davila7/claude-code-templates.
nanoAgentTeam/research-claw
ICML (International Conference on Machine Learning) paper formatting — activate when the user wants to submit to ICML, follow ICML template, or fix ICML format issues.
GPTomics/bioSkills
Maps query single-cell data onto reference atlases and transfers cell-type labels using scArches surgery (scVI/scANVI), Symphony, Azimuth, CellTypist, scPoli, popV, and foundation models, with…
Light0305/Light-skills
Verifies that every reference in a manuscript is real, correctly identified and actually supports its claim, and produces a citation registry for typesetting.
Light0305/Light-skills
Coordinates and recovers multi-stage Light research projects from a single passport file, with checkpoints, stale-work tracking and rerouting only when you approve.
Light0305/Light-skills
Builds an evidence-backed invention disclosure packet from a project or research result for attorney or patent-agent review, without giving legal advice.
Light0305/Light-skills
Audits, scaffolds and safely migrates research project folder structures, keeping existing repositories read-only until you approve exact moves from a plan.
Light0305/Light-skills
Prepares draft materials for a China software copyright registration from a real project: application worksheet, source deposit plan, operation manual and consistency checks.
Light0305/Light-skills
Evidence-based workflow for designing or modernizing a software system: current-state inventory, options, API and schema contracts, migration plans, ADRs and verification.
Categories
Light 科研主线第 7 步·结果分析:不描述好坏、解释「为什么」,把每条结论绑死到 claim + 证据强度,并防 p-hacking。. Light Result Analysis is an agent skill from Light0305/Light-skills.
Light Result Analysis fits situations like: tasks that involve Machine learning.
Run `npx skills add Light0305/Light-skills --skill light-result-analysis -a claude-code`. Or copy the skill folder (skills/light-result-analysis in Light0305/Light-skills) into .claude/skills/light-result-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Light0305/Light-skills --skill light-result-analysis -a codex`. Or copy the skill folder (skills/light-result-analysis in Light0305/Light-skills) into .agents/skills/light-result-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Light0305/Light-skills --skill light-result-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/light-result-analysis, .gemini/skills/light-result-analysis, .github/skills/light-result-analysis and .opencode/skills/light-result-analysis in your project.
Going by SKILL.md and its folder, Light Result Analysis needs Python and R for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Light Result Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Light Result Analysis: Lammps Deepmd (Hello-QM/catgo-LRG, 205 stars), scikit-survival Time-to-Event Modeling (davila7/claude-code-templates, 32k stars), Neuropixels Analysis (K-Dense-AI/scientific-agent-skills, 48k stars) and Molfeat (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Light0305 (a GitHub user) maintains it in Light0305/Light-skills, which has 641 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on July 6, 2026.
Source: Light0305/Light-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.