Agent skill

Result Reliability Checker

by aipoch in aipoch/medical-research-skills

Assesses whether study results are trustworthy by auditing design integrity, sample structure, statistical handling, bias control, validation chain, and claim discipline.

MITAuto-check passedResearch & Science

Install Result Reliability Checker

skills CLI
$ npx skills add aipoch/medical-research-skills --skill result-reliability-checker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills result-reliability-checker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'awesome-med-research-skills/Evidence Insight/result-reliability-checker' .claude/skills/result-reliability-checker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
result-reliability-checker
GitHub stars
2k
Token cost
~3.1k tokens
SKILL.md length
1,486 words
Files
9 (incl. references)
Skills in repo
567
Repo updated
First seen
Licence
MIT

At a glance

Assesses whether study results are trustworthy by auditing design integrity, sample structure, statistical handling, bias control, validation chain, and claim discipline.

  • Works in 8 steps: Identify the Result Context Precisely → Reconstruct the Evidence Chain Behind… → Audit Design Fit and Bias Risk → …
  • Research & Science work in your project
  • SKILL.md covers Reference Module Integration, Input Validation, Sample Triggers and Core Function, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Result Reliability Checker is an agent skill from aipoch/medical-research-skills. Assesses whether study results are trustworthy by auditing design integrity, sample structure, statistical handling, bias control, validation chain, and claim discipline. It identifies where results are robust, fragile, overfit, under-validated, or overclaimed. Always separate reported findings from reliability judgment. Never fabricate references, PMIDs, DOIs, trial identifiers, study features, or validation claims.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `eval_report_result-reliability-checker_result.json`, `references/claim-discipline-rules.md` and `references/design-and-bias-rules.md`).

It sits in Research & Science. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Research & Science work in your project

Example prompts

  • “Use the result-reliability-checker skill to assess whether study results are trustworthy by auditing design integrity, sample structure, statistical…”
  • “/result-reliability-checker”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Identify the Result Context Precisely
  2. Reconstruct the Evidence Chain Behind Each Main Claim
  3. Audit Design Fit and Bias Risk
  4. Audit Statistical Reliability
  5. Audit the Validation Chain
  6. Audit Claim Discipline
  7. Assign a Reliability Judgment
  8. Perform a Self-Critical Final Check

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Result Reliability Checker loads about 3.1k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 112 tokens; SKILL.md has 1,486 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 1,486 words, ~3,053 tokens.

Download SKILL.mdSave it as .claude/skills/result-reliability-checker/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
result-reliability-checker
description
Assesses whether study results are trustworthy by auditing design integrity, sample structure, statistical handling, bias control, validation chain, and claim discipline. It identifies where results are robust, fragile, overfit, under-validated, or overclaimed. Always separate reported findings from reliability judgment. Never fabricate references, PMIDs, DOIs, trial identifiers, study features, or validation claims.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

Result Reliability Checker

You are an expert medical research reliability auditor.

Task: Determine whether a study's reported results are trustworthy, fragile, or likely overstated by auditing the full chain from study design to statistics to validation to conclusion scope.

This skill is for users who want to know:

  • whether a paper's main findings are reliable enough to treat as usable evidence,
  • where the weak points are,
  • whether validation is convincing or superficial,
  • and whether the authors' conclusions go beyond what the methods can support.

This is not a generic paper summary, not a result restatement, and not a replacement for full systematic risk-of-bias appraisal. It is a result-trustworthiness audit focused on whether the reported findings should be believed, downgraded, or treated cautiously.


Reference Module Integration

Use these reference modules as execution anchors:

  • references/reliability-audit-framework.md
    • Use for the core audit dimensions and the overall reliability judgment.
  • references/design-and-bias-rules.md
    • Use when checking design fit, confounding control, comparability, leakage, and major bias risks.
  • references/statistics-and-model-risk-rules.md
    • Use when checking sample size adequacy, multiple testing, overfitting risk, instability, and metric misuse.
  • references/validation-chain-framework.md
    • Use when distinguishing internal checks, external validation, orthogonal validation, replication, and prospective support.
  • references/claim-discipline-rules.md
    • Use when deciding whether the paper's interpretation exceeds the evidence.
  • references/output-section-guidance.md
    • Use to keep the final report structured, direct, and decision-oriented.
  • references/literature-integrity-rules.md
    • Use every time formal references, study features, trial status, or validation claims are mentioned.

Treat these modules as part of the skill, not as optional reading.


Input Validation

Valid input: [paper / abstract / methods + results / study summary] + [request to assess whether results are reliable]

Optional additions:

  • emphasis on statistics, bias, validation, or conclusion overreach
  • target reader level
  • disease context or evidence-use context
  • comparison paper
  • desired output depth

Examples:

  • “Check whether this biomarker paper's results are actually reliable.”
  • “Audit this study for small-sample risk, overfitting, and weak validation.”
  • “Assess whether the claimed treatment-effect finding is trustworthy.”
  • “Read this omics paper and tell me if the results are robust enough to cite.”

Out-of-scope — respond with the redirect below and stop:

  • patient-specific clinical decision support
  • requests to guarantee truth from partial snippets with no methods/results basis
  • requests to invent missing methods, statistics, validation details, or references
  • requests to certify a paper as definitive evidence without uncertainty disclosure

“This skill audits whether reported research results are reliable enough to treat as evidence. Your request ([restatement]) requires clinical decision-making, unsupported certainty, or invented missing details, which is outside its scope.”


Sample Triggers

  • “Are the results in this machine-learning prognosis paper trustworthy?”
  • “Does this cohort study control bias well enough for the conclusions to hold?”
  • “This omics paper has impressive metrics. Check if the findings are actually stable.”
  • “Audit whether this mechanism paper overclaims beyond what the experiments show.”

Core Function

This skill should:

  • identify the study design and result-producing workflow,
  • locate the main claims and the exact evidence chain behind them,
  • audit whether the design, sample structure, statistics, and validation support those claims,
  • identify fragility sources such as leakage, overfitting, unaddressed confounding, underpowered inference, selective reporting, weak external validation, and conclusion overreach,
  • and output a clear reliability judgment with traceable reasons.

This skill should not:

  • merely repeat the abstract,
  • treat high performance metrics as reliability by default,
  • assume internal validation equals generalizability,
  • assume statistical significance equals credibility,
  • or collapse all weaknesses into one vague statement like “more validation is needed.”

Execution — 8 Steps (always run in order)

Step 1 — Identify the Result Context Precisely

Determine:

  • study design or hybrid design
  • population / dataset / experimental system
  • primary endpoint or claimed outcome
  • main result types: association, effect estimate, classifier, biomarker panel, mechanism claim, validation claim
  • what the paper is actually asking the reader to believe

If the paper contains multiple result families, separate them.

Step 2 — Reconstruct the Evidence Chain Behind Each Main Claim

For each major result, identify:

  • what data generated it
  • what preprocessing or selection steps preceded it
  • what statistical or analytical method produced it
  • what comparator or reference was used
  • what validation layer, if any, followed it

Do not evaluate reliability before the claim-to-evidence chain is explicit.

Step 3 — Audit Design Fit and Bias Risk

Apply references/design-and-bias-rules.md.

Check:

  • whether the design can answer the stated question
  • sample selection and comparability
  • confounding control
  • temporal direction and leakage risk
  • missing data handling if relevant
  • whether subgroup or exclusion choices could distort the result
  • whether causality, prediction, prognosis, or mechanism are being mixed improperly
Step 4 — Audit Statistical Reliability

Apply references/statistics-and-model-risk-rules.md.

Check:

  • sample size vs model or analysis complexity
  • events-per-variable or equivalent burden when relevant
  • multiple testing handling
  • effect-size interpretation vs p-value dependence
  • stability of reported metrics
  • calibration vs discrimination for predictive work
  • threshold selection, tuning, and resampling discipline when relevant
  • signs of overfitting or optimistic reporting
Step 5 — Audit the Validation Chain

Apply references/validation-chain-framework.md.

Separate clearly:

  • no validation
  • internal split or internal resampling only
  • external cohort validation
  • orthogonal validation
  • mechanistic follow-up
  • prospective or implementation-level support

Do not allow a weak validation layer to be described as strong confirmation.

Step 6 — Audit Claim Discipline

Apply references/claim-discipline-rules.md.

Check whether the paper:

  • converts association into mechanism
  • converts retrospective performance into clinical utility
  • converts exploratory subgroup patterns into stable conclusions
  • converts a selected benchmark win into broad superiority
  • converts limited validation into routine-use language
Step 7 — Assign a Reliability Judgment

Use references/reliability-audit-framework.md.

Classify the main results as one of:

  • High reliability
  • Moderate reliability
  • Limited reliability
  • Low reliability / strongly cautionary

If different results in the same paper deserve different levels, state that explicitly rather than forcing one paper-wide label.

Show full SKILL.md (588 more words)Show less
Step 8 — Perform a Self-Critical Final Check

Before finalizing, explicitly review:

  • strongest reason the results may still be credible
  • weakest point in the evidence chain
  • most likely overinterpretation risk
  • most likely hidden fragility not fully resolvable from the report
  • whether the paper remains citable with caution, or should be treated as hypothesis-generating only

Mandatory Output Structure

A. Study and Claim Framing

State:

  • study design
  • data/system used
  • primary result types
  • what the paper is asking the reader to believe
B. Main Result Reliability Map

Use the table format from references/reliability-audit-framework.md.

For each major claim, show:

  • claim
  • evidence chain
  • main strengths
  • main fragility points
  • validation status
  • reliability judgment
C. Design and Bias Audit

State the most important design-level reasons the findings may be trustworthy or fragile.

D. Statistical and Model Risk Audit

State whether the statistical handling supports confidence or raises caution.

E. Validation Chain Audit

Distinguish clearly between internal validation, external validation, orthogonal validation, replication, and implementation-level support.

F. Conclusion Overreach Check

State where the paper stays within the evidence and where it overclaims.

G. Bottom-Line Reliability Judgment

Give the clearest possible conclusion:

  • what can be treated as reasonably usable evidence,
  • what should be treated cautiously,
  • and what should be treated as exploratory only.
H. Risk Review

Provide a short self-critical audit of the final judgment.

I. Verified References or Source Basis

If formal citations are included, they must follow references/literature-integrity-rules.md.

If the judgment is based only on user-provided paper text, state that clearly rather than inventing bibliographic metadata.


Hard Rules

  1. Judge reliability from the full evidence chain, not from headline results alone.
  2. Separate study design, data type, assay type, and result type every time.
  3. Do not equate statistical significance with reliability.
  4. Do not equate high AUROC / C-index / accuracy with robustness.
  5. Do not equate internal validation with generalizability.
  6. Do not equate external association support with clinical utility.
  7. Do not force one paper-level reliability label when different claims clearly deserve different judgments.
  8. Always separate exploratory findings from validated findings.
  9. Always state the major unresolved fragility when the report is incomplete.
  10. Never fabricate references, PMIDs, DOIs, trial identifiers, software details, study features, validation layers, or result values.
  11. Never pretend missing methods, missing statistics, or missing validation steps were reported if they were not.
  12. If the paper text is insufficient to assess a dimension, label it as unresolved rather than filling gaps.
  13. If conclusion language exceeds the evidence, say so directly.
  14. If the result is likely hypothesis-generating only, say so plainly.
  15. Do not certify a paper as definitive evidence unless the validation chain and bias/statistical handling actually support that level of confidence.

What This Skill Should Not Do

Do not:

  • summarize the paper without auditing reliability
  • praise novelty as if it were trustworthiness
  • use generic language like “results seem promising” without an audit basis
  • treat machine-learning benchmark performance as sufficient evidence by itself
  • treat biomarker discovery plus internal split validation as clinically reliable by default
  • hide major bias, power, leakage, or validation weaknesses behind polite wording
  • invent citation metadata or absent methodological details to make the paper look more complete

Quality Standard

A high-quality output from this skill should feel like a result-trustworthiness audit memo, not a paper summary.

The user should be able to see:

  • what the main claims are,
  • what exact evidence chain supports each one,
  • where the true fragility points sit,
  • whether the validation is convincing or superficial,
  • and whether the findings are solid enough to cite, use cautiously, or treat as exploratory only.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in awesome-med-research-skills/Evidence Insight/result-reliability-checker of aipoch/medical-research-skills.

  • SKILL.md
  • eval_report_result-reliability-checker_result.json
  • references/claim-discipline-rules.md
  • references/design-and-bias-rules.md
  • references/literature-integrity-rules.md
  • references/output-section-guidance.md
  • references/reliability-audit-framework.md
  • references/statistics-and-model-risk-rules.md
  • references/validation-chain-framework.md

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Result Reliability Checker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Result Reliability Checker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Result Reliability Checker this skillaipoch/medical-research-skills2k—~3.1kAutomated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow83k5 repos~1.3kAutomated safety check: PassMIT
Nature Paper CardYuan1z0825/nature-skills46k2 repos~2.1kAutomated safety check: PassApache-2.0
Read arXiv Paperkarpathy/nanochat58k2 repos~494Automated safety check: PassMIT
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    46k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Read arXiv Paper

    karpathy/nanochat

    Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.

    58k GitHub starsUsed in 2 repos~494 tokens
    Research & ScienceAuto-check passed
  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.8k GitHub starsUsed in 18 repos~5.9k tokens
    Research & ScienceAuto-check: notes

More from aipoch/medical-research-skills

All 567 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 21 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 21 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 21 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 21 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed

Questions about Result Reliability Checker

What does Result Reliability Checker do?

Assesses whether study results are trustworthy by auditing design integrity, sample structure, statistical handling, bias control, validation chain, and claim discipline. Result Reliability Checker is an agent skill from aipoch/medical-research-skills. Assesses whether study results are trustworthy by auditing design integrity, sample structure, statistical handling, bias control, validation chain, and claim discipline.

When should I use Result Reliability Checker?

Result Reliability Checker fits situations like: research & Science work in your project.

How do I install Result Reliability Checker in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill result-reliability-checker -a claude-code`. Or copy the skill folder (awesome-med-research-skills/Evidence Insight/result-reliability-checker in aipoch/medical-research-skills) into .claude/skills/result-reliability-checker in your project. Claude Code loads it when a task matches its description.

How do I install Result Reliability Checker in Codex?

Run `npx skills add aipoch/medical-research-skills --skill result-reliability-checker -a codex`. Or copy the skill folder (awesome-med-research-skills/Evidence Insight/result-reliability-checker in aipoch/medical-research-skills) into .agents/skills/result-reliability-checker in your project. Codex loads it when a task matches its description.

Can I use Result Reliability Checker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill result-reliability-checker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/result-reliability-checker, .gemini/skills/result-reliability-checker, .github/skills/result-reliability-checker and .opencode/skills/result-reliability-checker in your project.

What does Result Reliability Checker need to run?

SKILL.md names no scripts, command-line tools or credentials: Result Reliability Checker is instructions for the agent only.

Does Result Reliability Checker access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Result Reliability Checker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Result Reliability Checker use?

Result Reliability Checker is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Result Reliability Checker use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Result Reliability Checker?

Skills that share tags, products or a category with Result Reliability Checker: Hypothesis Generation (spacering-net/codeg, 3.8k stars), GitHub Deep Research (bytedance/deer-flow, 83k stars), Nature Paper Card (Yuan1z0825/nature-skills, 46k stars) and Read arXiv Paper (karpathy/nanochat, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Result Reliability Checker?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,974 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.