Agent skill

Evidence Walkthrough

by PKU-YuanGroup in PKU-YuanGroup/OpenAI4S

Run the reference end-to-end research pass — fixed database query, local analysis, versioned artifacts with lineage, then an exported evidence package that verifies in a clean environment.

MITAuto-check passed

Install Evidence Walkthrough

skills CLI
$ npx skills add PKU-YuanGroup/OpenAI4S --skill evidence-walkthrough -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PKU-YuanGroup/OpenAI4S evidence-walkthrough --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evidence-walkthrough .claude/skills/evidence-walkthrough && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evidence-walkthrough
GitHub stars
622
Token cost
~1.7k tokens
SKILL.md length
546 words
Files
3
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Run the reference end-to-end research pass — fixed database query, local analysis, versioned artifacts with lineage, then an exported evidence package that verifies in a clean environment.

  • Works in 4 steps: Retrieve, and record what you retrieved → Analyse → Produce artifacts, declaring their inputs → …
  • SKILL.md covers Fixed inputs, Workflow, What to check before calling… and Offline and reproducibility
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Evidence Walkthrough is an agent skill from PKU-YuanGroup/OpenAI4S. Run the reference end-to-end research pass — fixed database query, local analysis, versioned artifacts with lineage, then an exported evidence package that verifies in a clean environment. Use as the first-run demonstration, as a benchmark case, or when a result must be handed to someone who was not there when it ran.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `README.md` and `README_zh.md`).

The repository describes itself as: Open-source AI agent for scientific research. Analyze data in Python/R with Claude, GPT, Gemini, and more. The licence is MIT.

Example prompts

  • “/evidence-walkthrough”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Retrieve, and record what you retrieved
  2. Analyse
  3. Produce artifacts, declaring their inputs
  4. Export and verify

What it can do on your machine

Read from SKILL.md and the folder at commit 4a72e87. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evidence Walkthrough loads about 1.7k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PKU-YuanGroup/OpenAI4S at commit 4a72e87, republished under its MIT licence (© PKU-YuanGroup). 546 words, ~1,730 tokens.

Download SKILL.mdSave it as .claude/skills/evidence-walkthrough/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evidence-walkthrough
description
Run the reference end-to-end research pass — fixed database query, local analysis, versioned artifacts with lineage, then an exported evidence package that verifies in a clean environment. Use as the first-run demonstration, as a benchmark case, or when a result must be handed to someone who was not there when it ran.
origin
openai4s
category
research-workflow

Evidence walkthrough

The reference pass a result has to survive: query → analyse → artifacts with lineage → an evidence package a stranger can verify.

Its point is not the science, which is deliberately small. Its point is that every step leaves evidence, and the package at the end can be checked by someone who does not trust this machine — a reviewer, a colleague, or you on a different laptop in six months.

Fixed inputs

Use these exact accessions. They are fixed so two runs are comparable and so this doubles as a benchmark case; changing them makes a run incomparable to every previous one.

python
ACCESSIONS = ["P69905", "P68871", "P02042", "P02100"]   # human haemoglobin subunits

Workflow

1. Retrieve, and record what you retrieved
python
import json

records = []
for accession in ACCESSIONS:
    hit = host.science.search("uniprot", accession, limit=1)
    records.append(hit)

# Check before saving, not after: an artifact written with three quarters of
# its evidence missing is already wrong by the time anyone can query it.
for accession, record in zip(ACCESSIONS, records):
    envelope = record["provenance"]
    assert accession in envelope["request_url"], accession
    assert envelope["retrieved_at"] and envelope["response_sha256"]

host.write_file("raw_uniprot.json", json.dumps(records, indent=2))
raw = host.save_artifact(
    "raw_uniprot.json",
    # EVERY retrieval, not the first one. This file is the evidence for four
    # independent requests, and `records[0]["provenance"]` describes exactly
    # one of them — the other three accessions would then sit inside an
    # artifact that claims to preserve their evidence while carrying no
    # request URL, no retrieval time and no response hash for them.
    source={
        "kind": "aggregate",
        "database": "uniprot",
        "queries": ACCESSIONS,
        "sources": [record["provenance"] for record in records],
    },
)                                                # -> {"version_id": ...}

Save the raw response before analysing it. The analysis is a claim; the raw response is the evidence for it, and a claim whose evidence was never written down cannot be rechecked later.

source is the other half. Every host.science.search result carries a provenance envelope naming the database, the exact request, the moment it was fetched and a hash of the bytes that came back. Pass it and the artifact can answer when was this true and was it the same data — without it a saved file records what you have but not what it is evidence of, and a rerun that quietly returned something different is indistinguishable from one that did not.

One artifact, four retrievals, so four envelopes. The aggregate above is the form to use whenever a file is assembled from more than one request: a single envelope covers a single request, and attaching one of four is worse than attaching none — the artifact then looks provenanced while three quarters of it is unaccounted for. Read it back and confirm every accession is there:

python
attached = json.loads(host.query(
    # `my_artifact_versions`, not `artifact_versions`: the base table is closed to
    # agent SQL by a SQLite authorizer, and this view is the same rows confined to
    # this session's scope. Reading the base table returned every project's.
    "SELECT source FROM my_artifact_versions WHERE version_id = ?",
    [raw["version_id"]],
)[0]["source"])
covered = {envelope["request_url"] for envelope in attached["sources"]}
missing = [a for a in ACCESSIONS if not any(a in url for url in covered)]
assert not missing, f"no retrieval provenance attached for {missing}"
2. Analyse

Keep it to what the raw file supports. Length, mass, and sequence composition are properties of the record; anything requiring a source you did not save is a claim you cannot back.

python
import json, collections

rows = []
for record in records:
    entry = (record.get("results") or [{}])[0]
    sequence = (entry.get("sequence") or {}).get("value", "")
    rows.append({
        "accession": entry.get("primaryAccession"),
        "name": entry.get("proteinDescription", {}).get(
            "recommendedName", {}).get("fullName", {}).get("value"),
        "length": len(sequence),
        "top_residues": collections.Counter(sequence).most_common(3),
    })
Show full SKILL.md (225 more words)Show less
3. Produce artifacts, declaring their inputs
python
host.write_file("summary.json", json.dumps(rows, indent=2))
summary = host.save_artifact("summary.json", input_version_ids=[raw["version_id"]])

input_version_ids is the lineage edge. Without it the summary is a file that appeared from nowhere; with it, anyone reading the artifact can walk back to the exact bytes it was derived from. Declare it on every derived artifact, including figures.

python
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt

fig, ax = plt.subplots(figsize=(6, 3.2))
ax.bar([r["accession"] for r in rows], [r["length"] for r in rows])
ax.set_ylabel("residues")
ax.set_title("Haemoglobin subunit lengths")
fig.tight_layout()
fig.savefig("lengths.png", dpi=150)
plt.close(fig)

host.save_artifact("lengths.png", input_version_ids=[raw["version_id"]])
4. Export and verify

Export the session package from the UI (or GET /api/v1/frames/<id>/session/export), then verify it the way a recipient would — with no daemon involved:

openai4s verify-package <session>.openai4s-session.zip

A pass means every listed file matches its recorded hash and the manifest matches its own digest. It does not establish who produced the package; that needs a signature, which this format does not carry. Say "verified intact", not "verified authentic".

What to check before calling it done

  • Every derived artifact declares input_version_ids. A missing edge is the difference between a result and an anecdote.
  • The raw retrieval is saved as its own artifact, not just parsed in memory.
  • verify-package exits 0 on the exported package.
  • Numbers in your summary can each be traced to a saved file.

Offline and reproducibility

The retrieval step needs the network. Everything after it is deterministic given the same raw file, so a benchmark run should fix the raw artifact and replay from step 2 — that separates "the analysis changed" from "the upstream database changed", which are different failures and only one of them is yours.

© PKU-YuanGroup, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/evidence-walkthrough of PKU-YuanGroup/OpenAI4S.

  • SKILL.md
  • README.md
  • README_zh.md

Open the folder on GitHubat commit 4a72e87

Compare with similar skills

Evidence Walkthrough next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evidence Walkthrough compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evidence Walkthrough this skillPKU-YuanGroup/OpenAI4S622—~1.7kAutomated safety check: PassMIT
Fixalirezarezvani/claude-skills28k1 repos~765Automated safety check: PassMIT
Fix Issuepytorch/pytorch104k—~2.3kAutomated safety check: PassCustom licence
Orch Fix Defectaffaan-m/ECC277k1 repos~414Automated safety check: PassMIT
Logic Fix Allsickn33/agentic-awesome-skills47k1 repos~1.3kAutomated safety check: PassMIT
Fixdavepoon/buildwithclaude3.6k—~14kAutomated safety check: NotesMIT

Similar skills

  • Fix

    alirezarezvani/claude-skills

    Fix failing or flaky Playwright tests. An agent skill from alirezarezvani/claude-skills.

    28k GitHub starsUsed in 1 repo~765 tokens
    Testing & QAAuto-check passed
  • Fix Issue

    pytorch/pytorch

    Fix bugs reported in PyTorch GitHub issues by reproducing, root-causing, and implementing a fix in the local working tree.

    104k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Orch Fix Defect

    affaan-m/ECC

    Orchestrate fixing a bug — reproduce it as a failing regression test, fix to green, review, and gated commit — by delegating each phase to the matching ECC agent.

    277k GitHub starsUsed in 1 repo~414 tokens
    Testing & QAAuto-check passed
  • Logic Fix All

    sickn33/agentic-awesome-skills

    Autonomous repository-wide audit-and-fix pipeline: health → review → locate/explain → fix → diff-verify → iterate until clean.

    47k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Fix

    davepoon/buildwithclaude

    Get fix intelligence for a vulnerability and propose concrete remediation for the current repository

    3.6k GitHub stars~14k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check: notes
  • Fix Dependabot

    remotion-dev/remotion

    Official

    Fix a Dependabot PR by updating all monorepo instances of the dependency, running bun install, and pushing

    63k GitHub stars~509 tokensUpdated yesterday
    DevelopmentAuto-check passed

More from PKU-YuanGroup/OpenAI4S

All 17 skills in this repo
  • Single Cell Rna Analysis

    PKU-YuanGroup/OpenAI4S

    Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…

    622 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Bioprobench

    PKU-YuanGroup/OpenAI4S

    Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

    622 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Reaction Atom Mapping

    PKU-YuanGroup/OpenAI4S

    Map atoms and changed bonds for a complete reaction with RXNMapper.

    622 GitHub stars~1.2k tokensUpdated 2 days ago
    Auto-check passed
  • Reaction Forward Prediction

    PKU-YuanGroup/OpenAI4S

    Predict ranked products from reactants and reagents with ReactionT5v2-forward; use for outcome prediction or round-trip recovery.

    622 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Reaction Yield Estimation

    PKU-YuanGroup/OpenAI4S

    Estimate yield for a fully specified reactant/reagent/product record with ReactionT5v2-yield.

    622 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Rfdiffusion

    PKU-YuanGroup/OpenAI4S

    Generate de novo protein backbones with RFdiffusion for protein-target binders, hotspot-conditioned interfaces, motif scaffolding, partial diffusion, or symmetric assemblies.

    622 GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check passed

Questions about Evidence Walkthrough

What does Evidence Walkthrough do?

Run the reference end-to-end research pass — fixed database query, local analysis, versioned artifacts with lineage, then an exported evidence package that verifies in a clean environment. Evidence Walkthrough is an agent skill from PKU-YuanGroup/OpenAI4S. Run the reference end-to-end research pass — fixed database query, local analysis, versioned artifacts with lineage, then an exported evidence package that verifies in a clean environment.

How do I install Evidence Walkthrough in Claude Code?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill evidence-walkthrough -a claude-code`. Or copy the skill folder (skills/evidence-walkthrough in PKU-YuanGroup/OpenAI4S) into .claude/skills/evidence-walkthrough in your project. Claude Code loads it when a task matches its description.

How do I install Evidence Walkthrough in Codex?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill evidence-walkthrough -a codex`. Or copy the skill folder (skills/evidence-walkthrough in PKU-YuanGroup/OpenAI4S) into .agents/skills/evidence-walkthrough in your project. Codex loads it when a task matches its description.

Can I use Evidence Walkthrough in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PKU-YuanGroup/OpenAI4S --skill evidence-walkthrough -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evidence-walkthrough, .gemini/skills/evidence-walkthrough, .github/skills/evidence-walkthrough and .opencode/skills/evidence-walkthrough in your project.

What does Evidence Walkthrough need to run?

SKILL.md names no scripts, command-line tools or credentials: Evidence Walkthrough is instructions for the agent only. Our summary lists: Python 3.

Does Evidence Walkthrough access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evidence Walkthrough safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evidence Walkthrough use?

Evidence Walkthrough is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evidence Walkthrough use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evidence Walkthrough?

Skills that share tags, products or a category with Evidence Walkthrough: Fix (alirezarezvani/claude-skills, 28k stars), Fix Issue (pytorch/pytorch, 104k stars), Orch Fix Defect (affaan-m/ECC, 277k stars) and Logic Fix All (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evidence Walkthrough?

PKU-YuanGroup (a GitHub organization) maintains it in PKU-YuanGroup/OpenAI4S, which has 622 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.

Source: PKU-YuanGroup/OpenAI4S on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.