Agent skill

Healthcare Eval Harness

by affaan-m in affaan-m/ECC

用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。

MITAuto-check passedAI & LLM Engineering

Install Healthcare Eval Harness

skills CLI
$ npx skills add affaan-m/ECC --skill healthcare-eval-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC healthcare-eval-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/zh-CN/skills/healthcare-eval-harness .claude/skills/healthcare-eval-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
healthcare-eval-harness
GitHub stars
276k
Token cost
~1.4k tokens
SKILL.md length
120 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。

  • Tasks that involve LLM evaluation
  • SKILL.md covers 使用场景, 工作原理 and 示例
  • Calls npx and jq
  • Tasks that involve Clinical and healthcare research

What it does

Healthcare Eval Harness is an agent skill from affaan-m/ECC. 用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation and Clinical and healthcare research. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM evaluation
  • Tasks that involve Clinical and healthcare research

Example prompts

  • “/healthcare-eval-harness”

Requirements

  • Node.js

What it can do on your machine

Read from SKILL.md and the folder at commit 4eb71d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Healthcare Eval Harness loads about 1.4k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 120 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 4eb71d9, republished under its MIT licence (© affaan-m). 120 words, ~1,411 tokens.

Download SKILL.mdSave it as .claude/skills/healthcare-eval-harness/SKILL.md (or your agent's skills folder).
name
healthcare-eval-harness
description
用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。
origin
Health1 Super Speciality Hospitals — contributed by Dr. Keyur Patel
version
1.0.0

医疗评估框架 — 患者安全验证

医疗应用部署的自动化验证系统。单个严重故障将阻止部署。患者安全不容妥协。

注意: 示例使用 Jest 作为参考测试运行器。请根据您的框架(Vitest、pytest、PHPUnit 等)调整命令——测试类别和通过阈值与框架无关。

使用场景

  • 部署任何 EMR/EHR 应用之前
  • 修改 CDSS 逻辑(药物相互作用、剂量验证、评分)之后
  • 更改涉及患者数据的数据库模式之后
  • 修改身份验证或访问控制之后
  • 配置医疗应用 CI/CD 流水线期间
  • 解决临床模块合并冲突之后

工作原理

评估框架按顺序运行五个测试类别。前三个(CDSS 准确性、PHI 暴露、数据完整性)是严重关卡,要求 100% 通过率——单个故障即阻止部署。其余两个(临床工作流、集成)是高优先级关卡,要求 95% 以上通过率。

每个类别对应一个 Jest 测试路径模式。CI 流水线使用 --bail(首次失败即停止)运行严重关卡,并使用 --coverage --coverageThreshold 强制执行覆盖率阈值。

评估类别

1. CDSS 准确性(严重 — 要求 100%)

测试所有临床决策支持逻辑:药物相互作用对(双向)、剂量验证规则、临床评分与发布规范的对比、无假阴性、无静默故障。

bash
npx jest --testPathPattern='tests/cdss' --bail --ci --coverage

2. PHI 暴露(严重 — 要求 100%)

测试受保护健康信息泄露:API 错误响应、控制台输出、URL 参数、浏览器存储、跨机构隔离、未认证访问、服务角色密钥缺失。

bash
npx jest --testPathPattern='tests/security/phi' --bail --ci

3. 数据完整性(严重 — 要求 100%)

测试临床数据安全:锁定就诊记录、审计追踪条目、级联删除保护、并发编辑处理、无孤立记录。

bash
npx jest --testPathPattern='tests/data-integrity' --bail --ci

4. 临床工作流(高优先级 — 要求 95% 以上)

测试端到端流程:就诊生命周期、模板渲染、用药集、药物/诊断搜索、处方 PDF、红色警报。

bash
tmp_json=$(mktemp)
npx jest --testPathPattern='tests/clinical' --ci --json --outputFile="$tmp_json" || true
total=$(jq '.numTotalTests // 0' "$tmp_json")
passed=$(jq '.numPassedTests // 0' "$tmp_json")
if [ "$total" -eq 0 ]; then
  echo "No clinical tests found" >&2
  exit 1
fi
rate=$(echo "scale=2; $passed * 100 / $total" | bc)
echo "Clinical pass rate: ${rate}% ($passed/$total)"

5. 集成合规性(高优先级 — 要求 95% 以上)

测试外部系统:HL7 消息解析(v2.x)、FHIR 验证、实验室结果映射、格式错误消息处理。

bash
tmp_json=$(mktemp)
npx jest --testPathPattern='tests/integration' --ci --json --outputFile="$tmp_json" || true
total=$(jq '.numTotalTests // 0' "$tmp_json")
passed=$(jq '.numPassedTests // 0' "$tmp_json")
if [ "$total" -eq 0 ]; then
  echo "No integration tests found" >&2
  exit 1
fi
rate=$(echo "scale=2; $passed * 100 / $total" | bc)
echo "Integration pass rate: ${rate}% ($passed/$total)"
通过/失败矩阵
类别阈值失败时操作
CDSS 准确性100%阻止部署
PHI 暴露100%阻止部署
数据完整性100%阻止部署
临床工作流95% 以上警告,允许经审查后部署
集成95% 以上警告,允许经审查后部署
CI/CD 集成
yaml
name: Healthcare Safety Gate
on: [push, pull_request]

jobs:
  safety-gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: '20'
      - run: npm ci

      # CRITICAL gates — 100% required, bail on first failure
      - name: CDSS Accuracy
        run: npx jest --testPathPattern='tests/cdss' --bail --ci --coverage --coverageThreshold='{"global":{"branches":80,"functions":80,"lines":80}}'

      - name: PHI Exposure Check
        run: npx jest --testPathPattern='tests/security/phi' --bail --ci

      - name: Data Integrity
        run: npx jest --testPathPattern='tests/data-integrity' --bail --ci

      # HIGH gates — 95%+ required, custom threshold check
      # HIGH gates — 95%+ required
      - name: Clinical Workflows
        run: |
          TMP_JSON=$(mktemp)
          npx jest --testPathPattern='tests/clinical' --ci --json --outputFile="$TMP_JSON" || true
          TOTAL=$(jq '.numTotalTests // 0' "$TMP_JSON")
          PASSED=$(jq '.numPassedTests // 0' "$TMP_JSON")
          if [ "$TOTAL" -eq 0 ]; then
            echo "::error::No clinical tests found"; exit 1
          fi
          RATE=$(echo "scale=2; $PASSED * 100 / $TOTAL" | bc)
          echo "Pass rate: ${RATE}% ($PASSED/$TOTAL)"
          if (( $(echo "$RATE < 95" | bc -l) )); then
            echo "::warning::Clinical pass rate ${RATE}% below 95%"
          fi

      - name: Integration Compliance
        run: |
          TMP_JSON=$(mktemp)
          npx jest --testPathPattern='tests/integration' --ci --json --outputFile="$TMP_JSON" || true
          TOTAL=$(jq '.numTotalTests // 0' "$TMP_JSON")
          PASSED=$(jq '.numPassedTests // 0' "$TMP_JSON")
          if [ "$TOTAL" -eq 0 ]; then
            echo "::error::No integration tests found"; exit 1
          fi
          RATE=$(echo "scale=2; $PASSED * 100 / $TOTAL" | bc)
          echo "Pass rate: ${RATE}% ($PASSED/$TOTAL)"
          if (( $(echo "$RATE < 95" | bc -l) )); then
            echo "::warning::Integration pass rate ${RATE}% below 95%"
          fi
反模式
  • 跳过 CDSS 测试,因为"上次通过了"
  • 将严重关卡阈值设为低于 100%
  • 在严重测试套件中使用 --no-bail
  • 在集成测试中模拟 CDSS 引擎(必须测试真实逻辑)
  • 安全关卡为红色时仍允许部署
  • 在 CDSS 套件中运行测试时不使用 --coverage

示例

示例 1:本地运行所有严重关卡
bash
npx jest --testPathPattern='tests/cdss' --bail --ci --coverage && \
npx jest --testPathPattern='tests/security/phi' --bail --ci && \
npx jest --testPathPattern='tests/data-integrity' --bail --ci
示例 2:检查高优先级关卡通过率
bash
tmp_json=$(mktemp)
npx jest --testPathPattern='tests/clinical' --ci --json --outputFile="$tmp_json" || true
jq '{
  passed: (.numPassedTests // 0),
  total: (.numTotalTests // 0),
  rate: (if (.numTotalTests // 0) == 0 then 0 else ((.numPassedTests // 0) / (.numTotalTests // 1) * 100) end)
}' "$tmp_json"
# Expected: { "passed": 21, "total": 22, "rate": 95.45 }
示例 3:评估报告
## 医疗评估:2026-03-27 [commit abc1234]

### 患者安全:通过

| 类别 | 测试数 | 通过 | 失败 | 状态 |
|----------|-------|------|------|--------|
| CDSS 准确性 | 39 | 39 | 0 | 通过 |
| PHI 暴露 | 8 | 8 | 0 | 通过 |
| 数据完整性 | 12 | 12 | 0 | 通过 |
| 临床工作流 | 22 | 21 | 1 | 95.5% 通过 |
| 集成 | 6 | 6 | 0 | 通过 |

### 覆盖率:84%(目标:80%以上)
### 结论:可安全部署

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/zh-CN/skills/healthcare-eval-harness of affaan-m/ECC.

Open the folder on GitHubat commit 4eb71d9

Compare with similar skills

Healthcare Eval Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Healthcare Eval Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Healthcare Eval Harness this skillaffaan-m/ECC276k—~1.4kAutomated safety check: PassMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Skill Conductorsmixs/skill-conductor179—~6.6kAutomated safety check: PassMIT
Evaluate RAGai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.0
Evalagentevals-dev/agentevals163—~904Automated safety check: PassApache-2.0
Synthetic Eval Data Generatorai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Skill Conductor

    smixs/skill-conductor

    Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor.

    179 GitHub stars~6.6k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed
  • Eval

    agentevals-dev/agentevals

    Evaluate and score agent behavior against a golden reference.

    163 GitHub stars~904 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 16 days ago
    AI & LLM EngineeringAuto-check passed
  • Code Model Evaluation Harness

    Orchestra-Research/AI-Research-SKILLs

    Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.

    13k GitHub starsUsed in 4 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Questions about Healthcare Eval Harness

What does Healthcare Eval Harness do?

用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。. Healthcare Eval Harness is an agent skill from affaan-m/ECC.

When should I use Healthcare Eval Harness?

Healthcare Eval Harness fits situations like: tasks that involve LLM evaluation; tasks that involve Clinical and healthcare research.

How do I install Healthcare Eval Harness in Claude Code?

Run `npx skills add affaan-m/ECC --skill healthcare-eval-harness -a claude-code`. Or copy the skill folder (docs/zh-CN/skills/healthcare-eval-harness in affaan-m/ECC) into .claude/skills/healthcare-eval-harness in your project. Claude Code loads it when a task matches its description.

How do I install Healthcare Eval Harness in Codex?

Run `npx skills add affaan-m/ECC --skill healthcare-eval-harness -a codex`. Or copy the skill folder (docs/zh-CN/skills/healthcare-eval-harness in affaan-m/ECC) into .agents/skills/healthcare-eval-harness in your project. Codex loads it when a task matches its description.

Can I use Healthcare Eval Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill healthcare-eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/healthcare-eval-harness, .gemini/skills/healthcare-eval-harness, .github/skills/healthcare-eval-harness and .opencode/skills/healthcare-eval-harness in your project.

What does Healthcare Eval Harness need to run?

Going by SKILL.md and its folder, Healthcare Eval Harness needs the command-line tools its instructions call (npx and jq). Our summary lists: Node.js.

Does Healthcare Eval Harness access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Healthcare Eval Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Healthcare Eval Harness use?

Healthcare Eval Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Healthcare Eval Harness use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Healthcare Eval Harness?

Skills that share tags, products or a category with Healthcare Eval Harness: Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars), Skill Conductor (smixs/skill-conductor, 179 stars), Evaluate RAG (ai-evals-course/evals-skills, 1.5k stars) and Eval (agentevals-dev/agentevals, 163 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Healthcare Eval Harness?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,111 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 10, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.