Agent skill

Semantic Consistency Auditor

by aipoch in aipoch/medical-research-skills

Use semantic consistency auditor for academic writing workflows that need structured execution, explicit assumptions, and clear output boundaries.

MITAuto-check passedResearch & Science

Install Semantic Consistency Auditor

skills CLI
$ npx skills add aipoch/medical-research-skills --skill semantic-consistency-auditor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills semantic-consistency-auditor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Academic Writing/semantic-consistency-auditor' .claude/skills/semantic-consistency-auditor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
semantic-consistency-auditor
GitHub stars
2k
Token cost
~3.1k tokens
SKILL.md length
1,012 words
Files
5 (incl. scripts, references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

Use semantic consistency auditor for academic writing workflows that need structured execution, explicit assumptions, and clear output boundaries.

  • Works in 2 steps: BERTScore → COMET (Cross-lingual Optimized Metric…
  • Tasks that involve Scientific writing
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 17 more sections
  • Runs Python scripts from its folder; calls python

What it does

Semantic Consistency Auditor is an agent skill from aipoch/medical-research-skills. Use semantic consistency auditor for academic writing workflows that need structured execution, explicit assumptions, and clear output boundaries.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/audit-reference.md`, `scripts/main.py` and `semantic-consistency-auditor_audit_result_v2.json`).

It sits in Research & Science, covering Scientific writing. It works with Python. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Tasks that involve Scientific writing

Example prompts

  • “/semantic-consistency-auditor”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. BERTScore
  2. COMET (Cross-lingual Optimized Metric for Evaluation of Translation)

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Semantic Consistency Auditor loads about 3.1k tokens when it runs, and up to ~3.3k if it reads all its reference files. Until then it costs about 44 tokens; SKILL.md has 1,012 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 1,012 words, ~3,087 tokens.

Download SKILL.mdSave it as .claude/skills/semantic-consistency-auditor/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
semantic-consistency-auditor
description
Use semantic consistency auditor for academic writing workflows that need structured execution, explicit assumptions, and clear output boundaries.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

Skill: Semantic Consistency Auditor

ID: 212
Name: semantic-consistency-auditor
Description: Introduces BERTScore and COMET algorithms to evaluate the semantic consistency between AI-generated clinical notes and expert gold standards from the "semantic entailment" level.

When to Use

  • Use this skill when the task needs Use semantic consistency auditor for academic writing workflows that need structured execution, explicit assumptions, and clear output boundaries.
  • Use this skill for academic writing tasks that require explicit assumptions, bounded scope, and a reproducible output format.
  • Use this skill when you need a documented fallback path for missing inputs, execution errors, or partial evidence.

Key Features

  • Scope-focused workflow aligned to: Use semantic consistency auditor for academic writing workflows that need structured execution, explicit assumptions, and clear output boundaries.
  • Packaged executable path(s): scripts/main.py.
  • Reference material available in references/ for task-specific guidance.
  • Structured execution path designed to keep outputs consistent and reviewable.

Dependencies

See ## Prerequisites above for related details.

  • Python: 3.10+. Repository baseline for current packaged skills.
  • bert_score: unspecified. Declared in requirements.txt.
  • comet: unspecified. Declared in requirements.txt.
  • dataclasses: unspecified. Declared in requirements.txt.
  • numpy: unspecified. Declared in requirements.txt.
  • torch: unspecified. Declared in requirements.txt.
  • yaml: unspecified. Declared in requirements.txt.

Example Usage

See ## Usage above for related details.

bash
cd "20260318/scientific-skills/Academic Writing/semantic-consistency-auditor"
python -m py_compile scripts/main.py
python scripts/main.py --help

Example run plan:

  1. Confirm the user input, output path, and any required config values.
  2. Edit the in-file CONFIG block or documented parameters if the script uses fixed settings.
  3. Run python scripts/main.py with the validated inputs.
  4. Review the generated output and return the final artifact with any assumptions called out.

Implementation Details

See ## Workflow above for related details.

  • Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
  • Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
  • Primary implementation surface: scripts/main.py.
  • Reference guidance: references/ contains supporting rules, prompts, or checklists.
  • Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
  • Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.

Quick Check

Use this command to verify that the packaged script entry point can be parsed before deeper execution.

bash
python -m py_compile scripts/main.py

Audit-Ready Commands

Use these concrete commands for validation. They are intentionally self-contained and avoid placeholder paths.

bash
python -m py_compile scripts/main.py
python scripts/main.py --help

Workflow

  1. Confirm the user objective, required inputs, and non-negotiable constraints before doing detailed work.
  2. Validate that the request matches the documented scope and stop early if the task would require unsupported assumptions.
  3. Use the packaged script path or the documented reasoning path with only the inputs that are actually available.
  4. Return a structured result that separates assumptions, deliverables, risks, and unresolved items.
  5. If execution fails or inputs are incomplete, switch to the fallback path and state exactly what blocked full completion.

Overview

Semantic Consistency Auditor is a medical AI evaluation tool used to assess the semantic consistency between AI-generated clinical notes and expert-written gold standards from a semantic level. This tool is not limited to traditional string matching or bag-of-words models, but uses deep learning models to understand semantic entailment relationships, capable of identifying expressions with different wording but similar meaning.

Algorithms

1. BERTScore

BERTScore uses pre-trained BERT model contextual embeddings to calculate similarity between candidate text and reference text:

  • Precision: How much semantics in the candidate text is covered by the reference text
  • Recall: How much semantics in the reference text is covered by the candidate text
  • F1 Score: Harmonic mean of Precision and Recall
2. COMET (Cross-lingual Optimized Metric for Evaluation of Translation)

COMET is a neural network-based evaluation metric originally used for machine translation evaluation, applicable to semantic entailment tasks:

  • Uses XLM-RoBERTa encoder to capture deep semantics
  • Outputs semantic consistency scores between 0-1
  • Gives high scores to semantically equivalent but differently expressed text
Show full SKILL.md (396 more words)Show less

Installation

text

# Create virtual environment (recommended)
python -m venv venv
source venv/bin/activate  # Linux/Mac

# Or venv\Scripts\activate  # Windows

# Install dependencies
pip install bertscore comet-ml transformers torch

Configuration

Configure in ~/.openclaw/skills/semantic-consistency-auditor/config.yaml:

yaml

# BERTScore Configuration
bertscore:
  model: "microsoft/deberta-xlarge-mnli"  # Or "bert-base-chinese" for Chinese
  lang: "zh"  # Language code: zh, en, etc.
  rescale_with_baseline: true
  device: "auto"  # auto, cpu, cuda

# COMET Configuration
comet:
  model: "Unbabel/wmt22-comet-da"  # COMET model
  batch_size: 8
  device: "auto"

# Evaluation Thresholds
thresholds:
  bertscore_f1: 0.85
  comet_score: 0.75
  semantic_consistency: 0.80  # Comprehensive score threshold

Usage

Command Line
text

# Evaluate single case pair
python scripts/main.py \
  --ai-generated "Patient presented with fever for 3 days, highest temperature 39°C, accompanied by cough." \
  --gold-standard "Patient chief complaint of fever for 3 days, highest temperature 39°C, accompanied by cough symptoms." \
  --output results.json

# Batch evaluation from JSON file
python scripts/main.py \
  --input-file batch_cases.json \
  --output results.json \
  --format detailed

# Use specific model
python scripts/main.py \
  --ai-generated "..." \
  --gold-standard "..." \
  --bert-model "bert-base-chinese" \
  --comet-model "Unbabel/wmt20-comet-da"
Python API
python
from semantic_consistency_auditor import SemanticConsistencyAuditor

# Initialize evaluator
auditor = SemanticConsistencyAuditor(
    bert_model="microsoft/deberta-xlarge-mnli",
    comet_model="Unbabel/wmt22-comet-da",
    lang="zh"
)

# Evaluate single case
result = auditor.evaluate(
    ai_text="Patient presented with fever for 3 days...",
    gold_text="Patient chief complaint of fever for 3 days..."
)

print(f"BERTScore F1: {result['bertscore']['f1']:.4f}")
print(f"COMET Score: {result['comet']['score']:.4f}")
print(f"Consistency: {result['consistency']:.4f}")
print(f"Passed: {result['passed']}")

# Batch evaluation
results = auditor.evaluate_batch([
    {"ai": "...", "gold": "..."},
    {"ai": "...", "gold": "..."}
])

Input Format

Single Case (Command Line)

Pass text directly through --ai-generated and --gold-standard parameters.

Batch Evaluation File (JSON)
json
[
  {
    "case_id": "CASE001",
    "ai_generated": "Patient presented with fever for 3 days, highest temperature 39°C, accompanied by cough.",
    "gold_standard": "Patient chief complaint of fever for 3 days, highest temperature 39°C, accompanied by cough symptoms.",
    "metadata": {
      "department": "Respiratory",
      "disease_type": "Upper respiratory infection"
    }
  },
  {
    "case_id": "CASE002",
    "ai_generated": "...",
    "gold_standard": "..."
  }
]

Output Format

Summary Mode
json
{
  "overall": {
    "total_cases": 100,
    "passed_cases": 85,
    "pass_rate": 0.85,
    "avg_bertscore_f1": 0.8923,
    "avg_comet_score": 0.8234,
    "avg_consistency": 0.8579
  },
  "thresholds": {
    "bertscore_f1": 0.85,
    "comet_score": 0.75,
    "semantic_consistency": 0.80
  }
}
Detailed Mode
json
{
  "cases": [
    {
      "case_id": "CASE001",
      "ai_generated": "Patient presented with fever for 3 days...",
      "gold_standard": "Patient chief complaint of fever for 3 days...",
      "metrics": {
        "bertscore": {
          "precision": 0.9123,
          "recall": 0.8934,
          "f1": 0.9028
        },
        "comet": {
          "score": 0.8234,
          "system_score": 0.8156
        },
        "semantic_consistency": 0.8631
      },
      "passed": true,
      "details": {
        "semantic_gaps": [],
        "matched_concepts": ["fever for 3 days", "temperature 39°C", "cough"]
      }
    }
  ],
  "summary": { ... }
}

Error Handling

  • If required inputs are missing, state exactly which fields are missing and request only the minimum additional information.
  • If the task goes outside the documented scope, stop instead of guessing or silently widening the assignment.
  • If scripts/main.py fails, report the failure point, summarize what still can be completed safely, and provide a manual fallback.
  • Do not fabricate files, citations, data, search results, or execution outcomes.

Performance Notes

  • BERTScore: First run will download model (approximately 400MB-1GB)
  • COMET: First run will download model (approximately 500MB-1.5GB)
  • GPU Acceleration: Significantly improves evaluation speed in CUDA environment
  • Batch Processing: Recommended for batch evaluation to fully utilize GPU parallel capability

References

  1. Zhang et al. "BERTScore: Evaluating Text Generation with BERT" ICLR 2020
  2. Rei et al. "COMET: A Neural Framework for MT Evaluation" EMNLP 2020
  3. Medical Record Standardization Evaluation Guidelines (National Health Commission)

Changelog

  • v1.0.0 (2026-02-06): Initial version, supports dual-algorithm evaluation with BERTScore and COMET

Prerequisites

text

# Python dependencies
pip install -r requirements.txt

Evaluation Criteria

Success Metrics
  • Successfully executes main functionality
  • Output meets quality standards
  • Handles edge cases gracefully
  • Performance is acceptable
Test Cases
  1. Basic Functionality: Standard input → Expected output
  2. Edge Case: Invalid input → Graceful error handling
  3. Performance: Large dataset → Acceptable processing time

Output Requirements

Every final response should make these items explicit when they are relevant:

  • Objective or requested deliverable
  • Inputs used and assumptions introduced
  • Workflow or decision path
  • Core result, recommendation, or artifact
  • Constraints, risks, caveats, or validation needs
  • Unresolved items and next-step checks

Input Validation

This skill accepts requests that match the documented purpose of semantic-consistency-auditor and include enough context to complete the workflow safely.

Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:

semantic-consistency-auditor only handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.

References

Response Template

Use the following fixed structure for non-trivial requests:

  1. Objective
  2. Inputs Received
  3. Assumptions
  4. Workflow
  5. Deliverable
  6. Risks and Limits
  7. Next Checks

If the request is simple, you may compress the structure, but still keep assumptions and limits explicit when they affect correctness.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in scientific-skills/Academic Writing/semantic-consistency-auditor of aipoch/medical-research-skills.

  • SKILL.md
  • references/audit-reference.md
  • requirements.txt
  • scripts/main.py
  • semantic-consistency-auditor_audit_result_v2.json

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Semantic Consistency Auditor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Semantic Consistency Auditor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Semantic Consistency Auditor this skillaipoch/medical-research-skills2k—~3.1kAutomated safety check: PassMIT
Nature-Style Scientific FiguresYuan1z0825/nature-skills47k—~2.9kAutomated safety check: PassApache-2.0
Citation ManagementK-Dense-AI/claude-scientific-writer2.4k2 repos~3.9kAutomated safety check: NotesMIT
Autonomous Researchfedericodeponte/opendraft507—~8.2kAutomated safety check: PassApache-2.0
Hugging Face Paper Publisherhuggingface/skills11k4 repos~4.2kAutomated safety check: PassApache-2.0
NSFC Abstract Writerhuangwb8/ChineseResearchLaTeX2.9k1 repos~1.3kAutomated safety check: PassMIT

Similar skills

  • Nature-Style Scientific Figures

    Yuan1z0825/nature-skills

    Creates, revises, audits and exports manuscript-ready scientific figures in Python or R, and routes AI-generated graphical abstracts to a separate workflow.

    47k GitHub stars~2.9k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Citation Management

    K-Dense-AI/claude-scientific-writer

    Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.

    2.4k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Autonomous Research

    federicodeponte/opendraft

    An 18-agent pipeline that turns one topic line into a drafted research paper, literature review, or thesis chapter.

    507 GitHub stars~8.2k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed
  • Official

    Indexes research papers on the Hugging Face Hub from arXiv, links them to models and datasets, claims authorship and generates markdown research articles from templates.

    11k GitHub starsUsed in 4 repos~4.2k tokens
    Research & ScienceAuto-check passed
  • NSFC Abstract Writer

    huangwb8/ChineseResearchLaTeX

    Writes Chinese and English abstracts for NSFC grant applications, with a recommended title and five alternatives, within set character limits.

    2.9k GitHub starsUsed in 1 repo~1.3k tokens
    Research & ScienceAuto-check passed
  • Survey Paper Generator

    dair-ai/dair-academy-plugins

    Builds a single-file HTML survey paper on an AI or ML topic from a research bundle the agent curates, with prose and SVG figures written by Kimi K2.6.

    614 GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check: notes

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 22 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 22 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 22 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 22 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 22 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 22 days ago
    Auto-check passed

Works with

Questions about Semantic Consistency Auditor

What does Semantic Consistency Auditor do?

Use semantic consistency auditor for academic writing workflows that need structured execution, explicit assumptions, and clear output boundaries. Semantic Consistency Auditor is an agent skill from aipoch/medical-research-skills. Use semantic consistency auditor for academic writing workflows that need structured execution, explicit assumptions, and clear output boundaries.

When should I use Semantic Consistency Auditor?

Semantic Consistency Auditor fits situations like: tasks that involve Scientific writing.

How do I install Semantic Consistency Auditor in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill semantic-consistency-auditor -a claude-code`. Or copy the skill folder (scientific-skills/Academic Writing/semantic-consistency-auditor in aipoch/medical-research-skills) into .claude/skills/semantic-consistency-auditor in your project. Claude Code loads it when a task matches its description.

How do I install Semantic Consistency Auditor in Codex?

Run `npx skills add aipoch/medical-research-skills --skill semantic-consistency-auditor -a codex`. Or copy the skill folder (scientific-skills/Academic Writing/semantic-consistency-auditor in aipoch/medical-research-skills) into .agents/skills/semantic-consistency-auditor in your project. Codex loads it when a task matches its description.

Can I use Semantic Consistency Auditor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill semantic-consistency-auditor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/semantic-consistency-auditor, .gemini/skills/semantic-consistency-auditor, .github/skills/semantic-consistency-auditor and .opencode/skills/semantic-consistency-auditor in your project.

What does Semantic Consistency Auditor need to run?

Going by SKILL.md and its folder, Semantic Consistency Auditor needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Semantic Consistency Auditor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Semantic Consistency Auditor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Semantic Consistency Auditor use?

Semantic Consistency Auditor is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Semantic Consistency Auditor use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 180 tokens, read only when the agent opens those files.

What are the alternatives to Semantic Consistency Auditor?

Skills that share tags, products or a category with Semantic Consistency Auditor: Nature-Style Scientific Figures (Yuan1z0825/nature-skills, 47k stars), Citation Management (K-Dense-AI/claude-scientific-writer, 2.4k stars), Autonomous Research (federicodeponte/opendraft, 507 stars) and Hugging Face Paper Publisher (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Semantic Consistency Auditor?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,978 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.