Legal Test Builder Patrick Munro
lawve-ai/awesome-legal-skills
Builds a high-fidelity interactive legal assessment as a single self-contained HTML artifact.
A comprehensive auditor for any agent skill — including Manus, OpenClaw/ClawHub, Claude, LobeHub, or custom SKILL.md-based skills.
$ npx skills add aipoch/medical-research-skills --skill skill-auditor -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aipoch/medical-research-skills skill-auditor --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skill-auditor .claude/skills/skill-auditor && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skill-auditor" agent skill from https://github.com/aipoch/medical-research-skills/tree/main/skill-auditor into .claude/skills/skill-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-auditor", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aipoch/medical-research-skills/tree/main/skill-auditorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aipoch/medical-research-skills --skill skill-auditor -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aipoch/medical-research-skills skill-auditor --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skill-auditor .agents/skills/skill-auditor && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skill-auditor" agent skill from https://github.com/aipoch/medical-research-skills/tree/main/skill-auditor into .agents/skills/skill-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-auditor", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aipoch/medical-research-skills --skill skill-auditor -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aipoch/medical-research-skills skill-auditor --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skill-auditor .cursor/skills/skill-auditor && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skill-auditor" agent skill from https://github.com/aipoch/medical-research-skills/tree/main/skill-auditor into .cursor/skills/skill-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-auditor", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aipoch/medical-research-skills.git --path skill-auditor--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aipoch/medical-research-skills --skill skill-auditor -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aipoch/medical-research-skills skill-auditor --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skill-auditor .gemini/skills/skill-auditor && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skill-auditor" agent skill from https://github.com/aipoch/medical-research-skills/tree/main/skill-auditor into .gemini/skills/skill-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-auditor", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aipoch/medical-research-skills skill-auditorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aipoch/medical-research-skills --skill skill-auditor -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skill-auditor .github/skills/skill-auditor && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skill-auditor" agent skill from https://github.com/aipoch/medical-research-skills/tree/main/skill-auditor into .github/skills/skill-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-auditor", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aipoch/medical-research-skills --skill skill-auditor -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aipoch/medical-research-skills skill-auditor --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skill-auditor .opencode/skills/skill-auditor && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skill-auditor" agent skill from https://github.com/aipoch/medical-research-skills/tree/main/skill-auditor into .opencode/skills/skill-auditor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-auditor", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skill-auditorA comprehensive auditor for any agent skill — including Manus, OpenClaw/ClawHub, Claude, LobeHub, or custom SKILL.md-based skills.
Skill Auditor is an agent skill from aipoch/medical-research-skills. A comprehensive auditor for any agent skill — including Manus, OpenClaw/ClawHub, Claude, LobeHub, or custom SKILL.md-based skills. Use this skill whenever a user wants to evaluate, audit, review, score, or quality-check an agent skill before publishing, updating, or deploying. Covers two hard veto gates (structural redlines + research integrity redlines), static quality scoring across 25 criteria (ISO 25010 + OpenSSF + Agent), dynamic test input generation, multi-mode execution testing, multi-layer output…
Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts and reference files (for example `README.md`, `references/basic_evaluation.md` and `references/basic_veto.md`).
It sits in Legal & Compliance, covering Contract review, Quizzes and assessments and Data analysis. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skill Auditor loads about 6.9k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 254 tokens; SKILL.md has 2,455 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 2,455 words, ~6,865 tokens.
.claude/skills/skill-auditor/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.This skill provides a standardized, end-to-end process for auditing any agent skill — from structural integrity to live functional performance. It combines two independent veto gates, static analysis across 25 criteria, dynamic execution, five-category specialized scoring, and human review into a single coherent workflow.
Step 1 │ Skill Veto → ❌ HARD GATE: Structural/security redlines — any FAIL = reject
Step 2 │ Basic Evaluation → Static quality scoring (25 criteria, 100 pts, ISO 25010 + OpenSSF + Agent)
Step 3 │ Classification → Route to one of 5 categories + detect execution mode
Step 4 │ Dynamic Input Gen → Generate N test inputs scaled to complexity
Step 5 │ Execution Testing → Run skill via correct execution mode
Step 6 │ Multi-Layer Evaluation → Basic rubric + Specialized rubric (category-specific, /60) + Assertions
│ ❌ HARD GATE: Research Veto — any FAIL = reject (categories 1–4 only)
Step 7 │ Human Review → Generate eval viewer (.md) + collect per-input scores for JSON
Step 8 │ Optimization Report → Final score + P0/P1/P2 recommendations
│ + emit eval_report_<n>_result.json for frontend visualizationTwo hard rejection gates run at Steps 1 and 6. Both are mandatory and cannot be skipped. All other steps run sequentially. Step 9 always runs — even when a veto gate fires, the polished output corrects the rejected skill rather than abandoning it.
All audit output must be written in English, regardless of the language used in the user's request or in the submitted skill.
This applies to every artifact produced by this skill:
.md file (Step 7)If the user communicates in another language, Claude may briefly acknowledge the request in that language, but must then conduct and present the full audit in English.
HARD GATE. Any FAIL = immediate rejection. Do not proceed to Step 2.
Read the target skill's SKILL.md and any bundled scripts. Check all four dimensions:
→ Full criteria: references/basic_veto.md
| Dimension | Immediate Rejection Triggers |
|---|---|
| T1. Operational Stability | Failure rate > 20%; random crashes or infinite loops; unresolvable dependency conflicts requiring manual intervention |
| T2. Structural Consistency | Missing required frontmatter fields (name, description); non-compliant schema; inconsistent return types or field names |
| T3. Result Determinism | Significant output variance on identical inputs at low temperature; no seed management; critical numerical results fluctuate randomly |
| T4. System Security | Direct execution of raw user-provided strings (eval/exec); no input filtering; prompt injection vectors present in scripts or instructions |
► If any T1–T4 dimension is FAIL: stop immediately. Output the rejection report below and do not continue.
SKILL VETO — REJECTED
══════════════════════════════════
Skill: <n>
Reason: Failed structural redline check
T1. Stability : PASS / FAIL — <reason>
T2. Contract : PASS / FAIL — <reason>
T3. Determinism : PASS / FAIL — <reason>
T4. Security : PASS / FAIL — <reason>
This skill must not be deployed. Fix all FAIL dimensions before resubmitting.
══════════════════════════════════Read the skill's full SKILL.md and all bundled files. Score each of the 25 criteria from 0–4.
→ Full rubric with per-level descriptions: references/basic_evaluation.md
| # | Category (Framework) | Criteria | Max |
|---|---|---|---|
| 1 | Functional Suitability (ISO 25010) | Completeness, Correctness, Appropriateness | 12 |
| 2 | Reliability (ISO 25010) | Fault Tolerance, Error Reporting, Recoverability | 12 |
| 3 | Performance & Context (ISO 25010 + Agent) | Token Cost, Execution Efficiency | 8 |
| 4 | Agent Usability (Shneiderman · Gerhardt-Powals) | Learnability, Consistency, Feedback Design, Error Prevention | 16 |
| 5 | Human Usability (Tognazzini · Norman) | Discoverability, Forgiveness | 8 |
| 6 | Security (ISO 25010 + OpenSSF) | Credential Safety, Input Validation, Data Safety | 12 |
| 7 | Maintainability (ISO 25010) | Modularity, Modifiability, Testability | 12 |
| 8 | Agent-Specific (Novel) | Trigger Precision, Progressive Disclosure, Composability, Idempotency, Escape Hatches | 20 |
Basic subtotal: __ / 100
Read the skill's description frontmatter and ## When to Use section.
→ Full category definitions: references/classification.md
| # | Category | Typical Skills |
|---|---|---|
| 1 | Evidence Insight | Search strategy builders, database scouts, critical appraisal tools, evidence synthesizers |
| 2 | Protocol Design | Experimental design generators, study-type advisors, statistical power planners, validation strategists |
| 3 | Data Analysis | R/Python code generators, bioinformatics pipelines, statistical modeling tools, ML workflows |
| 4 | Academic Writing | SCI manuscript writers, abstract generators, methods/discussion drafters, cover-letter tools |
| 5 | Other (General / Non-Research) | All skills that do not fall into categories 1–4 |
Research Veto scope: Categories 1–4 are subject to the Research Veto hard gate in Step 6. Category 5 is exempt.
Inspect the skill to determine how it is meant to be invoked.
| Mode | Indicators | How to Run in Step 5 |
|---|---|---|
| A: Direct | Only SKILL.md instructions, no scripts | Follow SKILL.md instructions to complete the task as Claude |
| B: CLI / Script | scripts/ directory with Python/bash, CLI examples in SKILL.md | Execute via bash: python scripts/xxx.py <args> |
| C: API | API endpoint patterns, fetch/curl usage in SKILL.md | Simulate or call the API as documented |
| D: Hybrid | Both instructions and scripts/API | Run script for deterministic parts; Claude for reasoning/generation parts |
Record the detected mode. It will be used in Step 5.
Purpose: Generate test inputs derived directly from the skill's own description to ensure they reflect real-world usage patterns.
| Complexity Level | Criteria | Test Input Count (N) |
|---|---|---|
| Simple | Single task type, narrow scope, < 3 reference files, no branching workflow | 3 inputs |
| Moderate | 2–3 task types, some branching, 3–5 reference files, moderate scope | 5 inputs |
| Complex | Multiple task types, branching logic, 5+ reference files, broad or specialized scope | 7 inputs |
Declare: Complexity: [Simple / Moderate / Complex] → Generating N inputs
Use this distribution based on N:
| Slot | Type | Always Include? |
|---|---|---|
| Input 1 | Canonical / happy path | ✅ Always |
| Input 2 | Variant A (different valid use case) | ✅ Always |
| Input 3 | Edge / boundary | ✅ Always |
| Input 4 | Variant B (third central use case) | If N ≥ 5 |
| Input 5 | Stress / complex / multi-part | If N ≥ 5 |
| Input 6 | Scope boundary (slightly outside) | If N = 7 |
| Input 7 | Adversarial / ambiguous | If N = 7 |
Format each as a realistic user message. Do not include expected answers.
Output:
GENERATED TEST INPUTS
═══════════════════════════════════════
Skill: <n> | Category: <1–5 + label> | Mode: <A/B/C/D>
Complexity: <level> → Generating <N> inputs
Input 1 (Canonical) : <prompt>
Input 2 (Variant A) : <prompt>
Input 3 (Edge) : <prompt>
[Input 4–7 if applicable]
═══════════════════════════════════════Run the skill on each of the N test inputs using the detected execution mode.
Load the skill's SKILL.md. Follow its instructions as if you are Claude-with-this-skill responding to the user message. Complete the task in full.
python scripts/<script_name>.py "<input_text_or_path>"Capture stdout/stderr. If execution fails, record the error and continue.
Follow the API usage pattern documented in the skill. Construct the request, execute, capture response. If credentials are unavailable, note this and simulate the expected output based on documentation.
Run the script/API component first. Pass its output to Claude for the reasoning/generation component.
For each input:
─── Input [N] ─────────────────────────
Mode : <A/B/C/D>
Input : <prompt>
Output :
<full output>
Status : COMPLETED / ERROR / PARTIAL
Notes : <anomalies, scope violations, unexpected behaviors>
────────────────────────────────────────Evaluate all N outputs across three parallel layers.
→ Full criteria: references/basic_evaluation.md
For each output, score four aggregate dimensions (0–10 each):
Per-output basic score: /40
Apply the rubric corresponding to the category from Step 3:
| Category | Reference File | Max |
|---|---|---|
| 1 — Evidence Insight | references/specialized_evaluation_literature.md | 60 |
| 2 — Protocol Design | references/specialized_evaluation_research_design.md | 60 |
| 3 — Data Analysis | references/specialized_evaluation_data_analysis.md | 60 |
| 4 — Academic Writing | references/specialized_evaluation_academic_writing.md | 60 |
| 5 — Other | references/specialized_evaluation_other.md | 60 |
Per-output specialized score: /60
For each output, write and evaluate 3–5 boolean assertions derived from the skill's stated purpose.
Assertion types:
Format:
Assertions for Input [N]:
[PASS/FAIL] <assertion statement> — <brief justification>
[PASS/FAIL] <assertion statement> — <brief justification>
[PASS/FAIL] <assertion statement> — <brief justification>
Assertion pass rate: X/YGATE: Any safety or scope assertion FAIL on 2+ outputs → Flag for rejection review.
Per-output total = Basic Score (Layer 1, /40) + Specialized Score (Layer 2, /60) = /100
Execution avg score = Mean of all N output totals
Assertion summary = Total passed / Total assertionsHARD GATE. Run after all N outputs are evaluated. Any FAIL = immediate rejection. Skip entirely for Category 5 (Other).
→ Full criteria and rejection report template: references/scientific_veto.md
| Dimension | Immediate Rejection Triggers |
|---|---|
| M1. Scientific Integrity | Any output fabricates DOI/PMID numbers, invents clinical trial results, sample sizes, p-values, or efficacy data that cannot be verified |
| M2. Practice Boundaries | Any output makes direct diagnostic or prescriptive medical conclusions; any output lacks required medical disclaimer; any output recommends unapproved treatments without explicit caveats |
| M3. Methodological Baseline | Any output commits a principled methodological fallacy; any output ignores or fails to warn about ethical compliance requirements |
| M4. Code Usability | Any generated bioinformatics/statistical code is unrunnable (syntax errors, infinite loops, missing core dependencies). Mark N/A for categories 1 & 4 if no code is generated. |
► If any M1–M4 dimension is FAIL: stop. Output the Research Veto rejection report. Do not proceed to Steps 7–8.
Generate a structured Markdown review document for human inspection.
# Eval Viewer — <Skill Name>
Generated: <date>
## Summary Table
| Input | Type | Basic /40 | Specialized /60 | Total /100 | Assertions | Status |
|---|---|---|---|---|---|---|
| 1 | Canonical | __ | __ | __ | X/Y PASS | ✅/⚠️/❌ |
...
**Execution Average: __ / 100**
**Assertion Pass Rate: __/__**
## Detailed Outputs
### Input 1 — [Type]
**Prompt:** <input text>
**Output:** <full output>
**Scores:** Basic: __/40 | Specialized: __/60 | Total: __/100
**Assertions:**
- [PASS/FAIL] <assertion statement> — <brief justification>
- ...Output location. Save the eval viewer next to the audited skill so it stays with the artifact it describes:
<audited_skill_path>/eval_viewer_<skill_name>.mdWhere <audited_skill_path> is the directory containing the skill's SKILL.md (the path the user passed in). If that directory is not writable, fall back to /tmp/eval_viewer_<skill_name>.md and note the fallback in the final report.
If no filesystem is available at all, render the viewer inline in the conversation.
Note for reviewer: Check ⚠️ and ❌ rows first. Patterns across 2+ outputs indicate structural skill issues.
Final Score = (Static Score × 0.4) + (Execution Avg × 0.6)→ Full scoring thresholds: references/scoring_rubric.md
| Score | Grade | Recommendation |
|---|---|---|
| 85–100 | ⭐ Production Ready | Deploy publicly |
| 75–84 | ✅ Limited Release | Throttled / monitored rollout |
| 60–74 | ⚠️ Beta Only | Internal / greylist only |
| < 60 | ❌ Reject | Do not deploy |
| Priority | Criteria | Action |
|---|---|---|
| P0 — Blocker | Any veto FAIL, safety assertion FAIL, score < 60 | Must fix before any deployment |
| P1 — Major | Score 60–74, repeated assertion failures, Layer 1 or 2 avg < 7/10 | Fix before production release |
| P2 — Minor | Score 75–84, isolated output weaknesses, style/format issues | Address before full scale |
For each issue found, output:
[P0/P1/P2] <Issue Title>
Observed in: Input(s) [N, M, ...]
Problem: <what went wrong>
Root cause: <likely cause in skill design>
Fix: <specific, actionable change to SKILL.md or scripts>══════════════════════════════════════════════════
SKILL AUDIT REPORT
══════════════════════════════════════════════════
Skill Name : <n>
Category : <label>
Execution Mode : <A/B/C/D>
Complexity : <Simple/Moderate/Complex> (N=<n> inputs)
Audited On : <date>
── STEP 1: Structural Veto ───────────────────────
Stability : PASS / FAIL
Contract : PASS / FAIL
Determinism : PASS / FAIL
Security : PASS / FAIL
── STEP 2: Static Evaluation (25 criteria) ───────
Functional Suitability : __/12
Reliability : __/12
Performance/Context : __/8
Agent Usability : __/16
Human Usability : __/8
Security : __/12
Maintainability : __/12
Agent-Specific : __/20
Static Subtotal : __/100
── STEP 3: Classification ────────────────────────
Category : <label>
Execution Mode : <A / B / C / D>
── STEP 4: Test Inputs ───────────────────────────
[N inputs listed with type labels]
── STEP 5: Execution Summary ─────────────────────
Input 1: [COMPLETED/PARTIAL/ERROR] — <note>
...
── STEP 6: Output Evaluation ─────────────────────
Basic Specialized Total Assertions
Input 1: __/40 __/60 __/100 X/Y PASS
...
Execution Avg : __/100
Total Assertion Pass Rate : __/__
[Research Veto — Evidence Insight / Protocol Design / Data Analysis / Academic Writing only; N/A for Other]
Scientific Integrity : PASS / FAIL / N/A
Practice Boundaries : PASS / FAIL / N/A
Methodological Ground : PASS / FAIL / N/A
Code Usability : PASS / FAIL / N/A
── STEP 7: Outputs ───────────────────────────────
<audited_skill_path>/eval_viewer_<n>.md : SAVED ✅
<audited_skill_path>/eval_report_<n>_result.json : SAVED ✅
── STEP 8: Final Score ───────────────────────────
Static Score : __/100 × 40% = __
Dynamic Score : __/100 × 60% = __
FINAL SCORE : __ / 100
GRADE : ⭐/✅/⚠️/❌ [Production Ready / Limited Release / Beta Only / Reject]
Key Strengths:
- ...
Optimization Recommendations:
[P0] ...
[P1] ...
[P2] ...
══════════════════════════════════════════════════⚠️ STRICT SCHEMA COMPLIANCE — NOT OPTIONAL. The JSON report is consumed by downstream tooling (skill-evaluator, the frontend viewer, automated regression dashboards). Any deviation from the schema below — a renamed key, a missing field, a string where an object is expected — breaks that tooling silently. Before emitting the JSON, you MUST:
- Re-read
references/report_json_schema.mdin full (not just the summary table below).- Build the JSON by copying the structure from the schema's complete example, then filling in values — do not invent your own key names or nesting.
- Run every item in the schema's Pre-Emit Checklist (§ "Pre-Emit Checklist" in
report_json_schema.md) before writing the file. Treat each unchecked item as a blocker, not a warning.Common drift to watch for — these have been observed in past audits and each one breaks the contract:
- Using
static_score.totalinstead ofstatic_score.subtotal.- Using
dynamic_score.execution_avg_scoreinstead ofdynamic_score.execution_avg.- Writing
veto_gates.research_veto.scientific_integrity: "PASS"(a bare string) instead of{ "result": "PASS", "detail": "..." }.- Omitting the top-level
gatefield insideskill_vetoandresearch_veto.- Per-input keys:
ninstead ofindex, missinglabel/status/status_flag/note.- Each static category value being a bare integer instead of
{ "score": N, "max": M, "note": "..." }.final.final_score/final.weighted_static/final.weighted_dynamicinstead offinal.score/final.static_weighted/final.dynamic_weighted.- Missing
final.grade_symbol.If a section of the schema's example does not match what you are about to emit, the example is canonical — change your output, not the schema.
→ Full schema + complete example: references/report_json_schema.md
Output location. Save the JSON next to the audited skill, using the same directory rule as the eval viewer:
<audited_skill_path>/eval_report_<skill_name>_result.jsonIf the audited skill directory already contains a prior eval_report_<skill_name>_result.json, overwrite it — the latest audit supersedes earlier ones. If that directory is not writable, fall back to /tmp/eval_report_<skill_name>_result.json and note the fallback in the final report.
JSON top-level nodes (all 7 required) — abbreviated reminder; the schema file is canonical:
| Node | Key rules |
|---|---|
meta | evaluator_version: "skill-auditor@1.0"; includes skill_name, description, evaluated_on, category, execution_mode, complexity, n_inputs |
veto_gates | skill_veto: top-level gate + stability, contract, determinism, security — no T-prefixes. research_veto: top-level applicable, gate + four dimensions, each as { result, detail } — no M-prefixes |
static_score | subtotal, max, categories. The 8 category keys must be exact and un-prefixed: functional_suitability, reliability, performance_context, agent_usability, human_usability, security, maintainability, agent_specific. Each value is an object { score, max, note }, never a bare integer. |
dynamic_score | execution_avg (not execution_avg_score), max, assertion_pass_rate: { passed, total }, inputs[]. Each input uses index (not n) and includes label, status, status_flag, note, basic, specialized, total, assertions_passed, assertions_total, and a full assertions[] array of { text, result, note } objects. |
final | static_weighted, dynamic_weighted, score (not final_score), max, grade, grade_symbol, deployable, veto_override |
key_strengths | plain-string array, 2–5 entries |
recommendations | P0 → P1 → P2 sorted, each with priority, title, observed_in, problem, root_cause, fix |
Pre-emit checklist (MUST verify all of these before writing the file — see full list in report_json_schema.md § Pre-Emit Checklist):
static_score.categories has exactly 8 keys; each value is { score, max, note }; all scores within 0–maxstatic_score.subtotal equals the sum of all 8 category score valuesdynamic_score.inputs has exactly N objects (matching meta.n_inputs); each has an assertions arrayassertions array has 3–5 entries (cardinality constraint — count it)assertions_passed equals count of "PASS" in its assertions arraybasic + specialized = totalkey_strengths has 2–5 entries (cardinality constraint — count it)research_veto.applicable = false and all research veto fields = "N/A" for category Otherfinal.veto_override = true if any gate is FAIL; final.deployable = false in that casefinal.grade and final.grade_symbol consistent with final.score per the threshold tablerecommendations sorted P0 → P1 → P2 (empty array allowed)This skill accepts: a SKILL.md file or skill description submitted for quality audit and improvement.
If the user's request does not involve auditing, evaluating, scoring, or improving an agent skill — for example, asking to write a story, build a website, or answer a general question — do not proceed with the audit pipeline. Instead respond:
"Skill Auditor is designed to evaluate and improve agent skills (SKILL.md files). Your request appears to be outside this scope. Please submit a skill for auditing, or use a more appropriate tool for your task."
Language note: The user may submit requests or skills in any language. Always produce the full audit output in English. See Language Policy above.
| File | Used In | Gate? |
|---|---|---|
references/basic_veto.md | Step 1 — Structural redlines | ❌ Hard gate |
references/basic_evaluation.md | Step 2 (static scoring) + Step 6 Layer 1 | — |
references/classification.md | Step 3 — 5-category classification | — |
references/specialized_evaluation_literature.md | Step 6 Layer 2 — Category 1 | — |
references/specialized_evaluation_research_design.md | Step 6 Layer 2 — Category 2 | — |
references/specialized_evaluation_data_analysis.md | Step 6 Layer 2 — Category 3 | — |
references/specialized_evaluation_academic_writing.md | Step 6 Layer 2 — Category 4 | — |
references/specialized_evaluation_other.md | Step 6 Layer 2 — Category 5 | — |
references/scientific_veto.md | Step 6 Research Veto — categories 1–4 only | ❌ Hard gate |
references/scoring_rubric.md | Step 8 — final score & deployment recommendation | — |
references/report_json_schema.md | Step 7 (data collection) + Step 8 (JSON output) | — |
scripts/evaluate_skill.py (structural pre-checks)Scene Override additions — Based on findings from the audit of differential-expression-analysis, three systematic biases were identified in basic_evaluation.md when applied to scientific computing + agent-first skills. Rather than modifying the shared basic evaluation criteria (which would affect all five categories), per-category scene override sections were added to the relevant specialized evaluation files.
Files modified:
references/specialized_evaluation_data_analysis.md — Added Scene Override section covering Fault Tolerance (2.1), Forgiveness (5.2), and Recoverability (2.3)references/specialized_evaluation_research_design.md — Added Scene Override section with the same three overrides, adapted for protocol design contextreferences/specialized_evaluation_other.md — Added Execution Mode Awareness note directing auditors to apply Category 3 overrides when the skill operates in agent-first Mode B/C/D contextRationale: The three affected basic evaluation criteria assume (1) human direct CLI operation and (2) general-purpose software tools. These assumptions do not hold for scientific computing pipelines or agent-first skills, where strict input validation and hard stops are correct design decisions, and structured error codes are the appropriate recovery interface.
© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 13 other files (scripts, references) in skill-auditor of aipoch/medical-research-skills.
Open the folder on GitHubat commit 686e09d
Skill Auditor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skill Auditor this skillaipoch/medical-research-skills | 2k | — | ~6.9k | Automated safety check: Pass | MIT | |
| Legal Test Builder Patrick Munrolawve-ai/awesome-legal-skills | 836 | — | ~2.3k | Automated safety check: Pass | AGPL-3.0 | |
| Weak Agent Testkklimuk/docx-cli | 216 | — | ~6.1k | Automated safety check: Notes | MIT | |
| Data Flow Mapperrevfactory/harness-100 | 1.3k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Respol Data Analysisbrycewang-stanford/Awesome-Journal-Skills | 1.2k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Respol Tables Figuresbrycewang-stanford/Awesome-Journal-Skills | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT |
lawve-ai/awesome-legal-skills
Builds a high-fidelity interactive legal assessment as a single self-contained HTML artifact.
kklimuk/docx-cli
Run the weak-agent adversarial test harness against docx-cli.
revfactory/harness-100
A data flow mapping tool that systematically maps personal information processing flows and identifies risk points.
brycewang-stanford/Awesome-Journal-Skills
A skill your agent uses when executing and stress-testing the empirical analysis for a Research Policy (RP) manuscript — building bibliometric/patent variables, running estimation or qualitative…
brycewang-stanford/Awesome-Journal-Skills
A skill your agent uses when exhibits are the bottleneck for a Research Policy (RP) manuscript — designing tables and figures (regression tables, patent/bibliometric maps, event-study plots, case…
evolsb/claude-legal-skill
Review legal contracts, NDAs, employment agreements, SaaS terms, and M&A documents.
aipoch/medical-research-skills
Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…
aipoch/medical-research-skills
Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.
aipoch/medical-research-skills
Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.
aipoch/medical-research-skills
A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.
aipoch/medical-research-skills
Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…
aipoch/medical-research-skills
Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.
A comprehensive auditor for any agent skill — including Manus, OpenClaw/ClawHub, Claude, LobeHub, or custom SKILL.md-based skills. Skill Auditor is an agent skill from aipoch/medical-research-skills.md-based skills.
Skill Auditor fits situations like: A user wants to evaluate; quality-check an agent skill before publishing; A user says audit my skill; evaluate my skill.
Run `npx skills add aipoch/medical-research-skills --skill skill-auditor -a claude-code`. Or copy the skill folder (skill-auditor in aipoch/medical-research-skills) into .claude/skills/skill-auditor in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aipoch/medical-research-skills --skill skill-auditor -a codex`. Or copy the skill folder (skill-auditor in aipoch/medical-research-skills) into .agents/skills/skill-auditor in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill skill-auditor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-auditor, .gemini/skills/skill-auditor, .github/skills/skill-auditor and .opencode/skills/skill-auditor in your project.
Going by SKILL.md and its folder, Skill Auditor needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Skill Auditor is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.9k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Skill Auditor: Legal Test Builder Patrick Munro (lawve-ai/awesome-legal-skills, 836 stars), Weak Agent Test (kklimuk/docx-cli, 216 stars), Data Flow Mapper (revfactory/harness-100, 1.3k stars) and Respol Data Analysis (brycewang-stanford/Awesome-Journal-Skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,974 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.
Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.