Agent skill

Eval Leakage Audit

by AlexZio00 in AlexZio00/sovereign-skills

Audits whether a verification (eval/metric/experiment/holdout) actually secures independent external ground truth, or whether the designer, the model, and the scorer are just confirming each other…

MITAuto-check passedMarketing & SEO

Install Eval Leakage Audit

skills CLI
$ npx skills add AlexZio00/sovereign-skills --skill eval-leakage-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AlexZio00/sovereign-skills eval-leakage-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AlexZio00/sovereign-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/eval-leakage-audit .claude/skills/eval-leakage-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-leakage-audit
GitHub stars
140
Token cost
~4.9k tokens
SKILL.md length
2,576 words
Files
4
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Audits whether a verification (eval/metric/experiment/holdout) actually secures independent external ground truth, or whether the designer, the model, and the scorer are just confirming each other…

  • Works in 5 steps: Identify the target verification… → Ask the core question: does independent… → Apply all 28 patterns below to the… → …
  • Tasks that involve A/B testing
  • SKILL.md covers Purpose, Trigger, Workflow and Invariants (never violate), plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Leakage Audit is an agent skill from AlexZio00/sovereign-skills. Audits whether a verification (eval/metric/experiment/holdout) actually secures independent external ground truth, or whether the designer, the model, and the scorer are just confirming each other in a circle — via a 28-pattern taxonomy. Read-only. Use before trusting any 'how we'll know it worked' — A/B tests, holdouts, scores, validation — especially when a result feels too clean or self-confirming. 한국어: '이 검증 순환논리 아닌지 봐줘', '이 평가 편파적이야?', '이 벤치마크 셀프체크야?'.

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `.claude-plugin/plugin.json`, `agents/openai.yaml` and `fixtures/golden-case-01.md`).

It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: 20 production-grade skills for AI coding agents — setup, scope, discipline, code review, security, session management, governance, ops, and quality audits (eval-leakage… The licence is MIT.

When your agent uses it

  • Tasks that involve A/B testing

Example prompts

  • “how we”
  • “Use the eval-leakage-audit skill to audit whether a verification (eval/metric/experiment/holdout) actually secures independent external ground…”
  • “/eval-leakage-audit”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Identify the target verification (eval/metric/experiment/holdout/"how will we know it worked"). Name its components — what plays the model…
  2. Ask the core question: does independent external ground truth actually enter the loop?
  3. Apply all 28 patterns below to the target and report only the ones that actually fire (don't list patterns that didn't fire)
  4. For every pattern that fires, propose a concrete fix aimed at restoring independence.
  5. Self-check this audit itself against patterns 3–5: is this auditor grading a bucket it drew itself? Is the verifier actually the designer?

What it can do on your machine

Read from SKILL.md and the folder at commit c062683. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Leakage Audit loads about 4.9k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 2,576 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AlexZio00/sovereign-skills at commit c062683, republished under its MIT licence (© AlexZio00). 2,576 words, ~4,895 tokens.

Download SKILL.mdSave it as .claude/skills/eval-leakage-audit/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
eval-leakage-audit
description
Audits whether a verification (eval/metric/experiment/holdout) actually secures independent external ground truth, or whether the designer, the model, and the scorer are just confirming each other in a circle — via a 28-pattern taxonomy. Read-only. Use before trusting any 'how we'll know it worked' — A/B tests, holdouts, scores, validation — especially when a result feels too clean or self-confirming. 한국어: '이 검증 순환논리 아닌지 봐줘', '이 평가 편파적이야?', '이 벤치마크 셀프체크야?'.
skill_type
analysis
tools
Read, Grep, Glob
disallowed-tools
Edit, Write, NotebookEdit, Bash
user-invocable
true
concurrency_profile.read_only
true
concurrency_profile.concurrency_safe
true
concurrency_profile.destructive
none
state_footprint
stateless
not_for
Post-code-change checks (tests pass, diff scope, side effects) -> verification (different target: verification=the code change itself, this skill=the…

Eval Leakage Audit — Verification Circularity Audit

Purpose

When some verification (eval/metric/experiment/holdout) gives you confidence that something "worked," this skill checks whether that confidence actually comes from independent external evidence, or whether the people who designed the verification and the model that scores it are just confirming each other (circularity) — via 28 concrete patterns. A passing gate is not proof of quality — it only proves what it was designed to check, and this skill is the executable tool that actually tests that principle in practice.

Dominant variable: does this verification actually receive independent external ground truth, or are the designer, the model, and the scorer mistaking self-confirmation for a result?

Discard if: the target has no evaluation/verification/benchmark concept at all (pure code refactor, doc changes, etc.).

Trigger

  • "audit this eval for leakage", "check if this benchmark is circular", "is this evaluation biased?"
  • "이 검증 순환논리 아닌지 봐줘", "이 평가 편파적이야?", "이 벤치마크 셀프체크야?"
  • /eval-leakage-audit

Workflow

  1. Identify the target verification (eval/metric/experiment/holdout/"how will we know it worked"). Name its components — what plays the model role, what plays the scorer role, what plays the designer role, and which dataset is involved.
  2. Ask the core question: does independent external ground truth actually enter the loop?
  3. Apply all 28 patterns below to the target and report only the ones that actually fire (don't list patterns that didn't fire):
    1. Recall, not reason — the answer was replayed from something already known, not actually derived
    2. Wrong null hypothesis — the ablation only strips the surface label while the actual leaking signal stays in place
    3. Shared hallucination — two components confirm each other and dress up the circularity as a number
    4. Tautology — the scorer grades the bucket it drew itself (precise term: checker_overfit)
    5. Verifier = designer — an unreproducible, undisclosed recipe is passed off as a holdout
    6. Shared-pool bias — train/holdout come from the same labeler pool, so the same bias enters both sides
    7. Frame injection — the question itself hints at the answer
    8. Demand characteristics — the subject being measured knows it's being measured and behaves differently as a result
    9. Dual-fail-flag — when two independent subjects (models/implementations) fail on exactly the same hidden case, suspect a defect in the scorer itself before blaming the subjects. Independent failures coinciding by chance is unlikely — a match more likely points to a shared cause (a scorer bug, or an error in the hidden case itself)
    10. Asymmetric-baseline self-falsification — if the metric itself uses a biased, asymmetric baseline statistic, the null expectation isn't 50%, producing false positives. Before reporting a result, self-falsify the metric with a symmetric check
    11. Evidence-burn — a fixture a model arm has already observed is spent: don't reuse it afterward as independent evidence, a holdout, or a replication fixture. Generate a new variant for the next round, or retire the fixture
    12. Ungraded grader — trusting the scorer/answer-key/rubric itself without self-verifying it first. Before trusting it, run 4 gates: (1) reference-pass — a known-correct implementation scores full marks; (2) buggy-baseline-fail — a deliberately flawed implementation scores in the expected low-to-mid band (too low or too high signals a miscalibrated grader); (3) mutation-kill — the grader actually catches ≥3 plausible wrong answers; (4) dual-fail-flag — if both arms of a comparison fail on the exact same case, treat it as a grader/fixture defect signal, not a candidate failure
    13. Ceiling task — misreading a benchmark saturated at full marks for every candidate (zero discriminative power) as "no difference." A single k=1 pilot showing both arms scoring ≥95% is too thin a sample to confirm a ceiling — flag it only as ceiling_suspected and require 1-2 decisive branching probes (k≥2, or reproduction under a different condition/task variant) before promoting it to confirmed ceiling_detected. Adding more fixtures of the same kind won't by itself recover discriminative power once a ceiling is actually confirmed. Respond by switching task families, promoting process-layer metrics (cost, verification cadence), or adding those branching probes. Post-hoc discriminative-power indicators: ceiling_rate, score_sd, bucket_entropy, winner_flip_rate
    14. Respawn masking — in systems with respawn/reset logic (games, simulations, state machines), scoring by an instant state snapshot lets a respawn disguise failure as a pass (for example: a character dies, auto-respawns, and the snapshot moment shows only alive — the death vanishes). Score by session-wide deterministic invariants instead (for example: death/reset event count = 0, no cumulative resource loss) rather than a state snapshot.
    15. Pseudo-replication — counting multiple probe cells drawn from the same arm as independent sample size n artificially inflates the sample and overstates statistical significance. Distinguish effective independent units (true independent observations) from raw probe cells (repeated measures within the same arm) and do not count the latter toward n.
    16. Stimulus calibration gap — the counterpart to pattern #12 (ungraded grader): even a perfect grader produces meaningless results if the test stimulus itself (question, scenario, prompt) never actually elicits the intended behavior. Run a calibration pass beforehand confirming the reference implementation actually triggers the target behavior for that stimulus.
    17. Unaudited cost-saving skips — leaving cost/time-saving skipped checks unaudited indefinitely lets not-checked quietly harden into no-problem. Even without full re-verification, periodically audit the skipped set via a deterministic minority sample (for example: sha256-minimum hashing).
    18. Goodhart co-evolution in self-improving loops — in a self-improving harness, if the loop designs the very scorer that grades its own improvements, the gate can drift lenient across iterations with no single discrete failure to point to — the loop is quietly reshaping the measure around its own output rather than the measure holding still. Guard with two layers together: fixed anchor tasks the loop never designed and never sees in advance (user-picked, undisclosed), plus a judge the loop doesn't control. Periodically re-check the self-improvement gate for lenient drift instead of trusting a one-time calibration. Ship a deliberately-broken fixture alongside every new capability so the gate's ability to actually fail something stays exercised, not assumed. Credit a harness improvement only against a fresh held-out task, never against the task that produced the improvement in the first place.
    19. Success provenance gap — when an agent lands on the "correct" answer, not distinguishing whether it followed the intended reasoning from the context it was given (authorized context) versus simply picked up the target value in passing during evaluation or retrieval (acquired target) lets one accuracy number hide two different sources of "success." Re-run the same items under three conditions — CLEAN (no target value exposed), GOLD (the correct target value exposed), SHAM (a decoy wrong value exposed) — and measure the GOLD–SHAM gap: a large gap signals the accuracy is coming from value dependence rather than reasoning. [borrowed from arXiv 2607.24054]
    20. Lenient-judge-mode non-disclosure — a self-LLM judge defaults to a lenient grading prompt (partial credit, near-match acceptance, etc.) and the headline number gets reported without ever disclosing that lenient mode was on. Without a strict-mode re-score for comparison, there's no way to tell whether the headline reflects real performance or just grading generosity. Before citing a headline number, confirm whether the grading mode (strict/lenient) is even stated. [borrowed from OpenViking benchmark analysis]
    21. Hardest-category denominator exclusion — a headline accuracy is presented as "M correct out of N total" while the hardest or most adversarial category has quietly been dropped from the denominator (N) before the division. If the exclusion lives only in a footnote or appendix — or isn't disclosed at all — readers can't tell whether the headline means "performance on the easy subset" or "performance overall." Demand per-category numerator/denominator breakdowns, and where a category was excluded, also check that category's own standalone accuracy. [borrowed from OpenViking benchmark analysis]
    22. Underpowered null reported as no effect — before interpreting "no effect," "marginal improvement," or any small effect, compute the minimum detectable effect (MDE) from the sample size n and the observed variance. If the MDE is more than 2x the observed effect, label the result INCONCLUSIVE(underpowered) and state in one line the n needed to detect that effect — do not report it as "no effect." For event-level metrics (hit/miss transitions), compute from the number of events. When citing an external benchmark figure, record its "M tasks x t repeats" allocation; if the repeat count is not disclosed, mark it unverified (repeats undisclosed). [borrowed from arXiv 2609.01519, 2608.20290, 2609.29140 (converging)]
    23. Degenerate strategy passes — before trusting a threshold or metric, run constant strategies (always reject, always pass, empty output) through the same scorer. If any of them clears the pass line, the metric rules nothing out and must be redesigned. A "zero violations + zero adoptions" signal is a post-hoc symptom; this is the pre-adoption check. [borrowed from arXiv 2608.27167]
    24. Retry-inflated headline — check whether the headline success rate includes retries, best-of-k, or an automatic recovery loop. If it does, ask for the first-attempt pass rate (pass@1) separately; if none is given, label the figure an "upper bound." Same disclosure family as #20 and #21. [borrowed from arXiv 2609.00050]
    25. Non-common item sets across compared runs — if ground-truth filtering or preprocessing leaves a different set of items for each compared subject (model, setting, version), recompute on the intersection of common items. If that is impossible, label the comparison "not comparable across runs." [borrowed from arXiv 2608.28623]
    26. Bundled-change attribution — before crediting an improvement to one factor, list everything that changed in the same round (instructions, schema, examples, scripts). If no comparison changed only one of them, report the effect as the effect of the whole bundle. [borrowed from arXiv 2609.01519]
    27. Stability-only validity — if the only evidence for a judge's or gate's quality is agreement rate or test-retest consistency, check whether responsiveness was also measured: does the verdict actually flip on a minimal pair of inputs that differ by a meaning-changing edit? If not measured, the judge cannot be told apart from a constant judge. [borrowed from arXiv 2608.24419]
    28. Selective-gate attribution — when claiming the value of a conditionally applied gate (selective revision, retry, escalation), check that it was compared against an "always apply" baseline as well as against "no gate (initial)". Reporting the gain over the initial state alone credits the revision procedure's own gain to the gate. Also ask for the 4-cell transition table (C->C, C->W, W->C, W->W) so correct answers destroyed (C->W) are not hidden inside a net gain. [borrowed from arXiv 2609.35832]
  4. For every pattern that fires, propose a concrete fix aimed at restoring independence.
  5. Self-check this audit itself against patterns 3–5: is this auditor grading a bucket it drew itself? Is the verifier actually the designer?
Show full SKILL.md (856 more words)Show less

Stratification-claim substantiation check: when the target claims it "compared stratified" (to guard against Simpson's-paradox-style aggregation bias, where a pooled comparison can reverse the direction seen within each subgroup), don't accept that claim as substantiated until 6 fields are recorded — (1) overall effect, (2) per-stratum effect, (3) direction consistency (same sign across strata), (4) effect dispersion (spread across strata), (5) minimum cell count (is each stratum's sample large enough to judge), (6) multiple-comparison correction (e.g. an FDR q-value). If even one is missing, the "we stratified" claim is unverifiable — the general principle is sound, but without these substantiation fields it's an empty assertion. (This is a statistical-substantiation gate, not one of the 28 circularity patterns — a separate axis.)

Reviewer Independence Honest 4-Label check: when the target verification leans on multiple reviewers/judges to claim independence, don't accept "independently reviewed" at face value. Vendor/model-family diversity is a necessary but not sufficient condition for independent — a reviewer on a different vendor that still consumes the same evidence (same dataset, same collection recipe) and applies the same methodology as the subject is vendor-diverse but not procedurally independent. Label the actual state as one of 4: independent (at least one reviewer confirmed to run on a different vendor/model family from the subject being judged, and that reviewer's evidence and methodology — data source, review procedure — are also independent of the subject's, not merely rubber-stamping the same inputs), same_vendor (every reviewer shares the subject's vendor/model family, or a different-vendor reviewer only reuses the subject's own evidence/methodology without independent evidence or procedure — either case is the same mind reviewing itself, just wearing a different label), unverified (the reviewer's vendor/model identity, or the independence of its evidence/methodology, can't be proven either way — don't default this to independent just because a vendor difference is visible), or unavailable (no reviewer was actually present). If the reviewer pool falls short of quorum and only same-vendor (or vendor-diverse-but-not-procedurally-independent) reviewers remain, keep them rather than reporting no signal at all — but the same_vendor label has to stay attached and visible; quorum survives on honest labeling, not on quietly upgrading the label to look cleaner than it is. (This is a source-diversity honesty gate, not one of the 28 circularity patterns — a separate axis, closest in spirit to pattern #9's dual-fail-flag.)

Invariants (never violate)

  1. Read-only — never redesign the verification or rewrite the experiment. Only name the leak points and the fixes. Violation → the audit and the redesign blur together and the original experiment's intent gets corrupted.
  2. Report only patterns that fired — don't mechanically list all 28; report only the ones actually backed by evidence. Violation → a laundry list dressed up as checklist completion, i.e. an unlabeled score dressed up as rigor.
  3. Blinding the output doesn't cure a leaking collection recipe — don't report safety just because outputs are hidden. Violation → mistaking surface-level blinding for real independence.
  4. The final report converges on one root cause, not a laundry list — even if multiple patterns fire, converge them into a single root cause. Violation → an unprioritized list is not an actionable report.

Scope Boundary

DoesDoes NOT
[READ] Identify verification components (model/scorer/designer/dataset)Redesign the experiment/benchmark or edit code
[READ] Apply the 28-pattern check and report only the ones that firedFormally list patterns that didn't fire
[READ] Propose an independence fix for each fired patternImplement the fix itself (propose only)
[READ] Self-check the audit itself against patterns 3–5Declare "clean" without a self-check

Rationalization Table

RationalizationCounter
"We hid the output values, so it's independent now"Violates Invariant 3. Output blinding and collection-recipe independence are different problems
"Let's show we checked all 28"Violates Invariant 2. Listing unfired patterns is a laundry list — completion theater without evidence
"This looks independent enough"Violates the Gate≠Oracle principle. "Looks similar" is a feeling, not evidence — judge only by which of the 28 patterns actually fired
"The auditor designed this experiment too, but it's fine"Exactly the self-check Workflow step 5 calls for — this is precisely the Verifier=designer case (#5) (corrected: this is not Invariant 4, which is the single-root-cause convergence rule and is unrelated to self-checking)
"Both sides failed, so both implementations are bad"Violates Dual-fail-flag (#9). If two independent implementations fail on exactly the same hidden case, check the scorer for a defect first — don't default to blaming the subjects

Output

Regression-check golden case: fixtures/golden-case-01.md -- after editing this skill, confirm that on this case only #3, #4, #5, #9, #20 and #21 fire (patterns #22-#28 are not exercised by it). The verdict is LLM-made, so this is a calibration aid, not a deterministic self-test.

In the conversation: identify components → fired patterns (name + evidence + fix) → converge to one root cause → self-check result (applying patterns 3–5).

Error Recovery

FailureRecovery
input_errorIf the verification target is unclear, ask "which part of this request is the eval?" before proceeding
missing_dataIf a component (model/scorer/designer/dataset) lacks information, state "cannot confirm" — never guess

Truthful Reporting

  1. No mock deception: report only patterns with actual evidence as "fired"; stay silent on the rest (don't list them).
  2. No silent brokenness: if the findings don't converge to a single root cause (several patterns are equally significant), state that explicitly too.

© AlexZio00, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in eval-leakage-audit of AlexZio00/sovereign-skills.

  • SKILL.md
  • .claude-plugin/plugin.json
  • agents/openai.yaml
  • fixtures/golden-case-01.md

Open the folder on GitHubat commit c062683

Compare with similar skills

Eval Leakage Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Leakage Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Leakage Audit this skillAlexZio00/sovereign-skills140—~4.9kAutomated safety check: PassMIT
Ab Testingcoreyhaines31/marketingskills54k3 repos~3.1kAutomated safety check: PassMIT
AnalyticsNexus-JPF/note-companion8707 repos~2.2kAutomated safety check: PassMIT
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab Test Analyzeririnabuht12-oss/marketing-skills4.1k—~1.4kAutomated safety check: PassNone
Ab Test Store Listingappeeky/aso-skills2.2k—~1.8kAutomated safety check: PassMIT

Similar skills

  • Ab Testing

    coreyhaines31/marketingskills

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

    54k GitHub starsUsed in 3 repos~3.1k tokens
    Marketing & SEOAuto-check passed
  • Analytics

    Nexus-JPF/note-companion

    When the user wants to set up, improve, or audit analytics tracking and measurement.

    870 GitHub starsUsed in 7 repos~2.2k tokens
    Marketing & SEOAuto-check passed
  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    4.1k GitHub stars~1.4k tokensUpdated 17 days ago
    Marketing & SEOAuto-check passed
  • Ab Test Store Listing

    appeeky/aso-skills

    When the user wants to A/B test App Store product page elements to improve conversion rate.

    2.2k GitHub stars~1.8k tokensUpdated 4 days ago
    Marketing & SEOAuto-check passed
  • Ab Test Setup

    freekmurze/dotfiles

    When the user wants to plan, design, or implement an A/B test or experiment.

    1k GitHub starsUsed in 14 repos~1.8k tokens
    Marketing & SEOAuto-check passed

More from AlexZio00/sovereign-skills

All 17 skills in this repo
  • Project Overview

    AlexZio00/sovereign-skills

    A skill your agent uses when the user wants a deterministic cross-project status map generated from registered projects' session handoffs.

    140 GitHub stars~2.4k tokensUpdated 2 days ago
    Auto-check passed
  • Scope

    AlexZio00/sovereign-skills

    Scope definition before implementation — two modes. An agent skill from AlexZio00/sovereign-skills.

    140 GitHub stars~4k tokensUpdated 2 days ago
    Auto-check passed
  • Project Init

    AlexZio00/sovereign-skills

    Interview-based project setup — generates CLAUDE.md, ROADMAP, .gitignore, .env.example from scratch.

    140 GitHub stars~4.1k tokensUpdated 2 days ago
    Auto-check: notes
  • Collab Audit

    AlexZio00/sovereign-skills

    This skill should be used when the user types /collab-audit or requests AI collaboration diagnosis.

    140 GitHub stars~8k tokensUpdated 2 days ago
    Auto-check passed
  • Doc Drift

    AlexZio00/sovereign-skills

    A skill your agent uses when the user wants to audit the memory and documents Claude Code loads into context — CLAUDE.md (user global + project + nested), MEMORY.md, @imports, .claude/skills…

    140 GitHub stars~6.2k tokensUpdated 2 days ago
    Auto-check passed
  • Session Checkpoint

    AlexZio00/sovereign-skills

    A skill your agent uses when saving session state before context compaction, switching tasks, or ending a session.

    140 GitHub stars~14k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Eval Leakage Audit

What does Eval Leakage Audit do?

Audits whether a verification (eval/metric/experiment/holdout) actually secures independent external ground truth, or whether the designer, the model, and the scorer are just confirming each other…. Eval Leakage Audit is an agent skill from AlexZio00/sovereign-skills. Audits whether a verification (eval/metric/experiment/holdout) actually secures independent external ground truth, or whether the designer, the model, and the scorer are just confirming each other in a circle — via a 28-pattern taxonomy.

When should I use Eval Leakage Audit?

Eval Leakage Audit fits situations like: tasks that involve A/B testing.

How do I install Eval Leakage Audit in Claude Code?

Run `npx skills add AlexZio00/sovereign-skills --skill eval-leakage-audit -a claude-code`. Or copy the skill folder (eval-leakage-audit in AlexZio00/sovereign-skills) into .claude/skills/eval-leakage-audit in your project. Claude Code loads it when a task matches its description.

How do I install Eval Leakage Audit in Codex?

Run `npx skills add AlexZio00/sovereign-skills --skill eval-leakage-audit -a codex`. Or copy the skill folder (eval-leakage-audit in AlexZio00/sovereign-skills) into .agents/skills/eval-leakage-audit in your project. Codex loads it when a task matches its description.

Can I use Eval Leakage Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlexZio00/sovereign-skills --skill eval-leakage-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-leakage-audit, .gemini/skills/eval-leakage-audit, .github/skills/eval-leakage-audit and .opencode/skills/eval-leakage-audit in your project.

What does Eval Leakage Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Leakage Audit is instructions for the agent only. Our summary lists: Python 3.

Does Eval Leakage Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Leakage Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Leakage Audit use?

Eval Leakage Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Leakage Audit use?

About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Leakage Audit?

Skills that share tags, products or a category with Eval Leakage Audit: Ab Testing (coreyhaines31/marketingskills, 54k stars), Analytics (Nexus-JPF/note-companion, 870 stars), Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars) and Ab Test Analyzer (irinabuht12-oss/marketing-skills, 4.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Leakage Audit?

AlexZio00 (a GitHub user) maintains it in AlexZio00/sovereign-skills, which has 140 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.

Source: AlexZio00/sovereign-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.