Agent skill

Jev Eval

by OneWave-AI in OneWave-AI/claude-skills

Build and run a labelled eval set for a System One model (Jev, Von, or any typed-decision config), then sweep criteria wordings and thresholds against it.

MITAuto-check passed

Install Jev Eval

skills CLI
$ npx skills add OneWave-AI/claude-skills --skill jev-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OneWave-AI/claude-skills jev-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OneWave-AI/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/jev-eval .claude/skills/jev-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
jev-eval
GitHub stars
323
Used in
1 other repo
Token cost
~1.6k tokens
SKILL.md length
685 words
Files
2 (incl. scripts)
Skills in repo
70
Repo updated
First seen
Licence
MIT

At a glance

Build and run a labelled eval set for a System One model (Jev, Von, or any typed-decision config), then sweep criteria wordings and thresholds against it.

  • Works in 5 steps: Build the set → Sweep wordings, not just models → Sweep thresholds for every noul → …
  • A Jev/Von classification is wrong
  • SKILL.md covers Run it, 1. Build the set, 2. Sweep wordings, not just… and 3. Sweep thresholds for every…, plus 3 more sections
  • Runs Python scripts from its folder; calls python

What it does

Jev Eval is an agent skill from OneWave-AI/claude-skills. Build and run a labelled eval set for a System One model (Jev, Von, or any typed-decision config), then sweep criteria wordings and thresholds against it. Use when a Jev/Von classification is wrong or unreliable, when choosing between the hosted API and a local open model, when tuning noul thresholds, or before shipping any typed-decision feature. Produces an accuracy-by-wording matrix and a calibrated threshold.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/sweep.py`).

The repository describes itself as: 200+ production-ready Claude Code skills for sales, marketing, design, engineering, and AI agent architecture. Built and maintained by OneWave AI. The licence is MIT.

When your agent uses it

  • A Jev/Von classification is wrong
  • Choosing between the hosted API and a local open model
  • Tuning noul thresholds
  • Before shipping any typed-decision feature

Example prompts

  • “/jev-eval”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Build the set
  2. Sweep wordings, not just models
  3. Sweep thresholds for every noul
  4. Score the confidence gate honestly
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit fc5b785. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Jev Eval loads about 1.6k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 685 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from OneWave-AI/claude-skills at commit fc5b785, republished under its MIT licence (© OneWave-AI). 685 words, ~1,634 tokens.

Download SKILL.mdSave it as .claude/skills/jev-eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
jev-eval
description
Build and run a labelled eval set for a System One model (Jev, Von, or any typed-decision config), then sweep criteria wordings and thresholds against it. Use when a Jev/Von classification is wrong or unreliable, when choosing between the hosted API and a local open model, when tuning noul thresholds, or before shipping any typed-decision feature. Produces an accuracy-by-wording matrix and a calibrated threshold.

Evaluating a typed-decision config

The eval set is the product. A System One model's accuracy is dominated by how the question was written, and the failure mode is silent — it returns a confident, type-valid, wrong answer. Without labels you cannot tell a bad question from a bad model.

Measured: rewriting the criteria moved an open model from 4/15 to 14/15 on 15 records. No model change. Then the same comparison at 150 records put that model at 61% overall against Jev's 97% — the 15-record read was an artifact of a small, easy set. Both facts are the point: wording swings results, and small sets lie about which way.

Run it

bash
python ~/.claude/skills/jev-eval/scripts/sweep.py labelled.json configs.json \
    --backend jev|von --question <name>

labelled.json is [{"id","state","truth"}]. configs.json maps a config name to {"instructions", "criteria"} — a dict of options makes it a choice, a list of levels makes it a score, omitting it makes it a noul. The script reads the Jev key from Keychain (typesafe-api-key), prints accuracy per config, labels the spread ROBUST or FRAGILE, sweeps thresholds for nouls, and scores the confidence gate.

Real output, 50 records of agent shell-command risk, only the backend changed:

config                         accuracy   ms/rec
A terse one-liners               45/50       410      <- Jev
B rich criteria                  49/50       415
C rich + exclusions              46/50       418
D deliberately lazy              44/50       411
spread: 44/50 to 49/50   (ROBUST - wording is not load-bearing)

A terse one-liners                9/50        66      <- Von, same configs
B rich criteria                  23/50       121
C rich + exclusions              15/50       129
D deliberately lazy              22/50        53
spread: 9/50 to 23/50    (FRAGILE - and the ceiling is still not usable)

Read the ceiling before the spread. A FRAGILE model whose best config is 23/50 is not a wording problem you can write your way out of — it is the wrong model for that question.

1. Build the set

50 records minimum, 200+ before shipping. Pull from the real stream, not synthetic data.

  • Include the boring middle, not just clean examples and dramatic edge cases
  • Include records with broken/missing metadata — that is where classifiers fail
  • Label by reading the record, before any model runs. Never label from model output
  • Store as JSON with the raw record plus a truth field
json
[{"id":"D-1994","state":"...full record text...","truth":"inbound_prospect"}]

If you cannot label a record confidently yourself, the model cannot either — either drop it or fix the question so the answer is determinate.

2. Sweep wordings, not just models

The core move. Write 3–4 genuinely different criteria configs and run all of them:

  • A — one terse line per option (what everyone writes first)
  • B — 3–4 sentences per option with concrete examples
  • C — B plus an explicit default and explicit exclusions ("ONLY when…")
  • D — deliberately lazy, four or five words, as a floor test

Report accuracy per config per model:

config                    JEV       VON      LAYA     (lead triage, n=50)
A terse one-liners      47/50     34/50     21/50
B rich criteria         49/50     28/50     15/50
C rich + exclusions     48/50     24/50     19/50
D deliberately lazy     48/50     22/50     24/50

Read the floor first, then the spread. The floor is "can this model do the job at all if I phrase it badly"; the spread is "how much will maintaining it cost me." Measured on the command task, every hosted model floors at 82-92% (Haiku 46-48, GPT-4.1-mini 45-50, Jev 44-49, GPT-5-mini 41-49) while Von floors at 9/50 and Laya at 18/50. Note that Jev is not more wording-robust than a small LLM — it swings the same ten points. What you buy is the floor, not immunity.

Note what the leads column does NOT show: a clean "richer is better" gradient. Von's best config here is the terse one. Whatever moves an open model's numbers is sensitivity to surface form, not comprehension, so do not assume your next criteria rewrite improves it — re-run the set.

Show full SKILL.md (179 more words)Show less

3. Sweep thresholds for every noul

Never ship 0.5. Sweep and read the curve:

python
for t in [0.5,0.6,0.7,0.75,0.8,0.85,0.9,0.95]:
    tp = sum(p>=t and y     for p,y in z); fp = sum(p>=t and not y for p,y in z)
    fn = sum(p< t and y     for p,y in z); tn = sum(p< t and not y for p,y in z)
    print(f"{t}  P={tp/max(tp+fp,1):.2f}  R={tp/max(tp+fn,1):.2f}  acc={(tp+tn)/len(z):.2f}")

Different models have different floors. Jev's noul sat at 0.2–0.5 on records that were plainly clean, where Claude went to 0.0 — so its real cut was ~0.85. A threshold tuned on one model does not transfer to another. Re-sweep when you switch.

4. Score the confidence gate honestly

Two numbers, always together:

python
errs   = [r for r in rows if r.pred != r.truth]
rights = [r for r in rows if r.pred == r.truth]
for g in [0.5,0.7,0.9]:
    caught    = sum(r.conf <  g for r in errs)      # errors the gate escalates
    escalated = sum(r.conf <  g for r in rows)      # total volume escalated
    print(f"gate {g}: catches {caught}/{len(errs)} errors, escalates {escalated/len(rows):.0%} of volume")

A gate catching 10/11 errors while escalating 93% of volume is not a working gate — it is a slow path with extra steps. Good calibration without good accuracy buys nothing.

5. Report

  • accuracy per config per model, with the spread called out
  • chosen threshold per noul, with the sweep that justified it
  • gate: errors caught and volume escalated
  • projected latency and cost per 1k at production volume
  • explicit statement of eval-set size and what it does not cover

Small sets lie. 15 records where two models both score 100% distinguishes nothing — say so rather than implying the tie is meaningful.

jev-integrate is the wiring workflow this feeds. jev-audit finds candidates worth evaluating.

© OneWave-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in jev-eval of OneWave-AI/claude-skills.

  • SKILL.md
  • scripts/sweep.py

Open the folder on GitHubat commit fc5b785

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in OneWave-AI/claude-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Jev Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Jev Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Jev Eval this skillOneWave-AI/claude-skills3231 repos~1.6kAutomated safety check: PassMIT
Eval Harnessaffaan-m/ECC275k—~2.2kAutomated safety check: PassMIT
Evalalirezarezvani/claude-skills28k1 repos~618Automated safety check: PassMIT
Eval Harnessaffaan-m/ECC275k1 repos~1.7kAutomated safety check: PassMIT
Eval-Driven Development Harnessaffaan-m/ECC275k—~1.5kAutomated safety check: PassMIT
Form Labelsthedaviddias/Front-End-Checklist74k—~565Automated safety check: PassMIT

Similar skills

  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    275k GitHub stars~2.2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub starsUsed in 1 repo~618 tokens
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

    275k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    275k GitHub stars~1.5k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Form Labels

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing rendered HTML, interactive components, or design-system patterns related to Associate labels with form controls.

    74k GitHub stars~565 tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Avoid Eval

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing scripts, client components, bundles, or runtime behavior related to Never use eval() or unsafe dynamic code execution.

    74k GitHub stars~554 tokensUpdated 2 days ago
    Auto-check passed

More from OneWave-AI/claude-skills

All 70 skills in this repo
  • CRM Data Cleanup

    OneWave-AI/claude-skills

    Finds duplicate and junk records in a CRM CSV export with fuzzy matching, normalizes fields and writes a reviewable merge plan plus import-ready files without touching the live CRM.

    323 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Design Export Repair

    OneWave-AI/claude-skills

    Repairs broken decks and PDFs exported from Claude Design or similar AI deck generators: clipped text, wrong fonts and corrupted .pptx package structure.

    323 GitHub stars~2.6k tokensUpdated 6 days ago
    Auto-check passed
  • Bi Measure Builder

    OneWave-AI/claude-skills

    Writes, explains, debugs, and optimizes BI calculations - Power BI / Fabric DAX measures and calculated columns, Tableau calculated fields (FIXED/INCLUDE/EXCLUDE LOD expressions, table…

    323 GitHub stars~2.2k tokensUpdated 6 days ago
    Auto-check passed
  • Bookkeeping Close

    OneWave-AI/claude-skills

    Categorizes transactions, reconciles bank and card statements to the ledger, works a month-end checklist and prepares a close package, without ever forcing a balance.

    323 GitHub stars~2k tokensUpdated 6 days ago
    Auto-check passed
  • CSV and Excel Merger

    OneWave-AI/claude-skills

    Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

    323 GitHub stars~1.6k tokensUpdated 6 days ago
    Auto-check passed
  • Sec Filing Puller

    OneWave-AI/claude-skills

    Pulls financial statement numbers for US public companies straight from SEC EDGAR's free official XBRL APIs (companyfacts, companyconcept, frames, submissions) into a cited table.

    323 GitHub stars~2.1k tokensUpdated 6 days ago
    Auto-check passed

Questions about Jev Eval

What does Jev Eval do?

Build and run a labelled eval set for a System One model (Jev, Von, or any typed-decision config), then sweep criteria wordings and thresholds against it. Jev Eval is an agent skill from OneWave-AI/claude-skills. Build and run a labelled eval set for a System One model (Jev, Von, or any typed-decision config), then sweep criteria wordings and thresholds against it.

When should I use Jev Eval?

Jev Eval fits situations like: A Jev/Von classification is wrong; choosing between the hosted API and a local open model; tuning noul thresholds; before shipping any typed-decision feature.

How do I install Jev Eval in Claude Code?

Run `npx skills add OneWave-AI/claude-skills --skill jev-eval -a claude-code`. Or copy the skill folder (jev-eval in OneWave-AI/claude-skills) into .claude/skills/jev-eval in your project. Claude Code loads it when a task matches its description.

How do I install Jev Eval in Codex?

Run `npx skills add OneWave-AI/claude-skills --skill jev-eval -a codex`. Or copy the skill folder (jev-eval in OneWave-AI/claude-skills) into .agents/skills/jev-eval in your project. Codex loads it when a task matches its description.

Can I use Jev Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OneWave-AI/claude-skills --skill jev-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/jev-eval, .gemini/skills/jev-eval, .github/skills/jev-eval and .opencode/skills/jev-eval in your project.

What does Jev Eval need to run?

Going by SKILL.md and its folder, Jev Eval needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Jev Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Jev Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Jev Eval use?

Jev Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Jev Eval use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Jev Eval?

Skills that share tags, products or a category with Jev Eval: Eval Harness (affaan-m/ECC, 275k stars), Eval (alirezarezvani/claude-skills, 28k stars), Eval Harness (affaan-m/ECC, 275k stars) and Eval-Driven Development Harness (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Jev Eval?

OneWave-AI (a GitHub organization) maintains it in OneWave-AI/claude-skills, which has 323 GitHub stars. The repository holds 70 skills in this directory. The repository was last updated on October 2, 2026.

Source: OneWave-AI/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.