Agent skill

Building With Jev

by notque in notque/vexjoy-agent

Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.

MITAuto-check: notes

Install Building With Jev

skills CLI
$ npx skills add notque/vexjoy-agent --skill building-with-jev -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install notque/vexjoy-agent building-with-jev --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/meta/building-with-jev .claude/skills/building-with-jev && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
building-with-jev
GitHub stars
439
Token cost
~11k tokens
SKILL.md length
6,131 words
Files
11 (incl. references)
Skills in repo
61
Repo updated
First seen
Licence
MIT

At a glance

Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.

  • Works in 7 steps: List phases. Read the skill's SKILL.md.… → Classify each phase into four columns:… → Convert judgment phases. Each judgment… → …
  • SKILL.md covers Reference Loading Table, Read the live docs, Transports and The three tiers, plus 11 more sections
  • Calls python3; reaches api.typesafe.ai; needs TYPESAFE_API_KEY and AI_GATEWAY_API_KEY

What it does

Building With Jev is an agent skill from notque/vexjoy-agent. Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.

Its SKILL.md is about 11k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `references/composition-patterns.md`, `references/composition-positions.md` and `references/decision-card.md`).

The repository describes itself as: VexJoy AI Agent with Jev Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop. The licence is MIT.

Example prompts

  • “/building-with-jev”

Requirements

  • A credential in TYPESAFE_API_KEY
  • A credential in AI_GATEWAY_API_KEY
  • Pre-approved tools (allowed-tools): Read, Edit, Write, Bash, Glob, Grep

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. List phases. Read the skill's SKILL.md. Write each phase on one row.
  2. Classify each phase into four columns: deterministic (program), judgment (Jev), generation (LLM), or orchestration (dispatch/coordination).
  3. Convert judgment phases. Each judgment becomes one or more Jev questions: Noul for gates, Choice for classification/routing, Score for…
  4. Keep deterministic phases in code. Regex scans, file reads, grep, counts, averages, formatting stay as programs.
  5. Isolate generation. If any phase requires new text (rewrite, diagnosis, plan), that phase keeps an LLM. The LLM receives all prior Jev…
  6. Write the policy function. A pure function policy(assessment) -> action with named thresholds is the dissolved skill's contract. It…
  7. Prove agreement. Run the Jev program on hand-labeled examples. Match or exceed the skill's accuracy before deleting the SKILL.md.

What it can do on your machine

Read from SKILL.md and the folder at commit 5218674. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Edit
    • Write
    • Bash
    • Glob
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.typesafe.ai

    Also links to:

    • docs.typesafe.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • TYPESAFE_API_KEY
    • AI_GATEWAY_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Building With Jev loads about 11k tokens when it runs, and up to ~29k if it reads all its reference files. Until then it costs about 30 tokens; SKILL.md has 6,131 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~11k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~29k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Edit, Write, Bash, Glob, Grep

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from notque/vexjoy-agent at commit 5218674, republished under its MIT licence (© notque). 6,131 words, ~10,724 tokens.

Download SKILL.mdSave it as .claude/skills/building-with-jev/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
building-with-jev
description
Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.
allowed-tools
Read, Edit, Write, Bash, Glob, Grep
user_invocable
false
routing.force_route
true
routing.triggers
jev, typesafe, noul, choice question, score question, system one, write a jev question, jev answers wrong, low confidence, jev criteria, jev state, confidence…
routing.not_for
Running the browser harness end to end (use browser-jev-automation) or routing a request (use do). This skill is for designing and fixing the Jev calls inside…
routing.pairs_with
browser-jev-automation, toolkit, do
routing.complexity
Complex
routing.category
meta

Building with Jev

Jev reads one state, answers every question in the request independently and in parallel, and returns a probability distribution over answers you defined. A head cannot read another head's answer: parallel heads share evidence, not reasoning. For one state, maximize independent heads that can change a decision or action, subject to their token cost and the 64,000-token request budget; omit noise heads. Code owns control flow, arithmetic, policy, and every serial dependency; Jev owns the snap judgment. It does not reason in steps, count, do arithmetic, or generate text. Use this skill to design the questions, fit the state, compose answers in code, wire the call into a hook or script, and fix a call that answers wrong.

Before writing or changing any Jev request, apply Jev production rules and tick its pre-ship checklist. It sets request size (2.5–4k tokens via Gateway until measured), per-run budget (~50k tokens), screen-then-detail above ~50 items, sending each stage at once with an instance cap of floor(0.25 × 250,000 / tokens_per_request), retries by status code, eval pacing, caching, logging, and the order to measure failures. The rest of this skill is the method for designing questions; those rules govern how requests are sent. Its four facts come first: use Vercel AI Gateway; too much context is the most common failure, so split an oversized request into as many small requests as it takes (the toolkit transports do this automatically); a small request is the first diagnostic; rate limits are normal, so retry them.

Prefer direct judgments over supplied evidence. Bounded action selection is valid when candidates and decision evidence are supplied. If answering requires an intermediate result that changes later evidence or candidates, code must resolve that dependency before a later Jev request.

Reference Loading Table

SignalLoad These FilesWhy
request or response shape, instruction objects, criteria objects, reading score/probabilities/confidencereferences/primitives.mdFull API shape and answer semantics
writing or rewriting instructions, criteria, levels, options, examplesreferences/question-design.mdQuestion rules with before/after pairs
max_tokens_exceeded, large inputs, batching, truncation, untrusted text in statereferences/state-and-budget.mdFitting stages, batching, bounds, adversarial state
any production Jev call: request size, run budget, parallelism, retries, caching, logging, failure diagnosisskills/shared-patterns/jev-production-lessons.mdThe concrete rules and checklist; read before shipping
429, 529, Gateway 503, rate limits, tokens per second, concurrency, fan-out size, retries, eval pacing, "works locally but fails in production"references/state-and-budget.md (Rate limits), references/improve-and-calibrate.md (Service failure or wrong design)Per-second pricing, pacing, backoff numbers, and the diagnosis order
fan-out, confidence gates, composite scores, taxonomy walks, cascades, second requests, long horizon, multi-step, agent loop, beam, branching, checkpointreferences/composition-patterns.mdDocs patterns plus ours, with script paths as worked examples
hooks, reader/storage/action, fail modes, persistence, calibration storereferences/integration-lifecycle.mdWhere a call lives and what happens when Jev is down
wrong answers, low confidence, clustered scores, revision discipline, known debt, baseline, OOD, holdoutreferences/improve-and-calibrate.mdSymptom table and labeled-example loop
dissolving a skill, replacing an LLM with Jev, three-tier classificationreferences/dissolving-a-skill.mdMethod, phase table, worked example
decision surface, card, gate design, threshold, what numbers mean, failure behavior, versionsreferences/decision-card.mdDecision card template: fields every gate must define before code ships
position of a judgment, operand, gate, post-judge, selector, verifier, logical operators, dissolve a skill phasereferences/composition-positions.md11 positions a judgment can occupy relative to a function, mapped to our scripts, with the walk-the-positions procedure
iteratively improve a skill, rubric, prompt, policy, or other artifact with broad Jev feedbackreferences/iteration-with-jev.mdControlled A/B iteration, wide independent question batteries, variance checks, stopping rules, and avoiding optimization artifacts
model limitations, version changes, or a failure-mode auditreferences/question-design.md, references/state-and-budget.md, references/primitives.md, references/improve-and-calibrate.mdLiteral wording, arithmetic and dates, indirection, state filtering, hostile state, structural invariants, generation boundaries, and labeled retesting

Read the live docs

The TypeSafe docs are the source of truth for the API, SDKs, models, limits, and prices. Read them as part of the task; this skill carries our build procedure and measured lessons.

  • Start at the documentation index. Append .md to a page path for Markdown.
  • Read the jaggedness page for the exact model version you deploy. Pin that version; when it changes, reread the page and rerun the labeled set before reusing thresholds.
  • Before you write an integration, read the API page, the page for each primitive you use, and the closest cookbook. A cookbook often shows a better decomposition than a plain classifier.
  • The typesafe:typesafe-ai skill lists the design patterns the docs cover (route and fill arguments, select instead of generate, rerank, feature discovery, verify and escalate). Load it when you explore what to build.
  • Treat thresholds and results in cookbooks as examples to test on your data.
  • Read the documented rate limits on the models page, not just the request limits. On 2026-09-22 they were 64,000 tokens per request, 32,000 for state plus the longest question, 250,000 input tokens per second, and 1,200 requests per minute, "subject to dynamic adjustment". scripts/jev_limits.py holds these numbers; update it when the page changes.

Transports

A program reaches Jev through one of two transports. Use Vercel AI Gateway (JEV_TRANSPORT=vercel); the direct API is for measurement only. Benchmark and price on the transport production uses; their latencies and error codes differ.

TransportHowErrors to back off onNotes
Direct APIPOST https://api.typesafe.ai/... with TYPESAFE_API_KEY; scripts/jev_router_common.py429, 529Official SDKs retry with backoff by default.
Vercel AI GatewayAI SDK experimental_evaluate with a plain model string, or scripts/jev_vercel.py; key is AI_GATEWAY_API_KEY (or OIDC on Vercel)429, 503, 529The Gateway passes the payload to Jev verbatim. It reports upstream rate limiting or overload as HTTP 503 GatewayInternalServerError ("Service temporarily unavailable"), and it uses the same 503 for fast transient failures whose rate grew with tokens per request in our measurements (2026-09-22: ~1.6k tokens 0/8, ~3k 1/8, ~5k 3/8, ~13k 5/8, one request at a time). A 503 is a retry signal, not proof of an outage, a bad payload, or rate limiting; measure which before redesigning. The Gateway question type "boolean" is Noul. Read providerMetadata.gateway.routing on failures.

Both transports share the same account-level rate limits. A direct-API probe that succeeds says nothing about the Gateway path, and the reverse; a direct-key 401/402 says nothing about a Gateway-routed app.

The three tiers

Three things run this toolkit: deterministic programs, Jev, and LLMs. Apply the lowest tier that can do the job.

TierWhenExamples
1. ProgramThe answer is computablesearch, parse, count, diff, validate, run a command, regex, build, test
2. JevThe answer is a judgment over evidence in handclassify, gate, score, triage, verify, choose from a fixed set, decide to escalate
3. LLMThe output is a new artifactwrite code, draft prose, produce a plan, diagnose a novel problem, synthesize across sources

An LLM call in a hook, gate, router, or review is a defect unless the output is generative. A Score, a Choice, or a yes/no decision is never generative. The narrow exception is a bounded review of Jev-referred residuals: the reviewer receives a frozen, source-bound evidence bundle and returns a fixed answer, never new prose or a replacement pipeline. Use it only when a held-out benchmark shows its marginal quality lift justifies its referral rate, marginal cost, and added wall time. When you catch an LLM doing a job Jev can do, replace it.

Jev bills input tokens: the state plus the full text of every question. Output is free. Questions in one request run concurrently, so batching avoids serial call latency, but every question still consumes tokens and shared request budget. Measure actual p50/p95 wall time and accumulated model call time on the workload; do not rely on a universal latency promise. An LLM costs far more per call, takes seconds, and can rationalize a wrong answer. The toolkit metric is LLM calls per request; Jev programs exist to drive it toward zero, with benchmarked residual review as the stated exception.

Programs produce the evidence. Jev judges it. The LLM acts on those judgments creatively, receiving tier 1 and 2 findings as prior_results, not re-judging them. The only exception is the bounded residual-review stage above. When all phases of a skill are tier 1 and 2, the skill dissolves into a Jev program and no LLM runs at all.

Tier 1 goes first on every unit; tier 2 receives the residual tier 1 leaves undecided. That is what "program first" means in practice: a rule the data supports is written in code and scored before any question is written.

Pick the shape

Name the shape of the problem first. The shape decides what code does, what Jev does, and how many requests a run costs.

ShapeSignsBuild
Decide from historylabeled outcomes exist; signals are computable from dataCode builds a correlation table and writes rules for the sure units. Jev judges the residual the rules leave undecided.
One document, many propertiesreview a file, grade a draft, check a diffOne request per document: the state once, every independent, action-changing question once. Stages are code thresholds over that one answer set. A second request carries only evidence the first lacked.
Pick from known optionsroute a request, classify an error, choose a templateCode produces the candidates. A cheap wide Choice ranks them; a second Choice reranks the shortlist with full detail; a confidence gate decides act, confirm, or hand off.
Many items, same questionrank comments, filter tool results, triage filesCode decides the obvious ends. The middle goes in one request as short per-item Nouls. Code counts and sums. Past about 50 items, or when one run would spend more than about 50,000 tokens, build a cascade instead: stage 1 asks one short fit Noul per item over compact state; stage 2 asks the full question set for the top survivors only. Do not fan the full question set out over every item in parallel.
Event streamsomething to check on every tool call, reply, or commitBuild it as an on-demand command. Promote it to a hook after the four conditions in step 11.
Select, then copyextract a value, pick a source span, recover structureCode finds the candidate values or spans. Jev selects the intended one. Code copies or normalizes it. No text is generated.
New text neededwrite, rewrite, plan, diagnoseFirst check whether "Select, then copy" fits. When it does not, an LLM writes. Jev grades the result against a rubric that has its own labeled set.

Build procedure

Build one Jev system at a time. A system is finished when it has labeled cases, a measured score, a measured cost per run, and an action that uses the answer. Start the next system after that.

Evidence of value is a labeled run. Unit tests with fake Jev answers show that the code runs; a labeled run shows that the grading is right.

StepTierDoExit gate
1. State the decision-Write one sentence: the decision, the unit it applies to (a match, a diff hunk, a prompt), and the action code takes on each answer.A person can label one unit by hand in under a minute.
2. Build the grader1Collect human-confirmed labeled units (x, y) with label provenance. Freeze fixture copies and hash both fixtures and rubric. Split train/dev from an untouched group-disjoint heldout. Provisional or agent labels are diagnostics, never action-promoting ground truth.score(predictions) runs on dev and prints the majority-class baseline; the heldout is sealed.
3. Discover signals1Run SQL or Python over train. For every computable signal, record accuracy against y, count, and the same per slice. Start from existing analytics code. Keep every signal; the table decides.A correlation table sorted by accuracy, with counts.
4. Write the policy1Turn the table into rules: rule(x) -> (action, sure). The strongest signal decides; a near-certain signal overrides. Score the rules on dev.The rules and their dev score are row one of the run log. The residual (every unit where sure is false) is counted.
5. Design the request1+2Build state for residual units only: correlated signals, bounded, labeled, arithmetic done in code, plus the rules' verdict and why it was unsure. Write one atomic question per judgment, worded from the table. Match the primitive to the action. Put every independent question about one state in one request. A question whose evidence/options depend on another answer is a second request after code builds the new state.The decision card is filled in (references/decision-card.md).
6. Price the run1Run the program on a ten-word input: the billed tokens are the fixed floor, your question text. Compute calls per run = units x calls per unit x rounds, tokens per call, tokens per run, referrals per run, and worst-case retry sends. Then price it per second: dump every request one run sends to JSON and run python3 scripts/jev-budget-check.py --payload run.json --concurrency C --concurrent-runs N --attempts A. Pick the request size with python3 scripts/jev-size-probe.py --payload run.json on the production transport. Price any eval the same way with --eval-cases. State all numbers.The numbers are ones you would approve and the budget check says ok: peak tokens per second and requests per minute stay under 25% of the documented limits with retries and concurrent users counted. When the floor exceeds the typical state, shorten the questions first. When tokens per run exceed about 50,000, redesign as a cascade before tuning anything else.
7. Smoke run2Run the three-unit set, then the dev sample.calls_failed is zero, every answer parses, and python3 scripts/jev-cost-report.py --since 1h matches the step 6 estimate.
8. Score1On the same dev set, report the rules alone, Jev on the residual, and the combined system, per slice, with Brier and a calibration curve. Count false positives beside recall. Run judge variance once over frozen rows.The combined score and its cost per run are in the run log.
9. Improve1+2First separate code errors and service failures (HTTP errors, timeouts) from wrong answers, by reading the exact state, questions, candidates, and answers of each miss. Then classify the wrong answers (state_lacked_evidence, criteria_ambiguous, wrong_primitive, label_noise). Change one state, instruction, criterion, or policy lever at a time. State changes must add needed decision evidence, not decorative context. Re-score. Keep the change when the combined score climbs and every slice holds.Each variant is logged with score and cost.
10. Report1Score the untouched heldout once after selection. Report p50/p95 wall time, accumulated model call time, throughput, and (when used) referral rate, marginal reviewer lift, cost, and time. Before later tuning, create a new independent heldout.One heldout number and workload metrics, reported beside the dev number.
11. Integrate1+2Ship an on-demand command with a reader, storage, and an action. Log whether each answer changed the action. Promote to a hook when four conditions hold: code decides the obvious cases first; the labeled set shows the answers are right; the answer distribution is meaningfully non-constant; something acts on the answer. Run a new hook in shadow mode first, and promote one hook at a time.A day of use shows the cost report and the action-changed rate you expected.
12. Next system-Start step 1 for the next decision.-

Step 9 levers, in search order: evidence in state; decomposition (one Score into several Nouls); criteria wording; thresholds; few-shot examples in state. Evidence comes first because the other levers work only on a signal that is present. Retune thresholds from stored probabilities, which costs zero calls. Derive a gate threshold from action costs, t = C_FP / (C_FP + C_FN), select it on one split, and report on another.

Cost model. Cost = calls x input tokens per call. Input tokens = state + the text of every question, with its criteria and examples. Output is free. Parallel questions reduce wall time relative to serial sends but do not make question text free. Fill a request with independent heads only while each has decision/action value and the total fits its 64,000-token budget; sequence only a head whose evidence or candidates are derived from a prior answer. Measure wall time separately from accumulated model call time (the sum of attempt durations): concurrency can lower the former while leaving the latter high. Throughput is units or KB divided by the chosen clock; label the clock. Three numbers govern a run:

NumberTargetReach it by
Sends per stateone per runone request per unit; stages as code thresholds; a second request only for new evidence
Fixed floor per callbelow the typical state sizeone- or two-line questions; what, not_for, and examples only where labeled misses call for them
Firing ratematches how often the answer changes an actionon-demand commands first; hooks after step 11's conditions
Peak tokens per secondunder 25% of the documented 250,000 (about 60,000), with retries and concurrent users countedfewer tokens per run (a cascade instead of full detail for every unit); an instance-wide in-flight cap sized from the budget so simultaneous runs cannot burst together; jittered backoff for 429/529
Tokens per answer, retries includedrequest size near the minimum of size / success_rate(size) on the production transportmeasure the transient failure rate at several request sizes (references/improve-and-calibrate.md), then pack requests to that size

Per-request fit is not enough. Every request can sit far under 64,000 tokens while one run still spends the whole per-second limit: 20 requests of 15,000 tokens sent together is 300,000 tokens in about a second. Retries then multiply it, because every failed request resends its full state. A design that works for one test query fails for real users, and an eval of 80 such runs spends millions of tokens in minutes. Price tokens per run and per second in step 6, not only tokens per request.

Measure repeatability before iterating. Run repeated frozen requests and measure answer variance and decision flips on the workload. Keep thresholds away from where answers cluster, and establish this noise floor before comparing variants. Treat cache behavior and circuit-breaker behavior as implementation details to verify in the current runner rather than performance guarantees.

Telemetry is part of the system. call_jev logs every call: script name, session id, input tokens, question count, payload hash, cached or not, error. An evidence-pipeline runner also persists the evidence-bundle ID/version, prompt/question version, model/version, attempt number, retry reason, deadline/cap, response, accounting, and final keep/refer/action outcome. Preserve source rows and provenance through joins: a relationship label is not permission to merge identities. Read the cost report after every multi-call run and compare it with the step 6 estimate.

Graders see only what you send. Send the richest available output and the evidence itself: command output, file content, stored answers. Tune on one label set and report on another.

Action-changing gates stay conservative. An unavailable, missing, or invalid Jev answer is unknown: exclude it from quality scores, retain its failure receipt, and never reinterpret it as no, pass, or permission to act. Keep any action-changing selector in shadow mode until human-confirmed, disjoint-heldout results show that its action improves the intended outcome. Confidence measures concentration, not authority: it cannot authorize an action or override source evidence, permissions, or deterministic safety rules. An evidence question needs a supplied source-evidence ledger; plausibility and apparent intent are not source evidence.

Sanity floors (majority class, the single strongest signal) prove the pipeline is wired. The bar is higher: the combined system climbs across iterations, calibration holds on the residual, and the test set agrees once. Spend scales with the residual, so a good policy keeps each round to hundreds of calls.

The grader decides how far the procedure goes. With outcomes that already happened (a result, a merged PR, a finished run), the loop runs unattended. With hand labels, it runs until the labels are used up; then the next step is more labels.

The same procedure replaces a skill: the skill's phases supply the signals and questions, hand labels are the grader, and the policy function replaces its gates (see "Dissolving a skill"). Systems compose: one system's decision is another's signal. Deterministic driver: scripts/jev-harness.py (loop, variance, sweep).

Primitives

PrimitiveAsk whenReturnsCode acts with
Noulclean yes/no; the probability is the signalnoul in [0, 1]; no confidenceif noul > t
Choiceone of a known unordered setchoice, probabilities, confidencea branch per option
Scorea position on a spectrum you can describe in stepsscore, legend, probabilities, confidencethreshold, rank, or round

score is the probability-weighted mean of level numbers (0-based), not a picked level. A 1.0 can be certainty on level 1 or a 0/2 split; read probabilities when the distinction matters. Threshold, rank, or round it; never interpolate a quantity from it. A Noul at 0.5 means unsure, not "medium"; distance from 0.5 is its confidence. Do not carry a threshold tuned on one primitive to another, and do not expect P(noul) and 1 - P(not noul) to agree. Full shapes: references/primitives.md.

confidence on a Choice or Score measures how concentrated the distribution is. It does not say the workflow is right, and it is not permission to act. Several acceptable options also spread probability, so low confidence on a harmless preference choice is fine. Set thresholds from your labeled data and the cost of each action.

Question rules

  • State the exact condition. Jev reads scoping words and negations literally. When you find yourself explaining what you meant after a miss, that explanation is the missing half of the instruction.
  • One narrow, coherent judgment per question. Split dimensions that are useful on their own; keep together a relationship that is the thing being judged (does this reply answer this question). Atomic does not mean one sentence: a bounded action choice or a reading in context is one judgment. No double negatives, no multi-hop questions.
  • The question ID is your key and is not sent to the model. Put the full meaning in instructions.
  • When Jev selects from candidates that code produced, check coverage first: Jev cannot pick a value that is not in the list.
  • Name the state path in backticks: `ticket.messages[0].text`.
  • Criteria and instruction ask the same thing in the same direction. A Noul whose true side describes "no" degrades.
  • Criteria encode the hard cases. Jev handles the obvious ones alone. what, not_for, examples per Choice option; true/false with what and examples for a subtle Noul.
  • Score levels: 2 to 10, each a standalone situation, one dimension, no numerals and no "worse than the previous". Give a rare extreme its own level.
  • Choice: add other or none when the list may not cover the input. Examples are concrete instances ("charged twice"), not descriptions of instances.
  • Instructions accept a string or an object (question, focus, inspect, note, compare, field). Pass schemas and rows as JSON, never serialized into a string.
  • Ask many specific questions, not one broad one. Put them in one call so the state is billed once. Every question's text is billed too: write each in one or two lines, and add what, not_for, and examples only where labeled misses show the question needs them.
Show full SKILL.md (2,319 more words)Show less

State rules

  • State is evidence, not instructions. No coaching in state; rules go in instructions and criteria.
  • Send only what the questions need. Irrelevant detail lowers accuracy and hides which input caused a miss.
  • Bound every field with a named constant; keep the tail; note omitted characters; label sections ([Request], [Diff], [Prior Assessment]).
  • Convert numbers to words or buckets. Compute dates, durations, counts, and sums in code. Jev does not count: one Noul per item, sum in code.
  • A request holds 64,000 tokens: the state plus every question. The state plus the longest single question must stay under 32,000. The account also has a per-second limit (250,000 input tokens per second, 1,200 requests per minute on 2026-09-22) shared by every request, retry, user, and eval. Check the models page for current limits. The state is billed again in every request, so fill each request with as many questions as fit before you start a second one. Fit state in stages. Every stage that calls Jev needs fitting, not just the first.
  • Keep observed facts and inferred values in separate, labeled fields. Check that the state is still current before you act on an answer about it.
  • Jev does not treat state as hostile. Text in state can steer answers. Apply skills/shared-patterns/untrusted-content-handling.md, name in criteria what counts, and run adversarial and self-describing test cases before deployment.

Composition patterns

PatternShapeWorked example
Speculative fan-outevery branch's questions in one call, each stating its own premise ("if this is a refund request, ..."); heads are independent and cannot see one another's answers; code ignores unused heads and their uncertaintyscripts/jev-browser-decide.py
Confidence-gated routinga floor below which nothing acts and a higher bar for high-stakes actions; paths act / confirm / hand off. Select both thresholds from labeled data and action costsscripts/jev-route.py
Composite scoringone Score per dimension, normalize by len(criteria) - 1, weights in code. Weighted sums suit preferences that offset one another; an "any serious violation" rule needs its own Noul per conditionreferences/composition-patterns.md
Intent routingChoice for intent plus complexity Score, both confidence-gatedscripts/jev-route.py
Taxonomy walkone Choice per level; each option's criteria is its trimmed subtree; follow several branches when closereferences/composition-patterns.md
Multi-Noul decompositionsplit a compound goal into one Noul per clause; combine in codereferences/composition-patterns.md
Cascade plus verificationone wide request per unit; code thresholds pick survivors. Send a second request when the first answer is needed to fetch evidence, build new state, or decide the next options; it carries only what the first lackedreferences/composition-patterns.md
Bounded residual reviewJev handles most units, a fixed-answer reviewer checks benchmarked referrals; runner sends the same source-bound bundle with immutable provenance; code accepts only the declared answer schemareferences/composition-patterns.md
Deterministic pre-filterprograms decide the obvious ends; Jev judges the middlescripts/jev-compact.py
History injectionrecent actions as "already taken, do not repeat"scripts/jev-browser-agent.py
Checkpoint searchcaller sets subgoals a few steps apart; Jev beam-searches between them with progress-comparison Choices; LLM only on a near-tie, missing answer, or stalled subgoal; a small ledger replaces raw historyscripts/jev_search.py

One screen each, with the code shape: references/composition-patterns.md.

Integration lifecycle

Start every program as an on-demand command. Promote it to a hook after it meets the four conditions in step 11 of the build procedure. Promote one hook at a time and read the cost report after a day of use.

Every integration has a reader (runs Jev), storage (findings persist somewhere read), and an action (something changes behavior). Missing any part wastes the call. Thread prior assessments into later calls as bounded, labeled evidence. Fail open for advisory checks; fail to warn for safety checks; never fail to block when Jev is unavailable. Validate every response with validate_jev_response before acting. Details: references/integration-lifecycle.md.

Improve a program

This is step 9 of the build procedure. Find the failing question on labeled data before changing anything.

Promote lessons deliberately. Experimental Jevmaxxing is hypothesis discovery, not guidance. Add a durable rule only when it has a clear mechanism, representative labeled evidence, a stated boundary or counterexample, and a measured improvement to an action, cost, or quality decision against a baseline. Otherwise leave the observation out; prune copied lore that cannot meet this standard.

SymptomLikely causeFix
Wrong with high confidenceliteral readingstate the exact condition; put the boundary case in criteria
Low-confidence Choiceoptions overlap or none fitsadd what, not_for, examples; add other
Low-confidence Scorelevels overlap, two dimensions, thin staterewrite levels as situations; split; add the missing field
Scores cluster mid-scalelevels are degrees or numeralsdescribe a situation per level; drop numerals
Extremes look alikeno level for the extremeadd one
Noul near 0.5vague conditiondefine it; add true/false examples
Accuracy falls with input sizeirrelevant statefilter in code; send fields, not blobs
Count, sum, date errorsJev doing arithmeticmove it to code; per-item Nouls
Nested or negated questions failindirectionask directly; split into two literal questions
Answer follows text in statestate steeringseparate trusted policy from untrusted text; require source evidence in code; adversarial tests; thresholds may abstain or refer, not authorize
Rewording trades one error for anotherone question, several propertiessplit into atomic questions
Answers right, decision wrongpolicychange weights or thresholds in code, not questions
Slow or costlysequential callsmerge into one request
One query works; real use or an eval fails with 429/529/503one run spends too much of the per-second limit; retries multiply itprice with jev-budget-check.py; cascade; cap concurrency; jittered backoff

Operational rules:

  • Store every answer's probabilities with the payload hash; retune thresholds from stored answers, which costs zero calls.
  • Measure judge variance over frozen rows before trusting a judge; gate only on a judge whose answers hold steady between runs.
  • Derive a gate threshold from action costs t = C_FP/(C_FP+C_FN), select on one split, report on another, re-measure when the data shifts.
  • Run a new gate in shadow mode (log the action it would take) until replayed fixtures pass, then enforce.
  • Let a domain rule veto an action regardless of model confidence (permit != confidence).
  • Grant "done" only to a post-execution probe (test, build, exit code); the probe result is what decides.

Rules: change one or two questions per revision; judge on labeled data, not confidence alone; keep the answer space stable once code depends on it; general rules in criteria, specific names only in examples. Full table and the known-debt note: references/improve-and-calibrate.md.

Dissolving a skill into a Jev program

A dissolution is the build procedure with the skill as the request. The method:

  1. List phases. Read the skill's SKILL.md. Write each phase on one row.
  2. Classify each phase into four columns: deterministic (program), judgment (Jev), generation (LLM), or orchestration (dispatch/coordination).
  3. Convert judgment phases. Each judgment becomes one or more Jev questions: Noul for gates, Choice for classification/routing, Score for severity/quality. Write criteria for the hard cases.
  4. Keep deterministic phases in code. Regex scans, file reads, grep, counts, averages, formatting stay as programs.
  5. Isolate generation. If any phase requires new text (rewrite, diagnosis, plan), that phase keeps an LLM. The LLM receives all prior Jev decisions as prior_results and does not re-judge.
  6. Write the policy function. A pure function policy(assessment) -> action with named thresholds is the dissolved skill's contract. It replaces the skill's gates.
  7. Prove agreement. Run the Jev program on hand-labeled examples. Match or exceed the skill's accuracy before deleting the SKILL.md.

Worked example. references/dissolving-a-skill.md walks one skill through the method: phase table, Jev question set, and policy function.

Checklist

  • This is the only Jev system under construction; the previous one has labeled cases, a score, a cost per run, and an action.
  • The fixed floor, sends per state, and calls per run are measured; each state is sent once per run.
  • The expected call count was computed before launch and matches the cost report after.
  • scripts/jev-budget-check.py says ok for one run at the planned concurrency, attempts, and concurrent users, and any eval was priced with --eval-cases before it ran.
  • A run over about 50 units or 50,000 tokens is a cascade: a cheap wide stage over every unit, full detail only for survivors.
  • Independent requests in a stage are sent together, not in waves; an instance-wide in-flight cap derived from the per-second budget bounds simultaneous runs. Rate answers (429/529, repeated 503s) back off exponentially with jitter (base at least 0.5 s) and honor Retry-After; a lone fast Gateway 503 retries after about 100 ms. Every run has a retry budget.
  • A deterministic policy over the signals is written and scored first; Jev receives the residual it leaves undecided.
  • Each question asks one property a person could answer in a second.
  • The primitive matches how code uses the answer.
  • Instructions state the exact condition and name state paths in backticks.
  • Criteria agree with the instruction and point the same way; hard cases are encoded.
  • Score levels are standalone situations with no numerals; Choices that may not cover the input have other.
  • Score uses a criteria list (2–10 level descriptions), never min/max. Choice uses a criteria map with what/not_for/examples.
  • Code does all counting, arithmetic, and date comparison.
  • State holds only what questions need, every field bounded by a named constant, sections labeled.
  • Every stage that calls Jev fits state and batches questions; calls_failed is zero on a labeled run.
  • All independent questions on one state travel in one request; serial dependencies are explicit second requests built by code.
  • Responses are validated; a pure policy function decides; thresholds sit in the policy.
  • Reader, storage, and action all exist; assessments persist with full distributions, prompt/model versions, attempts, and final actions.
  • Entity or linkage systems preserve original rows and provenance; relationship labels and identity merges are separate actions.
  • Workload reports distinguish wall time from accumulated model call time and label throughput's clock.
  • Any non-Jev residual reviewer is fixed-answer, source-bound, capped, and justified by a held-out marginal benchmark.
  • score is read as a weighted mean; code that rounds says so.
  • Adversarial and self-describing inputs are in the test set.
  • Labeled examples back every revision; the model version is pinned or the jaggedness page rechecked.
  • Fixtures and rubrics are hashed; labels record human/provisional provenance, and provisional labels never promote an action.
  • An untouched group-disjoint heldout is used once after selection; later tuning starts with a new independent heldout.
  • Missing, invalid, and unavailable answers are stored as unknown with separate failure receipts, never scored as pass or no.
  • Any action-changing selector remains shadow-only until human-confirmed, disjoint-heldout evidence shows the action is useful.
  • Evidence questions receive an explicit source-evidence ledger; confidence cannot override evidence or permission constraints.
  • Every script passes a validation probe: imports without error, --help exits 0, and a live call with representative input returns valid JSON with source != "error".
  • Hooks that grade agent output read the richest available text (task-notification result, not just the last assistant message).
  • Hooks that check grounding receive verifiable evidence (stored Jev answers, tool output summaries), not just file paths.
  • In development, run on a small representative sample before the full dataset. A bad question wastes every call.
  • Read the cost report after every multi-call run. Investigate scripts with no successful calls, duplicate payload hashes, or any unexpected failures.

Error handling

Error: HTTP 422 on the request

  • Cause: wrong schema. Choice needs criteria as a map; Noul uses criteria.true/criteria.false; Score uses a criteria list. Keys such as options, min, max are not part of the API.
  • Solution: match references/primitives.md; validate against the live API, not a mocked test.

Error: max_tokens_exceeded, or a request that fails because it carries too much context

  • Cause: state plus questions exceed a limit at some stage. Through the Gateway this is the most common failure, and it recurs well below the documented 64,000 tokens: large requests fail while small ones return fine.
  • Solution: check with a ~100-token request; if that returns, the cause is size. jev_transport.evaluate splits oversized requests into as many small ones as it takes, each with the same state; in other code, do the same with jev_limits.split_and_run. Shrink the state when it alone passes the target. Fit state in stages, cap items per call, and count failures per stage. See references/state-and-budget.md and fact 2 in the production rules.

Error: HTTP 429 or 529 (direct), or 503 through Vercel AI Gateway

  • Cause: rate limit (tokens per second or requests per minute) or an overloaded service. Through the Gateway these arrive as 503 GatewayInternalServerError. The most common cause in our programs is the program itself: one run, or an eval of many runs, sending more tokens per second than the account allows, or requests large enough that transient failures are frequent (through the Gateway, the 503 rate grew with tokens per request).
  • Solution: first price the run with scripts/jev-budget-check.py. If it is over 25% of a limit, fix the design (cascade, fewer tokens per run, a concurrency cap), not the retry loop. Then retry with exponential backoff and equal jitter (jev_limits.backoff_delay: base 0.5 s, doubling, capped, half random) so parallel requests that failed together do not retry together, and treat Retry-After as a floor. Cap attempts per request and retries per run; when the budget is spent, stop sending and return what finished. Persist every failed and retried attempt with its reason; retries count in workload cost and latency. Short fixed delays (100 ms, 200 ms) across many parallel requests make a retry storm that keeps the limit tripped.
  • Diagnose before blaming the payload or the service: see "Service failure or wrong design" in references/improve-and-calibrate.md.

Error: HTTP 401 or 402

  • Cause: bad key or exhausted credits. A retry never succeeds.
  • Solution: the breaker in call_jev stops further sends. Code outside call_jev (a sandboxed plugin) stops after the first such status. Keep the API key on the server side; never ship it to a browser.

Error: valid response, wrong decision

  • Cause: policy reads score as a level index or thresholds on the wrong primitive.
  • Solution: read probabilities; keep thresholds in the policy; see the known-debt note in references/improve-and-calibrate.md.

© notque, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/meta/building-with-jev of notque/vexjoy-agent.

  • SKILL.md
  • references/composition-patterns.md
  • references/composition-positions.md
  • references/decision-card.md
  • references/dissolving-a-skill.md
  • references/improve-and-calibrate.md
  • references/integration-lifecycle.md
  • references/iteration-with-jev.md
  • references/primitives.md
  • references/question-design.md
  • references/state-and-budget.md

Open the folder on GitHubat commit 5218674

Compare with similar skills

Building With Jev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Building With Jev compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Building With Jev this skillnotque/vexjoy-agent439—~11kAutomated safety check: NotesMIT
Integration Testingthedaviddias/Front-End-Checklist74k—~514Automated safety check: PassMIT
Frontend API Integration Patternssickn33/agentic-awesome-skills47k1 repos~1.9kAutomated safety check: PassMIT
API Integrationsickn33/agentic-awesome-skills47k1 repos~1.3kAutomated safety check: PassMIT
Labarchive Integrationdavila7/claude-code-templates32k10 repos~2.3kAutomated safety check: PassMIT
Robius Matrix Integrationsickn33/agentic-awesome-skills47k2 repos~3.6kAutomated safety check: PassMIT

Similar skills

  • Integration Testing

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing CI coverage, automated checks, or test strategy related to Write integration tests for key workflows.

    74k GitHub stars~514 tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Frontend API Integration Patterns

    sickn33/agentic-awesome-skills

    Production-ready patterns for integrating frontend applications with backend APIs, including race condition handling, request cancellation, retry strategies, error normalization, and UI state…

    47k GitHub starsUsed in 1 repo~1.9k tokens
    Frontend & DesignAuto-check passed
  • API Integration

    sickn33/agentic-awesome-skills

    Designs event-driven architectures, webhook systems, API chaining flows, ETL pipelines, and integration patterns between services.

    47k GitHub starsUsed in 1 repo~1.3k tokens
    Backend & APIsAuto-check passed
  • Labarchive Integration

    davila7/claude-code-templates

    Electronic lab notebook API integration. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~2.3k tokens
    Backend & APIsAuto-check passed
  • Robius Matrix Integration

    sickn33/agentic-awesome-skills

    CRITICAL: Use for Matrix SDK integration with Makepad. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~3.6k tokens
    Backend & APIsAuto-check passed
  • API Integration Architect

    sickn33/agentic-awesome-skills

    Design, implement, debug, and optimize API integrations with expert-level patterns for REST, GraphQL, webhooks, and authentication flows.

    47k GitHub starsUsed in 1 repo~2.3k tokens
    Backend & APIsAuto-check passed

More from notque/vexjoy-agent

All 61 skills in this repo
  • Game Asset Generator

    notque/vexjoy-agent

    Deterministic palette/matrix pixel art (not AI). An agent skill from notque/vexjoy-agent.

    439 GitHub stars~2.3k tokensUpdated 6 days ago
    Auto-check: notes
  • PR Workflow

    notque/vexjoy-agent

    Pull request lifecycle: commit, codex review, sync, review, fix, status, cleanup, and PR mining.

    439 GitHub stars~2.8k tokensUpdated 6 days ago
    Auto-check: notes
  • Architecture Deepening

    notque/vexjoy-agent

    Improve architecture across modules by deepening interfaces.

    439 GitHub stars~3.3k tokensUpdated 6 days ago
    Auto-check: notes
  • Code Quality

    notque/vexjoy-agent

    Code quality: cleanup, linting, formatting, quality gates. An agent skill from notque/vexjoy-agent.

    439 GitHub stars~1.5k tokensUpdated 6 days ago
    Auto-check: notes
  • Codebase Analyzer

    notque/vexjoy-agent

    Statistical rule discovery from Go codebase patterns. An agent skill from notque/vexjoy-agent.

    439 GitHub stars~2k tokensUpdated 6 days ago
    Auto-check: notes
  • Comment Quality

    notque/vexjoy-agent

    Review and fix temporal references in code comments. An agent skill from notque/vexjoy-agent.

    439 GitHub stars~2k tokensUpdated 6 days ago
    Auto-check: notes

Questions about Building With Jev

What does Building With Jev do?

Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model. Building With Jev is an agent skill from notque/vexjoy-agent. Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.

How do I install Building With Jev in Claude Code?

Run `npx skills add notque/vexjoy-agent --skill building-with-jev -a claude-code`. Or copy the skill folder (skills/meta/building-with-jev in notque/vexjoy-agent) into .claude/skills/building-with-jev in your project. Claude Code loads it when a task matches its description.

How do I install Building With Jev in Codex?

Run `npx skills add notque/vexjoy-agent --skill building-with-jev -a codex`. Or copy the skill folder (skills/meta/building-with-jev in notque/vexjoy-agent) into .agents/skills/building-with-jev in your project. Codex loads it when a task matches its description.

Can I use Building With Jev in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add notque/vexjoy-agent --skill building-with-jev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/building-with-jev, .gemini/skills/building-with-jev, .github/skills/building-with-jev and .opencode/skills/building-with-jev in your project.

What does Building With Jev need to run?

Going by SKILL.md and its folder, Building With Jev needs the command-line tools its instructions call (python3) and credentials named TYPESAFE_API_KEY and AI_GATEWAY_API_KEY. Our summary lists: A credential in TYPESAFE_API_KEY; A credential in AI_GATEWAY_API_KEY. Its frontmatter pre-approves these tools: Read, Edit, Write, Bash, Glob, Grep.

Does Building With Jev access the network?

SKILL.md names 2 domains. In commands or code: api.typesafe.ai; the agent is likely to contact it when it follows the instructions. As links in the text: docs.typesafe.ai. This is read from the text; nothing was executed.

Is Building With Jev safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Building With Jev use?

Building With Jev is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Building With Jev use?

About 11k tokens (SKILL.md is roughly 43k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.

What are the alternatives to Building With Jev?

Skills that share tags, products or a category with Building With Jev: Integration Testing (thedaviddias/Front-End-Checklist, 74k stars), Frontend API Integration Patterns (sickn33/agentic-awesome-skills, 47k stars), API Integration (sickn33/agentic-awesome-skills, 47k stars) and Labarchive Integration (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Building With Jev?

notque (a GitHub user) maintains it in notque/vexjoy-agent, which has 439 GitHub stars. The repository holds 61 skills in this directory. The repository was last updated on October 3, 2026.

Source: notque/vexjoy-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.