Integration Testing
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing CI coverage, automated checks, or test strategy related to Write integration tests for key workflows.
Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.
$ npx skills add notque/vexjoy-agent --skill building-with-jev -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install notque/vexjoy-agent building-with-jev --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/meta/building-with-jev .claude/skills/building-with-jev && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "building-with-jev" agent skill from https://github.com/notque/vexjoy-agent/tree/main/skills/meta/building-with-jev into .claude/skills/building-with-jev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-with-jev", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/notque/vexjoy-agent/tree/main/skills/meta/building-with-jevType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add notque/vexjoy-agent --skill building-with-jev -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install notque/vexjoy-agent building-with-jev --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/meta/building-with-jev .agents/skills/building-with-jev && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "building-with-jev" agent skill from https://github.com/notque/vexjoy-agent/tree/main/skills/meta/building-with-jev into .agents/skills/building-with-jev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-with-jev", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add notque/vexjoy-agent --skill building-with-jev -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install notque/vexjoy-agent building-with-jev --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/meta/building-with-jev .cursor/skills/building-with-jev && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "building-with-jev" agent skill from https://github.com/notque/vexjoy-agent/tree/main/skills/meta/building-with-jev into .cursor/skills/building-with-jev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-with-jev", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/notque/vexjoy-agent.git --path skills/meta/building-with-jev--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add notque/vexjoy-agent --skill building-with-jev -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install notque/vexjoy-agent building-with-jev --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/meta/building-with-jev .gemini/skills/building-with-jev && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "building-with-jev" agent skill from https://github.com/notque/vexjoy-agent/tree/main/skills/meta/building-with-jev into .gemini/skills/building-with-jev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-with-jev", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install notque/vexjoy-agent building-with-jevInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add notque/vexjoy-agent --skill building-with-jev -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/meta/building-with-jev .github/skills/building-with-jev && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "building-with-jev" agent skill from https://github.com/notque/vexjoy-agent/tree/main/skills/meta/building-with-jev into .github/skills/building-with-jev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-with-jev", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add notque/vexjoy-agent --skill building-with-jev -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install notque/vexjoy-agent building-with-jev --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/notque/vexjoy-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/meta/building-with-jev .opencode/skills/building-with-jev && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "building-with-jev" agent skill from https://github.com/notque/vexjoy-agent/tree/main/skills/meta/building-with-jev into .opencode/skills/building-with-jev/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building-with-jev", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
building-with-jevWrite, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.
Building With Jev is an agent skill from notque/vexjoy-agent. Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.
Its SKILL.md is about 11k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `references/composition-patterns.md`, `references/composition-positions.md` and `references/decision-card.md`).
The repository describes itself as: VexJoy AI Agent with Jev Intelligent Routing - /do routes plain-English requests to the right specialist agent and gates the work with reviews, tests, and a learning loop. The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5218674. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadEditWriteBashGlobGrepFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.typesafe.aiAlso links to:
docs.typesafe.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
TYPESAFE_API_KEYAI_GATEWAY_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Building With Jev loads about 11k tokens when it runs, and up to ~29k if it reads all its reference files. Until then it costs about 30 tokens; SKILL.md has 6,131 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Edit, Write, Bash, Glob, GrepAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from notque/vexjoy-agent at commit 5218674, republished under its MIT licence (© notque). 6,131 words, ~10,724 tokens.
.claude/skills/building-with-jev/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.Jev reads one state, answers every question in the request independently and in parallel, and returns a probability distribution over answers you defined. A head cannot read another head's answer: parallel heads share evidence, not reasoning. For one state, maximize independent heads that can change a decision or action, subject to their token cost and the 64,000-token request budget; omit noise heads. Code owns control flow, arithmetic, policy, and every serial dependency; Jev owns the snap judgment. It does not reason in steps, count, do arithmetic, or generate text. Use this skill to design the questions, fit the state, compose answers in code, wire the call into a hook or script, and fix a call that answers wrong.
Before writing or changing any Jev request, apply Jev production rules and tick its pre-ship checklist. It sets request size (2.5–4k tokens via Gateway until measured), per-run budget (~50k tokens), screen-then-detail above ~50 items, sending each stage at once with an instance cap of floor(0.25 × 250,000 / tokens_per_request), retries by status code, eval pacing, caching, logging, and the order to measure failures. The rest of this skill is the method for designing questions; those rules govern how requests are sent. Its four facts come first: use Vercel AI Gateway; too much context is the most common failure, so split an oversized request into as many small requests as it takes (the toolkit transports do this automatically); a small request is the first diagnostic; rate limits are normal, so retry them.
Prefer direct judgments over supplied evidence. Bounded action selection is valid when candidates and decision evidence are supplied. If answering requires an intermediate result that changes later evidence or candidates, code must resolve that dependency before a later Jev request.
| Signal | Load These Files | Why |
|---|---|---|
request or response shape, instruction objects, criteria objects, reading score/probabilities/confidence | references/primitives.md | Full API shape and answer semantics |
| writing or rewriting instructions, criteria, levels, options, examples | references/question-design.md | Question rules with before/after pairs |
max_tokens_exceeded, large inputs, batching, truncation, untrusted text in state | references/state-and-budget.md | Fitting stages, batching, bounds, adversarial state |
| any production Jev call: request size, run budget, parallelism, retries, caching, logging, failure diagnosis | skills/shared-patterns/jev-production-lessons.md | The concrete rules and checklist; read before shipping |
| 429, 529, Gateway 503, rate limits, tokens per second, concurrency, fan-out size, retries, eval pacing, "works locally but fails in production" | references/state-and-budget.md (Rate limits), references/improve-and-calibrate.md (Service failure or wrong design) | Per-second pricing, pacing, backoff numbers, and the diagnosis order |
| fan-out, confidence gates, composite scores, taxonomy walks, cascades, second requests, long horizon, multi-step, agent loop, beam, branching, checkpoint | references/composition-patterns.md | Docs patterns plus ours, with script paths as worked examples |
| hooks, reader/storage/action, fail modes, persistence, calibration store | references/integration-lifecycle.md | Where a call lives and what happens when Jev is down |
| wrong answers, low confidence, clustered scores, revision discipline, known debt, baseline, OOD, holdout | references/improve-and-calibrate.md | Symptom table and labeled-example loop |
| dissolving a skill, replacing an LLM with Jev, three-tier classification | references/dissolving-a-skill.md | Method, phase table, worked example |
| decision surface, card, gate design, threshold, what numbers mean, failure behavior, versions | references/decision-card.md | Decision card template: fields every gate must define before code ships |
| position of a judgment, operand, gate, post-judge, selector, verifier, logical operators, dissolve a skill phase | references/composition-positions.md | 11 positions a judgment can occupy relative to a function, mapped to our scripts, with the walk-the-positions procedure |
| iteratively improve a skill, rubric, prompt, policy, or other artifact with broad Jev feedback | references/iteration-with-jev.md | Controlled A/B iteration, wide independent question batteries, variance checks, stopping rules, and avoiding optimization artifacts |
| model limitations, version changes, or a failure-mode audit | references/question-design.md, references/state-and-budget.md, references/primitives.md, references/improve-and-calibrate.md | Literal wording, arithmetic and dates, indirection, state filtering, hostile state, structural invariants, generation boundaries, and labeled retesting |
The TypeSafe docs are the source of truth for the API, SDKs, models, limits, and prices. Read them as part of the task; this skill carries our build procedure and measured lessons.
.md to a page path for Markdown.typesafe:typesafe-ai skill lists the design patterns the docs cover (route and fill arguments, select instead of generate, rerank, feature discovery, verify and escalate). Load it when you explore what to build.scripts/jev_limits.py holds these numbers; update it when the page changes.A program reaches Jev through one of two transports. Use Vercel AI Gateway (JEV_TRANSPORT=vercel); the direct API is for measurement only. Benchmark and price on the transport production uses; their latencies and error codes differ.
| Transport | How | Errors to back off on | Notes |
|---|---|---|---|
| Direct API | POST https://api.typesafe.ai/... with TYPESAFE_API_KEY; scripts/jev_router_common.py | 429, 529 | Official SDKs retry with backoff by default. |
| Vercel AI Gateway | AI SDK experimental_evaluate with a plain model string, or scripts/jev_vercel.py; key is AI_GATEWAY_API_KEY (or OIDC on Vercel) | 429, 503, 529 | The Gateway passes the payload to Jev verbatim. It reports upstream rate limiting or overload as HTTP 503 GatewayInternalServerError ("Service temporarily unavailable"), and it uses the same 503 for fast transient failures whose rate grew with tokens per request in our measurements (2026-09-22: ~1.6k tokens 0/8, ~3k 1/8, ~5k 3/8, ~13k 5/8, one request at a time). A 503 is a retry signal, not proof of an outage, a bad payload, or rate limiting; measure which before redesigning. The Gateway question type "boolean" is Noul. Read providerMetadata.gateway.routing on failures. |
Both transports share the same account-level rate limits. A direct-API probe that succeeds says nothing about the Gateway path, and the reverse; a direct-key 401/402 says nothing about a Gateway-routed app.
Three things run this toolkit: deterministic programs, Jev, and LLMs. Apply the lowest tier that can do the job.
| Tier | When | Examples |
|---|---|---|
| 1. Program | The answer is computable | search, parse, count, diff, validate, run a command, regex, build, test |
| 2. Jev | The answer is a judgment over evidence in hand | classify, gate, score, triage, verify, choose from a fixed set, decide to escalate |
| 3. LLM | The output is a new artifact | write code, draft prose, produce a plan, diagnose a novel problem, synthesize across sources |
An LLM call in a hook, gate, router, or review is a defect unless the output is generative. A Score, a Choice, or a yes/no decision is never generative. The narrow exception is a bounded review of Jev-referred residuals: the reviewer receives a frozen, source-bound evidence bundle and returns a fixed answer, never new prose or a replacement pipeline. Use it only when a held-out benchmark shows its marginal quality lift justifies its referral rate, marginal cost, and added wall time. When you catch an LLM doing a job Jev can do, replace it.
Jev bills input tokens: the state plus the full text of every question. Output is free. Questions in one request run concurrently, so batching avoids serial call latency, but every question still consumes tokens and shared request budget. Measure actual p50/p95 wall time and accumulated model call time on the workload; do not rely on a universal latency promise. An LLM costs far more per call, takes seconds, and can rationalize a wrong answer. The toolkit metric is LLM calls per request; Jev programs exist to drive it toward zero, with benchmarked residual review as the stated exception.
Programs produce the evidence. Jev judges it. The LLM acts on those judgments creatively, receiving tier 1 and 2 findings as prior_results, not re-judging them. The only exception is the bounded residual-review stage above. When all phases of a skill are tier 1 and 2, the skill dissolves into a Jev program and no LLM runs at all.
Tier 1 goes first on every unit; tier 2 receives the residual tier 1 leaves undecided. That is what "program first" means in practice: a rule the data supports is written in code and scored before any question is written.
Name the shape of the problem first. The shape decides what code does, what Jev does, and how many requests a run costs.
| Shape | Signs | Build |
|---|---|---|
| Decide from history | labeled outcomes exist; signals are computable from data | Code builds a correlation table and writes rules for the sure units. Jev judges the residual the rules leave undecided. |
| One document, many properties | review a file, grade a draft, check a diff | One request per document: the state once, every independent, action-changing question once. Stages are code thresholds over that one answer set. A second request carries only evidence the first lacked. |
| Pick from known options | route a request, classify an error, choose a template | Code produces the candidates. A cheap wide Choice ranks them; a second Choice reranks the shortlist with full detail; a confidence gate decides act, confirm, or hand off. |
| Many items, same question | rank comments, filter tool results, triage files | Code decides the obvious ends. The middle goes in one request as short per-item Nouls. Code counts and sums. Past about 50 items, or when one run would spend more than about 50,000 tokens, build a cascade instead: stage 1 asks one short fit Noul per item over compact state; stage 2 asks the full question set for the top survivors only. Do not fan the full question set out over every item in parallel. |
| Event stream | something to check on every tool call, reply, or commit | Build it as an on-demand command. Promote it to a hook after the four conditions in step 11. |
| Select, then copy | extract a value, pick a source span, recover structure | Code finds the candidate values or spans. Jev selects the intended one. Code copies or normalizes it. No text is generated. |
| New text needed | write, rewrite, plan, diagnose | First check whether "Select, then copy" fits. When it does not, an LLM writes. Jev grades the result against a rubric that has its own labeled set. |
Build one Jev system at a time. A system is finished when it has labeled cases, a measured score, a measured cost per run, and an action that uses the answer. Start the next system after that.
Evidence of value is a labeled run. Unit tests with fake Jev answers show that the code runs; a labeled run shows that the grading is right.
| Step | Tier | Do | Exit gate |
|---|---|---|---|
| 1. State the decision | - | Write one sentence: the decision, the unit it applies to (a match, a diff hunk, a prompt), and the action code takes on each answer. | A person can label one unit by hand in under a minute. |
| 2. Build the grader | 1 | Collect human-confirmed labeled units (x, y) with label provenance. Freeze fixture copies and hash both fixtures and rubric. Split train/dev from an untouched group-disjoint heldout. Provisional or agent labels are diagnostics, never action-promoting ground truth. | score(predictions) runs on dev and prints the majority-class baseline; the heldout is sealed. |
| 3. Discover signals | 1 | Run SQL or Python over train. For every computable signal, record accuracy against y, count, and the same per slice. Start from existing analytics code. Keep every signal; the table decides. | A correlation table sorted by accuracy, with counts. |
| 4. Write the policy | 1 | Turn the table into rules: rule(x) -> (action, sure). The strongest signal decides; a near-certain signal overrides. Score the rules on dev. | The rules and their dev score are row one of the run log. The residual (every unit where sure is false) is counted. |
| 5. Design the request | 1+2 | Build state for residual units only: correlated signals, bounded, labeled, arithmetic done in code, plus the rules' verdict and why it was unsure. Write one atomic question per judgment, worded from the table. Match the primitive to the action. Put every independent question about one state in one request. A question whose evidence/options depend on another answer is a second request after code builds the new state. | The decision card is filled in (references/decision-card.md). |
| 6. Price the run | 1 | Run the program on a ten-word input: the billed tokens are the fixed floor, your question text. Compute calls per run = units x calls per unit x rounds, tokens per call, tokens per run, referrals per run, and worst-case retry sends. Then price it per second: dump every request one run sends to JSON and run python3 scripts/jev-budget-check.py --payload run.json --concurrency C --concurrent-runs N --attempts A. Pick the request size with python3 scripts/jev-size-probe.py --payload run.json on the production transport. Price any eval the same way with --eval-cases. State all numbers. | The numbers are ones you would approve and the budget check says ok: peak tokens per second and requests per minute stay under 25% of the documented limits with retries and concurrent users counted. When the floor exceeds the typical state, shorten the questions first. When tokens per run exceed about 50,000, redesign as a cascade before tuning anything else. |
| 7. Smoke run | 2 | Run the three-unit set, then the dev sample. | calls_failed is zero, every answer parses, and python3 scripts/jev-cost-report.py --since 1h matches the step 6 estimate. |
| 8. Score | 1 | On the same dev set, report the rules alone, Jev on the residual, and the combined system, per slice, with Brier and a calibration curve. Count false positives beside recall. Run judge variance once over frozen rows. | The combined score and its cost per run are in the run log. |
| 9. Improve | 1+2 | First separate code errors and service failures (HTTP errors, timeouts) from wrong answers, by reading the exact state, questions, candidates, and answers of each miss. Then classify the wrong answers (state_lacked_evidence, criteria_ambiguous, wrong_primitive, label_noise). Change one state, instruction, criterion, or policy lever at a time. State changes must add needed decision evidence, not decorative context. Re-score. Keep the change when the combined score climbs and every slice holds. | Each variant is logged with score and cost. |
| 10. Report | 1 | Score the untouched heldout once after selection. Report p50/p95 wall time, accumulated model call time, throughput, and (when used) referral rate, marginal reviewer lift, cost, and time. Before later tuning, create a new independent heldout. | One heldout number and workload metrics, reported beside the dev number. |
| 11. Integrate | 1+2 | Ship an on-demand command with a reader, storage, and an action. Log whether each answer changed the action. Promote to a hook when four conditions hold: code decides the obvious cases first; the labeled set shows the answers are right; the answer distribution is meaningfully non-constant; something acts on the answer. Run a new hook in shadow mode first, and promote one hook at a time. | A day of use shows the cost report and the action-changed rate you expected. |
| 12. Next system | - | Start step 1 for the next decision. | - |
Step 9 levers, in search order: evidence in state; decomposition (one Score into several Nouls); criteria wording; thresholds; few-shot examples in state. Evidence comes first because the other levers work only on a signal that is present. Retune thresholds from stored probabilities, which costs zero calls. Derive a gate threshold from action costs, t = C_FP / (C_FP + C_FN), select it on one split, and report on another.
Cost model. Cost = calls x input tokens per call. Input tokens = state + the text of every question, with its criteria and examples. Output is free. Parallel questions reduce wall time relative to serial sends but do not make question text free. Fill a request with independent heads only while each has decision/action value and the total fits its 64,000-token budget; sequence only a head whose evidence or candidates are derived from a prior answer. Measure wall time separately from accumulated model call time (the sum of attempt durations): concurrency can lower the former while leaving the latter high. Throughput is units or KB divided by the chosen clock; label the clock. Three numbers govern a run:
| Number | Target | Reach it by |
|---|---|---|
| Sends per state | one per run | one request per unit; stages as code thresholds; a second request only for new evidence |
| Fixed floor per call | below the typical state size | one- or two-line questions; what, not_for, and examples only where labeled misses call for them |
| Firing rate | matches how often the answer changes an action | on-demand commands first; hooks after step 11's conditions |
| Peak tokens per second | under 25% of the documented 250,000 (about 60,000), with retries and concurrent users counted | fewer tokens per run (a cascade instead of full detail for every unit); an instance-wide in-flight cap sized from the budget so simultaneous runs cannot burst together; jittered backoff for 429/529 |
| Tokens per answer, retries included | request size near the minimum of size / success_rate(size) on the production transport | measure the transient failure rate at several request sizes (references/improve-and-calibrate.md), then pack requests to that size |
Per-request fit is not enough. Every request can sit far under 64,000 tokens while one run still spends the whole per-second limit: 20 requests of 15,000 tokens sent together is 300,000 tokens in about a second. Retries then multiply it, because every failed request resends its full state. A design that works for one test query fails for real users, and an eval of 80 such runs spends millions of tokens in minutes. Price tokens per run and per second in step 6, not only tokens per request.
Measure repeatability before iterating. Run repeated frozen requests and measure answer variance and decision flips on the workload. Keep thresholds away from where answers cluster, and establish this noise floor before comparing variants. Treat cache behavior and circuit-breaker behavior as implementation details to verify in the current runner rather than performance guarantees.
Telemetry is part of the system. call_jev logs every call: script name, session id, input tokens, question count, payload hash, cached or not, error. An evidence-pipeline runner also persists the evidence-bundle ID/version, prompt/question version, model/version, attempt number, retry reason, deadline/cap, response, accounting, and final keep/refer/action outcome. Preserve source rows and provenance through joins: a relationship label is not permission to merge identities. Read the cost report after every multi-call run and compare it with the step 6 estimate.
Graders see only what you send. Send the richest available output and the evidence itself: command output, file content, stored answers. Tune on one label set and report on another.
Action-changing gates stay conservative. An unavailable, missing, or invalid Jev answer is unknown: exclude it from quality scores, retain its failure receipt, and never reinterpret it as no, pass, or permission to act. Keep any action-changing selector in shadow mode until human-confirmed, disjoint-heldout results show that its action improves the intended outcome. Confidence measures concentration, not authority: it cannot authorize an action or override source evidence, permissions, or deterministic safety rules. An evidence question needs a supplied source-evidence ledger; plausibility and apparent intent are not source evidence.
Sanity floors (majority class, the single strongest signal) prove the pipeline is wired. The bar is higher: the combined system climbs across iterations, calibration holds on the residual, and the test set agrees once. Spend scales with the residual, so a good policy keeps each round to hundreds of calls.
The grader decides how far the procedure goes. With outcomes that already happened (a result, a merged PR, a finished run), the loop runs unattended. With hand labels, it runs until the labels are used up; then the next step is more labels.
The same procedure replaces a skill: the skill's phases supply the signals and questions, hand labels are the grader, and the policy function replaces its gates (see "Dissolving a skill"). Systems compose: one system's decision is another's signal. Deterministic driver: scripts/jev-harness.py (loop, variance, sweep).
| Primitive | Ask when | Returns | Code acts with |
|---|---|---|---|
| Noul | clean yes/no; the probability is the signal | noul in [0, 1]; no confidence | if noul > t |
| Choice | one of a known unordered set | choice, probabilities, confidence | a branch per option |
| Score | a position on a spectrum you can describe in steps | score, legend, probabilities, confidence | threshold, rank, or round |
score is the probability-weighted mean of level numbers (0-based), not a picked level. A 1.0 can be certainty on level 1 or a 0/2 split; read probabilities when the distinction matters. Threshold, rank, or round it; never interpolate a quantity from it. A Noul at 0.5 means unsure, not "medium"; distance from 0.5 is its confidence. Do not carry a threshold tuned on one primitive to another, and do not expect P(noul) and 1 - P(not noul) to agree. Full shapes: references/primitives.md.
confidence on a Choice or Score measures how concentrated the distribution is. It does not say the workflow is right, and it is not permission to act. Several acceptable options also spread probability, so low confidence on a harmless preference choice is fine. Set thresholds from your labeled data and the cost of each action.
instructions.`ticket.messages[0].text`.true side describes "no" degrades.what, not_for, examples per Choice option; true/false with what and examples for a subtle Noul.other or none when the list may not cover the input. Examples are concrete instances ("charged twice"), not descriptions of instances.question, focus, inspect, note, compare, field). Pass schemas and rows as JSON, never serialized into a string.what, not_for, and examples only where labeled misses show the question needs them.instructions and criteria.[Request], [Diff], [Prior Assessment]).skills/shared-patterns/untrusted-content-handling.md, name in criteria what counts, and run adversarial and self-describing test cases before deployment.| Pattern | Shape | Worked example |
|---|---|---|
| Speculative fan-out | every branch's questions in one call, each stating its own premise ("if this is a refund request, ..."); heads are independent and cannot see one another's answers; code ignores unused heads and their uncertainty | scripts/jev-browser-decide.py |
| Confidence-gated routing | a floor below which nothing acts and a higher bar for high-stakes actions; paths act / confirm / hand off. Select both thresholds from labeled data and action costs | scripts/jev-route.py |
| Composite scoring | one Score per dimension, normalize by len(criteria) - 1, weights in code. Weighted sums suit preferences that offset one another; an "any serious violation" rule needs its own Noul per condition | references/composition-patterns.md |
| Intent routing | Choice for intent plus complexity Score, both confidence-gated | scripts/jev-route.py |
| Taxonomy walk | one Choice per level; each option's criteria is its trimmed subtree; follow several branches when close | references/composition-patterns.md |
| Multi-Noul decomposition | split a compound goal into one Noul per clause; combine in code | references/composition-patterns.md |
| Cascade plus verification | one wide request per unit; code thresholds pick survivors. Send a second request when the first answer is needed to fetch evidence, build new state, or decide the next options; it carries only what the first lacked | references/composition-patterns.md |
| Bounded residual review | Jev handles most units, a fixed-answer reviewer checks benchmarked referrals; runner sends the same source-bound bundle with immutable provenance; code accepts only the declared answer schema | references/composition-patterns.md |
| Deterministic pre-filter | programs decide the obvious ends; Jev judges the middle | scripts/jev-compact.py |
| History injection | recent actions as "already taken, do not repeat" | scripts/jev-browser-agent.py |
| Checkpoint search | caller sets subgoals a few steps apart; Jev beam-searches between them with progress-comparison Choices; LLM only on a near-tie, missing answer, or stalled subgoal; a small ledger replaces raw history | scripts/jev_search.py |
One screen each, with the code shape: references/composition-patterns.md.
Start every program as an on-demand command. Promote it to a hook after it meets the four conditions in step 11 of the build procedure. Promote one hook at a time and read the cost report after a day of use.
Every integration has a reader (runs Jev), storage (findings persist somewhere read), and an action (something changes behavior). Missing any part wastes the call. Thread prior assessments into later calls as bounded, labeled evidence. Fail open for advisory checks; fail to warn for safety checks; never fail to block when Jev is unavailable. Validate every response with validate_jev_response before acting. Details: references/integration-lifecycle.md.
This is step 9 of the build procedure. Find the failing question on labeled data before changing anything.
Promote lessons deliberately. Experimental Jevmaxxing is hypothesis discovery, not guidance. Add a durable rule only when it has a clear mechanism, representative labeled evidence, a stated boundary or counterexample, and a measured improvement to an action, cost, or quality decision against a baseline. Otherwise leave the observation out; prune copied lore that cannot meet this standard.
| Symptom | Likely cause | Fix |
|---|---|---|
| Wrong with high confidence | literal reading | state the exact condition; put the boundary case in criteria |
| Low-confidence Choice | options overlap or none fits | add what, not_for, examples; add other |
| Low-confidence Score | levels overlap, two dimensions, thin state | rewrite levels as situations; split; add the missing field |
| Scores cluster mid-scale | levels are degrees or numerals | describe a situation per level; drop numerals |
| Extremes look alike | no level for the extreme | add one |
| Noul near 0.5 | vague condition | define it; add true/false examples |
| Accuracy falls with input size | irrelevant state | filter in code; send fields, not blobs |
| Count, sum, date errors | Jev doing arithmetic | move it to code; per-item Nouls |
| Nested or negated questions fail | indirection | ask directly; split into two literal questions |
| Answer follows text in state | state steering | separate trusted policy from untrusted text; require source evidence in code; adversarial tests; thresholds may abstain or refer, not authorize |
| Rewording trades one error for another | one question, several properties | split into atomic questions |
| Answers right, decision wrong | policy | change weights or thresholds in code, not questions |
| Slow or costly | sequential calls | merge into one request |
| One query works; real use or an eval fails with 429/529/503 | one run spends too much of the per-second limit; retries multiply it | price with jev-budget-check.py; cascade; cap concurrency; jittered backoff |
Operational rules:
t = C_FP/(C_FP+C_FN), select on one split, report on another, re-measure when the data shifts.Rules: change one or two questions per revision; judge on labeled data, not confidence alone; keep the answer space stable once code depends on it; general rules in criteria, specific names only in examples. Full table and the known-debt note: references/improve-and-calibrate.md.
A dissolution is the build procedure with the skill as the request. The method:
prior_results and does not re-judge.policy(assessment) -> action with named thresholds is the dissolved skill's contract. It replaces the skill's gates.Worked example. references/dissolving-a-skill.md walks one skill through the method: phase table, Jev question set, and policy function.
scripts/jev-budget-check.py says ok for one run at the planned concurrency, attempts, and concurrent users, and any eval was priced with --eval-cases before it ran.Retry-After; a lone fast Gateway 503 retries after about 100 ms. Every run has a retry budget.other.criteria list (2–10 level descriptions), never min/max. Choice uses a criteria map with what/not_for/examples.calls_failed is zero on a labeled run.score is read as a weighted mean; code that rounds says so.unknown with separate failure receipts, never scored as pass or no.--help exits 0, and a live call with representative input returns valid JSON with source != "error".Error: HTTP 422 on the request
criteria as a map; Noul uses criteria.true/criteria.false; Score uses a criteria list. Keys such as options, min, max are not part of the API.references/primitives.md; validate against the live API, not a mocked test.Error: max_tokens_exceeded, or a request that fails because it carries too much context
jev_transport.evaluate splits oversized requests into as many small ones as it takes, each with the same state; in other code, do the same with jev_limits.split_and_run. Shrink the state when it alone passes the target. Fit state in stages, cap items per call, and count failures per stage. See references/state-and-budget.md and fact 2 in the production rules.Error: HTTP 429 or 529 (direct), or 503 through Vercel AI Gateway
GatewayInternalServerError. The most common cause in our programs is the program itself: one run, or an eval of many runs, sending more tokens per second than the account allows, or requests large enough that transient failures are frequent (through the Gateway, the 503 rate grew with tokens per request).scripts/jev-budget-check.py. If it is over 25% of a limit, fix the design (cascade, fewer tokens per run, a concurrency cap), not the retry loop. Then retry with exponential backoff and equal jitter (jev_limits.backoff_delay: base 0.5 s, doubling, capped, half random) so parallel requests that failed together do not retry together, and treat Retry-After as a floor. Cap attempts per request and retries per run; when the budget is spent, stop sending and return what finished. Persist every failed and retried attempt with its reason; retries count in workload cost and latency. Short fixed delays (100 ms, 200 ms) across many parallel requests make a retry storm that keeps the limit tripped.references/improve-and-calibrate.md.Error: HTTP 401 or 402
call_jev stops further sends. Code outside call_jev (a sandboxed plugin) stops after the first such status. Keep the API key on the server side; never ship it to a browser.Error: valid response, wrong decision
score as a level index or thresholds on the wrong primitive.probabilities; keep thresholds in the policy; see the known-debt note in references/improve-and-calibrate.md.© notque, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (references) in skills/meta/building-with-jev of notque/vexjoy-agent.
Open the folder on GitHubat commit 5218674
Building With Jev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Building With Jev this skillnotque/vexjoy-agent | 439 | — | ~11k | Automated safety check: Notes | MIT | |
| Integration Testingthedaviddias/Front-End-Checklist | 74k | — | ~514 | Automated safety check: Pass | MIT | |
| Frontend API Integration Patternssickn33/agentic-awesome-skills | 47k | 1 repos | ~1.9k | Automated safety check: Pass | MIT | |
| API Integrationsickn33/agentic-awesome-skills | 47k | 1 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Labarchive Integrationdavila7/claude-code-templates | 32k | 10 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Robius Matrix Integrationsickn33/agentic-awesome-skills | 47k | 2 repos | ~3.6k | Automated safety check: Pass | MIT |
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing CI coverage, automated checks, or test strategy related to Write integration tests for key workflows.
sickn33/agentic-awesome-skills
Production-ready patterns for integrating frontend applications with backend APIs, including race condition handling, request cancellation, retry strategies, error normalization, and UI state…
sickn33/agentic-awesome-skills
Designs event-driven architectures, webhook systems, API chaining flows, ETL pipelines, and integration patterns between services.
davila7/claude-code-templates
Electronic lab notebook API integration. An agent skill from davila7/claude-code-templates.
sickn33/agentic-awesome-skills
CRITICAL: Use for Matrix SDK integration with Makepad. An agent skill from sickn33/agentic-awesome-skills.
sickn33/agentic-awesome-skills
Design, implement, debug, and optimize API integrations with expert-level patterns for REST, GraphQL, webhooks, and authentication flows.
notque/vexjoy-agent
Deterministic palette/matrix pixel art (not AI). An agent skill from notque/vexjoy-agent.
notque/vexjoy-agent
Pull request lifecycle: commit, codex review, sync, review, fix, status, cleanup, and PR mining.
notque/vexjoy-agent
Improve architecture across modules by deepening interfaces.
notque/vexjoy-agent
Code quality: cleanup, linting, formatting, quality gates. An agent skill from notque/vexjoy-agent.
notque/vexjoy-agent
Statistical rule discovery from Go codebase patterns. An agent skill from notque/vexjoy-agent.
notque/vexjoy-agent
Review and fix temporal references in code comments. An agent skill from notque/vexjoy-agent.
Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model. Building With Jev is an agent skill from notque/vexjoy-agent. Write, compose, integrate, and improve programs that call Jev, TypeSafe's System One judgment model.
Run `npx skills add notque/vexjoy-agent --skill building-with-jev -a claude-code`. Or copy the skill folder (skills/meta/building-with-jev in notque/vexjoy-agent) into .claude/skills/building-with-jev in your project. Claude Code loads it when a task matches its description.
Run `npx skills add notque/vexjoy-agent --skill building-with-jev -a codex`. Or copy the skill folder (skills/meta/building-with-jev in notque/vexjoy-agent) into .agents/skills/building-with-jev in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add notque/vexjoy-agent --skill building-with-jev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/building-with-jev, .gemini/skills/building-with-jev, .github/skills/building-with-jev and .opencode/skills/building-with-jev in your project.
Going by SKILL.md and its folder, Building With Jev needs the command-line tools its instructions call (python3) and credentials named TYPESAFE_API_KEY and AI_GATEWAY_API_KEY. Our summary lists: A credential in TYPESAFE_API_KEY; A credential in AI_GATEWAY_API_KEY. Its frontmatter pre-approves these tools: Read, Edit, Write, Bash, Glob, Grep.
SKILL.md names 2 domains. In commands or code: api.typesafe.ai; the agent is likely to contact it when it follows the instructions. As links in the text: docs.typesafe.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Building With Jev is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 11k tokens (SKILL.md is roughly 43k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Building With Jev: Integration Testing (thedaviddias/Front-End-Checklist, 74k stars), Frontend API Integration Patterns (sickn33/agentic-awesome-skills, 47k stars), API Integration (sickn33/agentic-awesome-skills, 47k stars) and Labarchive Integration (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
notque (a GitHub user) maintains it in notque/vexjoy-agent, which has 439 GitHub stars. The repository holds 61 skills in this directory. The repository was last updated on October 3, 2026.
Source: notque/vexjoy-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.