Agent skill

Building With Typesafe Jev

by aaddrick in aaddrick/building-with-typesafe-jev

A skill your agent uses when code needs a judgment about natural language that rules or regex cannot make (classify, route, triage, moderate, score, rank, match, dedupe, filter, extract, verify, or…

MITAuto-check: notesAI & LLM Engineering

Install Building With Typesafe Jev

skills CLI
$ npx skills add aaddrick/building-with-typesafe-jev --skill building-with-typesafe-jev -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aaddrick/building-with-typesafe-jev building-with-typesafe-jev --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aaddrick/building-with-typesafe-jev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/building-with-typesafe-jev .claude/skills/building-with-typesafe-jev && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
building-with-typesafe-jev
GitHub stars
136
Token cost
~3k tokens
SKILL.md length
1,826 words
Files
16
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when code needs a judgment about natural language that rules or regex cannot make (classify, route, triage, moderate, score, rank, match, dedupe, filter, extract, verify, or…

  • Works in 11 steps: One snap judgment per question. Pose… → Put every question for one state in one… → A second request only when the first… → …
  • Code needs a judgment about natural language that rules
  • SKILL.md covers Overview, Test against the live API when…, Pick the primitive and Rules, plus 3 more sections
  • Reaches docs.typesafe.ai; needs TYPESAFE_API_KEY

What it does

Building With Typesafe Jev is an agent skill from aaddrick/building-with-typesafe-jev. Use when code needs a judgment about natural language that rules or regex cannot make (classify, route, triage, moderate, score, rank, match, dedupe, filter, extract, verify, or gate an action), when brainstorming where AI could fit in an app, when an LLM call exists only to pick a label, score, or yes/no, or when code uses TypeSafe AI, Jev, System One, typesafe-sdk, or @typesafe-ai/sdk.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files (for example `agents/openai.yaml`, `api-reference.md` and `patterns.md`).

It sits in AI & LLM Engineering, covering Brainstorming. It works with Vercel AI SDK. The repository describes itself as: Unofficial skill that teaches coding agents to build with TypeSafe AI's Jev: typed decisions, calibrated confidence, and prior art from 150+ community projects. The licence is MIT.

When your agent uses it

  • Code needs a judgment about natural language that rules
  • Regex cannot make (classify
  • Gate an action)
  • Brainstorming where AI could fit in an app

Example prompts

  • “/building-with-typesafe-jev”

Requirements

  • Python 3
  • A credential in TYPESAFE_API_KEY

Workflow steps

11 steps, taken from the first numbered list in SKILL.md.

  1. One snap judgment per question. Pose something a knowledgeable person decides in a second. Split "angry AND wants refund" into two…
  2. Put every question for one state in one request, including speculative ones. They run in parallel and in isolation. Extra questions add…
  3. A second request only when the first answer is needed to build it: to fetch more state, to build new state, or to pick the next options…
  4. Question IDs are never sent to the model. Write the full question in instructions.
  5. Write the exact condition. Jev reads literally: it answers the question you wrote, not the one you meant. Scope words, negations, and…
  6. Structure state as a named JSON object. Point at parts with backticked paths: "Does ticket.messages[0].text request a refund?". Send only…
  7. Score levels describe situations, not degrees. Write "Broken feature, workaround exists", not "moderate" or "2". Each level is judged…
  8. Choice: give the full option list plus other / none_of_the_above. When two options get confused, make each value an object: {"what"…
  9. Phrase a Noul so that high = yes. Keep criteria aligned with the instruction. Never map true to a "no" meaning.
  10. Keep math, counting, dates, and exact lookups in code. To count, ask one Noul per candidate and sum. For dates, extract the parts with…
  11. Put questions, thresholds, and weights in one constants module. Those are what humans review and tune.

What it can do on your machine

Read from SKILL.md and the folder at commit a46856c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • docs.typesafe.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • TYPESAFE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Building With Typesafe Jev loads about 3k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 1,826 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:20
    u did. Do not go looking for the key in `.env` files, shell profiles, or key files. Point the user to the README section

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aaddrick/building-with-typesafe-jev at commit a46856c, republished under its MIT licence (© aaddrick). 1,826 words, ~2,983 tokens.

Download SKILL.mdSave it as .claude/skills/building-with-typesafe-jev/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
building-with-typesafe-jev
description
Use when code needs a judgment about natural language that rules or regex cannot make (classify, route, triage, moderate, score, rank, match, dedupe, filter, extract, verify, or gate an action), when brainstorming where AI could fit in an app, when an LLM call exists only to pick a label, score, or yes/no, or when code uses TypeSafe AI, Jev, System One, typesafe-sdk, or @typesafe-ai/sdk.

Building with TypeSafe Jev

Overview

Jev is a System One model: it takes a state plus named, typed questions and returns typed answers with calibrated probabilities. It never generates text. Code owns the workflow. Jev supplies narrow snap judgments. It is not a chat or coding LLM and cannot power a coding agent.

Before you design anything new, check prior art: open prior-art/INDEX.md. It maps intents (control loop, gate, rerank, stream filter, incremental, agent memory, LLM pairing, and more) to shape files. Each shape file has a code sketch, field lessons, and linked community projects. Or grep prior-art/ for a domain word. It also lists the known bad fits.

Exact request/response shapes, SDK signatures, limits, and errors: see api-reference.md. It ends with a map of which live docs page to read for which task. Patterns and the cookbook index: see patterns.md. The live docs win on conflict: fetch https://docs.typesafe.ai/llms.txt, and append .md to any page path. Do not guess field names. Check them there. If the docs cannot be reached, read the installed SDK's types, tell the user you did, and do not invent details that depend on the version.

Open-ended request ("where could AI help in this app?"): work backward from what the app should show, select, change, or hand off. Offer two or three directions from prior-art/INDEX.md and recommend one. Concrete request: pick the shape and build. Either way, keep the user's stack and scope, and add Jev only where code needs a judgment.

Test against the live API when you can

You get better results when you check a design against the real endpoint. A live call catches a wrong field name, and it shows when Jev reads a question differently than you meant. If TYPESAFE_API_KEY is set in your shell, run a small probe before the code reaches the project. One call costs a fraction of a cent (input tokens only, 4.2 cents per million). If the key is not set, design from api-reference.md and say that you did. Do not go looking for the key in .env files, shell profiles, or key files. Point the user to the README section "If the agent cannot see the key".

Rules for the key:

  • Do not print, echo, log, or commit it.
  • Do not hardcode it or pass api_key= in code. Let the SDK read TYPESAFE_API_KEY. Keep it server-side.
  • TYPESAFE_LOG_LEVEL=debug logs request bodies without redaction. Headers stay redacted.
  • A 401 or TypeSafeAuthenticationError means the key is wrong or rotated. Ask the user to replace it where they stored it, never in the chat.
  • Keep a task to about 10 test calls, and probe loops to about 8 workers or fewer. Rate limits are shared across the account.

Pick the primitive

Answer shapePrimitiveReturnsBranch on
One of an unordered setChoice (≤255 options, map)choice, probabilities, confidencechoice; gate with confidence
Position on a spectrum you can describeScore (2–10 levels, ordered list)score (0…n-1, can be fractional), legend, probabilities, confidencethreshold or sort score
Yes/noNoul (optional true/false criteria)noul = P(yes), no confidencenoul > threshold

A Noul of 0.5 means "unsure", never "medium". Degree questions ("how strong in Python") need a Score. Several labels that can apply at once need one Noul per label, not a Choice.

Rules

  1. One snap judgment per question. Pose something a knowledgeable person decides in a second. Split "angry AND wants refund" into two questions and combine them in code. Atomic does not mean trivial: picking one bounded action or reading one message in context is a single judgment. Do not split apart a relationship the question is about (does this passage support that claim?).
  2. Put every question for one state in one request, including speculative ones. They run in parallel and in isolation. Extra questions add almost no latency, but each one adds input tokens, so check usage.input_tokens on a real request. Write the premise into the text: "If this is a technical problem, how severe is it?". Code ignores the answers on branches it does not take.
  3. A second request only when the first answer is needed to build it: to fetch more state, to build new state, or to pick the next options (hierarchical walk, shortlist then full text).
  4. Question IDs are never sent to the model. Write the full question in instructions.
  5. Write the exact condition. Jev reads literally: it answers the question you wrote, not the one you meant. Scope words, negations, and implied conditions count at face value. When a wrong answer makes you explain what you really meant, that explanation is the missing half of the instruction. Put boundary cases in the criteria. Avoid double negatives and property-of-a-property hops. Name the state path instead.
  6. Structure state as a named JSON object. Point at parts with backticked paths: "Does `ticket.messages[0].text` request a refund?". Send only what the question needs. Irrelevant state lowers accuracy. Include metadata (plan, timestamps) only when a question refers to it.
  7. Score levels describe situations, not degrees. Write "Broken feature, workaround exists", not "moderate" or "2". Each level is judged alone, so "worse than the previous level" means nothing. Give each level one dimension.
  8. Choice: give the full option list plus other / none_of_the_above. When two options get confused, make each value an object: {"what": ..., "not_for": ..., "examples": [...]}. Use the same keys on every option.
  9. Phrase a Noul so that high = yes. Keep criteria aligned with the instruction. Never map true to a "no" meaning.
  10. Keep math, counting, dates, and exact lookups in code. To count, ask one Noul per candidate and sum. For dates, extract the parts with Choices (include not_stated) and compare in code. For extraction, find candidates with regex or an LLM, then let Jev pick among them. To count distinct items, split into candidates in code (sentences, lines). Ask two Nouls per candidate: "is this an X?" and "is this X different from those in items[0..i-1]?". Sum the answers in code. Turn hex colors, RGB triples, and raw codes into a computed number or a named bucket before Jev sees them. The same goes for a safety condition you can write exactly, when a miss cannot be undone: a denylist of destructive commands, a spending limit, a protected branch. Check it in code before Jev runs. Jev may make that decision stricter, never looser.
  11. Put questions, thresholds, and weights in one constants module. Those are what humans review and tune.
Show full SKILL.md (755 more words)Show less

Using the numbers

  • Only picking the best option? Use choice (the argmax). A confidence threshold is not needed.
  • Acting on it? Split into three bands: act / confirm or review / escalate to a human or a reasoning LLM. Scale each threshold to what a wrong action costs. A read-only action can take a lower bar than a money-moving one.
  • Escalate only on uncertainty that feeds an action. Report uncertainty on informational outputs, and ignore it on branches the code does not take.
  • Belongs to two categories? Route to choice, and notify a second option whose probability is above about 0.25. Do not force a single label.
  • Confidence of 1.0 is normal when all the probability sits on one option. It is not a bug.
  • Composite judgment? Normalize each Score with score / (len(criteria) - 1), weight the results in code, and keep the raw answers.
  • A fractional score works for thresholds and sorting only. Do not read an exact magnitude off it by interpolating between two levels.
  • A Choice always has a winner, even when nothing fits. Add an absolute Noul (exists, stated, fits) in the same request to decide whether to act at all.
  • Combine parts with the minimum (a date from its parts, a call from its arguments). Combine "something is wrong" flags with the maximum, not the average.
  • Confidence is not the top probability. For a Choice it is (n·top − 1)/(n − 1) over n options (measured: top 0.92 of 3, confidence 0.88), so it depends only on the top probability and the option count. Score confidence also counts how far the rest of the probability sits from the top level: spread onto a neighboring level costs little, and mass at the far end costs a lot. See api-reference.md. Neither is proof of correctness: typed output guarantees the interface, not the truth. Tune thresholds on labeled data from the user, and plot confidence against accuracy to check them. Treat cookbook numbers as examples only.
  • Every number comes back rounded to 0.01. Nouls, probabilities, and confidence tie often. Give any sort on them a tiebreaker.
  • Do not carry a threshold from a Noul to a Choice. Do not expect P(q) + P(not q) = 1 across separate questions.
  • Pin jev-1.13.0 when thresholds were tuned against it. Log response.model.

When an answer is wrong

Look at the exact state, questions, candidates, answers, and the code that combined them, next to what actually happened. Then name the cause before you change anything:

  • Missing evidence: the state lacked what the question needed, or a candidate was never offered.
  • Model error: the state held the answer. Rewrite the question (rule 5) and re-test on the same case, or route cases like it to review.
  • Code error: wrong path, inverted Noul, bad threshold, or wrong combination.
  • Service failure: a timeout, a 429, or a 5xx. Handle it with the SDK's RetryPolicy, not by rewording the question.

Known weak spots in jev-1.13: https://docs.typesafe.ai/model-jaggedness/jev-1.13.md.

Common mistakes

MistakeFix
One API call per questionPut all questions in one system_one call
Treating the top probability as confidenceUse .confidence. It is computed for you and is never higher than the top probability.
Guessing the score scaleIt is 0 to n-1, weighted by level probability. Python keys legend/probabilities by int; HTTP keys them by string.
Score.criteria as a dict keyed by intUse an ordered list (SDK ≥ 0.6.0)
Numeric or vague Score levelsUse concrete situations, one dimension each
"Analyze this and decide what to do"Split into atomic questions and combine in code
Asking Jev to count, do date math, or compare hex colorsDo it in code. Ask per-item Nouls.
Double negatives, or "does the thing it refers to have..."Ask directly about a named state path
Speculative question with no stated premise"If X, then ...?"
Hand-rolled 429 retry loopThe SDKs retry 429/5xx with backoff. Tune with RetryPolicy.
16+ worker thread pool on one keyUse about 8 workers or fewer. Above that, the endpoint rate-limits.
Holistic "is this record good?" judgeUse per-field checks, one flaw per Noul
API key in browser codeKeep it server-side. JS needs dangerouslyAllowBrowser for a reason.
Jev as the only check before a destructive or money-moving actionPut a code rule first: a known-bad pattern blocks whatever Jev says. Jev judges only what the rule does not decide.
Hostile text in stateJev does not treat state as adversarial. Keep untrusted text in its own named field. Point questions at it by path. Add an injection Noul. Never let a Jev answer alone authorize a side effect.

© aaddrick, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files in skills/building-with-typesafe-jev of aaddrick/building-with-typesafe-jev.

  • SKILL.md
  • agents/openai.yaml
  • api-reference.md
  • patterns.md
  • prior-art/INDEX.md
  • prior-art/agent-context-memory.md
  • prior-art/control-loops.md
  • prior-art/embedding-in-infrastructure.md
  • prior-art/gates.md
  • prior-art/incremental-realtime.md
  • prior-art/judges-and-evals.md
  • prior-art/llm-pairing.md
  • prior-art/ranking-and-matching.md
  • prior-art/research-and-features.md
  • prior-art/select-from-candidates.md
  • prior-art/stream-filters.md

Open the folder on GitHubat commit a46856c

Compare with similar skills

Building With Typesafe Jev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Building With Typesafe Jev compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Building With Typesafe Jev this skillaaddrick/building-with-typesafe-jev136—~3kAutomated safety check: NotesMIT
Neurolink Guidejuspay/neurolink144—~1.4kAutomated safety check: PassMIT
Repo Conventionsjuspay/neurolink144—~1.3kAutomated safety check: PassMIT
AI SDKvercel-labs/ai-facts16820 repos~1.2kAutomated safety check: PassNone
Mem0 Provider for Vercel AI SDKmem0ai/mem067k—~2.3kAutomated safety check: PassApache-2.0
Yao Meta Skillyaojingang/yao-meta-skill2.7k—~768Automated safety check: PassMIT

Similar skills

  • Neurolink Guide

    juspay/neurolink

    Guide for using the NeuroLink SDK and CLI. An agent skill from juspay/neurolink.

    144 GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Repo Conventions

    juspay/neurolink

    NeuroLink's review standards — the critical rules to enforce, what NOT to comment on, the security bar, hot paths.

    144 GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • AI SDK

    vercel-labs/ai-facts

    Official

    Answer questions about the AI SDK and help build AI-powered features.

    168 GitHub starsUsed in 20 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Adds persistent memory to Vercel AI SDK apps with the Mem0 provider, using a wrapped model or standalone retrieve and store utilities.

    67k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Yao Meta Skill

    yaojingang/yao-meta-skill

    Create, improve, or evaluate an existing skill from workflows, prompts, SOPs, scripts.

    2.7k GitHub stars~768 tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

Works with

Questions about Building With Typesafe Jev

What does Building With Typesafe Jev do?

A skill your agent uses when code needs a judgment about natural language that rules or regex cannot make (classify, route, triage, moderate, score, rank, match, dedupe, filter, extract, verify, or…. Building With Typesafe Jev is an agent skill from aaddrick/building-with-typesafe-jev. Use when code needs a judgment about natural language that rules or regex cannot make (classify, route, triage, moderate, score, rank, match, dedupe, filter, extract, verify, or gate an action), when brainstorming where AI could fit in an app, when an LLM call exists only to pick a label, score, or yes/no, or when code uses TypeSafe AI, Jev, System One, typesafe-sdk, or @typesafe-ai/sdk.

When should I use Building With Typesafe Jev?

Building With Typesafe Jev fits situations like: code needs a judgment about natural language that rules; regex cannot make (classify; gate an action); brainstorming where AI could fit in an app.

How do I install Building With Typesafe Jev in Claude Code?

Run `npx skills add aaddrick/building-with-typesafe-jev --skill building-with-typesafe-jev -a claude-code`. Or copy the skill folder (skills/building-with-typesafe-jev in aaddrick/building-with-typesafe-jev) into .claude/skills/building-with-typesafe-jev in your project. Claude Code loads it when a task matches its description.

How do I install Building With Typesafe Jev in Codex?

Run `npx skills add aaddrick/building-with-typesafe-jev --skill building-with-typesafe-jev -a codex`. Or copy the skill folder (skills/building-with-typesafe-jev in aaddrick/building-with-typesafe-jev) into .agents/skills/building-with-typesafe-jev in your project. Codex loads it when a task matches its description.

Can I use Building With Typesafe Jev in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aaddrick/building-with-typesafe-jev --skill building-with-typesafe-jev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/building-with-typesafe-jev, .gemini/skills/building-with-typesafe-jev, .github/skills/building-with-typesafe-jev and .opencode/skills/building-with-typesafe-jev in your project.

What does Building With Typesafe Jev need to run?

Going by SKILL.md and its folder, Building With Typesafe Jev needs credentials named TYPESAFE_API_KEY. Our summary lists: Python 3; A credential in TYPESAFE_API_KEY.

Does Building With Typesafe Jev access the network?

SKILL.md names 1 domain. In commands or code: docs.typesafe.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Building With Typesafe Jev safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Building With Typesafe Jev use?

Building With Typesafe Jev is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Building With Typesafe Jev use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Building With Typesafe Jev?

Skills that share tags, products or a category with Building With Typesafe Jev: Neurolink Guide (juspay/neurolink, 144 stars), Repo Conventions (juspay/neurolink, 144 stars), AI SDK (vercel-labs/ai-facts, 168 stars) and Mem0 Provider for Vercel AI SDK (mem0ai/mem0, 67k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Building With Typesafe Jev?

aaddrick (a GitHub user) maintains it in aaddrick/building-with-typesafe-jev, which has 136 GitHub stars. The repository was last updated on October 1, 2026.

Source: aaddrick/building-with-typesafe-jev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.