Official agent skill

LangGraph Decision Models

by langchain-ai in langchain-ai/langchain-skills

Routes LangGraph agents with typed decision models that return probabilities, and finds LLM calls that only exist to produce a routing decision.

OfficialMITAuto-check passedAI & LLM Engineering

Install LangGraph Decision Models

skills CLI
$ npx skills add langchain-ai/langchain-skills --skill langgraph-decision-models -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install langchain-ai/langchain-skills langgraph-decision-models --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/langchain-ai/langchain-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/config/skills/langgraph-decision-models .claude/skills/langgraph-decision-models && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langgraph-decision-models
GitHub stars
1.3k
Token cost
~2.3k tokens
SKILL.md length
830 words
Files
2 (incl. references)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Routes LangGraph agents with typed decision models that return probabilities, and finds LLM calls that only exist to produce a routing decision.

  • Replacing an LLM call that only decides which graph branch to take
  • SKILL.md covers Install and wire, The conversion pattern, Three traps when reading answers and Question wording dominates…, plus 4 more sections
  • Calls uv; reaches gateway.smith.langchain.com; needs TYPESAFE_API_KEY and LANGSMITH_API_KEY
  • Auditing an agent for model calls whose output is parsed into a route

What it does

A decision model answers typed questions about state and returns probabilities instead of prose, replacing the habit of prompting an LLM, parsing its text and branching on it. `Noul` asks a yes or no question and returns the probability of yes, `Choice` picks one label with the full distribution and a confidence, and `Score` grades against an ordered rubric and returns an expected value with a confidence. `TypeSafeClassifier` is a LangChain runnable, so it drops into a graph node, and up to 32 questions in one request are answered independently.

The skill advises using one where a node generates text just to extract a decision, such as routing, triage, filtering or per-item classification, and not where the output is the product or the judgment needs multi-step reasoning. It covers installing `langchain-typesafe`, which is alpha and should be pinned, wiring a base URL and API key so providers switch through constructor arguments, the conversion pattern of one request per item with a plain Python router, how to read answers, threshold design, three traps that cause silent misrouting, and a conversion playbook reference.

When your agent uses it

  • Replacing an LLM call that only decides which graph branch to take
  • Auditing an agent for model calls whose output is parsed into a route
  • Choosing probability thresholds for a classifier node
  • Wiring a decision model through the LangSmith Gateway

Example prompts

  • “Audit our support agent for LLM calls that only produce a routing decision and convert them.”
  • “Add a Noul check for whether the ticket is urgent, and route on its probability.”
  • “Why does my Choice-based router always take the first branch? Check how I read the answers.”

Requirements

  • Python with the `langchain-typesafe` package
  • A model endpoint reachable with a base URL and API key

What it can do on your machine

Read from SKILL.md and the folder at commit 16a992f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • gateway.smith.langchain.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • TYPESAFE_API_KEY
    • LANGSMITH_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LangGraph Decision Models loads about 2.3k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 830 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from langchain-ai/langchain-skills at commit 16a992f, republished under its MIT licence (© langchain-ai). 830 words, ~2,341 tokens.

Download SKILL.mdSave it as .claude/skills/langgraph-decision-models/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
langgraph-decision-models
description
INVOKE THIS SKILL when routing a LangGraph agent with a decision model (TypeSafe Jev, SemIf) instead of an LLM, or when auditing an existing agent for LLM calls that only produce a routing decision. Covers langchain-typesafe Noul/Choice/Score, reading answers correctly, threshold design, and LangSmith Gateway wiring.
<overview>
A **decision model** answers typed questions about state and returns probabilities instead of prose. It replaces the common pattern of prompting an LLM, parsing its text, and branching on the result.
  • Noul(instructions=...) — binary question, returns a probability of yes
  • Choice(instructions=..., criteria={...}) — picks one label, returns the full distribution plus confidence
  • Score(instructions=..., criteria=[...]) — grades against an ordered rubric, returns an expected value plus confidence

TypeSafeClassifier is a LangChain Runnable[ClassifierRequest, ClassifierResponse], so it drops into a node like any other runnable. Up to 32 questions share one request and are answered independently — your code combines them.

Reach for one when a node generates text only so you can parse a decision out of it: routing, triage, filtering, guardrails, or per-item classification over a batch.

Do not reach for one when the node's output is the product (summaries, drafts, code) or when the judgment needs multi-step reasoning. A decision model classifies; it does not think. </overview>


Install and wire

langchain-typesafe is alpha (0.0.1a3) and TypeSafeClassifier is marked @beta — pin it and expect churn.

bash
uv add langchain-typesafe

Three ways to reach a model. The classifier POSTs to {base_url}/v1/systemone with Authorization: Bearer {api_key}, so switching providers is constructor arguments only:

<ex-wiring>
<python>
```python
import os
from langchain_typesafe import TypeSafeClassifier

1. TypeSafe directly (Jev). Reads TYPESAFE_API_KEY when api_key is omitted.

classifier = TypeSafeClassifier(model="jev-latest")

2. SemIf, hosted on the LangSmith Gateway. Note: LangSmith key, not a TypeSafe key.

classifier = TypeSafeClassifier( model="semif-qwen3.5-4b", api_key=os.environ["LANGSMITH_API_KEY"], base_url="https://gateway.smith.langchain.com", )

3. Jev through the Gateway (BYOK). The typesafe/ prefix routes to a

TYPESAFE_API_KEY stored in LangSmith workspace secrets. Without that secret

every typesafe/* id returns 424 Failed Dependency -- before the model name is

even validated, so a 424 does not confirm the id is real.

classifier = TypeSafeClassifier( model="typesafe/jev-1.13.0", api_key=os.environ["LANGSMITH_API_KEY"], base_url="https://gateway.smith.langchain.com", )

</python>
</ex-wiring>

---

## The conversion pattern

Ask every question about a page/item in **one** request, put the typed response in state, and let a plain function route on it. The router is ordinary Python — testable without touching a network.

<ex-classify-and-route>
<python>
```python
from typing import TypedDict
from langchain_typesafe import ClassifierResponse, Noul, Score, TypeSafeClassifier
from langgraph.graph import StateGraph, START, END

QUESTIONS = {
    "relevant": Score(
        instructions="How relevant is this ticket to a billing problem?",
        criteria=["Unrelated.", "Possibly related.", "Directly about billing."],
    ),
    "angry": Noul(instructions="Is the customer expressing anger?"),
}

class State(TypedDict):
    text: str
    answers: ClassifierResponse
    route: str

classifier = TypeSafeClassifier(model="jev-latest")

def classify(state: State) -> dict:
    # One request, every question. They are answered independently.
    return {"answers": classifier.invoke(
        {"state": state["text"], "questions": QUESTIONS}
    )}

def route(state: State) -> str:
    a = state["answers"]
    if a.nouls["angry"].noul > 0.7:
        return "escalate"
    if a.scores["relevant"].score < 0.5:
        return "close"
    return "handle"
</python>
</ex-classify-and-route>

Reading answers: response.nouls[id].noul, response.choices[id].choice, response.scores[id].score. Each view is keyed by your question id; response.answers holds them all.


Three traps when reading answers

These cause silent misrouting, not exceptions.

1. Score.score is an expected value, not a level. It is a probability-weighted average over the rubric and is routinely fractional. score == 0 almost never fires — a "not responsive" item lands at 0.07, not 0. Always compare against a band.

python
if a.scores["relevant"].score < 0.5:   # correct
if a.scores["relevant"].score == 0:    # WRONG -- nearly never true

2. Confidence measures distribution shape, not correctness. On a Score, confidence reports how concentrated the rubric distribution is. An item sitting cleanly between two levels scores low confidence even when the model is entirely clear about it. A blanket confidence < X -> escalate rule therefore escalates items the model already decided. Gate on confidence only inside the ambiguous middle:

python
if score < NOT_RELEVANT:                          # decisive -- trust it
    return "close"
if score < RELEVANT or confidence < MIN_CONF:     # ambiguous -- escalate
    return "human_review"
return "handle"

3. Thresholds do not transfer between models. Calibration is part of the model. The same policy over the same items routes differently on Jev vs SemIf vs an LLM adapter. Re-tune thresholds whenever you change models, and pin the model id.


Question wording dominates accuracy

A loose question produces confident wrong answers, and no threshold fixes it. Use criteria to say what each outcome means, including what should not count.

In a measured case, "Is this a confidential communication with a lawyer?" scored a routine finance memo at 0.798. Rewriting it to name the actual test — written by or to a lawyer, with an explicit carve-out for finance and accounting content — moved the same page to 0.005 while a genuinely privileged page held at 0.991.

python
Noul(
    instructions=(
        "Was this written by or to a lawyer, or does it convey a lawyer's legal "
        "advice? Answer no for ordinary business or accounting discussion, even "
        "when the subject is litigation-sensitive."
    ),
    criteria=NoulCriteria(
        true="A named attorney is author or recipient, or it relays legal advice.",
        false="Business or accounting content with no attorney involved.",
    ),
)

Before blaming the model, rewrite the question and re-measure.


Show full SKILL.md (290 more words)Show less

Auditing an existing agent

To find where a decision model fits, look for these in the codebase — see references/conversion-playbook.md for the full walkthrough.

SignalWhat to look for
Generate-then-parseAn LLM call whose output is immediately regex'd, json.loads'd, or string-matched into a branch
Prompted classifiersPrompts containing "respond with one of", "answer yes or no", "rate from 1 to 5"
Sampling for costComments or configs that check only the first N items because checking all is too expensive
Brittle rulesKeyword lists or regexes standing in for semantic judgment
Re-reading contextThe same document re-sent to a model for each separate question

The last two matter most: cheap semantic judgments change what you can build, not just the bill. If evaluating every item became affordable, what would you stop sampling?


Expectations

Measured on a 24-item batch, identical LangGraph graph and routing policy, only the classifier swapped:

per itemtokens (6 items)notes
Jev 1.13.00.27s3,648 in / 318 outreports usage
SemIf 4B0.49snot reportedhosted on the Gateway
Claude Sonnet 52.87s7,930 in / 864 outvia structured output

Routing agreed on 4–5 of 6 items across engines; disagreements clustered on genuinely borderline items. Treat these as shape, not benchmarks — measure on your own workload.

If you compare against an LLM baseline, use method="json_schema" so the comparison is fair. LangChain's with_structured_output defaults to method="function_calling", which injects a tool schema into every request — 556/35 tokens versus 228/12 for the native output_config.format path on the same one-field probe.


Batching with Send

Classification is per-item and independent, so fan out with Send and let each item route on its own.

<ex-fan-out>
<python>
```python
from langgraph.types import Send

def fan_out(state): return [Send("classify_item", {"text": t}) for t in state["items"]]

builder.add_conditional_edges(START, fan_out, ["classify_item"])

</python>
</ex-fan-out>

Fan-out hides latency, so it flatters slow classifiers most: in the run above, Sonnet gained 8x from concurrency and Jev only 1.4x — yet Jev still finished first. Compare throughput, not the speedup multiple.

---

## Related skills

- **langgraph-fundamentals** — StateGraph, `Send`, `Command`, conditional edges
- **langgraph-human-in-the-loop** — `interrupt()` for the escalation branch above
- **langchain-middleware** — structured output when you need an LLM, not a classifier

© langchain-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in config/skills/langgraph-decision-models of langchain-ai/langchain-skills.

  • SKILL.md
  • references/conversion-playbook.md

Open the folder on GitHubat commit 16a992f

Compare with similar skills

LangGraph Decision Models next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LangGraph Decision Models compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LangGraph Decision Models this skilllangchain-ai/langchain-skills1.3k—~2.3kAutomated safety check: PassMIT
Tool Designagentailor/fullstack-langgraph-nextjs-agent132—~3.2kAutomated safety check: PassMIT
Deepagents Setup Configurationsoba-labs/langchain-agent-skills107—~1.9kAutomated safety check: PassMIT
Langgraph Testing Evaluationsoba-labs/langchain-agent-skills107—~2.3kAutomated safety check: PassMIT
Langchain Observabilityjeremylongshore/tons-of-skills-marketplace2.8k—~3.9kAutomated safety check: NotesMIT
LangSmith Trace DebuggingComposioHQ/awesome-claude-skills77k8 repos~2.7kAutomated safety check: PassNone

Similar skills

  • Tool Design

    agentailor/fullstack-langgraph-nextjs-agent

    Design and verify tools that AI agents can actually use — for any framework or language (MCP servers, LangChain/LangGraph, function-calling, raw JSON schema; TypeScript, Python, or otherwise).

    132 GitHub stars~3.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Deepagents Setup Configuration

    soba-labs/langchain-agent-skills

    Initialize, validate, and troubleshoot Deep Agents projects in Python or JavaScript using the deepagents package.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Langgraph Testing Evaluation

    soba-labs/langchain-agent-skills

    A skill your agent uses when you need to test or evaluate LangGraph/LangChain agents: writing unit or integration tests, generating test scaffolds, mocking LLM/tool behavior, running trajectory…

    107 GitHub stars~2.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Langchain Observability

    jeremylongshore/tons-of-skills-marketplace

    Wire LangSmith tracing and custom metric callbacks into a LangChain 1.0 chain or LangGraph 1.0 agent correctly — env-var spelling, subgraph propagation, per-tenant dimensions, cost and latency…

    2.8k GitHub stars~3.9k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • LangSmith Trace Debugging

    ComposioHQ/awesome-claude-skills

    Debugs LangChain and LangGraph agents by pulling recent execution traces with the langsmith-fetch CLI and reporting errors, tool calls, timings and token use.

    77k GitHub starsUsed in 8 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Example Agent

    GetBindu/Bindu

    Add a new self-contained example agent under examples/. An agent skill from GetBindu/Bindu.

    10k GitHub stars~1.1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes

More from langchain-ai/langchain-skills

All 22 skills in this repo
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Swarm Parallel Dispatch

    langchain-ai/langchain-skills

    Official

    Fans a list of independent items out to subagents in parallel, merges the results back into a table and supports retrying only the rows that failed.

    1.3k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Langgraph Human In The Loop

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when implementing human-in-the-loop patterns, pausing for approval, or handling errors in LangGraph.

    1.3k GitHub starsUsed in 1 repo~4.1k tokens
    Auto-check passed
  • Langgraph Persistence

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when your LangGraph needs to persist state, remember conversations, travel through history, or configure subgraph checkpointer scoping.

    1.3k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Deep Agents Core

    langchain-ai/langchain-skills

    Official

    Explains how to build agents with the Deep Agents framework: create_deep_agent, the built-in middleware, the harness, SKILL.md format and configuration options.

    1.3k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Deep Agents Orchestration

    langchain-ai/langchain-skills

    Official

    INVOKE THIS SKILL when using subagents, task planning, or human approval in Deep Agents.

    1.3k GitHub stars~3.4k tokensUpdated today
    Auto-check passed

Questions about LangGraph Decision Models

What does LangGraph Decision Models do?

Routes LangGraph agents with typed decision models that return probabilities, and finds LLM calls that only exist to produce a routing decision. A decision model answers typed questions about state and returns probabilities instead of prose, replacing the habit of prompting an LLM, parsing its text and branching on it. `Noul` asks a yes or no question and returns the probability of yes, `Choice` picks one label with the full distribution and a confidence, and `Score` grades against an ordered rubric and returns an expected value with a confidence.

When should I use LangGraph Decision Models?

LangGraph Decision Models fits situations like: replacing an LLM call that only decides which graph branch to take; auditing an agent for model calls whose output is parsed into a route; choosing probability thresholds for a classifier node; wiring a decision model through the LangSmith Gateway.

How do I install LangGraph Decision Models in Claude Code?

Run `npx skills add langchain-ai/langchain-skills --skill langgraph-decision-models -a claude-code`. Or copy the skill folder (config/skills/langgraph-decision-models in langchain-ai/langchain-skills) into .claude/skills/langgraph-decision-models in your project. Claude Code loads it when a task matches its description.

How do I install LangGraph Decision Models in Codex?

Run `npx skills add langchain-ai/langchain-skills --skill langgraph-decision-models -a codex`. Or copy the skill folder (config/skills/langgraph-decision-models in langchain-ai/langchain-skills) into .agents/skills/langgraph-decision-models in your project. Codex loads it when a task matches its description.

Can I use LangGraph Decision Models in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add langchain-ai/langchain-skills --skill langgraph-decision-models -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langgraph-decision-models, .gemini/skills/langgraph-decision-models, .github/skills/langgraph-decision-models and .opencode/skills/langgraph-decision-models in your project.

What does LangGraph Decision Models need to run?

Going by SKILL.md and its folder, LangGraph Decision Models needs the command-line tools its instructions call (uv) and credentials named TYPESAFE_API_KEY and LANGSMITH_API_KEY. Our summary lists: Python with the `langchain-typesafe` package; A model endpoint reachable with a base URL and API key.

Does LangGraph Decision Models access the network?

SKILL.md names 1 domain. In commands or code: gateway.smith.langchain.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is LangGraph Decision Models safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does LangGraph Decision Models use?

LangGraph Decision Models is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LangGraph Decision Models use?

About 2.3k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to LangGraph Decision Models?

Skills that share tags, products or a category with LangGraph Decision Models: Tool Design (agentailor/fullstack-langgraph-nextjs-agent, 132 stars), Deepagents Setup Configuration (soba-labs/langchain-agent-skills, 107 stars), Langgraph Testing Evaluation (soba-labs/langchain-agent-skills, 107 stars) and Langchain Observability (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LangGraph Decision Models?

langchain-ai (a GitHub organization, an official publisher) maintains it in langchain-ai/langchain-skills, which has 1,274 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 8, 2026.

Source: langchain-ai/langchain-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.