---
name: deciding-with-confidence
description: "Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds (is this tool call grounded, should the agent ask first, is the customer angry), and how severe or urgent it is on a rubric, each with a calibrated confidence. Runs Claude Haiku in the request and response shapes of OpenAI's Decisions API, for when no purpose-built decision model is at hand. Use for \"route this ticket\", \"triage these\", \"classify with a confidence\", \"give me a probability that\", \"rate the severity\", \"score against this rubric\", guardrail checks before a tool call, or \"a Decisions API without OpenAI\"."
metadata:
  version: 0.1.1
---

# Deciding with confidence

`scripts/decide.py` takes an OpenAI Decisions API request (`input` plus
`questions` of type `predicate`, `choice` or `score`) and returns its response
shape: a probability per option, the chosen value or score, and `confidence`.
Claude Haiku 5.5 answers. Its logits are not exposed, so the distribution is
estimated: each of k parallel samples states a probability per option with the
options in a different order, the samples are averaged, and the average is
temperature-scaled from a labelled eval. Disagreement between samples lowers
`confidence` and shows up in `diagnostics.agreement`.

Expect about 85% of what a purpose-built decision model gets: on the bundled
104-item eval, accuracy 0.846 against Jev's 0.894, with a calibration error
(ECE 0.041) that matches it. The whole accuracy gap was on eight-way intent
questions; yes/no and rubric questions tied. Disagreement between samples is
the strongest warning sign. `references/method.md` has the numbers.

## When NOT to use this skill

| Situation | Use instead |
|---|---|
| The label set is too large to list in a prompt (a taxonomy, hundreds of tags) | `hallucinating-labels` |
| The answer needs reasoning, multi-step work or a written explanation | a normal Claude call; this skill forbids explanation |
| You need extracted fields or free-form JSON | structured outputs on the Messages API |
| A real decision model is reachable (OpenAI `/v1/decisions`, Strands Decider, Jev) and latency matters | that model: 0.1–0.3 s against ~0.7–4 s here |
| Picking which Claude tier runs a subagent | `agent-routing` |
| The evidence is an image | not supported; the transports send text only |

Abandon a run when `diagnostics.agreement` is below 1 on most of a batch: the
questions or option descriptions are ambiguous, and more samples will not fix
them. Rewrite the options (distinct, observable criteria) before sampling more.

## Procedure

1. **Write the request** as JSON in the OpenAI shape. Give every question a
   unique `name`. Options need `description`s that say when each applies; add
   an `other` choice when the set is not exhaustive. Score `levels` go lowest
   first.

   ```json
   {"input": "I was charged twice for my order.",
    "questions": [
      {"type": "choice", "name": "department", "instructions": "Which department should handle this?",
       "choices": [{"value": "billing", "description": "Payments, invoices, refunds."},
                   {"value": "shipping", "description": "Delivery and tracking."},
                   {"value": "other", "description": "Anything else."}]},
      {"type": "predicate", "name": "angry", "instructions": "Is the customer expressing anger?"},
      {"type": "score", "name": "severity", "instructions": "How severe is this?",
       "levels": [{"label": "Cosmetic", "description": "No lost functionality."},
                  {"label": "Workaround", "description": "Fails, another way works."},
                  {"label": "Blocked", "description": "Fails with no workaround."}]}]}
   ```

2. **Pick the transport.** `--transport auto` (the default) takes the first that works:

   | Transport | Needs | Wall time per request |
   |---|---|---|
   | `api` | `pip install anthropic` and `ANTHROPIC_API_KEY` or `ANTHROPIC_AUTH_TOKEN` | ~0.7 s expected (the API time measured inside `claude -p`) |
   | `cli` | the `claude` CLI, logged in | ~4 s, measured |
   | `bedrock` | `pip install anthropic`, `AWS_REGION`, AWS credentials; explicit only | as `api` |

   With none of them, use the subagent path in step 3b.

3. **Run it.**

   a. With a transport:
   ```bash
   S=/mnt/skills/user/deciding-with-confidence/scripts   # or wherever the skill lives
   python3 $S/decide.py run request.json            # k=3 by default
   echo '{...}' | python3 $S/decide.py run -
   ```
   From Python: `sys.path.insert(0, S); import decide; decide.decide(req)` returns
   the response dict, or None when no sample was usable.

   b. In a Claude Code session, the plugin's `decider` subagent is the
   cheaper path: it is pinned to `claude-haiku-5-5` (the model the calibration
   was fitted on) and carries its instructions inline, so it answers without a
   tool call. Install the `deciding-with-confidence` plugin, or copy
   `agents/decider.md` to `.claude/agents/`; agent definitions load at session
   start. Then:
   ```bash
   python3 $S/decide.py emit request.json -k 3 --agent deciding-with-confidence:decider > batch.json
   ```
   Use the name the Agent tool lists: `deciding-with-confidence:decider` from
   the plugin, `decider` from `.claude/agents/`. Send each `samples[i].prompt`
   verbatim as an Agent call with that `subagent_type`, all in one message.
   Collect the reply texts into a JSON list in `replies.json`, then
   `python3 $S/decide.py aggregate batch.json replies.json`. Pool before acting:
   a single reply carries no agreement signal and no calibration.

   Without the agent installed, omit `--agent`: the system prompt then rides
   inside each prompt for a stock general-purpose subagent with
   `model: haiku`. That works and costs about 58K tokens a sample against 9K.

4. **Read the answer.** Predicates give `probability` (P true). Choices give
   `choice`, `probabilities` and `confidence`. Scores give `score` (the
   probability-weighted level index, so 1.4 sits between levels 1 and 2),
   `probabilities` and `confidence`. Every answer carries `diagnostics`:
   `agreement`, `spread`, the pre-calibration `raw` pool and `per_sample`.

5. **Act on it with thresholds, never on the top answer alone.**
   - `agreement < 1` (the samples' top answers differ): escalate. On the bundled
     eval, unanimous answers were 92% right and split ones 44%.
   - Top probability under 0.7: escalate. Together with the split rule that
     sent 23 of 104 items onward and left 6 wrong, assuming the escalation
     target got every escalated item right. The 0.7 was chosen on the same
     data, so treat it as a starting point.
   - Set the final thresholds from labelled examples of your own traffic,
     weighing the cost of a false positive against a false negative.

6. **Calibrate on your own labels** when you have 50 or more per question type:
   ```bash
   python3 $S/decide.py eval my_labelled.jsonl --out results.jsonl   # {id, input, questions, labels}
   python3 $S/decide.py calibrate results.jsonl --out my-cal.json
   export DECIDER_CALIBRATION=$PWD/my-cal.json
   ```
   The bundled `assets/calibration.json` fits only `choice` (T = 0.62, fitted
   through the `cli` transport). Predicate and score keep T = 1 until a type
   has 50 labels, because smaller sets swung between folds.

## Earned exceptions

| Rule | Banned when | Earned when |
|---|---|---|
| k=3 samples | latency or cost is the constraint | always the default: parallel samples cost no wall time, and k=1 loses the agreement signal |
| Bundled calibration | your transport or domain differs from the eval's (`cli`, banking and support text) | you have not yet labelled 50 items of your own |
| Escalate on `agreement < 1` | nothing stronger sits behind it | a stronger model or a human can take the split items |

## Common failure modes

- **`decide: no usable sample; first error: TypeError: "Could not resolve
  authentication method..."`** → the `api` transport found no key. Set
  `ANTHROPIC_API_KEY`, or pass `--transport cli`.
- **`decide: no transport`** → no key and no `claude` on PATH. Use step 3b.
- **`diagnostics.samples` below k** → some replies did not parse. Usually a
  transport timeout under heavy concurrency; lower `DECIDER_WORKERS` (default 8)
  or raise `DECIDER_TIMEOUT` (default 60 s).
- **Every answer near uniform** (confidence around 0.1–0.2 throughout) →
  options without descriptions, or descriptions that overlap. Rewrite them so
  each has distinct, observable criteria; more samples will not help.
- **A subagent reply with prose around the JSON** → `aggregate` takes the
  outermost `{...}` and copes. A reply with no JSON at all counts as a failed
  sample.
- **Transcript text showing up as evidence in the subagent path** → the
  `[no-context]` line was dropped from the prompt. Send the emitted prompt
  verbatim.

## Verification

Run the tests, then one live request:

```bash
python3 -m pytest /mnt/skills/user/deciding-with-confidence/tests -q    # 39 pass, no network
echo '{"input":"Nobody can log in since the 2pm deploy.","questions":[{"type":"predicate","name":"outage","instructions":"Is a service down for users?"}]}' \
  | python3 $S/decide.py run -
```

The live answer should have `probability` above 0.9 and `diagnostics.samples`
equal to 3. A **bad success** is an answer with `samples: 1` from a k=3 run:
it looks normal but has lost the agreement signal. Check `samples` when a
batch runs under load.

## Diagnosed failures

- **2026-10-09, first eval:** 14 of 520 samples were discarded because Haiku
  left the options it ruled out of 8-way choices. Missing keys now share the
  probability mass the sample left unassigned, and the re-run lost none. The
  prompt also asks for ruled-out options at 0.001.
- **2026-10-09:** `claude -p --json-schema` doubled API time (2.2 s against
  0.8 s) with a second turn and three times the output tokens. The `cli`
  transport sends the JSON contract in the system prompt instead.
- **2026-10-09:** the claude-workspace PreToolUse hook that appends
  parent-transcript chunks to Agent prompts would have put them in the
  evidence. `emit` appends `[no-context]`, which that hook honours.
