Agent skill

Deciding With Confidence

by oaustegard in oaustegard/claude-skills

Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds…

MITAuto-check passedEducation

Install Deciding With Confidence

skills CLI
$ npx skills add oaustegard/claude-skills --skill deciding-with-confidence -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills deciding-with-confidence --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/deciding-with-confidence .claude/skills/deciding-with-confidence && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deciding-with-confidence
GitHub stars
150
Token cost
~2.6k tokens
SKILL.md length
1,123 words
Files
10 (incl. scripts, references, assets)
Skills in repo
67
Repo updated
First seen
Licence
MIT

At a glance

Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds…

  • Works in 6 steps: Write the request as JSON in the OpenAI… → Pick the transport. --transport auto… → Run it. → …
  • Route this ticket
  • SKILL.md covers When NOT to use this skill, Procedure, Earned exceptions and Common failure modes, plus 2 more sections
  • Runs Python scripts from its folder; calls python3, pip and claude; needs ANTHROPIC_API_KEY and ANTHROPIC_AUTH_TOKEN

What it does

Deciding With Confidence is an agent skill from oaustegard/claude-skills. Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds (is this tool call grounded, should the agent ask first, is the customer angry), and how severe or urgent it is on a rubric, each with a calibrated confidence. Runs Claude Haiku in the request and response shapes of OpenAI's Decisions API, for when no purpose-built decision model is at hand. Use for "route this ticket"…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts, reference files and assets (for example `CHANGELOG.md`, `README.md` and `agents/decider.md`).

It sits in Education, covering Quizzes and assessments and Customer support. It works with OpenAI. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • Route this ticket
  • Classify with a confidence
  • Give me a probability that
  • Rate the severity

Example prompts

  • “route this ticket”
  • “triage these”
  • “classify with a confidence”
  • “/deciding-with-confidence”

Requirements

  • Python 3
  • A credential in ANTHROPIC_API_KEY
  • A credential in ANTHROPIC_AUTH_TOKEN

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Write the request as JSON in the OpenAI shape. Give every question a
  2. Pick the transport. --transport auto (the default) takes the first that works
  3. Run it.
  4. Read the answer. Predicates give probability (P true). Choices give
  5. Act on it with thresholds, never on the top answer alone.
  6. Calibrate on your own labels when you have 50 or more per question type

What it can do on your machine

Read from SKILL.md and the folder at commit 63d432e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY
    • ANTHROPIC_AUTH_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deciding With Confidence loads about 2.6k tokens when it runs, and up to ~4.1k if it reads all its reference files. Until then it costs about 186 tokens; SKILL.md has 1,123 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~186
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 63d432e, republished under its MIT licence (© oaustegard). 1,123 words, ~2,591 tokens.

Download SKILL.mdSave it as .claude/skills/deciding-with-confidence/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
deciding-with-confidence
description
Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds (is this tool call grounded, should the agent ask first, is the customer angry), and how severe or urgent it is on a rubric, each with a calibrated confidence. Runs Claude Haiku in the request and response shapes of OpenAI's Decisions API, for when no purpose-built decision model is at hand. Use for "route this ticket", "triage these", "classify with a confidence", "give me a probability that", "rate the severity", "score against this rubric", guardrail checks before a tool call, or "a Decisions API without OpenAI".
metadata.version
0.1.1

Deciding with confidence

scripts/decide.py takes an OpenAI Decisions API request (input plus questions of type predicate, choice or score) and returns its response shape: a probability per option, the chosen value or score, and confidence. Claude Haiku 5.5 answers. Its logits are not exposed, so the distribution is estimated: each of k parallel samples states a probability per option with the options in a different order, the samples are averaged, and the average is temperature-scaled from a labelled eval. Disagreement between samples lowers confidence and shows up in diagnostics.agreement.

Expect about 85% of what a purpose-built decision model gets: on the bundled 104-item eval, accuracy 0.846 against Jev's 0.894, with a calibration error (ECE 0.041) that matches it. The whole accuracy gap was on eight-way intent questions; yes/no and rubric questions tied. Disagreement between samples is the strongest warning sign. references/method.md has the numbers.

When NOT to use this skill

SituationUse instead
The label set is too large to list in a prompt (a taxonomy, hundreds of tags)hallucinating-labels
The answer needs reasoning, multi-step work or a written explanationa normal Claude call; this skill forbids explanation
You need extracted fields or free-form JSONstructured outputs on the Messages API
A real decision model is reachable (OpenAI /v1/decisions, Strands Decider, Jev) and latency mattersthat model: 0.1–0.3 s against ~0.7–4 s here
Picking which Claude tier runs a subagentagent-routing
The evidence is an imagenot supported; the transports send text only

Abandon a run when diagnostics.agreement is below 1 on most of a batch: the questions or option descriptions are ambiguous, and more samples will not fix them. Rewrite the options (distinct, observable criteria) before sampling more.

Procedure

  1. Write the request as JSON in the OpenAI shape. Give every question a unique name. Options need descriptions that say when each applies; add an other choice when the set is not exhaustive. Score levels go lowest first.

    json
    {"input": "I was charged twice for my order.",
     "questions": [
       {"type": "choice", "name": "department", "instructions": "Which department should handle this?",
        "choices": [{"value": "billing", "description": "Payments, invoices, refunds."},
                    {"value": "shipping", "description": "Delivery and tracking."},
                    {"value": "other", "description": "Anything else."}]},
       {"type": "predicate", "name": "angry", "instructions": "Is the customer expressing anger?"},
       {"type": "score", "name": "severity", "instructions": "How severe is this?",
        "levels": [{"label": "Cosmetic", "description": "No lost functionality."},
                   {"label": "Workaround", "description": "Fails, another way works."},
                   {"label": "Blocked", "description": "Fails with no workaround."}]}]}
  2. Pick the transport. --transport auto (the default) takes the first that works:

    TransportNeedsWall time per request
    apipip install anthropic and ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN~0.7 s expected (the API time measured inside claude -p)
    clithe claude CLI, logged in~4 s, measured
    bedrockpip install anthropic, AWS_REGION, AWS credentials; explicit onlyas api

    With none of them, use the subagent path in step 3b.

  3. Run it.

    a. With a transport:

    bash
    S=/mnt/skills/user/deciding-with-confidence/scripts   # or wherever the skill lives
    python3 $S/decide.py run request.json            # k=3 by default
    echo '{...}' | python3 $S/decide.py run -

    From Python: sys.path.insert(0, S); import decide; decide.decide(req) returns the response dict, or None when no sample was usable.

    b. In a Claude Code session, the plugin's decider subagent is the cheaper path: it is pinned to claude-haiku-5-5 (the model the calibration was fitted on) and carries its instructions inline, so it answers without a tool call. Install the deciding-with-confidence plugin, or copy agents/decider.md to .claude/agents/; agent definitions load at session start. Then:

    bash
    python3 $S/decide.py emit request.json -k 3 --agent deciding-with-confidence:decider > batch.json

    Use the name the Agent tool lists: deciding-with-confidence:decider from the plugin, decider from .claude/agents/. Send each samples[i].prompt verbatim as an Agent call with that subagent_type, all in one message. Collect the reply texts into a JSON list in replies.json, then python3 $S/decide.py aggregate batch.json replies.json. Pool before acting: a single reply carries no agreement signal and no calibration.

    Without the agent installed, omit --agent: the system prompt then rides inside each prompt for a stock general-purpose subagent with model: haiku. That works and costs about 58K tokens a sample against 9K.

  4. Read the answer. Predicates give probability (P true). Choices give choice, probabilities and confidence. Scores give score (the probability-weighted level index, so 1.4 sits between levels 1 and 2), probabilities and confidence. Every answer carries diagnostics: agreement, spread, the pre-calibration raw pool and per_sample.

  5. Act on it with thresholds, never on the top answer alone.

    • agreement < 1 (the samples' top answers differ): escalate. On the bundled eval, unanimous answers were 92% right and split ones 44%.
    • Top probability under 0.7: escalate. Together with the split rule that sent 23 of 104 items onward and left 6 wrong, assuming the escalation target got every escalated item right. The 0.7 was chosen on the same data, so treat it as a starting point.
    • Set the final thresholds from labelled examples of your own traffic, weighing the cost of a false positive against a false negative.
  6. Calibrate on your own labels when you have 50 or more per question type:

    bash
    python3 $S/decide.py eval my_labelled.jsonl --out results.jsonl   # {id, input, questions, labels}
    python3 $S/decide.py calibrate results.jsonl --out my-cal.json
    export DECIDER_CALIBRATION=$PWD/my-cal.json

    The bundled assets/calibration.json fits only choice (T = 0.62, fitted through the cli transport). Predicate and score keep T = 1 until a type has 50 labels, because smaller sets swung between folds.

Show full SKILL.md (375 more words)Show less

Earned exceptions

RuleBanned whenEarned when
k=3 sampleslatency or cost is the constraintalways the default: parallel samples cost no wall time, and k=1 loses the agreement signal
Bundled calibrationyour transport or domain differs from the eval's (cli, banking and support text)you have not yet labelled 50 items of your own
Escalate on agreement < 1nothing stronger sits behind ita stronger model or a human can take the split items

Common failure modes

  • decide: no usable sample; first error: TypeError: "Could not resolve authentication method..." → the api transport found no key. Set ANTHROPIC_API_KEY, or pass --transport cli.
  • decide: no transport → no key and no claude on PATH. Use step 3b.
  • diagnostics.samples below k → some replies did not parse. Usually a transport timeout under heavy concurrency; lower DECIDER_WORKERS (default 8) or raise DECIDER_TIMEOUT (default 60 s).
  • Every answer near uniform (confidence around 0.1–0.2 throughout) → options without descriptions, or descriptions that overlap. Rewrite them so each has distinct, observable criteria; more samples will not help.
  • A subagent reply with prose around the JSON → aggregate takes the outermost {...} and copes. A reply with no JSON at all counts as a failed sample.
  • Transcript text showing up as evidence in the subagent path → the [no-context] line was dropped from the prompt. Send the emitted prompt verbatim.

Verification

Run the tests, then one live request:

bash
python3 -m pytest /mnt/skills/user/deciding-with-confidence/tests -q    # 39 pass, no network
echo '{"input":"Nobody can log in since the 2pm deploy.","questions":[{"type":"predicate","name":"outage","instructions":"Is a service down for users?"}]}' \
  | python3 $S/decide.py run -

The live answer should have probability above 0.9 and diagnostics.samples equal to 3. A bad success is an answer with samples: 1 from a k=3 run: it looks normal but has lost the agreement signal. Check samples when a batch runs under load.

Diagnosed failures

  • 2026-10-09, first eval: 14 of 520 samples were discarded because Haiku left the options it ruled out of 8-way choices. Missing keys now share the probability mass the sample left unassigned, and the re-run lost none. The prompt also asks for ruled-out options at 0.001.
  • 2026-10-09: claude -p --json-schema doubled API time (2.2 s against 0.8 s) with a second turn and three times the output tokens. The cli transport sends the JSON contract in the system prompt instead.
  • 2026-10-09: the claude-workspace PreToolUse hook that appends parent-transcript chunks to Agent prompts would have put them in the evidence. emit appends [no-context], which that hook honours.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references, assets) in deciding-with-confidence of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • agents/decider.md
  • assets/calibration.json
  • assets/eval.jsonl
  • assets/system-prompt.md
  • references/method.md
  • scripts/decide.py
  • tests/test_decide.py

Open the folder on GitHubat commit 63d432e

Compare with similar skills

Deciding With Confidence next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deciding With Confidence compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deciding With Confidence this skilloaustegard/claude-skills150—~2.6kAutomated safety check: PassMIT
Woo AI Smokewoocommerce/woocommerce-ios358—~7.4kAutomated safety check: NotesGPL-2.0
Promptfoo Evaluationdaymade/claude-code-skills1.4k—~3kAutomated safety check: PassMIT
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch67k—~2kAutomated safety check: PassMIT
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch67k—~2.1kAutomated safety check: PassMIT
Claude Certification Tutorrohitg00/ai-engineering-from-scratch67k—~3kAutomated safety check: PassMIT

Similar skills

  • Woo AI Smoke

    woocommerce/woocommerce-ios

    Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

    358 GitHub stars~7.4k tokensUpdated yesterday
    EducationAuto-check: notes
  • Promptfoo Evaluation

    daymade/claude-code-skills

    Configures and runs LLM evaluation using Promptfoo framework.

    1.4k GitHub stars~3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    67k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    67k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Claude Certification Tutor

    rohitg00/ai-engineering-from-scratch

    Guides a learner through one of four independent Claude certification tracks with onboarding, lessons, practice labs, mock exams and remediation.

    67k GitHub stars~3k tokensUpdated today
    EducationAuto-check passed
  • Generate Verifiers Env

    adithya-s-k/FineEnvs

    Builds a Verifiers (PrimeIntellect) variant of an RL environment.

    461 GitHub starsUsed in 1 repo~2.3k tokens
    EducationAuto-check passed

More from oaustegard/claude-skills

All 67 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Review Before Shipping

    oaustegard/claude-skills

    Has a fresh-context adversary attack a blog post, recommendation, analysis brief or piece of code before you ship it, using a profile suited to that kind of artifact.

    150 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Deciding With Confidence

What does Deciding With Confidence do?

Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds…. Deciding With Confidence is an agent skill from oaustegard/claude-skills. Routes, triages, flags and rates a piece of text with a probability for every option: which department or queue a ticket goes to, which intent a message expresses, whether a yes/no condition holds (is this tool call grounded, should the agent ask first, is the customer angry), and how severe or urgent it is on a rubric, each with a calibrated confidence.

When should I use Deciding With Confidence?

Deciding With Confidence fits situations like: route this ticket; classify with a confidence; give me a probability that; rate the severity.

How do I install Deciding With Confidence in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill deciding-with-confidence -a claude-code`. Or copy the skill folder (deciding-with-confidence in oaustegard/claude-skills) into .claude/skills/deciding-with-confidence in your project. Claude Code loads it when a task matches its description.

How do I install Deciding With Confidence in Codex?

Run `npx skills add oaustegard/claude-skills --skill deciding-with-confidence -a codex`. Or copy the skill folder (deciding-with-confidence in oaustegard/claude-skills) into .agents/skills/deciding-with-confidence in your project. Codex loads it when a task matches its description.

Can I use Deciding With Confidence in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill deciding-with-confidence -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deciding-with-confidence, .gemini/skills/deciding-with-confidence, .github/skills/deciding-with-confidence and .opencode/skills/deciding-with-confidence in your project.

What does Deciding With Confidence need to run?

Going by SKILL.md and its folder, Deciding With Confidence needs Python for the scripts in its folder, the command-line tools its instructions call (python3, pip and claude) and credentials named ANTHROPIC_API_KEY and ANTHROPIC_AUTH_TOKEN. Our summary lists: Python 3; A credential in ANTHROPIC_API_KEY; A credential in ANTHROPIC_AUTH_TOKEN.

Does Deciding With Confidence access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Deciding With Confidence safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Deciding With Confidence use?

Deciding With Confidence is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deciding With Confidence use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Deciding With Confidence?

Skills that share tags, products or a category with Deciding With Confidence: Woo AI Smoke (woocommerce/woocommerce-ios, 358 stars), Promptfoo Evaluation (daymade/claude-code-skills, 1.4k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 67k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 67k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deciding With Confidence?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 67 skills in this directory. The repository was last updated on October 10, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.