Agent skill

Hallucinating Labels

by oaustegard in oaustegard/claude-skills

Assign items to a CLOSED label vocabulary that is too large to put in a prompt — product taxonomies, category hierarchies, tag vocabularies, routing tables, ICD/SIC-style code lists.

MITAuto-check passedAI & LLM Engineering

Install Hallucinating Labels

skills CLI
$ npx skills add oaustegard/claude-skills --skill hallucinating-labels -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills hallucinating-labels --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/hallucinating-labels .claude/skills/hallucinating-labels && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hallucinating-labels
GitHub stars
150
Token cost
~2.8k tokens
SKILL.md length
1,548 words
Files
5 (incl. scripts)
Skills in repo
66
Repo updated
First seen
Licence
MIT

At a glance

Assign items to a CLOSED label vocabulary that is too large to put in a prompt — product taxonomies, category hierarchies, tag vocabularies, routing tables, ICD/SIC-style code lists.

  • Classify these into our taxonomy
  • SKILL.md covers Check the boundary first, Procedure, Anchor the prompt on REGISTER,… and Batch 40 per call. Never one…, plus 5 more sections
  • Runs Python scripts from its folder; calls python3
  • Tag these against the existing tag list

What it does

Hallucinating Labels is an agent skill from oaustegard/claude-skills. Assign items to a CLOSED label vocabulary that is too large to put in a prompt — product taxonomies, category hierarchies, tag vocabularies, routing tables, ICD/SIC-style code lists. A cheap model writes the label it thinks the vocabulary would use, and an embedder snaps that writing onto the nearest legal value, so the schema is never transmitted and the output is always in-vocabulary. Use for "classify these into our taxonomy", "tag these against the existing tag list", "map these queries to categories", "the…

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts (for example `CHANGELOG.md`, `README.md` and `scripts/snap.py`).

It sits in AI & LLM Engineering, covering Structured output and tool calling. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • Classify these into our taxonomy
  • Tag these against the existing tag list
  • Map these queries to categories
  • The enum is too big to send

Example prompts

  • “classify these into our taxonomy”
  • “tag these against the existing tag list”
  • “map these queries to categories”
  • “/hallucinating-labels”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit cf49d47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • softwaredoug.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hallucinating Labels loads about 2.8k tokens when it runs. Until then it costs about 209 tokens; SKILL.md has 1,548 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~209
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit cf49d47, republished under its MIT licence (© oaustegard). 1,548 words, ~2,807 tokens.

Download SKILL.mdSave it as .claude/skills/hallucinating-labels/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
hallucinating-labels
description
Assign items to a CLOSED label vocabulary that is too large to put in a prompt — product taxonomies, category hierarchies, tag vocabularies, routing tables, ICD/SIC-style code lists. A cheap model writes the label it thinks the vocabulary would use, and an embedder snaps that writing onto the nearest legal value, so the schema is never transmitted and the output is always in-vocabulary. Use for "classify these into our taxonomy", "tag these against the existing tag list", "map these queries to categories", "the enum is too big to send", or a Literal/enum that hits a provider cap. NOT for a vocabulary that fits in a prompt — structured output measured 0.701 acc@1 there against this pattern's 0.564. NOT for open-ended labelling with no fixed vocabulary, and not for ranked retrieval over documents (bm25).
metadata.version
0.1.0

hallucinating-labels

Ask a cheap model to write a plausible label for the item. Snap that label onto the real vocabulary with an embedder. The model never sees the label set.

Doug Turnbull's pattern (softwaredoug.com, 2026-08-10), with the two prompt and boundary corrections that measurement produced.

Check the boundary first

If the whole vocabulary fits in a prompt, do not use this skill. Ship the label list and ask for a constrained choice. Measured on WANDS (860 labels, 468 queries, one gold label each, gemini-3.5-flash-lite):

approachacc@1acc@3input tokens/item
structured output, all 860 labels shipped0.7010.7445,265
this skill0.5640.6906
embed the item directly, no model0.4170.5640

Shipping the vocabulary is 14 points more accurate and 880× more expensive. Take the accuracy unless the tokens are the problem. The tokens are the problem when the vocabulary does not fit, when a provider enum cap rejects it, or when per-call cost at volume dominates — a 5,000-label vocabulary is roughly 30k tokens on every single call.

This skill still beats every model-free baseline by a wide margin, so it is the right tool whenever shipping the vocabulary is off the table.

Procedure

1. Write the vocabulary to a file, one label per line, and index it once.

bash
python3 scripts/snap.py build --vocab categories.txt --out .snap-index.pkl

Default backend is tfidf — sklearn only, no download. Pass --backend minilm when sentence-transformers and a ~90 MB download are available and the items share no wording with the labels; it scored 0.564 to tfidf's 0.528 on WANDS. Where items literally contain their own label words, tfidf wins outright (0.416 vs 0.356 on a memory-tag corpus).

2. Write the labels yourself, in batches of 40, using the register prompt below. Write them to a file, one per line, in the same order as the items.

3. Snap.

bash
python3 scripts/snap.py snap --index .snap-index.pkl --labels written.txt --k 3

Add --min-score 0.35 to get null instead of a bad snap, and --items items.txt --union for long items (see below). Output is JSON with the top-k legal labels per item.

4. Report the nulls and the low scores. A snap at cosine 0.18 is noise wearing a legal label. Never present one as a classification.

Anchor the prompt on REGISTER, not on novelty

This is the correction that matters most, and it is the opposite of what the source post's prompt says. Its prompt opens "create a novel, never-seen-before classification". That instruction is safe only with a model too weak to follow it.

Measured on the same 40 WANDS queries, MiniLM backend:

promptmodelacc@1acc@3
embed the query directly, no model—0.5000.650
"novel, never-seen-before"gemini-3.5-flash-lite0.5750.675
"novel, never-seen-before"Haiku 4.5 subagent0.1000.275
register-anchored (below)Haiku 4.5 subagent0.5250.750

Haiku obeyed. Asked for novelty it produced novelty — Hydraulic Styling Thrones, Weathered Branch-Frame Reflectors, Chromatic Comfort Accents — and scored a fifth of what doing nothing scores. Gemini flash-lite half-ignored the same instruction and wrote Salon & Styling Chairs, Rustic Wall Mirrors, which is what the snap needs. The pattern wants a novel instance in the vocabulary's register, and "never-seen-before" asks for novel wording. A better instruction-follower is worse at the badly-worded prompt.

The register prompt also beat the novelty prompt on Gemini across all 468 queries (0.564 vs 0.489 acc@1), so it is strictly the better wording. Use this:

You are writing entries for a {DOMAIN} vocabulary.

For each item below, write the label that this vocabulary WOULD file that item under. Write it the way the vocabulary writes labels — match the examples' register, length and wording exactly.

Do not worry about whether the label already exists. Write the obvious one. Do not invent novel or creative wording, do not use marketing adjectives, do not hedge, do not explain.

Examples of the register: {6-8 REAL LABELS FROM THE VOCABULARY}

Output one line per item, in the same order, formatted exactly as: <n>. <label>

ITEMS: {NUMBERED ITEMS}

The examples are load-bearing — they are how the register gets communicated. Draw 6-8 real labels from the vocabulary. They are not the vocabulary; sending eight labels is not sending five thousand.

Batch 40 per call. Never one item per call

Batching is free: 0.496/0.641 batched ×40 against 0.489/0.613 unbatched, at 1/17 the input tokens and 1/9 the wall-clock.

Write the labels yourself when the items are already in context — you are the cheap model here, and it costs one short generation. Delegate to a Haiku subagent only in batches, and only when the item list is long enough to be worth it. A subagent invocation carries a measured floor of ~32,500 tokens before it reads your prompt: a general-purpose Haiku subagent asked to output the single word ok, with zero tool calls, spent 32,539. Per item that floor is 813 tokens at batch 40 and 32,500 at batch 1.

Parse the model's numbered reply back by index, not by zipping positionally. When the model drops item 2 of 40, zipping shifts every later item onto its neighbour's label and nothing signals it. A dropped item is an empty label and then a null.

Show full SKILL.md (721 more words)Show less

Long items want --union, not a different pattern

The written label replaces direct embedding cleanly when item and label are the same kind of string — a WANDS query and a WANDS category are both short noun phrases. When the item is a 1,500-character document and the label is one word, the written label throws away most of the document, and the direct embedding still has it. The two are complementary.

Memory store, 1,273 tags, 250 documents of 300-2000 characters, mean 4.8 gold tags, tfidf:

arm@1@3@5
embed the document directly0.4160.6280.712
write 5 tags, novelty prompt0.2080.3520.424
write 5 tags, register prompt0.5080.7000.792
both, interleaved (--union)0.6720.8520.888

Note the middle two rows. With the wrong prompt this corpus says the pattern loses to doing nothing by 2x; with the right one it wins, and the union wins by a lot more. Long items amplify the register error rather than causing a separate problem — a distinctive vocabulary is exactly where novel wording lands furthest from anything legal.

Rule: item and label share a register → write labels and snap them. Item is a long document → do that and pass --items ... --union. Neither case is a reason to reach for the novelty prompt.

Picking the encoder, and why there is no local-LLM version

The encoder is the whole system when you cannot reach an API. Snapping the raw query, no model call anywhere, full WANDS set:

encoderint8 ONNXacc@1acc@3
all-MiniLM-L6-v223 MB0.4170.564
bge-small-en-v1.533 MB0.4270.583
gte-small33 MB0.4550.594
bge-base-en-v1.5109 MB0.4620.630

gte-small is the knee. bge-base buys +0.007 acc@1 for 3.3x the download.

Do not substitute a tiny local model for the label-writing half. Pleias Monad (57M) and Baguettotron (321M) both have onnx-community builds — 35 MB and 236 MB at q4f16 — so a wholly client-side pipeline packages fine. Neither earns its bytes. As writers they score 0.425 and 0.400 acc@1 against a 0.500 no-model control on the same 40 queries: they echo the query (smart coffee table → Smart coffee table) and bleed from the few-shot exemplars (chair and a half recliner → Chair & Recycling Bins). As likelihood rerankers over the encoder's top-10 — which asks them for no format compliance at all — they score 0.325 and 0.350 against the same 0.500, with the gold label present in that top-10 for 82.5% of queries.

The reason is the same one that makes the register prompt matter: what the cheap model contributes here is not reasoning but a prior over how taxonomies name things, learned from web-scale pretraining. A small model trained for reasoning has no retail-taxonomy prior, and reasoning does not substitute for one. In a browser, ship the encoder alone.

Failure modes

signalcausefix
Snapped labels are wrong but confident; written labels read like ad copynovelty-anchored prompt, obeyedswitch to the register prompt; read the written labels before blaming the snap — Hydraulic Styling Thrones is a prompt bug, not an embedder bug
Everything snaps to the same one or two labelsthe register examples are unrepresentative, or the vocabulary has one dominant stringdraw examples spanning the vocabulary's breadth
Scores cluster near 0.15items and labels share no surface wording--backend minilm, and --union if items are long
Item n onward all shifted by onepositional zipping of a reply with a dropped lineparse by the emitted index; verify counts match before snapping
Accuracy below the no-model controlnovelty-anchored prompt, or a long item without --unionfix the prompt first — it cost 30 points on one corpus and 7.5 on another; then add --union

Verify it beat doing nothing

Run the no-model control before shipping this anywhere. Snap the items directly (--labels items.txt, no written labels) and compare. On one of the two corpora measured here the control won by 2x under the wrong prompt. If you have no gold labels to score against, hand-check 20 items both ways — the control is one command and its absence is how this pattern gets adopted where it loses.

  • bm25 — ranked retrieval over documents. Different problem: no closed label set.
  • agent-routing — routing to a small named set, which fits in a prompt. Ship it instead.
  • muninn_utils.hypothetical_classifier — the same pattern as a Python API with Gemini wired in, for Muninn sessions.

Experiment, arms and artifacts: oaustegard/experiments/hypothetical-classification.

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in hallucinating-labels of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • scripts/snap.py
  • tests/test_snap.py

Open the folder on GitHubat commit cf49d47

Compare with similar skills

Hallucinating Labels next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hallucinating Labels compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hallucinating Labels this skilloaustegard/claude-skills150—~2.8kAutomated safety check: PassMIT
Planning With Filesjarrodwatts/claude-code-config1.1k5 repos~967Automated safety check: PassNone
Tool Use Data Synthesissunny-glow/Auto-BenchMax1.3k—~3.3kAutomated safety check: PassNone
Agent Harness ConstructionKartikLabhshetwar/mind-mentor1486 repos~500Automated safety check: PassApache-2.0
Prompt Engineering Patternswshobson/agents40k—~1.3kAutomated safety check: PassMIT
Model Benchmarkstheopenco/llmgateway1.7k—~1.1kAutomated safety check: NotesCustom licence

Similar skills

  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed
  • Tool Use Data Synthesis

    sunny-glow/Auto-BenchMax

    Synthesize training data for ANY tool-use / agentic benchmark, in ANY repo.

    1.3k GitHub stars~3.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agent Harness Construction

    KartikLabhshetwar/mind-mentor

    Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates.

    148 GitHub starsUsed in 6 repos~500 tokens
    AI & LLM EngineeringAuto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Model Benchmarks

    theopenco/llmgateway

    Run and report repository model or provider-mapping benchmarks.

    1.7k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Agent Prompt Quality Bar

    mastra-ai/mastra

    Universal quality bar and final audit rubric for any agent system prompt.

    29k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from oaustegard/claude-skills

All 66 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Review Before Shipping

    oaustegard/claude-skills

    Has a fresh-context adversary attack a blog post, recommendation, analysis brief or piece of code before you ship it, using a profile suited to that kind of artifact.

    150 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed

Questions about Hallucinating Labels

What does Hallucinating Labels do?

Assign items to a CLOSED label vocabulary that is too large to put in a prompt — product taxonomies, category hierarchies, tag vocabularies, routing tables, ICD/SIC-style code lists. Hallucinating Labels is an agent skill from oaustegard/claude-skills. Assign items to a CLOSED label vocabulary that is too large to put in a prompt — product taxonomies, category hierarchies, tag vocabularies, routing tables, ICD/SIC-style code lists.

When should I use Hallucinating Labels?

Hallucinating Labels fits situations like: classify these into our taxonomy; tag these against the existing tag list; map these queries to categories; the enum is too big to send.

How do I install Hallucinating Labels in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill hallucinating-labels -a claude-code`. Or copy the skill folder (hallucinating-labels in oaustegard/claude-skills) into .claude/skills/hallucinating-labels in your project. Claude Code loads it when a task matches its description.

How do I install Hallucinating Labels in Codex?

Run `npx skills add oaustegard/claude-skills --skill hallucinating-labels -a codex`. Or copy the skill folder (hallucinating-labels in oaustegard/claude-skills) into .agents/skills/hallucinating-labels in your project. Codex loads it when a task matches its description.

Can I use Hallucinating Labels in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill hallucinating-labels -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hallucinating-labels, .gemini/skills/hallucinating-labels, .github/skills/hallucinating-labels and .opencode/skills/hallucinating-labels in your project.

What does Hallucinating Labels need to run?

Going by SKILL.md and its folder, Hallucinating Labels needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Hallucinating Labels access the network?

SKILL.md names 1 domain. As links in the text: softwaredoug.com. This is read from the text; nothing was executed.

Is Hallucinating Labels safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Hallucinating Labels use?

Hallucinating Labels is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hallucinating Labels use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hallucinating Labels?

Skills that share tags, products or a category with Hallucinating Labels: Planning With Files (jarrodwatts/claude-code-config, 1.1k stars), Tool Use Data Synthesis (sunny-glow/Auto-BenchMax, 1.3k stars), Agent Harness Construction (KartikLabhshetwar/mind-mentor, 148 stars) and Prompt Engineering Patterns (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hallucinating Labels?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 66 skills in this directory. The repository was last updated on October 8, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.